Chengyue Jiang

dblp:256/1067 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
8since 2021 · last 2024
0000-0002-5665-9774ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Information extraction and text analysis · 47% Language models and text generation · 18% Trustworthy machine learning · 15%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
instruction tuning
1.522024
SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence Understanding · AAAI 2024
EcomGPT: Instruction-Tuning Large Language Models with Chain-of-Task Tasks for E-commerce · AAAI 2024
Natural language and speech › Information extraction and text analysis
entity typing
1.532024
Recall, Expand, and Multi-Candidate Cross-Encode: Fast and Accurate Ultra-Fine Entity Typing · ACL (1) 2023
Modeling Label Correlations for Ultra-Fine Entity Typing with Neural Pairwise Conditional Random Field · EMNLP 2022
SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence Understanding · AAAI 2024
Natural language and speech › Information extraction and text analysis › entity typing
ultra-fine entity typing
1.222023
Recall, Expand, and Multi-Candidate Cross-Encode: Fast and Accurate Ultra-Fine Entity Typing · ACL (1) 2023
Modeling Label Correlations for Ultra-Fine Entity Typing with Neural Pairwise Conditional Random Field · EMNLP 2022
Natural language and speech › Information extraction and text analysis
text classification
1.122023
Using Interpretation Methods for Model Enhancement · EMNLP 2023
Cold-Start and Interpretability: Turning Regular Expressions into Trainable Recurrent Neural Networks · EMNLP (1) 2020
Natural language and speech › Language models and text generation
large language model
0.812024
SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence Understanding · AAAI 2024
Machine learning › Trustworthy machine learning
interpretability
0.712023
Using Interpretation Methods for Model Enhancement · EMNLP 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology
ontological knowledge
0.712023
Do PLMs Know and Understand Ontological Knowledge? · ACL (1) 2023
Machine learning › Trustworthy machine learning › language model interpretability
pretrained language model probing
0.712023
Do PLMs Know and Understand Ontological Knowledge? · ACL (1) 2023
Machine learning › Trustworthy machine learning
rationale-based training
0.712023
Using Interpretation Methods for Model Enhancement · EMNLP 2023
Information retrieval
ranking
0.712023
Recall, Expand, and Multi-Candidate Cross-Encode: Fast and Accurate Ultra-Fine Entity Typing · ACL (1) 2023
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
conditional random field
0.612022
Modeling Label Correlations for Ultra-Fine Entity Typing with Neural Pairwise Conditional Random Field · EMNLP 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
neural-symbolic integration
0.512021
Neuralizing Regular Expressions for Slot Filling · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis
slot filling
0.512021
Neuralizing Regular Expressions for Slot Filling · EMNLP (1) 2021
Machine learning › Deep learning architectures and training
recurrent neural network
0.412020
Cold-Start and Interpretability: Turning Regular Expressions into Trainable Recurrent Neural Networks · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis
event extraction
0.212024
SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence Understanding · AAAI 2024
Machine learning › Transfer learning and domain adaptation
zero-shot transfer
0.212024
EcomGPT: Instruction-Tuning Large Language Models with Chain-of-Task Tasks for E-commerce · AAAI 2024
Machine learning › Transfer learning and domain adaptation
low-resource learning
0.212023
Using Interpretation Methods for Model Enhancement · EMNLP 2023
Logic in computer science › knowledge representation and reasoning
ontology entailment
0.212023
Do PLMs Know and Understand Ontological Knowledge? · ACL (1) 2023

Methods — techniques the papers use, named apart from their topics

instruction tuning · 1.5recall-expand-filter · 1.3probing · 1.3multi-candidate cross-encoding · 1.3logical reasoning evaluation · 1.3cross-encoder · 1.3recurrent neural network · 0.9data synthesis · 0.8chain-of-tasks · 0.8autoregressive model · 0.8
YearPublicationVenuePosition
2024 EcomGPT: Instruction-Tuning Large Language Models with Chain-of-Task Tasks for E-commerce
abstract
Recently, instruction-following Large Language Models (LLMs) , represented by ChatGPT, have exhibited exceptional performance in general Natural Language Processing (NLP) tasks. However, the unique characteristics of E-commerce data pose significant challenges to general LLMs. An LLM tailored specifically for E-commerce scenarios, possessing robust cross-dataset/task generalization capabilities, is a pressing necessity. To solve this issue, in this work, we proposed the first E-commerce instruction dataset EcomInstruct, with a total of 2.5 million instruction data. EcomInstruct scales up the data size and task diversity by constructing atomic tasks with E-commerce basic data types, such as product information, user reviews. Atomic tasks are defined as intermediate tasks implicitly involved in solving a final task, which we also call Chain-of-Task tasks. We developed EcomGPT with different parameter scales by training the backbone model BLOOMZ with the EcomInstruct. Benefiting from the fundamental semantic understanding capabilities acquired from the Chain-of-Task tasks, EcomGPT exhibits excellent zero-shot generalization capabilities. Extensive experiments and human evaluations demonstrate that EcomGPT outperforms ChatGPT in term of cross-dataset/task generalization on E-commerce tasks. The EcomGPT will be public at https://github.com/Alibaba-NLP/EcomGPT.
Yangning Li, Shirong Ma, Xiaobin Wang, Shen Huang, Chengyue Jiang, Hai-Tao Zheng 0002, Pengjun Xie, Fei Huang 0002, Yong Jiang 0005
AAAI5
2024 SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence Understanding
abstract
Large language models (LLMs) have shown impressive abilities for open-domain NLP tasks. However, LLMs are sometimes too footloose for natural language understanding (NLU) tasks which always have restricted output and input format. Their performances on NLU tasks are highly related to prompts or demonstrations and are shown to be poor at performing several representative NLU tasks, such as event extraction and entity typing. To this end, we present SeqGPT, a bilingual (i.e., English and Chinese) open-source autoregressive model specially enhanced for open-domain natural language understanding. We express all NLU tasks with two atomic tasks, which define fixed instructions to restrict the input and output format but still ``open'' for arbitrarily varied label sets. The model is first instruction-tuned with extremely fine-grained labeled data synthesized by ChatGPT and then further fine-tuned by 233 different atomic tasks from 152 datasets across various domains. The experimental results show that SeqGPT has decent classification and extraction ability, and is capable of performing language understanding tasks on unseen domains. We also conduct empirical studies on the scaling of data and model size as well as on the transfer across tasks. Our models are accessible at https://github.com/Alibaba-NLP/SeqGPT.
Tianyu Yu 0002, Chengyue Jiang, Chao Lou, Shen Huang, Xiaobin Wang, Wei Liu 0131, Jiong Cai, Yangning Li, Kewei Tu, Hai-Tao Zheng 0002, Ningyu Zhang 0001, Pengjun Xie, Fei Huang 0002, Yong Jiang 0005
AAAI2
2023 Recall, Expand, and Multi-Candidate Cross-Encode: Fast and Accurate Ultra-Fine Entity Typing
abstract
Ultra-fine entity typing (UFET) predicts extremely free-formed types (e.g., president, politician) of a given entity mention (e.g., Joe Biden) in context.State-of-the-art (SOTA) methods use the cross-encoder (CE) based architecture.CE concatenates a mention (and its context) with each type and feeds the pair into a pretrained language model (PLM) to score their relevance.It brings deeper interaction between the mention and the type to reach better performance but has to perform N (the type set size) forward passes to infer all the types of a single mention.CE is therefore very slow in inference when the type set is large (e.g., N = 10k for UFET).To this end, we propose to perform entity typing in a recall-expand-filter manner.The recall and expansion stages prune the large type set and generate K (typically much smaller than N ) most relevant type candidates for each mention.At the filter stage, we use a novel model called MCCE to concurrently encode and score all these K candidates in only one forward pass to obtain the final type prediction.We investigate different model options for each stage and conduct extensive experiments to compare each option, experiments show that our method reaches SOTA performance on UFET and is thousands of times faster than the CE-based architecture.We also found our method is very effective in fine-grained (130 types) and coarse-grained (9 types) entity typing.
Chengyue Jiang, Wenyang Hui, Yong Jiang 0005, Xiaobin Wang, Pengjun Xie, Kewei Tu
ACL (1)1
2023 Do PLMs Know and Understand Ontological Knowledge?
abstract
Ontological knowledge, which comprises classes and properties and their relationships, is integral to world knowledge.It is significant to explore whether Pretrained Language Models (PLMs) know and understand such knowledge.However, existing PLM-probing studies focus mainly on factual knowledge, lacking a systematic probing of ontological knowledge.In this paper, we focus on probing whether PLMs store ontological knowledge and have a semantic understanding of the knowledge rather than rote memorization of the surface form.To probe whether PLMs know ontological knowledge, we investigate how well PLMs memorize: (1) types of entities; (2) hierarchical relationships among classes and properties, e.g., Person is a subclass of Animal and Member of Sports Team is a subproperty of Member of ; (3) domain and range constraints of properties, e.g., the subject of Member of Sports Team should be a Person and the object should be a Sports Team.To further probe whether PLMs truly understand ontological knowledge beyond memorization, we comprehensively study whether they can reliably perform logical reasoning with given knowledge according to ontological entailment rules.Our probing results show that PLMs can memorize certain ontological knowledge and utilize implicit knowledge in reasoning.However, both the memorizing and reasoning performances are less than perfect, indicating incomplete knowledge and understanding.
Weiqi Wu, Chengyue Jiang, Yong Jiang 0005, Pengjun Xie, Kewei Tu
ACL (1)2
2023 COMBO: A Complete Benchmark for Open KG Canonicalization
abstract
Open knowledge graph (KG) consists of (subject, relation, object) triples extracted from millions of raw text.The subject and object noun phrases and the relation in open KG have severe redundancy and ambiguity and need to be canonicalized.Existing datasets for open KG canonicalization only provide gold entitylevel canonicalization for noun phrases.In this paper, we present COMBO, a Complete Benchmark for Open KG canonicalization.Compared with existing datasets, we additionally provide gold canonicalization for relation phrases, gold ontology-level canonicalization for noun phrases, as well as source sentences from which triples are extracted.We also propose metrics for evaluating each type of canonicalization.On the COMBO dataset, we empirically compare previously proposed canonicalization methods as well as a few simple baseline methods based on pretrained language models.We find that properly encoding the phrases in a triple using pretrained language models results in better relation canonicalization and ontology-level canonicalization of the noun phrase.We release our dataset, baselines, and evaluation scripts at
Chengyue Jiang, Yong Jiang 0005, Weiqi Wu, Pengjun Xie, Kewei Tu
EACL1
2023 Using Interpretation Methods for Model Enhancement
abstract
In the age of neural natural language processing, there are plenty of works trying to derive interpretations of neural models.Intuitively, when gold rationales exist during training, one can additionally train the model to match its interpretation with the rationales.However, this intuitive idea has not been fully explored.In this paper, we propose a framework of utilizing interpretation methods and gold rationales to enhance models.Our framework is very general in the sense that it can incorporate various interpretation methods.Previously proposed gradient-based methods can be shown as an instance of our framework.We also propose two novel instances utilizing two other types of interpretation methods, erasure/replace-based and extractor-based methods, for model enhancement.We conduct comprehensive experiments on a variety of tasks.Experimental results show that our framework is effective especially in low-resource settings in enhancing models with various interpretation methods, and our two newly-proposed methods outperform gradient-based methods in most settings.Code is available at https://github. com/Chord-Chen-30/UIMER.
Chengyue Jiang, Kewei Tu
EMNLP2
2022 Modeling Label Correlations for Ultra-Fine Entity Typing with Neural Pairwise Conditional Random Field
abstract
Ultra-fine entity typing (UFET) aims to predict a wide range of type phrases that correctly describe the categories of a given entity mention in a sentence.Most recent works infer each entity type independently, ignoring the correlations between types, e.g., when an entity is inferred as a president, it should also be a politician and a leader.To this end, we use an undirected graphical model called pairwise conditional random field (PCRF) to formulate the UFET problem, in which the type variables are not only unarily influenced by the input but also pairwisely relate to all the other type variables.We use various modern backbones for entity typing to compute unary potentials, and derive pairwise potentials from type phrase representations that both capture prior semantic information and facilitate accelerated inference.We use mean-field variational inference for efficient type inference on very large type sets and unfold it as a neural network module to enable end-to-end training.Experiments on UFET show that the Neural-PCRF consistently outperforms its backbones with little cost and results in a competitive performance against crossencoder based SOTA while being thousands of times faster.We also find Neural-PCRF effective on a widely used fine-grained entity typing dataset with a smaller type set.We pack Neural-PCRF as a network module that can be plugged onto multi-label type classifiers with ease and release it in github.com/modelscope/ adaseq/examples/NPCRF.
Chengyue Jiang, Yong Jiang 0005, Weiqi Wu, Pengjun Xie, Kewei Tu
EMNLP1
2021 Neuralizing Regular Expressions for Slot Filling
abstract
Neural models and symbolic rules such as regular expressions have their respective merits and weaknesses.In this paper, we study the integration of the two approaches for the slot filling task by converting regular expressions into neural networks.Specifically, we first convert regular expressions into a special form of finite-state transducers, then unfold its approximate inference algorithm as a bidirectional recurrent neural model that performs slot filling via sequence labeling.Experimental results show that our model has superior zero-shot and few-shot performance and stays competitive when there are sufficient training data.
Chengyue Jiang, Zijian Jin, Kewei Tu
EMNLP (1)1
2020 Cold-Start and Interpretability: Turning Regular Expressions into Trainable Recurrent Neural Networks
abstract
Neural networks can achieve impressive performance on many natural language processing applications, but they typically need large labeled data for training and are not easily interpretable.On the other hand, symbolic rules such as regular expressions are interpretable, require no training, and often achieve decent accuracy; but rules cannot benefit from labeled data when available and hence underperform neural networks in rich-resource scenarios.In this paper, we propose a type of recurrent neural networks called FA-RNNs that combine the advantages of neural networks and regular expression rules.An FA-RNN can be converted from regular expressions and deployed in zero-shot and cold-start scenarios.It can also utilize labeled data for training to achieve improved prediction accuracy.After training, an FA-RNN often remains interpretable and can be converted back into regular expressions.We apply FA-RNNs to text classification and observe that FA-RNNs significantly outperform previous neural approaches in both zeroshot and low-resource settings and remain very competitive in rich-resource settings.
Chengyue Jiang, Yinggong Zhao, Shanbo Chu, Libin Shen, Kewei Tu
EMNLP (1)1