VLDB 2026 Research / reviewers in the wild / expert
Chengguang Tang
dblp:264/5495
· DBLP profile ↗
8ranked-venue papers
0as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 40% Question answering and dialogue systems · 23% Information extraction and text analysis · 20% | |
| Databases, data mining, and information retrieval
1 paper |
Knowledge graphs · 100% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.9 | 1 | 2025 | Aligning Language Models Using Follow-up Likelihood as Reward Signal · AAAI 2025 |
Natural language and speech › Language models and text generation › alignment
preference alignment |
0.9 | 1 | 2025 | Aligning Language Models Using Follow-up Likelihood as Reward Signal · AAAI 2025 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.9 | 1 | 2025 | Aligning Language Models Using Follow-up Likelihood as Reward Signal · AAAI 2025 |
Natural language and speech › Question answering and dialogue systems › knowledge-grounded dialogue
document-grounded dialogue |
0.6 | 1 | 2022 | Layout-Aware Information Extraction for Document-Grounded Dialogue: Dataset, Method and Demonstration · ACM Multimedia 2022 |
Natural language and speech › Information extraction and text analysis › document analysis
document information extraction |
0.6 | 1 | 2022 | Layout-Aware Information Extraction for Document-Grounded Dialogue: Dataset, Method and Demonstration · ACM Multimedia 2022 |
Natural language and speech › Language models and text generation › pre-trained language model
knowledge-enhanced pre-trained language model |
0.6 | 1 | 2022 | DKPLM: Decomposable Knowledge-Enhanced Pre-trained Language Model for Natural Language Understanding · AAAI 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph |
0.6 | 1 | 2022 | DKPLM: Decomposable Knowledge-Enhanced Pre-trained Language Model for Natural Language Understanding · AAAI 2022 |
Natural language and speech › Language models and text generation › knowledge editing
knowledge injection into language models |
0.6 | 1 | 2022 | DKPLM: Decomposable Knowledge-Enhanced Pre-trained Language Model for Natural Language Understanding · AAAI 2022 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.6 | 1 | 2022 | DKPLM: Decomposable Knowledge-Enhanced Pre-trained Language Model for Natural Language Understanding · AAAI 2022 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.6 | 1 | 2022 | DKPLM: Decomposable Knowledge-Enhanced Pre-trained Language Model for Natural Language Understanding · AAAI 2022 |
Natural language and speech › Information extraction and text analysis › document analysis › document information extraction
visually rich document information extraction |
0.6 | 1 | 2022 | Layout-Aware Information Extraction for Document-Grounded Dialogue: Dataset, Method and Demonstration · ACM Multimedia 2022 |
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking |
0.5 | 1 | 2021 | Unsupervised Learning of Deterministic Dialogue Structure with Edge-Enhanced Graph Auto-Encoder · AAAI 2021 |
Natural language and speech › Question answering and dialogue systems › dialogue modeling
dialogue structure modeling |
0.5 | 1 | 2021 | Unsupervised Learning of Deterministic Dialogue Structure with Edge-Enhanced Graph Auto-Encoder · AAAI 2021 |
Knowledge graphs › link prediction
few-shot knowledge graph completion |
0.5 | 1 | 2021 | Relational Learning with Gated and Attentive Neighbor Aggregator for Few-Shot Knowledge Graph Completion · SIGIR 2021 |
Knowledge graphs
link prediction |
0.5 | 1 | 2021 | Relational Learning with Gated and Attentive Neighbor Aggregator for Few-Shot Knowledge Graph Completion · SIGIR 2021 |
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue |
0.4 | 1 | 2020 | Learning Low-Resource End-To-End Goal-Oriented Dialog for Fast and Reliable System Deployment · ACL 2020 |
Methods — techniques the papers use, named apart from their topics
preference data mining · 0.9follow-up likelihood · 0.9direct preference optimization · 0.9token-based language model · 0.6relational knowledge decoding · 0.6pseudo token representation · 0.6layout feature encoding · 0.6knowledge-aware long-tail entity detection · 0.6transh · 0.5response selection · 0.5meta-learning · 0.5graph autoencoder · 0.5gated attentive neighbor aggregator · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Aligning Language Models Using Follow-up Likelihood as Reward SignalabstractIn natural human-to-human conversations, participants often receive feedback signals from one another based on their follow-up reactions. These reactions can include verbal responses, facial expressions, changes in emotional state, and other non-verbal cues. Similarly, in human-machine interactions, the machine can leverage the user's follow-up utterances as feedback signals to assess whether it has appropriately addressed the user's request. Therefore, we propose using the likelihood of follow-up utterances as rewards to differentiate preferred responses from less favored ones, without relying on human or commercial LLM-based preference annotations. Our proposed reward mechanism, ``Follow-up Likelihood as Reward" (FLR), matches the performance of strong reward models trained on large-scale human or GPT-4 annotated data on 8 pairwise-preference and 4 rating-based benchmarks. Building upon the FLR mechanism, we propose to automatically mine preference data from the online generations of a base policy model. The preference data are subsequently used to boost the helpfulness of the base model through direct alignment from preference (DAP) methods, such as direct preference optimization (DPO). Lastly, we demonstrate that fine-tuning the language model that provides follow-up likelihood with natural language feedback significantly enhances FLR's performance on reward modeling benchmarks and effectiveness in aligning the base policy model's helpfulness. Chen Zhang 0055, Dading Chong, Feng Jiang 0007, Chengguang Tang, Anningzhe Gao, Guohua Tang, Haizhou Li 0001 |
AAAI | 4 |
| 2022 | DKPLM: Decomposable Knowledge-Enhanced Pre-trained Language Model for Natural Language UnderstandingabstractKnowledge-Enhanced Pre-trained Language Models (KEPLMs) are pre-trained models with relation triples injecting from knowledge graphs to improve language understanding abilities.Experiments show that our model outperforms other KEPLMs significantly over zero-shot knowledge probing tasks and multiple knowledge-aware language understanding tasks. To guarantee effective knowledge injection, previous studies integrate models with knowledge encoders for representing knowledge retrieved from knowledge graphs. The operations for knowledge retrieval and encoding bring significant computational burdens, restricting the usage of such models in real-world applications that require high inference speed. In this paper, we propose a novel KEPLM named DKPLM that decomposes knowledge injection process of the pre-trained language models in pre-training, fine-tuning and inference stages, which facilitates the applications of KEPLMs in real-world scenarios. Specifically, we first detect knowledge-aware long-tail entities as the target for knowledge injection, enhancing the KEPLMs' semantic understanding abilities and avoiding injecting redundant information. The embeddings of long-tail entities are replaced by ``pseudo token representations'' formed by relevant knowledge triples. We further design the relational knowledge decoding task for pre-training to force the models to truly understand the injected knowledge by relation triple reconstruction. Experiments show that our model outperforms other KEPLMs significantly over zero-shot knowledge probing tasks and multiple knowledge-aware language understanding tasks. We further show that DKPLM has a higher inference speed than other competing models due to the decomposing mechanism. Taolin Zhang 0001, Chengyu Wang 0001, Minghui Qiu, Chengguang Tang, Jun Huang 0007 |
AAAI | 5 |
| 2022 | Layout-Aware Information Extraction for Document-Grounded Dialogue: Dataset, Method and DemonstrationabstractBuilding document-grounded dialogue systems have received growing interest as documents convey a wealth of human knowledge and commonly exist in enterprises. Wherein, how to comprehend and retrieve information from documents is a challenging research problem. Previous work ignores the visual property of documents and treats them as plain text, resulting in incomplete modality. In this paper, we propose a Layout-aware document-level Information Extraction dataset, LIE, to facilitate the study of extracting both structural and semantic knowledge from visually rich documents (VRDs), so as to generate accurate responses in dialogue systems. LIE contains 62k annotations of three extraction tasks from 4,061 pages in product and official documents, becoming the largest VRD-based information extraction dataset to the best of our knowledge. We also develop benchmark methods that extend the token-based language model to consider layout features like humans. Empirical results show that layout is critical for VRD-based extraction, and system demonstration also verifies that the extracted knowledge can help locate the answers that users care about. Zhenyu Zhang 0006, Bowen Yu 0002, Haiyang Yu 0003, Tingwen Liu, Cheng Fu 0003, Chengguang Tang, Jian Sun 0021, Yongbin Li 0001 |
ACM Multimedia | 7 |
| 2021 | Unsupervised Learning of Deterministic Dialogue Structure with Edge-Enhanced Graph Auto-EncoderabstractIt is important for task-oriented dialogue systems to discover the dialogue structure (i.e. the general dialogue flow) from dialogue corpora automatically. Previous work models dialogue structure by extracting latent states for each utterance first and then calculating the transition probabilities among states. These two-stage methods ignore the contextual information when calculating the probabilities, which makes the transitions between the states ambiguous. This paper proposes a conversational graph (CG) to represent deterministic dialogue structure where nodes and edges represent the utterance and context information respectively. An unsupervised Edge-Enhanced Graph Auto-Encoder (EGAE) architecture is designed to model local-contextual and global-structural information for conversational graph learning. Furthermore, a self-supervised objective is introduced with the response selection task to guide the unsupervised learning of the dialogue structure. Experimental results on several public datasets demonstrate that the novel model outperforms several alternatives in aggregating utterances with similar semantics. The effectiveness of the learned dialogue structured is also verified by more than 5\% joint accuracy improvement in the downstream task of low resource dialogue state tracking. Yajing Sun, Yong Shan, Chengguang Tang, Yue Hu 0002, Yinpei Dai, Jing Yu 0007, Jian Sun 0021, Fei Huang 0002, Luo Si |
AAAI | 3 |
| 2021 | HORNET: Enriching Pre-trained Language Representations with Heterogeneous Knowledge SourcesabstractKnowledge-Enhanced Pre-trained Language Models (KEPLMs) improve the language understanding abilities of deep language models by leveraging the rich semantic knowledge from knowledge graphs, other than plain pre-training texts. However, previous efforts mostly use homogeneous knowledge (especially structured relation triples in knowledge graphs) to enhance the context-aware representations of entity mentions, whose performance may be limited by the coverage of knowledge graphs. Also, it is unclear whether these KEPLMs truly understand the injected semantic knowledge due to the "black-box'' training mechanism. In this paper, we propose a novel KEPLM named HORNET, which integrates Heterogeneous knowledge from various structured and unstructured sources into the Roberta NETwork and hence takes full advantage of both linguistic and factual knowledge simultaneously. Specifically, we design a hybrid attention heterogeneous graph convolution network (HaHGCN) to learn heterogeneous knowledge representations based on the structured relation triplets from knowledge graphs and the unstructured entity description texts. Meanwhile, we propose the explicit dual knowledge understanding tasks to help induce a more effective infusion of the heterogeneous knowledge, promoting our model for learning the complicated mappings from the knowledge graph embedding space to the deep context-aware embedding space and vice versa. Experiments show that our HORNET model outperforms various KEPLM baselines on knowledge-aware tasks including knowledge probing, entity typing and relation extraction. Our model also achieves substantial improvement over several GLUE benchmark datasets, compared to other KEPLMs. Taolin Zhang 0001, Zerui Cai, Chengyu Wang 0001, Peng Li 0056, Yang Li 0218, Minghui Qiu, Chengguang Tang, Jun Huang 0007 |
CIKM | 7 |
| 2021 | When Few-Shot Learning Meets Large-Scale Knowledge-Enhanced Pre-training: Alibaba at FewCLUE
Ziyun Xu, Chengyu Wang 0001, Peng Li 0056, Yang Li 0218, Boyu Hou, Minghui Qiu, Chengguang Tang, Jun Huang 0007 |
NLPCC (2) | 8 |
| 2021 | Relational Learning with Gated and Attentive Neighbor Aggregator for Few-Shot Knowledge Graph CompletionabstractAiming at expanding few-shot relations' coverage in knowledge graphs (KGs), few-shot knowledge graph completion (FKGC) has recently gained more research interests. Some existing models employ a few-shot relation's multi-hop neighbor information to enhance its semantic representation. However, noise neighbor information might be amplified when the neighborhood is excessively sparse and no neighbor is available to represent the few-shot relation. Moreover, modeling and inferring complex relations of one-to-many (1-N), many-to-one (N-1), and many-to-many (N-N) by previous knowledge graph completion approaches requires high model complexity and a large amount of training instances. Thus, inferring complex relations in the few-shot scenario is difficult for FKGC models due to limited training instances. In this paper, we propose a few-shot relational learning with global-local framework to address the above issues. At the global stage, a novel gated and attentive neighbor aggregator is built for accurately integrating the semantics of a few-shot relation's neighborhood, which helps filtering the noise neighbors even if a KG contains extremely sparse neighborhoods. For the local stage, a meta-learning based TransH (MTransH) method is designed to model complex relations and train our model in a few-shot learning fashion. Extensive experiments show that our model outperforms the state-of-the-art FKGC approaches on the frequently-used benchmark datasets NELL-One and Wiki-One. Compared with the strong baseline model MetaR, our model achieves 5-shot FKGC performance improvements of 8.0% on NELL-One and 2.8% on Wiki-One by the metric [email protected] Guanglin Niu, Yang Li 0218, Chengguang Tang, Ruiying Geng, Hao Wang 0005, Jian Sun 0021, Fei Huang 0002, Luo Si |
SIGIR | 3 |
| 2020 | Learning Low-Resource End-To-End Goal-Oriented Dialog for Fast and Reliable System DeploymentabstractExisting end-to-end dialog systems perform less effectively when data is scarce. To obtain an acceptable success in real-life online services with only a handful of training examples, both fast adaptability and reliable performance are highly desirable for dialog systems. In this paper, we propose the Meta-Dialog System (MDS), which combines the advantages of both meta-learning approaches and human-machine collaboration. We evaluate our methods on a new extended-bAbI dataset and a transformed MultiWOZ dataset for low-resource goal-oriented dialog learning. Experimental results show that MDS significantly outperforms non-meta-learning baselines and can achieve more than 90% per-turn accuracies with only 10 dialogs on the extended-bAbI dataset. Yinpei Dai, Hangyu Li 0003, Chengguang Tang, Yongbin Li 0001, Jian Sun 0021, Xiaodan Zhu 0001 |
ACL | 3 |