Chenguang Zhu 0001

dblp:48/7536-1 · DBLP profile ↗
← Back
45ranked-venue papers
6as first author
41since 2021 · last 2025
0000-0001-6955-8924ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 44 · 5 first-author · 40 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Law of the Weakest Link: Cross Capabilities of Large Language Models
abstract
The development and evaluation of Large Language Models (LLMs) have largely focused on individual capabilities. However, this overlooks the intersection of multiple abilities across different types of expertise that are often required for real-world tasks, which we term **cross capabilities**. To systematically explore this concept, we first define seven core individual capabilities and then pair them to form seven common cross capabilities, each supported by a manually constructed taxonomy. Building on these definitions, we introduce *CrossEval*, a benchmark comprising 1,400 human-annotated prompts, with 100 prompts for each individual and cross capability. To ensure reliable evaluation, we involve expert annotators to assess 4,200 model responses, gathering 8,400 human ratings with detailed explanations to serve as reference examples. Our findings reveal that current LLMs consistently exhibit the ``Law of the Weakest Link,'' where cross-capability performance is significantly constrained by the weakest component. Across 58 cross-capability scores from 17 models, 38 scores are lower than all individual capabilities, while 20 fall between strong and weak, but closer to the weaker ability. These results highlight LLMs' underperformance in cross-capability tasks, emphasizing the need to identify and improve their weakest capabilities as a key research priority. The code, benchmarks, and evaluations are available on our [project website](https://www.llm-cross-capabilities.org).
Ming Zhong 0005, Aston Zhang, Wenhan Xiong, Chenguang Zhu 0001, Zhengxing Chen, Chloe Bi, Mike Lewis, Sravya Popuri, Sharan Narang, Melanie Kambadur, Dhruv Mahajan 0001, Sergey Edunov, Jiawei Han 0001, Laurens van der Maaten
ICLR6
2025 Self-Generated Critiques Boost Reward Modeling for Language Models
abstract
Yue Yu, Zhengxing Chen, Aston Zhang, Liang Tan, Chenguang Zhu, Richard Yuanzhe Pang, Yundi Qian, Xuewei Wang, Suchin Gururangan, Chao Zhang, Melanie Kambadur, Dhruv Mahajan, Rui Hou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yue Yu 0009, Zhengxing Chen, Aston Zhang, Chenguang Zhu 0001, Richard Yuanzhe Pang, Yundi Qian, Suchin Gururangan, Melanie Kambadur, Dhruv Mahajan 0001
NAACL (Long Papers)5
2024 PEARL: Prompting Large Language Models to Plan and Execute Actions Over Long Documents
abstract
Simeng Sun, Yang Liu, Shuohang Wang, Dan Iter, Chenguang Zhu, Mohit Iyyer. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Simeng Sun, Yang Liu 0124, Shuohang Wang, Dan Iter, Chenguang Zhu 0001, Mohit Iyyer
EACL (1)5
2023 i-Code: An Integrative and Composable Multimodal Learning Framework
abstract
Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to one or two modalities. We present i-Code, a self-supervised pretraining framework where users may flexibly combine the modalities of vision, speech, and language into unified and general-purpose vector representations. In this framework, data from each modality are first given to pretrained single-modality encoders. The encoder outputs are then integrated with a multimodal fusion network, which uses novel merge- and co-attention mechanisms to effectively combine information from the different modalities. The entire system is pretrained end-to-end with new objectives including masked modality unit modeling and cross-modality contrastive learning. Unlike previous research using only video for pretraining, the i-Code framework can dynamically process single, dual, and triple-modality data during training and inference, flexibly projecting different combinations of modalities into a single representation space. Experimental results demonstrate how i-Code can outperform state-of-the-art techniques on five multimodal understanding tasks and single-modality benchmarks, improving by as much as 11% and demonstrating the power of integrative multimodal pretraining.
Ziyi Yang 0011, Yuwei Fang, Chenguang Zhu 0001, Reid Pryzant, Dongdong Chen 0001, Yu Shi 0001, Yichong Xu, Yao Qian, Mei Gao, Liyang Lu, Yujia Xie, Robert Gmyr, Noel Codella, Naoyuki Kanda, Bin Xiao 0004, Lu Yuan 0001, Takuya Yoshioka, Michael Zeng 0001, Xuedong Huang 0001
AAAI3
2023 UniSumm and SummZoo: Unified Model and Diverse Benchmark for Few-Shot Summarization
abstract
Yulong Chen, Yang Liu, Ruochen Xu, Ziyi Yang, Chenguang Zhu, Michael Zeng, Yue Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yulong Chen 0001, Yang Liu 0124, Ruochen Xu, Ziyi Yang 0011, Chenguang Zhu 0001, Michael Zeng 0001, Yue Zhang 0004
ACL (1)5
2023 APOLLO: A Simple Approach for Adaptive Pretraining of Language Models for Logical Reasoning
abstract
Soumya Sanyal, Yichong Xu, Shuohang Wang, Ziyi Yang, Reid Pryzant, Wenhao Yu, Chenguang Zhu, Xiang Ren. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Soumya Sanyal 0001, Yichong Xu, Shuohang Wang, Ziyi Yang 0011, Reid Pryzant, Wenhao Yu 0002, Chenguang Zhu 0001, Xiang Ren 0001
ACL (1)7
2023 Z-Code++: A Pre-trained Language Model Optimized for Abstractive Summarization
abstract
Pengcheng He, Baolin Peng, Song Wang, Yang Liu, Ruochen Xu, Hany Hassan, Yu Shi, Chenguang Zhu, Wayne Xiong, Michael Zeng, Jianfeng Gao, Xuedong Huang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Baolin Peng, Song Wang 0012, Yang Liu 0124, Ruochen Xu, Hany Hassan, Yu Shi 0001, Chenguang Zhu 0001, Wayne Xiong, Michael Zeng 0001, Jianfeng Gao 0001, Xuedong Huang 0001
ACL (1)8
2023 Unifying Vision, Text, and Layout for Universal Document Processing
abstract
We propose Universal Document Processing (UDOP), a foundation Document AI model which unifies text, image, and layout modalities together with varied task formats, including document understanding and generation. UDOP leverages the spatial correlation between textual content and document image to model image, text, and layout modalities with one uniform representation. With a novel Vision-Text-Layout Transformer, UDOP unifies pretraining and multi-domain downstream tasks into a prompt-based sequence generation scheme. UDOP is pretrained on both large-scale unlabeled document corpora using innovative self-supervised objectives and diverse labeled data. UDOP also learns to generate document images from text and layout modalities via masked image reconstruction. To the best of our knowledge, this is the first time in the field of document AI that one model simultaneously achieves high-quality neural document editing and content customization. Our method sets the state-of-the-art on 8 Document AI tasks, e.g., document understanding and QA, across diverse data domains like finance reports, academic papers, and web-sites. UDOP ranks first on the leaderboard of the Document Understanding Benchmark.11Code and models: https://github.com/microsoft/i-Code/tree/main/i-Code-Doc
Zineng Tang, Ziyi Yang 0011, Yuwei Fang, Yang Liu 0124, Chenguang Zhu 0001, Michael Zeng 0001, Cha Zhang, Mohit Bansal
CVPR6
2023 Improving Commonsense in Vision-Language Models via Knowledge Graph Riddles
abstract
This paper focuses on analyzing and improving the commonsense ability of recent popular vision-language (VL) models. Despite the great success, we observe that existing VL-models still lack commonsense knowledge/reasoning ability (e.g., “Lemons are sour”), which is a vital component towards artificial general intelligence. Through our analysis, we find one important reason is that existing large-scale VL datasets do not contain much commonsense knowledge, which motivates us to improve the commonsense of VL-models from the data perspective. Rather than collecting a new VL training dataset, we propose a more scalable strategy, i.e., “Data Augmentation with kNowledge graph linearization for CommonsensE capability” (DANCE). It can be viewed as one type of data augmentation technique, which can inject commonsense knowledge into existing VL datasets on the fly during training. More specifically, we leverage the commonsense knowledge graph (e.g., ConceptNet) and create variants of text description in VL datasets via bidirectional sub-graph sequentialization. For better commonsense evaluation, we further propose the first retrieval-based commonsense diagnostic benchmark. By conducting extensive experiments on some representative VL-models, we demonstrate that our DANCE technique is able to significantly improve the commonsense ability while maintaining the performance on vanilla retrieval tasks. The code and data are available at https://github.com/pleaseconnectwifi/DANCE.
Shuquan Ye, Yujia Xie, Dongdong Chen 0001, Yichong Xu, Lu Yuan 0001, Chenguang Zhu 0001, Jing Liao 0001
CVPR6
2023 The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT Interactions
abstract
Siru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu, Heng Ji, Jiawei Han. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Siru Ouyang, Shuohang Wang, Ming Zhong 0005, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu 0001, Heng Ji 0001, Jiawei Han 0001
EMNLP8
2023 Automatic Prompt Optimization with "Gradient Descent" and Beam Search
abstract
Large Language Models (LLMs) have shown impressive performance as general purpose agents, but their abilities remain highly dependent on prompts which are hand written with onerous trial-and-error effort.We propose a simple and nonparametric solution to this problem, Prompt Optimization with Textual Gradients (ProTeGi), which is inspired by numerical gradient descent to automatically improve prompts, assuming access to training data and an LLM API.The algorithm uses minibatches of data to form natural language "gradients" that criticize the current prompt, much like how numerical gradients point in the direction of error ascent.The natural language gradients are then "propagated" into the prompt by editing the prompt in the opposite semantic direction of the gradient.These gradient descent steps are guided by a beam search and bandit selection procedure which significantly improves algorithmic efficiency.Preliminary results across three benchmark NLP tasks and the novel problem of LLM jailbreak detection suggest that Automatic Prompt Optimization can outperform prior prompt editing techniques and improve an initial prompt's performance by up to 31%, by using data to rewrite vague task descriptions into more precise annotation instructions.1
Reid Pryzant, Dan Iter, Jerry Li 0001, Yin Tat Lee, Chenguang Zhu 0001, Michael Zeng 0001
EMNLP5
2023 Generate rather than Retrieve: Large Language Models are Strong Context Generators
Wenhao Yu 0002, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal 0001, Chenguang Zhu 0001, Michael Zeng 0001, Meng Jiang 0001
ICLR7
2023 Any-to-Any Generation via Composable Diffusion
abstract
We present Composable Diffusion (CoDi), a novel generative model capable of generating any combination of output modalities, such as language, image, video, or audio, from any combination of input modalities. Unlike existing generative AI systems, CoDi can generate multiple modalities in parallel and its input is not limited to a subset of modalities like text or image. Despite the absence of training datasets for many combinations of modalities, we propose to align modalities in both the input and output space. This allows CoDi to freely condition on any input combination and generate any group of modalities, even if they are not present in the training data. CoDi employs a novel composable generation strategy which involves building a shared multimodal space by bridging alignment in the diffusion process, enabling the synchronized generation of intertwined modalities, such as temporally aligned video and audio. Highly customizable and flexible, CoDi achieves strong joint-modality generation quality, and outperforms or is on par with the unimodal state-of-the-art for single-modality synthesis.
Zineng Tang, Ziyi Yang 0011, Chenguang Zhu 0001, Michael Zeng 0001, Mohit Bansal
NeurIPS3
2023 Knowledge-Augmented Methods for Natural Language Processing
abstract
Knowledge in NLP has been a rising trend especially after the advent of large-scale pre-trained models. Knowledge is critical to equip statistics-based models with common sense, logic and other external information. In this tutorial, we will introduce recent state-of-the-art works in applying knowledge in language understanding, language generation and commonsense reasoning.
Chenguang Zhu 0001, Yichong Xu, Xiang Ren 0001, Bill Y. Lin, Meng Jiang 0001, Wenhao Yu 0002
WSDM1
2023 MACSum: Controllable Summarization with Mixed Attributes
abstract
Abstract Controllable summarization allows users to generate customized summaries with specified attributes. However, due to the lack of designated annotations of controlled summaries, existing work has to craft pseudo datasets by adapting generic summarization benchmarks. Furthermore, most research focuses on controlling single attributes individually (e.g., a short summary or a highly abstractive summary) rather than controlling a mix of attributes together (e.g., a short and highly abstractive summary). In this paper, we propose MACSum, the first human-annotated summarization dataset for controlling mixed attributes. It contains source texts from two domains, news articles and dialogues, with human-annotated summaries controlled by five designed attributes (Length, Extractiveness, Specificity, Topic, and Speaker). We propose two simple and effective parameter-efficient approaches for the new task of mixed controllable summarization based on hard prompt tuning and soft prefix tuning. Results and analysis demonstrate that hard prompt models yield the best performance on most metrics and human evaluations. However, mixed-attribute control is still challenging for summarization tasks. Our dataset and code are available at https://github.com/psunlpgroup/MACSum.
Yusen Zhang 0001, Yang Liu 0124, Ziyi Yang 0011, Yuwei Fang, Yulong Chen 0001, Dragomir R. Radev, Chenguang Zhu 0001, Michael Zeng 0001, Rui Zhang 0037
Trans. Assoc. Comput. Linguistics7
2022 JAKET: Joint Pre-training of Knowledge Graph and Language Understanding
abstract
Knowledge graphs (KGs) contain rich information about world knowledge, entities, and relations. Thus, they can be great supplements to existing pre-trained language models. However, it remains a challenge to efficiently integrate information from KG into language modeling. And the understanding of a knowledge graph requires related context. We propose a novel joint pre-training framework, JAKET, to model both the knowledge graph and language. The knowledge module and language module provide essential information to mutually assist each other: the knowledge module produces embeddings for entities in text while the language module generates context-aware initial embeddings for entities and relations in the graph. Our design enables the pre-trained model to easily adapt to unseen knowledge graphs in new domains. Experiment results on several knowledge-aware NLP tasks show that our proposed framework achieves superior performance by effectively leveraging knowledge in language understanding.
Donghan Yu, Chenguang Zhu 0001, Yiming Yang 0002, Michael Zeng 0001
AAAI2
2022 DialogLM: Pre-trained Model for Long Dialogue Understanding and Summarization
abstract
Dialogue is an essential part of human communication and cooperation. Existing research mainly focuses on short dialogue scenarios in a one-on-one fashion. However, multi-person interactions in the real world, such as meetings or interviews, are frequently over a few thousand words. There is still a lack of corresponding research and powerful tools to understand and process such long dialogues. Therefore, in this work, we present a pre-training framework for long dialogue understanding and summarization. Considering the nature of long conversations, we propose a window-based denoising approach for generative pre-training. For a dialogue, it corrupts a window of text with dialogue-inspired noise, and guides the model to reconstruct this window based on the content of the remaining conversation. Furthermore, to process longer input, we augment the model with sparse attention which is combined with conventional attention in a hybrid manner. We conduct extensive experiments on five datasets of long dialogues, covering tasks of dialogue summarization, abstractive question answering and topic segmentation. Experimentally, we show that our pre-trained model DialogLM significantly surpasses the state-of-the-art models across datasets and tasks. Source code and all the pre-trained models are available on our GitHub repository (https://github.com/microsoft/DialogLM).
Ming Zhong 0005, Yang Liu 0124, Yichong Xu, Chenguang Zhu 0001, Michael Zeng 0001
AAAI4
2022 SummN: A Multi-Stage Summarization Framework for Long Input Dialogues and Documents
abstract
Yusen Zhang, Ansong Ni, Ziming Mao, Chen Henry Wu, Chenguang Zhu, Budhaditya Deb, Ahmed Awadallah, Dragomir Radev, Rui Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yusen Zhang 0001, Ansong Ni, Ziming Mao, Chen Henry Wu, Chenguang Zhu 0001, Budhaditya Deb, Ahmed Awadallah 0001, Dragomir R. Radev, Rui Zhang 0037
ACL (1)5
2022 Leveraging Visual Knowledge in Language Tasks: An Empirical Study on Intermediate Pre-training for Cross-Modal Knowledge Transfer
abstract
Pre-trained language models are still far from human performance in tasks that need understanding of properties (e.g.appearance, measurable quantity) and affordances of everyday objects in the real world since the text lacks such information due to reporting bias.In this work, we study whether integrating visual knowledge into a language model can fill the gap.We investigate two types of knowledge transfer: (1) text knowledge transfer using image captions that may contain enriched visual knowledge and (2) cross-modal knowledge transfer using both images and captions with vision-language training objectives.On 5 downstream tasks that may need visual knowledge to solve the problem, we perform extensive empirical comparisons over the presented objectives.Our experiments show that visual knowledge transfer can improve performance in both low-resource and fully supervised settings.1
Woojeong Jin 0001, Chenguang Zhu 0001, Jay Pujara, Xiang Ren 0001
ACL (1)3
2022 DYLE: Dynamic Latent Extraction for Abstractive Long-Input Summarization
abstract
Ziming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang, Rui Zhang, Tao Yu, Budhaditya Deb, Chenguang Zhu, Ahmed Awadallah, Dragomir Radev. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Ziming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang 0001, Rui Zhang 0037, Tao Yu 0009, Budhaditya Deb, Chenguang Zhu 0001, Ahmed Awadallah 0001, Dragomir R. Radev
ACL (1)8
2022 Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data
abstract
Shuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu, Siqi Sun, Ruochen Xu, Chenguang Zhu, Michael Zeng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Shuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu 0124, Ruochen Xu, Chenguang Zhu 0001, Michael Zeng 0001
ACL (1)7
2022 KG-FiD: Infusing Knowledge Graph in Fusion-in-Decoder for Open-Domain Question Answering
abstract
Donghan Yu, Chenguang Zhu, Yuwei Fang, Wenhao Yu, Shuohang Wang, Yichong Xu, Xiang Ren, Yiming Yang, Michael Zeng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Donghan Yu, Chenguang Zhu 0001, Yuwei Fang, Wenhao Yu 0002, Shuohang Wang, Yichong Xu, Xiang Ren 0001, Yiming Yang 0002, Michael Zeng 0001
ACL (1)2
2022 An Empirical Study of Training End-to-End Vision-and-Language Transformers
abstract
Vision-and-language (VL) pre-training has proven to be highly effective on various VL downstream tasks. While recent work has shown that fully transformer-based VL models can be more efficient than previous region-feature-based methods, their performance on downstream tasks often degrades significantly. In this paper, we present Meter, a Multimodal End-to-end TransformER framework, through which we investigate how to design and pre-train a fully transformer-based VL model in an end-to-end manner. Specifically, we dissect the model designs along multiple dimensions: vision encoders (e.g., CLIP-ViT, Swin transformer), text encoders (e.g., RoBERTa, De-BERTa), multimodal fusion module (e.g., merged attention vs. co-attention), architectural design (e.g., encoder-only vs. encoder-decoder), and pre-training objectives (e.g., masked image modeling). We conduct comprehensive experiments and provide insights on how to train a performant VL transformer. Meterachieves an accuracy of 77.64% on the VQAv2 test-std set using only 4M images for pre-training, surpassing the state-of-the-art region-feature-based model by 1.04%, and outperforming the previous best fully transformer-based model by 1.6%. Notably, when further scaled up, our best VQA model achieves an accuracy of 80.54%. Code and pre-trained models are released at https://github.com/zdou0830/METER.
Zi-Yi Dou, Yichong Xu, Zhe Gan, Shuohang Wang, Chenguang Zhu 0001, Pengchuan Zhang, Lu Yuan 0001, Nanyun Peng 0001, Zicheng Liu 0001, Michael Zeng 0001
CVPR7
2022 CLIP-Event: Connecting Text and Images with Event Structures
abstract
Vision-language (V+L) pretraining models have achieved great success in supporting multimedia applications by understanding the alignments between images and text. While existing vision-language pretraining models primarily focus on understanding objects in images or entities in text, they often ignore the alignment at the level of events and their argument structures. In this work, we propose a contrastive learning framework to enforce vision-language pretraining models to comprehend events and associated argument (participant) roles. To achieve this, we take advantage of text information extraction technologies to obtain event structural knowledge, and utilize multiple prompt functions to contrast difficult negative descriptions by manipulating event structures. We also design an event graph alignment loss based on optimal transport to capture event argument structures. In addition, we collect a large event-rich dataset (106,875 images) for pretraining, which provides a more challenging image retrieval benchmark to assess the understanding of complicated lengthy sentences11The data and code are publicly available for research purpose in https://github.com/limanling/clip-event.. Experiments show that our zero-shot CLIP-Event outperforms the state-of-the-art supervised model in argument extraction on Multimedia Event Extraction, achieving more than 5% absolute F-score gain in event extraction, as well as significant improvements on a variety of downstream tasks under zero-shot settings.
Manling Li, Ruochen Xu, Shuohang Wang, Luowei Zhou, Xudong Lin 0003, Chenguang Zhu 0001, Michael Zeng 0001, Heng Ji 0001, Shih-Fu Chang
CVPR6
2022 Retrieval Augmentation for Commonsense Reasoning: A Unified Approach
abstract
A common thread of retrieval-augmented methods in the existing literature focuses on retrieving encyclopedic knowledge, such as Wikipedia, which facilitates well-defined entity and relation spaces that can be modeled.However, applying such methods to commonsense reasoning tasks faces two unique challenges, i.e., the lack of a general large-scale corpus for retrieval and a corresponding effective commonsense retriever.In this paper, we systematically investigate how to leverage commonsense knowledge retrieval to improve commonsense reasoning tasks.We proposed a unified framework of Retrieval-Augmented Commonsense reasoning (called RACO), including a newly constructed commonsense corpus with over 20 million documents and novel strategies for training a commonsense retriever.We conducted experiments on four different commonsense reasoning tasks.Extensive evaluation results showed that our proposed RACO can significantly outperform other knowledgeenhanced method counterparts, achieving new SoTA performance on the CommonGen 1 and CREAK 2 leaderboards.Our code is available at https://github.com/wyu97/RACo.
Wenhao Yu 0002, Chenguang Zhu 0001, Zhihan Zhang 0001, Shuohang Wang, Zhuosheng Zhang 0001, Yuwei Fang, Meng Jiang 0001
EMNLP2
2022 Empowering Language Models with Knowledge Graph Reasoning for Open-Domain Question Answering
abstract
Answering open-domain questions requires world knowledge about in-context entities.As pre-trained Language Models (LMs) lack the power to store all required knowledge, external knowledge sources, such as knowledge graphs, are often used to augment LMs.In this work, we propose knOwledge REasOning empowered Language Model (OREOLM), which consists of a novel Knowledge Interaction Layer that can be flexibly plugged into existing Transformer-based LMs to interact with a differentiable Knowledge Graph Reasoning module collaboratively.In this way, LM guides KG to walk towards the desired answer, while the retrieved knowledge improves LM.By adopting OREOLM to RoBERTa and T5, we show significant performance gain, achieving state-of-art results in the Closed-Book setting.The performance enhancement is mainly from the KG reasoning's capacity to infer missing relational facts.In addition, OREOLM provides reasoning paths as rationales to interpret the model's decision.
Ziniu Hu, Yichong Xu, Wenhao Yu 0002, Shuohang Wang, Ziyi Yang 0011, Chenguang Zhu 0001, Kai-Wei Chang 0001, Yizhou Sun
EMNLP6
2022 Leveraging Locality in Abstractive Text Summarization
abstract
Neural attention models have achieved significant improvements on many natural language processing tasks.However, the quadratic memory complexity of the self-attention module with respect to the input length hinders their applications in long text summarization.Instead of designing more efficient attention modules, we approach this problem by investigating if models with a restricted context can have competitive performance compared with the memory-efficient attention models that maintain a global context by treating the input as a single sequence.Our model is applied to individual pages, which contain parts of inputs grouped by the principle of locality, during both the encoding and decoding stages.We empirically investigated three kinds of locality in text summarization at different levels of granularity, ranging from sentences to documents.Our experimental results show that our model has a better performance compared with strong baseline models with efficient attention modules, and our analysis provides further insights into our locality-aware modeling strategy.1
Yixin Liu 0003, Ansong Ni, Linyong Nan, Budhaditya Deb, Chenguang Zhu 0001, Ahmed Awadallah 0001, Dragomir R. Radev
EMNLP5
2022 ParaTag: A Dataset of Paraphrase Tagging for Fine-Grained Labels, NLG Evaluation, and Data Augmentation
abstract
Paraphrase identification has been formulated as a binary classification task to decide whether two sentences hold a paraphrase relationship.Existing paraphrase datasets only annotate a binary label for each sentence pair.However, after a systematical analysis of existing paraphrase datasets, we found that the degree of paraphrase cannot be well characterized by a single binary label.And the criteria of paraphrase are not even consistent within the same dataset.We hypothesize that such issues would limit the effectiveness of paraphrase models trained on these data.To this end, we propose a novel fine-grained paraphrase annotation schema that labels the minimum spans of tokens in a sentence that don't have the corresponding paraphrases in the other sentence.Under this setting, we frame paraphrasing as a sequence tagging task.We collect 30k sentence pairs in English with the new annotation schema, resulting in the ParaTag dataset.In addition to reporting baseline results on ParaTag using state-of-art language models, we show that ParaTag is especially useful for training an automatic scorer for language generation evaluation.Finally, we train a paraphrase generation model from ParaTag and achieve better data augmentation performance on the GLUE benchmark than other public paraphrasing datasets.1
Shuohang Wang, Ruochen Xu, Yang Liu 0124, Chenguang Zhu 0001, Michael Zeng 0001
EMNLP4
2022 A Unified Encoder-Decoder Framework with Entity Memory
abstract
Entities, as important carriers of real-world knowledge, play a key role in many NLP tasks.We focus on incorporating entity knowledge into an encoder-decoder framework for informative text generation.Existing approaches tried to index, retrieve, and read external documents as evidence, but they suffered from a large computational overhead.In this work, we propose an Encoder-Decoder framework with an entity Memory, namely EDMem.The entity knowledge is stored in the memory as latent representations, and the memory is pre-trained on Wikipedia along with encoder-decoder parameters.To precisely generate entity names, we design three decoding methods to constrain entity generation by linking entities in the memory.EDMem is a unified framework that can be used on various entity-intensive question answering and generation tasks.Extensive experimental results show that EDMem outperforms both memory-based auto-encoder models and non-memory encoder-decoder models.
Zhihan Zhang 0001, Wenhao Yu 0002, Chenguang Zhu 0001, Meng Jiang 0001
EMNLP3
2022 Towards a Unified Multi-Dimensional Evaluator for Text Generation
abstract
Multi-dimensional evaluation is the dominant paradigm for human evaluation in Natural Language Generation (NLG), i.e., evaluating the generated text from multiple explainable dimensions, such as coherence and fluency.However, automatic evaluation in NLG is still dominated by similarity-based metrics, and we lack a reliable framework for a more comprehensive evaluation of advanced models.In this paper, we propose a unified multi-dimensional evaluator UNIEVAL for NLG.We re-frame NLG evaluation as a Boolean Question Answering (QA) task, and by guiding the model with different questions, we can use one evaluator to evaluate from multiple dimensions.Furthermore, thanks to the unified Boolean QA format, we are able to introduce an intermediate learning phase that enables UNIEVAL to incorporate external knowledge from multiple related tasks and gain further improvement.Experiments on three typical NLG tasks show that UNIEVAL correlates substantially better with human judgments than existing metrics.Specifically, compared to the top-performing unified evaluators, UNIEVAL achieves a 23% higher correlation on text summarization, and over 43% on dialogue response generation.Also, UNIEVAL demonstrates a strong zero-shot learning ability for unseen evaluation dimensions and tasks.Source code, data and all pre-trained evaluators are available on our GitHub repository 1 . Generated Summary:Harry Kane is nominated for both the PFA player and young player of the season.The Spurs striker has been released from the awards ceremony on Sunday.The Tottenham striker features in a new animation.Reference Summary: Harry Kane has been in superb form for Tottenham this season.The 21-year-old has scored 30 goals in all competitions for Spurs.Kane also made his England debut and scored within two minutes.Document: Harry Kane's celebrations this season have always shown him to be an animated young man . . .Similarity-based Evaluators ROUGE-1: 0.44 ROUGE-2: 0.25 ROUGE-L: 0.42 BERTScore: 0.24 Single-dimensional Evaluators (predicted by two different evaluators (Deng et al., 2021)) Consistency: 0.87 Relevance: 0.74 Unified Evaluator (predicted by BARTScore, and the scoring range is negative infinity to 0) Precision: -5.45 Recall: -4.93 F1: -5.19
Ming Zhong 0005, Yang Liu 0005, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu 0003, Chenguang Zhu 0001, Heng Ji 0001, Jiawei Han 0001
EMNLP7
2022 Human Parity on CommonsenseQA: Augmenting Self-Attention with External Attention
abstract
Most of today's AI systems focus on using self-attention mechanisms and transformer architectures on large amounts of diverse data to achieve impressive performance gains. In this paper, we propose to augment the transformer architecture with an external attention mechanism to bring external knowledge and context to bear. By integrating external information into the prediction process, we hope to reduce the need for ever-larger models and increase the democratization of AI systems. We find that the proposed external attention mechanism can significantly improve the performance of existing AI systems, allowing practitioners to easily customize foundation AI models to many diverse downstream applications. In particular, we focus on the task of Commonsense Reasoning, demonstrating that the proposed external attention mechanism can augment existing transformer models and significantly improve the model's reasoning capabilities. The proposed system, Knowledgeable External Attention for commonsense Reasoning (KEAR), reaches human parity on the open CommonsenseQA research benchmark with an accuracy of 89.4% in comparison to the human accuracy of 88.9%.
Yichong Xu, Chenguang Zhu 0001, Shuohang Wang, Hao Cheng 0002, Xiaodong Liu 0003, Jianfeng Gao 0001, Michael Zeng 0001, Xuedong Huang 0001
IJCAI2
2022 REVIVE: Regional Visual Representation Matters in Knowledge-Based Visual Question Answering
abstract
This paper revisits visual representation in knowledge-based visual question answering (VQA) and demonstrates that using regional information in a better way can significantly improve the performance. While visual representation is extensively studied in traditional VQA, it is under-explored in knowledge-based VQA even though these two tasks share the common spirit, i.e., rely on visual input to answer the question. Specifically, we observe in most state-of-the-art knowledge-based VQA methods: 1) visual features are extracted either from the whole image or in a sliding window manner for retrieving knowledge, and the important relationship within/among object regions is neglected; 2) visual features are not well utilized in the final answering model, which is counter-intuitive to some extent. Based on these observations, we propose a new knowledge-based VQA method REVIVE, which tries to utilize the explicit information of object regions not only in the knowledge retrieval stage but also in the answering model. The key motivation is that object regions and inherent relationship are important for knowledge-based VQA. We perform extensive experiments on the standard OK-VQA dataset and achieve new state-of the-art performance, i.e., 58.0 accuracy, surpassing previous state-of-the-art method by a large margin (+3.6%). We also conduct detailed analysis and show the necessity of regional information in different framework components for knowledge-based VQA. Code is publicly available at https://github.com/yzleroy/REVIVE.
Yuanze Lin, Yujia Xie, Dongdong Chen 0001, Yichong Xu, Chenguang Zhu 0001, Lu Yuan 0001
NeurIPS5
2022 Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners
abstract
The goal of this work is to build flexible video-language models that can generalize to various video-to-text tasks from few examples. Existing few-shot video-language learners focus exclusively on the encoder, resulting in the absence of a video-to-text decoder to handle generative tasks. Video captioners have been pretrained on large-scale video-language datasets, but they rely heavily on finetuning and lack the ability to generate text for unseen tasks in a few-shot setting. We propose VidIL, a few-shot Video-language Learner via Image and Language models, which demonstrates strong performance on few-shot video-to-text tasks without the necessity of pretraining or finetuning on any video datasets. We use image-language models to translate the video content into frame captions, object, attribute, and event phrases, and compose them into a temporal-aware template. We then instruct a language model, with a prompt containing a few in-context examples, to generate a target output from the composed content. The flexibility of prompting allows the model to capture any form of text input, such as automatic speech recognition (ASR) transcripts. Our experiments demonstrate the power of language models in understanding videos on a wide variety of video-language tasks, including video captioning, video question answering, video caption retrieval, and video future event prediction. Especially, on video future event prediction, our few-shot model significantly outperforms state-of-the-art supervised models trained on large-scale video datasets.Code and processed data are publicly available for research purposes at https://github.com/MikeWangWZHL/VidIL.
Zhenhailong Wang, Manling Li, Ruochen Xu, Luowei Zhou, Jie Lei 0003, Xudong Lin 0003, Shuohang Wang, Ziyi Yang 0011, Chenguang Zhu 0001, Derek Hoiem, Shih-Fu Chang, Mohit Bansal, Heng Ji 0001
NeurIPS9
2021 RADDLE: An Evaluation Benchmark and Analysis Platform for Robust Task-oriented Dialog Systems
abstract
Baolin Peng, Chunyuan Li, Zhu Zhang, Chenguang Zhu, Jinchao Li, Jianfeng Gao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Baolin Peng, Chunyuan Li, Zhu (Drew) Zhang, Chenguang Zhu 0001, Jinchao Li, Jianfeng Gao 0001
ACL/IJCNLP (1)4
2021 Injecting Entity Types into Entity-Guided Text Generation
abstract
Recent successes in deep generative modeling have led to significant advances in natural language generation (NLG).Incorporating entities into neural generation models has demonstrated great improvements by assisting to infer the summary topic and to generate coherent content.To enhance the role of entity in NLG, in this paper, we aim to model the entity type in the decoding phase to generate contextual words accurately.We develop a novel NLG model to produce a target sequence based on a given list of entities.Our model has a multistep decoder that injects the entity types into the process of entity mention generation.Experiments on two public news datasets demonstrate type injection performs better than existing type embedding concatenation baselines.
Xiangyu Dong 0004, Wenhao Yu 0002, Chenguang Zhu 0001, Meng Jiang 0001
EMNLP (1)3
2021 Sentence-Permuted Paragraph Generation
abstract
Generating paragraphs of diverse contents is important in many applications.Existing generation models produce similar contents from homogenized contexts due to the fixed left-toright sentence order.Our idea is permuting the sentence orders to improve the content diversity of multi-sentence paragraph.We propose a novel framework PermGen whose objective is to maximize the expected log-likelihood of output paragraph distributions with respect to all possible sentence orders.PermGen uses hierarchical positional embedding and designs new procedures for both training phase and inference phase.Experiments on three paragraph generation benchmarks demonstrate Per-mGen generates more diverse outputs with a higher quality than existing models.
Wenhao Yu 0002, Chenguang Zhu 0001, Tong Zhao 0003, Zhichun Guo, Meng Jiang 0001
EMNLP (1)2
2021 Data Augmentation for Spoken Language Understanding via Pretrained Language Models
abstract
The training of spoken language understanding (SLU) models often faces the problem of data scarcity. In this paper, we put forward a data augmentation method using pretrained language models to boost the variability and accuracy of generated utterances. Furthermore, we investigate and propose solutions to two previously overlooked semi-supervised learning scenarios of data scarcity in SLU: i) Rich-in-Ontology: ontology information with numerous valid dialogue acts is given; ii) Rich-in-Utterance: a large number of unlabelled utterances are available. Empirical results show that our method can produce synthetic training data that boosts the performance of language understanding models in various scenarios.
Baolin Peng, Chenguang Zhu 0001, Michael Zeng 0001, Jianfeng Gao 0001
Interspeech2
2021 SPLAT: Speech-Language Joint Pre-Training for Spoken Language Understanding
abstract
Spoken language understanding (SLU) requires a model to analyze input acoustic signal to understand its linguistic content and make predictions.To boost the models' performance, various pre-training methods have been proposed to learn rich representations from large-scale unannotated speech and text.However, the inherent disparities between the two modalities necessitate a mutual analysis.In this paper, we propose a novel semisupervised learning framework, SPLAT, to jointly pre-train the speech and language modules.Besides conducting a self-supervised masked language modeling task on the two individual modules using unpaired speech and text, SPLAT aligns representations from the two modules in a shared latent space using a small amount of paired speech and text.Thus, during fine-tuning, the speech module alone can produce representations carrying both acoustic information and contextual semantic knowledge of an input acoustic signal.Experimental results verify the effectiveness of our approach on various SLU tasks.For example, SPLAT improves the previous stateof-the-art performance on the Spoken SQuAD dataset by more than 10%.
Yu-An Chung, Chenguang Zhu 0001, Michael Zeng 0001
NAACL-HLT2
2021 Enhancing Factual Consistency of Abstractive Summarization
abstract
Chenguang Zhu, William Hinthorn, Ruochen Xu, Qingkai Zeng, Michael Zeng, Xuedong Huang, Meng Jiang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Chenguang Zhu 0001, William Hinthorn, Ruochen Xu, Qingkai Zeng 0001, Michael Zeng 0001, Xuedong Huang 0001, Meng Jiang 0001
NAACL-HLT1
2021 MediaSum: A Large-scale Media Interview Dataset for Dialogue Summarization
abstract
This paper introduces MEDIASUM 1 , a largescale media interview dataset consisting of 463.6K transcripts with abstractive summaries.To create this dataset, we collect interview transcripts from NPR and CNN and employ the overview and topic descriptions as summaries.Compared with existing public corpora for dialogue summarization, our dataset is an order of magnitude larger and contains complex multi-party conversations from multiple domains.We conduct statistical analysis to demonstrate the unique positional bias exhibited in the transcripts of televised and radioed interviews.We also show that MEDIASUM can be used in transfer learning to improve a model's performance on other dialogue summarization tasks.
Chenguang Zhu 0001, Yang Liu 0124, Michael Zeng 0001
NAACL-HLT1
2021 Leveraging Lead Bias for Zero-shot Abstractive News Summarization
abstract
A typical journalistic convention in news articles is to deliver the most salient information in the beginning, also known as the lead bias. While this phenomenon can be exploited in generating a summary, it has a detrimental effect on teaching a model to discriminate and extract important information in general. We propose that this lead bias can be leveraged in our favor in a simple and effective way to pre-train abstractive news summarization models on large-scale unlabeled news corpora: predicting the leading sentences using the rest of an article. We collect a massive news corpus and conduct data cleaning and filtering via statistical analysis. We then apply self-supervised pre-training on this dataset to existing generation models BART and T5 for domain adaptation. Via extensive experiments on six benchmark datasets, we show that this approach can dramatically improve the summarization quality and achieve state-of-the-art results for zero-shot news summarization without any fine-tuning. For example, in the DUC2003 dataset, the ROUGE-1 score of BART increases 13.7% after the lead-bias pre-training. We deploy the model in Microsoft News and provide public APIs as well as a demo website for multi-lingual news summarization.
Chenguang Zhu 0001, Ziyi Yang 0011, Robert Gmyr, Michael Zeng 0001, Xuedong Huang 0001
SIGIR1
2019 Embedding Imputation with Grounded Language Information
abstract
Due to the ubiquitous use of embeddings as input representations for a wide range of natural language tasks, imputation of embeddings for rare and unseen words is a critical problem in language processing.Embedding imputation involves learning representations for rare or unseen words during the training of an embedding model, often in a post-hoc manner.In this paper, we propose an approach for embedding imputation which uses grounded information in the form of a knowledge graph.This is in contrast to existing approaches which typically make use of vector space properties or subword information.We propose an online method to construct a graph from grounded information and design an algorithm to map from the resulting graphical structure to the space of the pre-trained embeddings.Finally, we evaluate our approach on a range of rare and unseen word tasks across various domains and show that our model can learn better representations.For example, on the Card-660 task our method improves Pearson's and Spearman's correlation coefficients upon the stateof-the-art by 11% and 17.8% respectively using GloVe embeddings.
Ziyi Yang 0011, Chenguang Zhu 0001, Vin Sachidananda, Eric Darve
ACL (1)2
2019 Multi-task Learning for Natural Language Generation in Task-Oriented Dialogue
abstract
Chenguang Zhu, Michael Zeng, Xuedong Huang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Chenguang Zhu 0001, Michael Zeng 0001, Xuedong Huang 0001
EMNLP/IJCNLP (1)1
2019 SIM: A Slot-Independent Neural Model for Dialogue State Tracking
abstract
Dialogue state tracking is an important component in task-oriented dialogue systems to identify users' goals and requests as a dialogue proceeds.However, as most previous models are dependent on dialogue slots, the model complexity soars when the number of slots increases.In this paper, we put forward a slotindependent neural model (SIM) to track dialogue states while keeping the model complexity invariant to the number of dialogue slots.The model utilizes attention mechanisms between user utterance and system actions.SIM achieves state-of-the-art results on WoZ and DSTC2 tasks, with only 20% of the model size of previous models.
Chenguang Zhu 0001, Michael Zeng 0001, Xuedong Huang 0001
SIGdial1
2012 Information diffusion and external influence in networks
abstract
Social networks play a fundamental role in the diffusion of information. However, there are two different ways of how information reaches a person in a network. Information reaches us through connections in our social networks, as well as through the influence external out-of-network sources, like the mainstream media. While most present models of information adoption in networks assume information only passes from a node to node via the edges of the underlying network, the recent availability of massive online social media data allows us to study this process in more detail.
Seth A. Myers, Chenguang Zhu 0001, Jure Leskovec
KDD2