EDBT 2026 Demo / reviewers in the wild / expert
Shuohang Wang
dblp:173/5469
· DBLP profile ↗
45ranked-venue papers
9as first author
28since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 9 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video GenerationabstractThe current state-of-the-art video generative models can produce commercial-grade videos with highly realistic details. However, they still struggle to coherently present multiple sequential events in the stories specified by the prompts, which is foreseeable an essential capability for future long video generation scenarios. For example, top T2V generative models still fail to generate a video of the short simple story "how to put an elephant into a refrigerator." While existing detail-oriented benchmarks primarily focus on fine-grained metrics like aesthetic quality and spatial-temporal consistency, they fall short of evaluating models’ abilities to handle event-level story presentation. To address this gap, we introduce StoryEval, a story-oriented benchmark specifically designed to assess text-to-video (T2V) models’ story-completion capabilities. StoryEval features 423 prompts spanning 7 classes, each representing short stories composed of 2–4 consecutive events. We employ Vision-Language Models, such as GPT-4o and LLaVA-OV-Chat-72B, to verify the completion of each event in the generated videos, applying a unanimous voting method to enhance reliability. Our methods ensure high alignment with human evaluations, and the evaluation of 11 models reveals its challenge, with none exceeding an average story-completion rate of 50%. StoryEval provides a new benchmark for advancing T2V models and highlights the challenges and opportunities in developing next-generation solutions for coherent story-driven video generation. Project website is available at https://ypwang61.github.io/project/StoryEval."The universe is made of stories, not of atoms."— Muriel Rukeyser Xuehai He, Shuohang Wang, Simon S. Du, Yelong Shen |
CVPR | 6 |
| 2025 | Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long GenerationabstractRecent advances in language modeling have demonstrated the effectiveness of State Space Models (SSMs) for efficient sequence modeling. While hybrid architectures such as Samba and the decoder-decoder architecture, YOCO, have shown promising performance gains over Transformers, prior works have not investigated the efficiency potential of representation sharing between SSM layers. In this paper, we introduce the Gated Memory Unit (GMU), a simple yet effective mechanism for efficient memory sharing across layers. We apply it to create SambaY, a decoder-hybrid-decoder architecture that incorporates GMUs in the cross-decoder to share memory readout states from a Samba-based self-decoder. SambaY significantly enhances decoding efficiency, preserves linear pre-filling time complexity, and boosts long-context performance, all while eliminating the need for explicit positional encoding. Through extensive scaling experiments, we demonstrate that our model exhibits a significantly lower irreducible loss compared to a strong YOCO baseline, indicating superior performance scalability under large-scale compute regimes. Our largest model enhanced with Differential Attention, Phi4-mini-Flash-Reasoning, achieves significantly better performance than Phi4-mini-Reasoning on reasoning tasks such as Math500, AIME24/25, and GPQA Diamond without any reinforcement learning, while delivering up to 10× higher decoding throughput on 2K-length prompts with 32K generation length under the vLLM inference framework. We release our training codebase on open-source data at https://github.com/microsoft/ArchScale. Liliang Ren, Young Jin Kim 0006, Adam Atkinson, Zheng Zhan 0001, Jiankai Sun, Baolin Peng, Shuohang Wang, Hao Cheng 0002, Jianfeng Gao 0001, Weizhu Chen, Yelong Shen |
NeurIPS | 10 |
| 2025 | Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleabstractWe show that reinforcement learning with verifiable reward using one training example (1-shot RLVR) is effective in incentivizing the math reasoning capabilities of large language models (LLMs). Applying RLVR to the base model Qwen2.5-Math-1.5B, we identify a single example that elevates model performance on MATH500 from 36.0\% to 73.6\% (8.6\% improvement beyond format correction), and improves the average performance across six common mathematical reasoning benchmarks from 17.6\% to 35.7\% (7.0\% non-format gain). This result matches the performance obtained using the 1.2k DeepScaleR subset (MATH500: 73.6\%, average: 35.9\%), which contains the aforementioned example. Furthermore, RLVR with only two examples even slightly exceeds these results (MATH500: 74.8\%, average: 36.6\%). Similar substantial improvements are observed across various models (Qwen2.5-Math-7B, Llama3.2-3B-Instruct, DeepSeek-R1-Distill-Qwen-1.5B), RL algorithms (GRPO and PPO), and different math examples.
In addition, we identify some interesting phenomena during 1-shot RLVR, including cross-category generalization, increased frequency of self-reflection, and sustained test performance improvement even after the training accuracy has saturated, a phenomenon we term \textit{post-saturation generalization}.
Moreover, we verify that the effectiveness of 1-shot RLVR primarily arises from the policy gradient loss, distinguishing it from the "grokking" phenomenon.
We also show the critical role of promoting exploration (e.g., by incorporating entropy loss with an appropriate coefficient) in 1-shot RLVR training.
We also further discuss related observations about format correction, label robustness and prompt modification.
These findings can inspire future work on RLVR efficiency and encourage a re-examination of recent progress and the underlying mechanisms in RLVR.
Our code, models, and data are open source at https://github.com/ypwang61/One-Shot-RLVR. Liliang Ren, Baolin Peng, Hao Cheng 0002, Xuehai He, Jianfeng Gao 0001, Weizhu Chen, Shuohang Wang, Simon S. Du, Yelong Shen |
NeurIPS | 12 |
| 2025 | Routing Mamba: Scaling State Space Models with Mixture-of-Experts ProjectionabstractState Space Models (SSMs) offer remarkable performance gains in efficient sequence modeling, with constant per-step inference-time computation and memory complexity. Recent advances, such as Mamba, further enhance SSMs with input-dependent gating and hardware-aware implementations, positioning them as strong alternatives to Transformers for long sequence modeling. However, efficiently scaling the expressive power of SSMs, particularly with Mixture of Experts (MoE), remains challenging, as naive integration attempts often falter or degrade performance. In this work, we introduce Routing Mamba (RoM), a novel approach that scales SSM parameters using sparse mixtures of linear projection experts. By sharing routing decisions between projection layers and lightweight sub-modules within Mamba across experts, RoM leverages synergies among linear projection experts for effective and efficient sparse scaling of Mamba layers. At a scale of 1.3B active parameters (10B total) and 16K training sequence length, RoM achieves language modeling performance equivalent to a dense Mamba model requiring over 2.3$\times$ more active parameters, and demonstrates consistent perplexity across context lengths. Experimental results further show RoM effectively scales hybrid language models, yielding a 23% FLOPS saving compared to dense Mamba scaling for similar performance. We release our training codebase at https://github.com/zhanzheng8585/Routing-Mamba. Zheng Zhan 0001, Liliang Ren, Shuohang Wang, Yeyun Gong, Yanzhi Wang 0001, Yelong Shen |
NeurIPS | 3 |
| 2024 | PEARL: Prompting Large Language Models to Plan and Execute Actions Over Long DocumentsabstractSimeng Sun, Yang Liu, Shuohang Wang, Dan Iter, Chenguang Zhu, Mohit Iyyer. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Simeng Sun, Yang Liu 0124, Shuohang Wang, Dan Iter, Chenguang Zhu 0001, Mohit Iyyer |
EACL (1) | 3 |
| 2024 | SciAgent: Tool-augmented Language Models for Scientific ReasoningabstractYubo Ma, Zhibin Gou, Junheng Hao, Ruochen Xu, Shuohang Wang, Liangming Pan, Yujiu Yang, Yixin Cao, Aixin Sun. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yubo Ma, Zhibin Gou, Junheng Hao, Ruochen Xu, Shuohang Wang, Liangming Pan, Yujiu Yang 0001, Yixin Cao 0002, Aixin Sun |
EMNLP | 5 |
| 2023 | APOLLO: A Simple Approach for Adaptive Pretraining of Language Models for Logical ReasoningabstractSoumya Sanyal, Yichong Xu, Shuohang Wang, Ziyi Yang, Reid Pryzant, Wenhao Yu, Chenguang Zhu, Xiang Ren. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Soumya Sanyal 0001, Yichong Xu, Shuohang Wang, Ziyi Yang 0011, Reid Pryzant, Wenhao Yu 0002, Chenguang Zhu 0001, Xiang Ren 0001 |
ACL (1) | 3 |
| 2023 | G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentabstractThe quality of texts generated by natural language generation (NLG) systems is hard to measure automatically.Conventional referencebased metrics, such as BLEU and ROUGE, have been shown to have relatively low correlation with human judgments, especially for tasks that require creativity and diversity.Recent studies suggest using large language models (LLMs) as reference-free metrics for NLG evaluation, which have the benefit of being applicable to new tasks that lack human references.However, these LLM-based evaluators still have lower human correspondence than medium-size neural evaluators.In this work, we present G-EVAL, a framework of using large language models with chain-of-thoughts (CoT) and a form-filling paradigm, to assess the quality of NLG outputs.We experiment with two generation tasks, text summarization and dialogue generation.We show that G-EVAL with GPT-4 as the backbone model achieves a Spearman correlation of 0.514 with human on summarization task, outperforming all previous methods by a large margin.We also propose analysis on the behavior of LLM-based evaluators, and highlight the potential concern of LLM-based evaluators having a bias towards the LLM-generated texts. 1 Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu |
EMNLP | 4 |
| 2023 | The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT InteractionsabstractSiru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu, Heng Ji, Jiawei Han. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Siru Ouyang, Shuohang Wang, Ming Zhong 0005, Yizhu Jiao, Dan Iter, Reid Pryzant, Chenguang Zhu 0001, Heng Ji 0001, Jiawei Han 0001 |
EMNLP | 2 |
| 2023 | Generate rather than Retrieve: Large Language Models are Strong Context Generators
Wenhao Yu 0002, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal 0001, Chenguang Zhu 0001, Michael Zeng 0001, Meng Jiang 0001 |
ICLR | 3 |
| 2023 | Prompting GPT-3 To Be Reliable
Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jordan L. Boyd-Graber |
ICLR | 4 |
| 2023 | Sparse Modular Activation for Efficient Sequence ModelingabstractRecent hybrid models combining Linear State Space Models (SSMs) with self-attention mechanisms have demonstrated impressive results across a range of sequence modeling tasks. However, current approaches apply attention modules statically and uniformly to all elements in the input sequences, leading to sub-optimal quality-efficiency trade-offs. To address this limitation, we introduce Sparse Modular Activation (SMA), a general mechanism enabling neural networks to sparsely and dynamically activate sub-modules for sequence elements in a differentiable manner. Through allowing each element to skip non-activated sub-modules, SMA reduces computation and memory consumption of neural networks at both training and inference stages. To validate the effectiveness of SMA on sequence modeling, we design a novel neural architecture, SeqBoat, which employs SMA to sparsely activate a Gated Attention Unit (GAU) based on the state representations learned from an SSM. By constraining the GAU to only conduct local attention on the activated inputs, SeqBoat can achieve linear inference complexity with theoretically infinite attention span, and provide substantially better quality-efficiency trade-off than the chunking-based models. With experiments on a wide range of tasks, including long sequence modeling, speech classification and language modeling, SeqBoat brings new state-of-the-art results among hybrid models with linear complexity, and reveals the amount of attention needed for each task through the learned sparse activation patterns. Our code is publicly available at https://github.com/renll/SeqBoat. Liliang Ren, Shuohang Wang, Yichong Xu, ChengXiang Zhai |
NeurIPS | 3 |
| 2022 | Playing Lottery Tickets with Vision and LanguageabstractLarge-scale pre-training has recently revolutionized vision-and-language (VL) research. Models such as LXMERT and UNITER have significantly lifted the state of the art over a wide range of VL tasks. However, the large number of parameters in such models hinders their application in practice. In parallel, work on the lottery ticket hypothesis (LTH) has shown that deep neural networks contain small matching subnetworks that can achieve on par or even better performance than the dense networks when trained in isolation. In this work, we perform the first empirical study to assess whether such trainable subnetworks also exist in pre-trained VL models. We use UNITER as the main testbed (also test on LXMERT and ViLT), and consolidate 7 representative VL tasks for experiments, including visual question answering, visual commonsense reasoning, visual entailment, referring expression comprehension, image-text retrieval, GQA, and NLVR2. Through comprehensive analysis, we summarize our main findings as follows. (i) It is difficult to find subnetworks that strictly match the performance of the full model. However, we can find relaxed winning tickets at 50%-70% sparsity that maintain 99% of the full accuracy. (ii) Subnetworks found by task-specific pruning transfer reasonably well to the other tasks, while those found on the pre-training tasks at 60%/70% sparsity transfer universally, matching 98%/96% of the full accuracy on average over all the tasks. (iii) Besides UNITER, other models such as LXMERT and ViLT can also play lottery tickets. However, the highest sparsity we can achieve for ViLT is far lower than LXMERT and UNITER (30% vs. 70%). (iv) LTH also remains relevant when using other training methods (e.g., adversarial training). Zhe Gan, Yen-Chun Chen 0001, Tianlong Chen 0001, Yu Cheng 0001, Shuohang Wang, Jingjing Liu 0001, Zicheng Liu 0001 |
AAAI | 6 |
| 2022 | Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training DataabstractShuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu, Siqi Sun, Ruochen Xu, Chenguang Zhu, Michael Zeng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Shuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu 0124, Ruochen Xu, Chenguang Zhu 0001, Michael Zeng 0001 |
ACL (1) | 1 |
| 2022 | KG-FiD: Infusing Knowledge Graph in Fusion-in-Decoder for Open-Domain Question AnsweringabstractDonghan Yu, Chenguang Zhu, Yuwei Fang, Wenhao Yu, Shuohang Wang, Yichong Xu, Xiang Ren, Yiming Yang, Michael Zeng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Donghan Yu, Chenguang Zhu 0001, Yuwei Fang, Wenhao Yu 0002, Shuohang Wang, Yichong Xu, Xiang Ren 0001, Yiming Yang 0002, Michael Zeng 0001 |
ACL (1) | 5 |
| 2022 | An Empirical Study of Training End-to-End Vision-and-Language TransformersabstractVision-and-language (VL) pre-training has proven to be highly effective on various VL downstream tasks. While recent work has shown that fully transformer-based VL models can be more efficient than previous region-feature-based methods, their performance on downstream tasks often degrades significantly. In this paper, we present Meter, a Multimodal End-to-end TransformER framework, through which we investigate how to design and pre-train a fully transformer-based VL model in an end-to-end manner. Specifically, we dissect the model designs along multiple dimensions: vision encoders (e.g., CLIP-ViT, Swin transformer), text encoders (e.g., RoBERTa, De-BERTa), multimodal fusion module (e.g., merged attention vs. co-attention), architectural design (e.g., encoder-only vs. encoder-decoder), and pre-training objectives (e.g., masked image modeling). We conduct comprehensive experiments and provide insights on how to train a performant VL transformer. Meterachieves an accuracy of 77.64% on the VQAv2 test-std set using only 4M images for pre-training, surpassing the state-of-the-art region-feature-based model by 1.04%, and outperforming the previous best fully transformer-based model by 1.6%. Notably, when further scaled up, our best VQA model achieves an accuracy of 80.54%. Code and pre-trained models are released at https://github.com/zdou0830/METER. Zi-Yi Dou, Yichong Xu, Zhe Gan, Shuohang Wang, Chenguang Zhu 0001, Pengchuan Zhang, Lu Yuan 0001, Nanyun Peng 0001, Zicheng Liu 0001, Michael Zeng 0001 |
CVPR | 5 |
| 2022 | CLIP-Event: Connecting Text and Images with Event StructuresabstractVision-language (V+L) pretraining models have achieved great success in supporting multimedia applications by understanding the alignments between images and text. While existing vision-language pretraining models primarily focus on understanding objects in images or entities in text, they often ignore the alignment at the level of events and their argument structures. In this work, we propose a contrastive learning framework to enforce vision-language pretraining models to comprehend events and associated argument (participant) roles. To achieve this, we take advantage of text information extraction technologies to obtain event structural knowledge, and utilize multiple prompt functions to contrast difficult negative descriptions by manipulating event structures. We also design an event graph alignment loss based on optimal transport to capture event argument structures. In addition, we collect a large event-rich dataset (106,875 images) for pretraining, which provides a more challenging image retrieval benchmark to assess the understanding of complicated lengthy sentences11The data and code are publicly available for research purpose in https://github.com/limanling/clip-event.. Experiments show that our zero-shot CLIP-Event outperforms the state-of-the-art supervised model in argument extraction on Multimedia Event Extraction, achieving more than 5% absolute F-score gain in event extraction, as well as significant improvements on a variety of downstream tasks under zero-shot settings. Manling Li, Ruochen Xu, Shuohang Wang, Luowei Zhou, Xudong Lin 0003, Chenguang Zhu 0001, Michael Zeng 0001, Heng Ji 0001, Shih-Fu Chang |
CVPR | 3 |
| 2022 | Retrieval Augmentation for Commonsense Reasoning: A Unified ApproachabstractA common thread of retrieval-augmented methods in the existing literature focuses on retrieving encyclopedic knowledge, such as Wikipedia, which facilitates well-defined entity and relation spaces that can be modeled.However, applying such methods to commonsense reasoning tasks faces two unique challenges, i.e., the lack of a general large-scale corpus for retrieval and a corresponding effective commonsense retriever.In this paper, we systematically investigate how to leverage commonsense knowledge retrieval to improve commonsense reasoning tasks.We proposed a unified framework of Retrieval-Augmented Commonsense reasoning (called RACO), including a newly constructed commonsense corpus with over 20 million documents and novel strategies for training a commonsense retriever.We conducted experiments on four different commonsense reasoning tasks.Extensive evaluation results showed that our proposed RACO can significantly outperform other knowledgeenhanced method counterparts, achieving new SoTA performance on the CommonGen 1 and CREAK 2 leaderboards.Our code is available at https://github.com/wyu97/RACo. Wenhao Yu 0002, Chenguang Zhu 0001, Zhihan Zhang 0001, Shuohang Wang, Zhuosheng Zhang 0001, Yuwei Fang, Meng Jiang 0001 |
EMNLP | 4 |
| 2022 | Empowering Language Models with Knowledge Graph Reasoning for Open-Domain Question AnsweringabstractAnswering open-domain questions requires world knowledge about in-context entities.As pre-trained Language Models (LMs) lack the power to store all required knowledge, external knowledge sources, such as knowledge graphs, are often used to augment LMs.In this work, we propose knOwledge REasOning empowered Language Model (OREOLM), which consists of a novel Knowledge Interaction Layer that can be flexibly plugged into existing Transformer-based LMs to interact with a differentiable Knowledge Graph Reasoning module collaboratively.In this way, LM guides KG to walk towards the desired answer, while the retrieved knowledge improves LM.By adopting OREOLM to RoBERTa and T5, we show significant performance gain, achieving state-of-art results in the Closed-Book setting.The performance enhancement is mainly from the KG reasoning's capacity to infer missing relational facts.In addition, OREOLM provides reasoning paths as rationales to interpret the model's decision. Ziniu Hu, Yichong Xu, Wenhao Yu 0002, Shuohang Wang, Ziyi Yang 0011, Chenguang Zhu 0001, Kai-Wei Chang 0001, Yizhou Sun |
EMNLP | 4 |
| 2022 | ParaTag: A Dataset of Paraphrase Tagging for Fine-Grained Labels, NLG Evaluation, and Data AugmentationabstractParaphrase identification has been formulated as a binary classification task to decide whether two sentences hold a paraphrase relationship.Existing paraphrase datasets only annotate a binary label for each sentence pair.However, after a systematical analysis of existing paraphrase datasets, we found that the degree of paraphrase cannot be well characterized by a single binary label.And the criteria of paraphrase are not even consistent within the same dataset.We hypothesize that such issues would limit the effectiveness of paraphrase models trained on these data.To this end, we propose a novel fine-grained paraphrase annotation schema that labels the minimum spans of tokens in a sentence that don't have the corresponding paraphrases in the other sentence.Under this setting, we frame paraphrasing as a sequence tagging task.We collect 30k sentence pairs in English with the new annotation schema, resulting in the ParaTag dataset.In addition to reporting baseline results on ParaTag using state-of-art language models, we show that ParaTag is especially useful for training an automatic scorer for language generation evaluation.Finally, we train a paraphrase generation model from ParaTag and achieve better data augmentation performance on the GLUE benchmark than other public paraphrasing datasets.1 Shuohang Wang, Ruochen Xu, Yang Liu 0124, Chenguang Zhu 0001, Michael Zeng 0001 |
EMNLP | 1 |
| 2022 | Human Parity on CommonsenseQA: Augmenting Self-Attention with External AttentionabstractMost of today's AI systems focus on using self-attention mechanisms and transformer architectures on large amounts of diverse data to achieve impressive performance gains. In this paper, we propose to augment the transformer architecture with an external attention mechanism to bring external knowledge and context to bear. By integrating external information into the prediction process, we hope to reduce the need for ever-larger models and increase the democratization of AI systems. We find that the proposed external attention mechanism can significantly improve the performance of existing AI systems, allowing practitioners to easily customize foundation AI models to many diverse downstream applications. In particular, we focus on the task of Commonsense Reasoning, demonstrating that the proposed external attention mechanism can augment existing transformer models and significantly improve the model's reasoning capabilities. The proposed system, Knowledgeable External Attention for commonsense Reasoning (KEAR), reaches human parity on the open CommonsenseQA research benchmark with an accuracy of 89.4% in comparison to the human accuracy of 88.9%. Yichong Xu, Chenguang Zhu 0001, Shuohang Wang, Hao Cheng 0002, Xiaodong Liu 0003, Jianfeng Gao 0001, Michael Zeng 0001, Xuedong Huang 0001 |
IJCAI | 3 |
| 2022 | Language Models with Image Descriptors are Strong Few-Shot Video-Language LearnersabstractThe goal of this work is to build flexible video-language models that can generalize to various video-to-text tasks from few examples. Existing few-shot video-language learners focus exclusively on the encoder, resulting in the absence of a video-to-text decoder to handle generative tasks. Video captioners have been pretrained on large-scale video-language datasets, but they rely heavily on finetuning and lack the ability to generate text for unseen tasks in a few-shot setting. We propose VidIL, a few-shot Video-language Learner via Image and Language models, which demonstrates strong performance on few-shot video-to-text tasks without the necessity of pretraining or finetuning on any video datasets. We use image-language models to translate the video content into frame captions, object, attribute, and event phrases, and compose them into a temporal-aware template. We then instruct a language model, with a prompt containing a few in-context examples, to generate a target output from the composed content. The flexibility of prompting allows the model to capture any form of text input, such as automatic speech recognition (ASR) transcripts. Our experiments demonstrate the power of language models in understanding videos on a wide variety of video-language tasks, including video captioning, video question answering, video caption retrieval, and video future event prediction. Especially, on video future event prediction, our few-shot model significantly outperforms state-of-the-art supervised models trained on large-scale video datasets.Code and processed data are publicly available for research purposes at https://github.com/MikeWangWZHL/VidIL. Zhenhailong Wang, Manling Li, Ruochen Xu, Luowei Zhou, Jie Lei 0003, Xudong Lin 0003, Shuohang Wang, Ziyi Yang 0011, Chenguang Zhu 0001, Derek Hoiem, Shih-Fu Chang, Mohit Bansal, Heng Ji 0001 |
NeurIPS | 7 |
| 2021 | FILTER: An Enhanced Fusion Method for Cross-lingual Language UnderstandingabstractLarge-scale cross-lingual language models (LM), such as mBERT, Unicoder and XLM, have achieved great success in cross-lingual representation learning. However, when applied to zero-shot cross-lingual transfer tasks, most existing methods use only single-language input for LM finetuning, without leveraging the intrinsic cross-lingual alignment between different languages that proves essential for multilingual tasks. In this paper, we propose FILTER, an enhanced fusion method that takes cross-lingual data as input for XLM finetuning. Specifically, FILTER first encodes text input in the source language and its translation in the target language independently in the shallow layers, then performs cross-language fusion to extract multilingual knowledge in the intermediate layers, and finally performs further language-specific encoding. During inference, the model makes predictions based on the text input in the target language and its translation in the source language. For simple tasks such as classification, translated text in the target language shares the same label as the source language. However, this shared label becomes less accurate or even unavailable for more complex tasks such as question answering, NER and POS tagging. To tackle this issue, we further propose an additional KL-divergence self-teaching loss for model training, based on auto-generated soft pseudo-labels for translated text in the target language. Extensive experiments demonstrate that FILTER achieves new state of the art on two challenging multilingual multi-task benchmarks, XTREME and XGLUE. Yuwei Fang, Shuohang Wang, Zhe Gan, Jingjing Liu 0001 |
AAAI | 2 |
| 2021 | EarlyBERT: Efficient BERT Training via Early-bird Lottery TicketsabstractXiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan, Zhangyang Wang, Jingjing Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xiaohan Chen 0001, Yu Cheng 0001, Shuohang Wang, Zhe Gan, Zhangyang Wang, Jingjing Liu 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | UC2: Universal Cross-Lingual Cross-Modal Vision-and-Language Pre-TrainingabstractVision-and-language pre-training has achieved impressive success in learning multimodal representations between vision and language. To generalize this success to non-English languages, we introduce UC2, the first machine translation-augmented framework for cross-lingual cross-modal representation learning. To tackle the scarcity problem of multilingual captions for image datasets, we first augment existing English-only datasets with other languages via machine translation (MT). Then we extend the standard Masked Language Modeling and Image-Text Matching training objectives to multilingual setting, where alignment between different languages is captured through shared visual context (i.e., using image as pivot). To facilitate the learning of a joint embedding space of images and all languages of interest, we further propose two novel pre-training tasks, namely Masked Region-to-Token Modeling (MRTM) and Visual Translation Language Modeling (VTLM), leveraging MT-enhanced translated data. Evaluation on multilingual image-text retrieval and multilingual visual question answering benchmarks demonstrates that our proposed framework achieves new state of the art on diverse non-English benchmarks while maintaining comparable performance to monolingual pre-trained models on English tasks. Mingyang Zhou 0004, Luowei Zhou, Shuohang Wang, Yu Cheng 0001, Zhou Yu 0005, Jingjing Liu 0001 |
CVPR | 3 |
| 2021 | InfoBERT: Improving Robustness of Language Models from An Information Theoretic Perspective
Boxin Wang, Shuohang Wang, Yu Cheng 0001, Zhe Gan, Ruoxi Jia 0001, Bo Li 0026, Jingjing Liu 0001 |
ICLR | 2 |
| 2021 | LightningDOT: Pre-training Visual-Semantic Embeddings for Real-Time Image-Text RetrievalabstractSiqi Sun, Yen-Chun Chen, Linjie Li, Shuohang Wang, Yuwei Fang, Jingjing Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yen-Chun Chen 0001, Shuohang Wang, Yuwei Fang, Jingjing Liu 0001 |
NAACL-HLT | 4 |
| 2021 | The Elastic Lottery Ticket HypothesisabstractLottery Ticket Hypothesis (LTH) raises keen attention to identifying sparse trainable subnetworks, or winning tickets, which can be trained in isolation to achieve similar or even better performance compared to the full models. Despite many efforts being made, the most effective method to identify such winning tickets is still Iterative Magnitude-based Pruning (IMP), which is computationally expensive and has to be run thoroughly for every different network. A natural question that comes in is: can we “transform” the winning ticket found in one network to another with a different architecture, yielding a winning ticket for the latter at the beginning, without re-doing the expensive IMP? Answering this question is not only practically relevant for efficient “once-for-all” winning ticket finding, but also theoretically appealing for uncovering inherently scalable sparse patterns in networks. We conduct extensive experiments on CIFAR-10 and ImageNet, and propose a variety of strategies to tweak the winning tickets found from different networks of the same model family (e.g., ResNets). Based on these results, we articulate the Elastic Lottery Ticket Hypothesis (E-LTH): by mindfully replicating (or dropping) and re-ordering layers for one network, its corresponding winning ticket could be stretched (or squeezed) into a subnetwork for another deeper (or shallower) network from the same family, whose performance is nearly the same competitive as the latter’s winning ticket directly found by IMP. We have also extensively compared E-LTH with pruning-at-initialization and dynamic sparse training methods, as well as discussed the generalizability of E-LTH to different model families, layer types, and across datasets. Code is available at https://github.com/VITA-Group/ElasticLTH. Xiaohan Chen 0001, Yu Cheng 0001, Shuohang Wang, Zhe Gan, Jingjing Liu 0001, Zhangyang Wang |
NeurIPS | 3 |
| 2020 | Multi-Level Head-Wise Match and Aggregation in Transformer for Textual Sequence Matching
Shuohang Wang, Yunshi Lan, Yi Tay, Jing Jiang 0001, Jingjing Liu 0001 |
AAAI | 1 |
| 2020 | Multi-Fact Correction in Abstractive Text SummarizationabstractPre-trained neural abstractive summarization systems have dominated extractive strategies on news summarization performance, at least in terms of ROUGE.However, systemgenerated abstractive summaries often face the pitfall of factual inconsistency: generating incorrect facts with respect to the source text.To address this challenge, we propose Span-Fact, a suite of two factual correction models that leverages knowledge learned from question answering models to make corrections in system-generated summaries via span selection.Our models employ single or multimasking strategies to either iteratively or autoregressively replace entities in order to ensure semantic consistency w.r.t. the source text, while retaining the syntactic structure of summaries generated by abstractive summarization models.Experiments show that our models significantly boost the factual consistency of system-generated summaries without sacrificing summary quality in terms of both automatic metrics and human evaluation.* *Most of this work was done when the first author was an intern at Microsoft.CNNDM Source (CNN) About a quarter of a million Australian homes and businesses have no power after a "once in a decade" storm battered Sydney and nearby areas.About 4,500 people Yue Dong 0002, Shuohang Wang, Zhe Gan, Yu Cheng 0001, Jackie Chi Kit Cheung, Jingjing Liu 0001 |
EMNLP (1) | 2 |
| 2020 | Hierarchical Graph Network for Multi-hop Question AnsweringabstractIn this paper, we present Hierarchical Graph Network (HGN) for multi-hop question answering.To aggregate clues from scattered texts across multiple paragraphs, a hierarchical graph is created by constructing nodes on different levels of granularity (questions, paragraphs, sentences, entities), the representations of which are initialized with pre-trained contextual encoders.Given this hierarchical graph, the initial node representations are updated through graph propagation, and multihop reasoning is performed via traversing through the graph edges for each subsequent sub-task (e.g., paragraph selection, supporting facts extraction, answer prediction).By weaving heterogeneous nodes into an integral unified graph, this hierarchical differentiation of node granularity enables HGN to support different question answering sub-tasks simultaneously.Experiments on the HotpotQA benchmark demonstrate that the proposed model achieves new state of the art, outperforming existing multi-hop QA approaches. 1 Yuwei Fang, Zhe Gan, Rohit Pillai, Shuohang Wang, Jingjing Liu 0001 |
EMNLP (1) | 5 |
| 2020 | Contrastive Distillation on Intermediate Representations for Language Model CompressionabstractExisting language model compression methods mostly use a simple L 2 loss to distill knowledge in the intermediate representations of a large BERT model to a smaller one.Although widely used, this objective by design assumes that all the dimensions of hidden representations are independent, failing to capture important structural knowledge in the intermediate layers of the teacher network.To achieve better distillation efficacy, we propose Contrastive Distillation on Intermediate Representations (CODIR), a principled knowledge distillation framework where the student is trained to distill knowledge through intermediate layers of the teacher via a contrastive objective.By learning to distinguish positive sample from a large set of negative samples, CoDIR facilitates the student's exploitation of rich information in teacher's hidden layers.CoDIR can be readily applied to compress large-scale language models in both pretraining and finetuning stages, and achieves superb performance on the GLUE benchmark, outperforming state-of-the-art compression methods. 1 Zhe Gan, Yuwei Fang, Yu Cheng 0001, Shuohang Wang, Jingjing Liu 0001 |
EMNLP (1) | 5 |
| 2020 | Cross-Thought for Sentence Encoder Pre-trainingabstractIn this paper, we propose Cross-Thought, a novel approach to pre-training sequence encoder, which is instrumental in building reusable sequence embeddings for large-scale NLP tasks such as question answering.Instead of using the original signals of full sentences, we train a Transformer-based sequence encoder over a large set of short sequences, which allows the model to automatically select the most useful information for predicting masked words.Experiments on question answering and textual entailment tasks demonstrate that our pre-trained encoder can outperform state-of-the-art encoders trained with continuous sentence signals as well as traditional masked language modeling baselines.Our proposed approach also achieves new state of the art on HotpotQA (full-wiki setting) by improving intermediate information retrieval performance.1 Shuohang Wang, Yuwei Fang, Zhe Gan, Yu Cheng 0001, Jingjing Liu 0001, Jing Jiang 0001 |
EMNLP (1) | 1 |
| 2020 | T3: Tree-Autoencoder Constrained Adversarial Text Generation for Targeted AttackabstractAdversarial attacks against natural language processing systems, which perform seemingly innocuous modifications to inputs, can induce arbitrary mistakes to the target models.Though raised great concerns, such adversarial attacks can be leveraged to estimate the robustness of NLP models.Compared with the adversarial example generation in continuous data domain (e.g., image), generating adversarial text that preserves the original meaning is challenging since the text space is discrete and non-differentiable.To handle these challenges, we propose a target-controllable adversarial attack framework T3, which is applicable to a range of NLP tasks.In particular, we propose a tree-based autoencoder to embed the discrete text data into a continuous representation space, upon which we optimize the adversarial perturbation.A novel tree-based decoder is then applied to regularize the syntactic correctness of the generated text and manipulate it on either sentence (T3(SENT)) or word (T3(WORD)) level.We consider two most representative NLP tasks: sentiment analysis and question answering (QA).Extensive experimental results and human studies show that T3 generated adversarial texts can successfully manipulate the NLP models to output the targeted incorrect answer without misleading the human.Moreover, we show that the generated adversarial texts have high transferability which enables the black-box attacks in practice.Our work sheds light on an effective and general way to examine the robustness of NLP models.Our code is publicly available at Boxin Wang, Hengzhi Pei, Boyuan Pan, Qian Chen 0003, Shuohang Wang, Bo Li 0026 |
EMNLP (1) | 5 |
| 2019 | Simple and Effective Curriculum Pointer-Generator Networks for Reading Comprehension over Long NarrativesabstractYi Tay, Shuohang Wang, Anh Tuan Luu, Jie Fu, Minh C. Phan, Xingdi Yuan, Jinfeng Rao, Siu Cheung Hui, Aston Zhang. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Yi Tay, Shuohang Wang, Anh Tuan Luu, Jie Fu 0001, Minh C. Phan, Xingdi Yuan, Jinfeng Rao, Siu Cheung Hui, Aston Zhang |
ACL (1) | 2 |
| 2019 | Lightweight and Efficient Neural Natural Language Processing with Quaternion NetworksabstractMany state-of-the-art neural models for NLP are heavily parameterized and thus memory inefficient.This paper proposes a series of lightweight and memory efficient neural architectures for a potpourri of natural language processing (NLP) tasks.To this end, our models exploit computation using Quaternion algebra and hypercomplex spaces, enabling not only expressive inter-component interactions but also significantly (75%) reduced parameter size due to lesser degrees of freedom in the Hamilton product.We propose Quaternion variants of models, giving rise to new architectures such as the Quaternion attention Model and Quaternion Transformer.Extensive experiments on a battery of NLP tasks demonstrates the utility of proposed Quaternion-inspired models, enabling up to 75% reduction in parameter size without significant loss in performance. Yi Tay, Aston Zhang, Anh Tuan Luu, Jinfeng Rao, Shuai Zhang 0007, Shuohang Wang, Jie Fu 0001, Siu Cheung Hui |
ACL (1) | 6 |
| 2019 | Multi-hop Knowledge Base Question Answering with an Iterative Sequence Matching ModelabstractKnowledge Base Question Answering (KBQA) has attracted much attention and recently there has been more interest in multi-hop KBQA. In this paper, we propose a novel iterative sequence matching model to address several limitations of previous methods for multi-hop KBQA. Our method iteratively grows the candidate relation paths that may lead to answer entities. The method prunes away less relevant branches and incrementally assigns matching scores to the paths. Empirical results demonstrate that our method can significantly outperform existing methods on three different benchmark datasets. Yunshi Lan, Shuohang Wang, Jing Jiang 0001 |
ICDM | 2 |
| 2019 | Knowledge Base Question Answering with Topic UnitsabstractKnowledge base question answering (KBQA) is an important task in natural language processing. Existing methods for KBQA usually start with entity linking, which considers mostly named entities found in a question as the starting points in the KB to search for answers to the question. However, relying only on entity linking to look for answer candidates may not be sufficient. In this paper, we propose to perform topic unit linking where topic units cover a wider range of units of a KB. We use a generation-and-scoring approach to gradually refine the set of topic units. Furthermore, we use reinforcement learning to jointly learn the parameters for topic unit linking and answer candidate ranking in an end-to-end manner. Experiments on three commonly used benchmark datasets show that our method consistently works well and outperforms the previous state of the art on two datasets. Yunshi Lan, Shuohang Wang, Jing Jiang 0001 |
IJCAI | 2 |
| 2019 | Compositional De-Attention NetworksabstractAttentional models are distinctly characterized by their ability to learn relative importance, i.e., assigning a different weight to input values. This paper proposes a new quasi-attention that is compositional in nature, i.e., learning whether to \textit{add}, \textit{subtract} or \textit{nullify} a certain vector when learning representations. This is strongly contrasted with vanilla attention, which simply re-weights input tokens. Our proposed \textit{Compositional De-Attention} (CoDA) is fundamentally built upon the intuition of both similarity and dissimilarity (negative affinity) when computing affinity scores, benefiting from a greater extent of expressiveness. We evaluate CoDA on six NLP tasks, i.e. open domain question answering, retrieval/ranking, natural language inference, machine translation, sentiment analysis and text2code generation. We obtain promising experimental results, achieving state-of-the-art performance on several tasks/datasets. Yi Tay, Anh Tuan Luu, Aston Zhang, Shuohang Wang, Siu Cheung Hui |
NeurIPS | 4 |
| 2019 | Knowledge Base Question Answering With a Matching-Aggregation Model and Question-Specific Contextual RelationsabstractMaking use of knowledge bases to answer questions (KBQA) is a key direction in question answering systems. Researchers have developed a diverse range of methods to address this problem, but there are still some limitations with the existing methods. Specifically, the existing neural network-based methods for KBQA have not taken advantage of the recent “matching-aggregation” framework for the sequence matching, and when representing a candidate answer entity, they may not choose the most useful context of the candidate for matching. In this paper, we explore the use of a “matching-aggregation” framework to match candidate answers with questions. We further make use of question-specific contextual relations to enhance the representations of candidate answer entities. Our complete method is able to achieve state-of-the-art performance on two benchmark datasets: WebQuestions and SimpleQuestions. Yunshi Lan, Shuohang Wang, Jing Jiang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | R3: Reinforced Ranker-Reader for Open-Domain Question AnsweringabstractIn recent years researchers have achieved considerable success applying neural network methods to question answering (QA). These approaches have achieved state of the art results in simplified closed-domain settings such as the SQuAD (Rajpurkar et al. 2016) dataset, which provides a pre-selected passage, from which the answer to a given question may be extracted. More recently, researchers have begun to tackle open-domain QA, in which the model is given a question and access to a large corpus (e.g., wikipedia) instead of a pre-selected passage (Chen et al. 2017a). This setting is more complex as it requires large-scale search for relevant passages by an information retrieval component, combined with a reading comprehension model that “reads” the passages to generate an answer to the question. Performance in this setting lags well behind closed-domain performance. In this paper, we present a novel open-domain QA system called Reinforced Ranker-Reader (R3), based on two algorithmic innovations. First, we propose a new pipeline for open-domain QA with a Ranker component, which learns to rank retrieved passages in terms of likelihood of extracting the ground-truth answer to a given question. Second, we propose a novel method that jointly trains the Ranker along with an answer-extraction Reader model, based on reinforcement learning. We report extensive experimental results showing that our method significantly improves on the state of the art for multiple open-domain QA datasets. Shuohang Wang, Mo Yu, Tim Klinger, Wei Zhang 0057, Shiyu Chang, Gerald Tesauro, Bowen Zhou 0002, Jing Jiang 0001 |
AAAI | 1 |
| 2018 | Evidence Aggregation for Answer Re-Ranking in Open-Domain Question Answering
Shuohang Wang, Mo Yu, Jing Jiang 0001, Wei Zhang 0057, Shiyu Chang, Tim Klinger, Gerald Tesauro, Murray Campbell |
ICLR (Poster) | 1 |
| 2017 | A Compare-Aggregate Model for Matching Text Sequences
Shuohang Wang, Jing Jiang 0001 |
ICLR (Poster) | 1 |
| 2017 | Machine Comprehension Using Match-LSTM and Answer Pointer
Shuohang Wang, Jing Jiang 0001 |
ICLR (Poster) | 1 |
| 2016 | Learning Natural Language Inference with LSTMabstractNatural language inference (NLI) is a fundamentally important task in natural language processing that has many applications.The recently released Stanford Natural Language Inference (SNLI) corpus has made it possible to develop and evaluate learning-centered methods such as deep neural networks for natural language inference (NLI).In this paper, we propose a special long short-term memory (LSTM) architecture for NLI.Our model builds on top of a recently proposed neural attention model for NLI but is based on a significantly different idea.Instead of deriving sentence embeddings for the premise and the hypothesis to be used for classification, our solution uses a match-LSTM to perform wordby-word matching of the hypothesis with the premise.This LSTM is able to place more emphasis on important word-level matching results.In particular, we observe that this LSTM remembers important mismatches that are critical for predicting the contradiction or the neutral relationship label.On the SNLI corpus, our model achieves an accuracy of 86.1%, outperforming the state of the art. Shuohang Wang, Jing Jiang 0001 |
HLT-NAACL | 1 |