Wenhan Xiong

dblp:203/8542 · DBLP profile ↗
← Back
26ranked-venue papers
10as first author
15since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 10 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Law of the Weakest Link: Cross Capabilities of Large Language Models
abstract
The development and evaluation of Large Language Models (LLMs) have largely focused on individual capabilities. However, this overlooks the intersection of multiple abilities across different types of expertise that are often required for real-world tasks, which we term **cross capabilities**. To systematically explore this concept, we first define seven core individual capabilities and then pair them to form seven common cross capabilities, each supported by a manually constructed taxonomy. Building on these definitions, we introduce *CrossEval*, a benchmark comprising 1,400 human-annotated prompts, with 100 prompts for each individual and cross capability. To ensure reliable evaluation, we involve expert annotators to assess 4,200 model responses, gathering 8,400 human ratings with detailed explanations to serve as reference examples. Our findings reveal that current LLMs consistently exhibit the ``Law of the Weakest Link,'' where cross-capability performance is significantly constrained by the weakest component. Across 58 cross-capability scores from 17 models, 38 scores are lower than all individual capabilities, while 20 fall between strong and weak, but closer to the weaker ability. These results highlight LLMs' underperformance in cross-capability tasks, emphasizing the need to identify and improve their weakest capabilities as a key research priority. The code, benchmarks, and evaluations are available on our [project website](https://www.llm-cross-capabilities.org).
Ming Zhong 0005, Aston Zhang, Wenhan Xiong, Chenguang Zhu 0001, Zhengxing Chen, Chloe Bi, Mike Lewis, Sravya Popuri, Sharan Narang, Melanie Kambadur, Dhruv Mahajan 0001, Sergey Edunov, Jiawei Han 0001, Laurens van der Maaten
ICLR5
2025 Scene-LLM: Extending Language Model for 3D Visual Reasoning
Xilun Chen 0002, Yixin Nie, Wenhan Xiong
WACV5
2024 Prompting Large Language Models with Speech Recognition Abilities
abstract
Large language models (LLMs) have proven themselves highly flexible, able to solve a wide range of generative tasks, such as abstractive summarization and open-ended question answering. In this paper we extend the capabilities of LLM by directly attaching a small audio encoder allowing it to perform speech recognition. By directly prepending a sequence of audio embeddings to the text token embeddings, the LLM can be converted to an automatic speech recognition (ASR) system, and be used in the exact same manner as its textual counterpart. Experiments on Multilingual LibriSpeech (MLS) show that incorporating a conformer encoder into the open sourced LLaMA-7B allows it to outperform monolingual baselines by 18% relatively in WER and perform multilingual speech recognition, despite LLaMA being trained overwhelmingly on English text. Furthermore, we perform ablation studies to investigate whether the LLM can be completely frozen during training to maintain its original capabilities, scaling up the audio encoder, and increasing the audio encoder striding to generate fewer embeddings. The results from these studies show that multilingual ASR is possible even when the LLM is frozen, or when strides of almost 1 second are used in the audio encoder opening up the possibility for LLMs to operate on long-form audio.
Yassir Fathullah, Chunyang Wu, Egor Lakomkin, Junteng Jia, Yuan Shangguan, Ke Li 0023, Jinxi Guo, Wenhan Xiong, Jay Mahadeokar, Ozlem Kalinli, Christian Fügen, Mike Seltzer
ICASSP8
2024 LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models
abstract
Chi Han, Qifan Wang, Hao Peng, Wenhan Xiong, Yu Chen, Heng Ji, Sinong Wang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Chi Han, Qifan Wang 0001, Hao Peng 0009, Wenhan Xiong, Yu Chen 0022, Heng Ji 0001, Sinong Wang
NAACL-HLT4
2024 Effective Long-Context Scaling of Foundation Models
abstract
Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, Hao Ma. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Wenhan Xiong, Igor Molybog, Prajjwal Bhargava, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, Hao Ma 0001
NAACL-HLT1
2024 FLAME : Factuality-Aware Alignment for Large Language Models
abstract
Alignment is a procedure to fine-tune pre-trained large language models (LLMs) to follow natural language instructions and serve as helpful AI assistants. We have observed, however, that the conventional alignment process fails to enhance the factual accuracy of LLMs, and often leads to the generation of more false facts (i.e., *hallucination*). In this paper, we study how to make the LLM alignment process more factual, by first identifying factors that lead to hallucination in both alignment steps: supervised fine-tuning (SFT) and reinforcement learning (RL). In particular, we find that training the LLM on new or unfamiliar knowledge can encourage hallucination. This makes SFT less factual as it trains on human-labeled data that may be novel to the LLM. Furthermore, reward functions used in standard RL often inadequately capture factuality and favor longer and more detailed responses, which inadvertently promote hallucination. Based on these observations, we propose *FactuaLity-aware AlignMEnt*, comprised of *factuality-aware SFT* and *factuality-aware RL* through direct preference optimization. Experiments show that our proposed *FLAME* guides LLMs to output more factual responses while maintaining their instruction-following capability.
Sheng-Chieh Lin, Luyu Gao, Barlas Oguz, Wenhan Xiong, Jimmy Lin, Scott Yih, Xilun Chen 0002
NeurIPS4
2024 Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
abstract
The quadratic complexity and weak length extrapolation of Transformers limits their ability to scale to long sequences, and while sub-quadratic solutions like linear attention and state space models exist, they empirically underperform Transformers in pretraining efficiency and downstream task accuracy. We introduce MEGALODON, an neural architecture for efficient sequence modeling with unlimited context length. MEGALODON inherits the architecture of MEGA (exponential moving average with gated attention), and further introduces multiple technical components to improve its capability and stability, including complex exponential moving average (CEMA), timestep normalization layer, normalized attention mechanism and pre-norm with two-hop residual configuration. In a controlled head-to-head comparison with LLAMA2, MEGALODON achieves better efficiency than Transformer in the scale of 7 billion parameters and 2 trillion training tokens. MEGALODON reaches a training loss of 1.70, landing mid-way between LLAMA2-7B (1.75) and LLAMA2-13B (1.67). This result is robust throughout a wide range of benchmarks, where MEGALODON consistently outperforms Transformers across different tasks, domains, and modalities.
Xuezhe Ma, Wenhan Xiong, Beidi Chen, Lili Yu, Hao Zhang 0025, Jonathan May, Luke Zettlemoyer, Omer Levy, Chunting Zhou
NeurIPS3
2023 Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language Compositionality
abstract
Contrastively trained vision-language models have achieved remarkable progress in vision and language representation learning.However, recent research has highlighted severe limitations of these models in their ability to perform compositional reasoning over objects, attributes, and relations.Scene graphs have emerged as an effective way to understand images compositionally.These are graphstructured semantic representations of images that contain objects, their attributes, and relations with other objects in a scene.In this work, we consider the scene graph parsed from text as a proxy for the image scene graph and propose a graph decomposition and augmentation framework along with a coarse-to-fine contrastive learning objective between images and text that aligns sentences of various complexities to the same image.We also introduce novel negative mining techniques in the scene graph space for improving attribute binding and relation understanding.Through extensive experiments, we demonstrate the effectiveness of our approach that significantly improves attribute binding, relation understanding, systematic generalization, and productivity on multiple recently proposed benchmarks (For example, improvements up to 18% for systematic generalization, 16.5% for relation understanding over a strong baseline), while achieving similar or better performance than CLIP on various general multimodal tasks.
Harman Singh, Pengchuan Zhang, Qifan Wang 0001, Wenhan Xiong, Jingfei Du, Yu Chen 0022
EMNLP5
2023 Multi-Head State Space Model for Speech Recognition
Yassir Fathullah, Chunyang Wu, Yuan Shangguan, Junteng Jia, Wenhan Xiong, Jay Mahadeokar, Chunxi Liu, Yangyang Shi, Ozlem Kalinli, Mike Seltzer, Mark J. F. Gales
INTERSPEECH5
2022 SCROLLS: Standardized CompaRison Over Long Language Sequences
abstract
Uri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, Omer Levy. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Uri Shaham 0002, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta 0001, Wenhan Xiong, Mor Geva, Jonathan Berant, Omer Levy
EMNLP8
2022 Boosted Dense Retriever
abstract
Patrick Lewis, Barlas Oguz, Wenhan Xiong, Fabio Petroni, Scott Yih, Sebastian Riedel. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Patrick S. H. Lewis, Barlas Oguz, Wenhan Xiong, Fabio Petroni, Scott Yih, Sebastian Riedel 0001
NAACL-HLT3
2022 Simple Local Attentions Remain Competitive for Long-Context Tasks
abstract
Wenhan Xiong, Barlas Oguz, Anchit Gupta, Xilun Chen, Diana Liskovich, Omer Levy, Scott Yih, Yashar Mehdad. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Wenhan Xiong, Barlas Oguz, Anchit Gupta, Xilun Chen 0002, Diana Liskovich, Omer Levy, Scott Yih, Yashar Mehdad
NAACL-HLT1
2021 Progressively Pretrained Dense Corpus Index for Open-Domain Question Answering
abstract
Commonly used information retrieval methods such as TF-IDF in open-domain question answering (QA) systems are insufficient to capture deep semantic matching that goes beyond lexical overlaps.Some recent studies consider the retrieval process as maximum inner product search (MIPS) using dense question and paragraph representations, achieving promising results on several informationseeking QA datasets.However, the pretraining of the dense vector representations is highly resource-demanding, e.g., requires a very large batch size and lots of training steps.In this work, we propose a sample-efficient method to pretrain the paragraph encoder.First, instead of using heuristically created pseudo questionparagraph pairs for pretraining, we use an existing pretrained sequence-to-sequence model to build a strong question generator that creates high-quality pretraining data.Second, we propose a simple progressive pretraining algorithm to ensure the existence of effective negative samples in each batch.Across three opendomain QA datasets, our method consistently outperforms a strong dense retrieval baseline that uses 6 times more computation for training.On two of the datasets, our method achieves more than 4-point absolute improvement in terms of answer exact match.
Wenhan Xiong, Hong Wang 0023, William Yang Wang
EACL1
2021 Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval
Wenhan Xiong, Xiang Li 0069, Srinivasan Iyer 0001, Jingfei Du, Patrick S. H. Lewis, William Yang Wang, Yashar Mehdad, Scott Yih, Sebastian Riedel 0001, Douwe Kiela, Barlas Oguz
ICLR1
2021 Unsupervised Multi-hop Question Answering by Question Generation
abstract
Liangming Pan, Wenhu Chen, Wenhan Xiong, Min-Yen Kan, William Yang Wang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Liangming Pan, Wenhu Chen, Wenhan Xiong, Min-Yen Kan, William Yang Wang
NAACL-HLT3
2020 Pretrained Encyclopedia: Weakly Supervised Knowledge-Pretrained Language Model
Wenhan Xiong, Jingfei Du, William Yang Wang, Veselin Stoyanov
ICLR1
2020 SafeRoute: Learning to Navigate Streets Safely in an Urban Environment
abstract
Recent studies show that 85% of women have changed their traveled routes to avoid harassment and assault. Despite this, current mapping tools do not empower users with information to take charge of their personal safety. We propose SafeRoute, a novel solution to the problem of navigating cities and avoiding street harassment and crime. Unlike other street navigation applications, SafeRoute introduces a new type of path generation via deep reinforcement learning. This enables us to successfully optimize for multi-criteria path-finding and incorporate representation learning within our framework. Our agent learns to pick favorable streets to create a safe and short path with a reward function that incorporates safety and efficiency. Given access to recent crime reports in many urban cities, we train our model for experiments in Boston, New York, and San Francisco. We test our model on areas of these cities, specifically the populated downtown regions with high foot traffic. We evaluate SafeRoute and successfully improve over state-of-the-art methods by up to 17% in local average distance from crimes while decreasing path length by up to 7%.
Sharon Levy, Wenhan Xiong, Elizabeth M. Belding, William Yang Wang
ACM Trans. Intell. Syst. Technol.2
2019 Self-Supervised Learning for Contextualized Extractive Summarization
abstract
Existing models for extractive summarization are usually trained from scratch with a crossentropy loss, which does not explicitly capture the global context at the document level.In this paper, we aim to improve this task by introducing three auxiliary pre-training tasks that learn to capture the document-level context in a self-supervised fashion.Experiments on the widely-used CNN/DM dataset validate the effectiveness of the proposed auxiliary tasks.Furthermore, we show that after pretraining, a clean model with simple building blocks is able to outperform previous state-ofthe-art that are carefully designed.1
Hong Wang 0023, Xin Wang 0061, Wenhan Xiong, Mo Yu, Shiyu Chang, William Yang Wang
ACL (1)3
2019 TWEETQA: A Social Media Focused Question Answering Dataset
abstract
With social media becoming increasingly popular on which lots of news and real-time events are reported, developing automated question answering systems is critical to the effective-ness of many applications that rely on real-time knowledge. While previous datasets have concentrated on question answering (QA) for formal text like news and Wikipedia, we present the first large-scale dataset for QA over social media data. To ensure that the tweets we collected are useful, we only gather tweets used by journalists to write news articles. We then ask human annotators to write questions and answers upon these tweets. Unlike otherQA datasets like SQuAD in which the answers are extractive, we allow the answers to be abstractive. We show that two recently proposed neural models that perform well on formal texts are limited in their performance when applied to our dataset. In addition, even the fine-tuned BERT model is still lagging behind human performance with a large margin. Our results thus point to the need of improved QA systems targeting social media text.
Wenhan Xiong, Jiawei Wu 0003, Hong Wang 0023, Vivek Kulkarni, Mo Yu, Shiyu Chang, William Yang Wang
ACL (1)1
2019 Improving Question Answering over Incomplete KBs with Knowledge-Aware Reader
abstract
We propose a new end-to-end question answering model, which learns to aggregate answer evidence from an incomplete knowledge base (KB) and a set of retrieved text snippets.Under the assumptions that the structured KB is easier to query and the acquired knowledge can help the understanding of unstructured text, our model first accumulates knowledge of entities from a question-related KB subgraph; then reformulates the question in the latent space and reads the texts with the accumulated entity knowledge at hand.The evidence from KB and texts are finally aggregated to predict answers.On the widely-used KBQA benchmark WebQSP, our model achieves consistent improvements across settings with different extents of KB incompleteness. 1
Wenhan Xiong, Mo Yu, Shiyu Chang, William Yang Wang
ACL (1)1
2019 Learning to Learn and Predict: A Meta-Learning Approach for Multi-Label Classification
abstract
Jiawei Wu, Wenhan Xiong, William Yang Wang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jiawei Wu 0003, Wenhan Xiong, William Yang Wang
EMNLP/IJCNLP (1)2
2018 Look Before You Leap: Bridging Model-Free and Model-Based Reinforcement Learning for Planned-Ahead Vision-and-Language Navigation
Xin Wang 0061, Wenhan Xiong, Hongmin Wang, William Yang Wang
ECCV (16)2
2018 One-Shot Relational Learning for Knowledge Graphs
abstract
Knowledge graphs (KGs) are the key components of various natural language processing applications.To further expand KGs' coverage, previous studies on knowledge graph completion usually require a large number of training instances for each relation.However, we observe that long-tail relations are actually more common in KGs and those newly added relations often do not have many known triples for training.In this work, we aim at predicting new facts under a challenging setting where only one training instance is available.We propose a one-shot relational learning framework, which utilizes the knowledge extracted by embedding models and learns a matching metric by considering both the learned embeddings and one-hop graph structures.Empirically, our model yields considerable performance improvements over existing embedding models, and also eliminates the need of retraining the embedding models when dealing with newly added relations. 1
Wenhan Xiong, Mo Yu, Shiyu Chang, William Yang Wang
EMNLP1
2018 Scheduled Policy Optimization for Natural Language Communication with Intelligent Agents
abstract
We investigate the task of learning to interpret natural language instructions by jointly reasoning with visual observations and language inputs. Unlike current methods which start with learning from demonstrations (LfD) and then use reinforcement learning (RL) to fine-tune the model parameters, we propose a novel policy optimization algorithm which can dynamically schedule demonstration learning and RL. The proposed training paradigm provides efficient exploration and generalization beyond existing methods. Comparing to existing ensemble models, the best single model based on our proposed method tremendously decreases the execution error by 55% on a block-world environment. To further illustrate the exploration strategy of our RL algorithm, our paper includes systematic studies on the evolution of policy entropy during training.
Wenhan Xiong, Mo Yu, Shiyu Chang, William Yang Wang
IJCAI1
2018 Variational Knowledge Graph Reasoning
abstract
Wenhu Chen, Wenhan Xiong, Xifeng Yan, William Yang Wang. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Wenhu Chen, Wenhan Xiong, Xifeng Yan, William Yang Wang
NAACL-HLT2
2017 DeepPath: A Reinforcement Learning Method for Knowledge Graph Reasoning
abstract
We study the problem of learning to reason in large scale knowledge graphs (KGs).More specifically, we describe a novel reinforcement learning framework for learning multi-hop relational paths: we use a policy-based agent with continuous states based on knowledge graph embeddings, which reasons in a KG vector space by sampling the most promising relation to extend its path.In contrast to prior work, our approach includes a reward function that takes the accuracy, diversity, and efficiency into consideration.Experimentally, we show that our proposed method outperforms a path-ranking based algorithm and knowledge graph embedding methods on Freebase and Never-Ending Language Learning datasets.1
Wenhan Xiong, Thien Hoang, William Yang Wang
EMNLP1