VLDB 2026 Research / reviewers in the wild / expert
Xilun Chen 0002
dblp:96/10207-2
· DBLP profile ↗
18ranked-venue papers
6as first author
11since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense RetrieversabstractLarge language models (LLMs) have demonstrated strong effectiveness and robustness when fine-tuned as dense retrievers.However, their large parameter size presents significant computational challenges at inference time.While smaller retrievers offer better efficiency, they often fail to generalize effectively with limited supervised fine-tuning data.In this work, we introduce DRAMA, a training framework that leverages LLMs to train smaller generalizable dense retrievers.In particular, we adopt pruned LLMs as the backbone and train on diverse LLM-augmented data in a single-stage contrastive learning setup.Experiments show that DRAMA offers better multilingual and long-context capabilities than traditional encoder-based retrievers, and achieves strong performance across multiple tasks and languages.1 * Equal contribution.† Work done while at Meta. 1 Code and checkpoints will be available at https://github. com/facebookresearch/dpr-scale/tree/main/drama. Xueguang Ma, Xi Victoria Lin, Barlas Oguz, Jimmy Lin, Scott Yih, Xilun Chen 0002 |
ACL (1) | 6 |
| 2025 | Scene-LLM: Extending Language Model for 3D Visual Reasoning
Xilun Chen 0002, Yixin Nie, Wenhan Xiong |
WACV | 3 |
| 2024 | Few-Shot Data Synthesis for Open Domain Multi-Hop Question AnsweringabstractFew-shot learning for open domain multi-hop question answering typically relies on the incontext learning capability of large language models (LLMs).While powerful, these LLMs usually contain tens or hundreds of billions of parameters, making them rather inefficient at inference time.To improve performance of smaller language models, we propose a data synthesis framework for multi-hop question answering that requires less than 10 humanannotated question answer pairs.Our framework depends only on rich, naturally-occurring relationships among documents and is built upon the data generation functions parameterized by LLMs and prompts.We synthesize millions of multi-hop questions and claims to finetune language models, evaluated on popular benchmarks for multi-hop question answering and fact verification.Empirically, our approach improves model performance significantly, allowing the finetuned models to be competitive with GPT-3.5 based approaches while being almost one-third the size in parameter count.What is the elevation range for the area that the eastern sector of the Colorado orogeny extends into?Query1: the eastern section of the Colorado orogeny Query2: the elevation range for the High Plains Unanswerable Question generation Query generation Query verification Question: What is the elevation range for … Query: the eastern section of the Colorado … Observation: 1: The Colorado orogeny, or Colorado orogen, was an orogeny … … Answer: 1,800 to 7,000 ft Question answering 1,800 to 7,000 ft Randomly sample document pairsThe causes of World War II are debated, but contributing factors included the Second Italo-Ethiopian War, Spanish Civil War, Second Sino-Japanese War, Soviet-Japanese border conflicts, the rise of fascism in Europe, and European tensions in the aftermath of World War I.The Second Italo-Ethiopian War, also referred to as the Second Italo-Abyssinian War, was a war of aggression which was fought between Italy and Ethiopia from October 1935 to February 1937. Events occurred in sequenceThe Colorado orogeny, or Colorado orogen, was an orogeny in Colorado and surrounding areas which was a part of the development of the ancestral Rockies.The eastern sector extends into the High Plains and is called the Central Plains orogeny.The High Plains are a subregion of the Great Plains.From east to west, the High Plains rise in elevation from around 1,800 to 7,000 ft (550 to 2,130 m). Extra geographical informationNew York, often called New York City or NYC, is the most populous city in the United States.With a 2020 population of 8,804,190 distributed over 300.46 square miles (778.2 km 2 ), the city is the most densely populated major city in the United States.NYC is more than twice as populous as Los Angeles, the nation's second-largest city.Los Angeles, often referred to by its initials L.A., officially the City of Los Angeles, is the most populous city in the U.S. state of California. Extra demographic information Mingda Chen, Xilun Chen 0002, Scott Yih |
EACL (1) | 2 |
| 2024 | RA-DIT: Retrieval-Augmented Dual Instruction TuningabstractRetrieval-augmented language models (RALMs) improve performance by accessing long-tail and up-to-date knowledge from external data stores, but are challenging to build. Existing approaches require either expensive retrieval-specific modifications to LM pre-training or use post-hoc integration of the data store that leads to suboptimal performance. We introduce Retrieval-Augmented Dual Instruction Tuning (RA-DIT), a lightweight fine-tuning methodology that provides a third option by retrofitting any LLM with retrieval capabilities. Our approach operates in two distinct fine-tuning steps: (1) one updates a pre-trained LM to better use retrieved information, while (2) the other updates the retriever to return more relevant results, as preferred by the LM. By fine-tuning over tasks that require both knowledge utilization and contextual awareness, we demonstrate that each stage yields significant performance improvements, and using both leads to additional gains. Our best model, RA-DIT 65B, achieves state-of-the-art performance across a range of knowledge-intensive zero- and few-shot learning benchmarks, significantly outperforming existing in-context RALM approaches by up to +8.9% in 0-shot setting and +1.4% in 5-shot setting on average. Xi Victoria Lin, Xilun Chen 0002, Mingda Chen, Maria Lomeli, Richard James 0001, Pedro Rodríguez 0001, Jacob Kahn, Gergely Szilvasy, Mike Lewis, Luke Zettlemoyer, Scott Yih |
ICLR | 2 |
| 2024 | Nearest Neighbor Speculative Decoding for LLM Generation and AttributionabstractLarge language models (LLMs) often hallucinate and lack the ability to provide attribution for their generations. Semi-parametric LMs, such as kNN-LM, approach these limitations by refining the output of an LM for a given prompt using its nearest neighbor matches in a non-parametric data store. However, these models often exhibit slow inference speeds and produce non-fluent texts. In this paper, we introduce Nearest Neighbor Speculative Decoding (NEST), a novel semi-parametric language modeling approach that is capable of incorporating real-world text spans of arbitrary length into the LM generations and providing attribution to their sources. NEST performs token-level retrieval at each inference step to compute a semi-parametric mixture distribution and identify promising span continuations in a corpus. It then uses an approximate speculative decoding procedure that accepts a prefix of the retrieved span or generates a new token. NEST significantly enhances the generation quality and attribution rate of the base LM across a variety of knowledge-intensive tasks, surpassing the conventional kNN-LM method and performing competitively with in-context retrieval augmentation. In addition, NEST substantially improves the generation speed, achieving a 1.8x speedup in inference time when applied to Llama-2-Chat 70B. Code will be released at https://github.com/facebookresearch/NEST/tree/main. Minghan Li 0002, Xilun Chen 0002, Ari Holtzman, Beidi Chen, Jimmy Lin, Scott Yih, Xi Victoria Lin |
NeurIPS | 2 |
| 2024 | FLAME : Factuality-Aware Alignment for Large Language ModelsabstractAlignment is a procedure to fine-tune pre-trained large language models (LLMs) to follow natural language instructions and serve as helpful AI assistants.
We have observed, however, that the conventional alignment process fails to enhance the factual accuracy of LLMs, and often leads to the generation of more false facts (i.e., *hallucination*).
In this paper, we study how to make the LLM alignment process more factual, by first identifying factors that lead to hallucination in both alignment steps: supervised fine-tuning (SFT) and reinforcement learning (RL).
In particular, we find that training the LLM on new or unfamiliar knowledge can encourage hallucination.
This makes SFT less factual as it trains on human-labeled data that may be novel to the LLM.
Furthermore, reward functions used in standard RL often inadequately capture factuality and favor longer and more detailed responses, which inadvertently promote hallucination.
Based on these observations, we propose *FactuaLity-aware AlignMEnt*, comprised of *factuality-aware SFT* and *factuality-aware RL* through direct preference optimization.
Experiments show that our proposed *FLAME* guides LLMs to output more factual responses while maintaining their instruction-following capability. Sheng-Chieh Lin, Luyu Gao, Barlas Oguz, Wenhan Xiong, Jimmy Lin, Scott Yih, Xilun Chen 0002 |
NeurIPS | 7 |
| 2023 | CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector RetrievalabstractMinghan Li, Sheng-Chieh Lin, Barlas Oguz, Asish Ghoshal, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, Xilun Chen. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Minghan Li 0002, Sheng-Chieh Lin, Barlas Oguz, Asish Ghoshal, Jimmy Lin, Yashar Mehdad, Scott Yih, Xilun Chen 0002 |
ACL (1) | 8 |
| 2023 | Hierarchical Video-Moment Retrieval and Step-CaptioningabstractThere is growing interest in searching for information from large video corpora. Prior works have studied relevant tasks, such as text-based video retrieval, moment retrieval, video summarization, and video captioning in isolation, without an end-to-end setup that can jointly search from video corpora and generate summaries. Such an end-to-end setup would allow for many interesting applications, e.g., a text-based search that finds a relevant video from a video corpus, extracts the most relevant moment from that video, and segments the moment into important steps with captions. To address this, we present the HIREST (HIerarchical REtrieval and STep-captioning) dataset and propose a new benchmark that covers hierarchical information retrieval and visual/textual stepwise summarization from an instructional video corpus. Hirest consists of 3.4K text-video pairs from an instructional video dataset, where 1.1 K videos have annotations of moment spans relevant to text query and breakdown of each moment into key instruction steps with caption and timestamps (totaling 8.6K step captions). Our hierarchical benchmark consists of video retrieval, moment retrieval, and two novel moment segmentation and step captioning tasks. In moment segmentation, models break down a video moment into instruction steps and identify start-end boundaries. In step captioning, models generate a textual summary for each step. We also present starting point task-specific and end-to-end joint baseline models for our new benchmark. While the baseline models show some promising results, there still exists large room for future improvement by the community.11code and data: https://github.com/j-min/HiREST Abhaysinh Zala, Jaemin Cho 0001, Satwik Kottur, Xilun Chen 0002, Barlas Oguz, Yashar Mehdad, Mohit Bansal |
CVPR | 4 |
| 2022 | Simple Local Attentions Remain Competitive for Long-Context TasksabstractWenhan Xiong, Barlas Oguz, Anchit Gupta, Xilun Chen, Diana Liskovich, Omer Levy, Scott Yih, Yashar Mehdad. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Wenhan Xiong, Barlas Oguz, Anchit Gupta, Xilun Chen 0002, Diana Liskovich, Omer Levy, Scott Yih, Yashar Mehdad |
NAACL-HLT | 4 |
| 2021 | Muppet: Massive Multi-task Representations with Pre-FinetuningabstractWe propose pre-finetuning, an additional largescale learning stage between language model pre-training and fine-tuning.Pre-finetuning is massively multi-task learning (around 50 datasets, over 4.8 million total labeled examples), and is designed to encourage learning of representations that generalize better to many different tasks.We show that prefinetuning consistently improves performance for pretrained discriminators (e.g.RoBERTa) and generation models (e.g.BART) on a wide range of tasks (sentence prediction, commonsense reasoning, MRC, etc.), while also significantly improving sample efficiency during fine-tuning.We also show that large-scale multi-tasking is crucial; pre-finetuning can hurt performance when few tasks are used up until a critical point (usually above 15) after which performance improves linearly in the number of tasks. Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen 0002, Luke Zettlemoyer, Sonal Gupta |
EMNLP (1) | 4 |
| 2021 | Learning Better Structured Representations Using Low-rank Adaptive Label Smoothing
Asish Ghoshal, Xilun Chen 0002, Sonal Gupta, Luke Zettlemoyer, Yashar Mehdad |
ICLR | 2 |
| 2020 | Low-Resource Domain Adaptation for Compositional Task-Oriented Semantic ParsingabstractTask-oriented semantic parsing is a critical component of virtual assistants, which is responsible for understanding the user's intents (set reminder, play music, etc.).Recent advances in deep learning have enabled several approaches to successfully parse more complex queries (Gupta et al., 2018;Rongali et al., 2020), but these models require a large amount of annotated training data to parse queries on new domains (e.g.reminder, music).In this paper, we focus on adapting taskoriented semantic parsers to low-resource domains, and propose a novel method that outperforms a supervised neural model at a 10-fold data reduction.In particular, we identify two fundamental factors for low-resource domain adaptation: better representation learning and better training techniques.Our representation learning uses BART (Lewis et al., 2020) to initialize our model which outperforms encoder-only pre-trained representations used in previous work.Furthermore, we train with optimization-based meta-learning (Finn et al., 2017) to improve generalization to lowresource domains.This approach significantly outperforms all baseline methods in the experiments on a newly collected multi-domain taskoriented semantic parsing dataset (TOPv2 1 ). Xilun Chen 0002, Asish Ghoshal, Yashar Mehdad, Luke Zettlemoyer, Sonal Gupta |
EMNLP (1) | 1 |
| 2019 | Multi-Source Cross-Lingual Model Transfer: Learning What to ShareabstractModern NLP applications have enjoyed a great boost utilizing neural networks models.Such deep neural models, however, are not applicable to most human languages due to the lack of annotated training data for various NLP tasks.Cross-lingual transfer learning (CLTL) is a viable method for building NLP models for a low-resource target language by leveraging labeled data from other (source) languages.In this work, we focus on the multilingual transfer setting where training data in multiple source languages is leveraged to further boost target language performance.Unlike most existing methods that rely only on language-invariant features for CLTL, our approach coherently utilizes both languageinvariant and language-specific features at instance level.Our model leverages adversarial networks to learn language-invariant features, and mixture-of-experts models to dynamically exploit the similarity between the target language and each individual source language 1 .This enables our model to learn effectively what to share between various languages in the multilingual setup.Moreover, when coupled with unsupervised multilingual embeddings, our model can operate in a zero-resource setting where neither target language training data nor cross-lingual resources are available.Our model achieves significant performance gains over prior art, as shown in an extensive set of experiments over multiple text classification and sequence tagging tasks including a large-scale industry dataset. Xilun Chen 0002, Ahmed Awadallah 0001, Hany Hassan, Wei Wang 0238, Claire Cardie |
ACL (1) | 1 |
| 2018 | Unsupervised Multilingual Word EmbeddingsabstractMultilingual Word Embeddings (MWEs) represent words from multiple languages in a single distributional vector space.Unsupervised MWE (UMWE) methods acquire multilingual embeddings without cross-lingual supervision, which is a significant advantage over traditional supervised approaches and opens many new possibilities for low-resource languages.Prior art for learning UMWEs, however, merely relies on a number of independently trained Unsupervised Bilingual Word Embeddings (UBWEs) to obtain multilingual embeddings.These methods fail to leverage the interdependencies that exist among many languages.To address this shortcoming, we propose a fully unsupervised framework for learning MWEs 1 that directly exploits the relations between all language pairs.Our model substantially outperforms previous approaches in the experiments on multilingual word translation and cross-lingual word similarity.In addition, our model even beats supervised approaches trained with cross-lingual resources. Xilun Chen 0002, Claire Cardie |
EMNLP | 1 |
| 2018 | Multinomial Adversarial Networks for Multi-Domain Text ClassificationabstractMany text classification tasks are known to be highly domain-dependent.Unfortunately, the availability of training data can vary drastically across domains.Worse still, for some domains there may not be any annotated data at all.In this work, we propose a multinomial adversarial network 1 (MAN) to tackle this real-world problem of multi-domain text classification (MDTC) in which labeled data may exist for multiple domains, but in insufficient amounts to train effective classifiers for one or more of the domains.We provide theoretical justifications for the MAN framework, proving that different instances of MANs are essentially minimizers of various f-divergence metrics (Ali and Silvey, 1966) among multiple probability distributions.MANs are thus a theoretically sound generalization of traditional adversarial networks that discriminate over two distributions.More specifically, for the MDTC task, MAN learns features that are invariant across multiple domains by resorting to its ability to reduce the divergence among the feature distributions of each domain.We present experimental results showing that MANs significantly outperform the prior art on the MDTC task.We also show that MANs achieve state-of-the-art performance for domains with no labeled data. Xilun Chen 0002, Claire Cardie |
NAACL-HLT | 1 |
| 2018 | Adversarial Deep Averaging Networks for Cross-Lingual Sentiment ClassificationabstractIn recent years great success has been achieved in sentiment classification for English, thanks in part to the availability of copious annotated resources. Unfortunately, most languages do not enjoy such an abundance of labeled data. To tackle the sentiment classification problem in low-resource languages without adequate annotated data, we propose an Adversarial Deep Averaging Network (ADAN 1 ) to transfer the knowledge learned from labeled data on a resource-rich source language to low-resource languages where only unlabeled data exist. ADAN has two discriminative branches: a sentiment classifier and an adversarial language discriminator. Both branches take input from a shared feature extractor to learn hidden representations that are simultaneously indicative for the classification task and invariant across languages. Experiments on Chinese and Arabic sentiment classification demonstrate that ADAN significantly outperforms state-of-the-art systems. Xilun Chen 0002, Yu Sun 0020, Ben Athiwaratkun, Claire Cardie, Kilian Q. Weinberger |
Trans. Assoc. Comput. Linguistics | 1 |
| 2017 | A Rectangle Mining Method for Understanding the Semantics of Financial TablesabstractFinancial statements report crucial information in tables with complex semantic structure, which are desirable, yet challenging, to interpret automatically. For example, in such tables a row of data cells is often explained by the headers of other rows. In a departure from prior art, we propose a rectangle mining framework for understanding complex tables, which considers rectangular regions rather than individual cells or pairs of cells in a table. We instantiate this framework with ReMine, an algorithm for extracting row header semantics of table, and show that it significantly outperforms prior pair-wise classification approaches on two datasets: (i) a set of manually labeled financial tables from multiple companies, and (ii) the ICDAR 2013 Table Competition dataset. Xilun Chen 0002, Laura Chiticariu, Marina Danilevsky, Alexandre V. Evfimievski, Prithviraj Sen |
ICDAR | 1 |
| 2013 | Multi-Domain Adaptation for SMT Using Multi-Task LearningabstractDomain adaptation for SMT usually adapts models to an individual specific domain.However, it often lacks some correlation among different domains where common knowledge could be shared to improve the overall translation quality.In this paper, we propose a novel multi-domain adaptation approach for SMT using Multi-Task Learning (MTL), with in-domain models tailored for each specific domain and a general-domain model shared by different domains.The parameters of these models are tuned jointly via MTL so that they can learn general knowledge more accurately and exploit domain knowledge better.Our experiments on a largescale English-to-Chinese translation task validate that the MTL-based adaptation approach significantly and consistently improves the translation quality compared to a non-adapted baseline.Furthermore, it also outperforms the individual adaptation of each specific domain. Lei Cui 0001, Xilun Chen 0002, Dongdong Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001 |
EMNLP | 2 |