EDBT 2026 Demo / reviewers in the wild / expert
Liang Wang 0046
dblp:56/4499-46
· DBLP profile ↗
21ranked-venue papers
10as first author
17since 2021 · last 2026
0000-0003-4664-7136ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 10 first-author · 15 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal EmbeddingsabstractMultimodal embedding models, built upon causal Vision Language Models (VLMs), have shown promise in various tasks.However, current approaches face three limitations: causal attention in VLM backbones is suboptimal for embedding tasks; scalability issues due to reliance on high-quality labeled paired data for contrastive learning; and limited diversity in training objectives and data.To address these issues, we propose MoCa, a two-stage framework for transforming pre-trained VLMs into bidirectional multimodal embedding models.The first stage, Modality-aware Continual Pre-training, introduces a joint reconstruction objective that simultaneously denoises interleaved texts and images, enhancing bidirectional context-aware reasoning.The second stage, Heterogeneous Contrastive Fine-tuning, leverages diverse, semantically rich multimodal data beyond simple image-caption pairs to enhance generalization and alignment.Our method addresses the stated limitations by introducing bidirectional attention through continual pre-training, scaling effectively with massive unlabeled datasets via joint reconstruction objectives, and utilizing diverse multimodal data for enhanced representation robustness.Experiments demonstrate that MoCa consistently improves performance across MMEB and ViDoRe-v2 benchmarks, achieving new state-of-the-arts, and exhibits strong scalability with both model size and training data on MMEB.We have released the model weights and data on our project page https://haon-chen.github.io/MoCa/. Haonan Chen 0005, Yuping Luo, Liang Wang 0046, Nan Yang 0002, Furu Wei, Zhicheng Dou |
ACL (1) | 4 |
| 2026 | Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM HallucinationsabstractWen Luo, Guangyue Peng, Wei Li, Shaohang Wei, Feifan Song, Liang Wang, Nan Yang, Xingxing Zhang, Jing Jin, Furu Wei, Houfeng Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wen Luo 0001, Guangyue Peng, Wei Li 0101, Shaohang Wei, Feifan Song 0001, Liang Wang 0046, Nan Yang 0002, Xingxing Zhang 0002, Furu Wei, Houfeng Wang |
ACL (1) | 6 |
| 2025 | Generative Representational Instruction TuningabstractAll text-based language problems can be reduced to either generation or embedding. Current models only perform well at one or the other. We introduce generative representational instruction tuning (GRIT) whereby a large language model is trained to handle both generative and embedding tasks by distinguishing between them through instructions. Compared to other open models, our resulting GritLM-7B is among the top models on the Massive Text Embedding Benchmark (MTEB) and outperforms various models up to its size on a range of generative tasks. By scaling up further, GritLM-8x7B achieves even stronger generative performance while still being among the best embedding models. Notably, we find that GRIT matches training on only generative or embedding data, thus we can unify both at no performance loss. Among other benefits, the unification via GRIT speeds up Retrieval-Augmented Generation (RAG) by > 60% for long documents, by no longer requiring separate retrieval and generation models. Models, code, etc. are freely available at https://github.com/ContextualAI/gritlm. Niklas Muennighoff, Hongjin Su, Liang Wang 0046, Nan Yang 0002, Furu Wei, Tao Yu 0009, Amanpreet Singh, Douwe Kiela |
ICLR | 3 |
| 2025 | Little Giants: Synthesizing High-Quality Embedding Data at ScaleabstractHaonan Chen, Liang Wang, Nan Yang, Yutao Zhu, Ziliang Zhao, Furu Wei, Zhicheng Dou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Haonan Chen 0005, Liang Wang 0046, Nan Yang 0002, Yutao Zhu 0001, Ziliang Zhao 0001, Furu Wei, Zhicheng Dou |
NAACL (Long Papers) | 2 |
| 2025 | Chain-of-Retrieval Augmented GenerationabstractThis paper introduces an approach for training o1-like RAG models that retrieve and reason over relevant information step by step before generating the final answer. Conventional RAG methods usually perform a single retrieval step before the generation process, which limits their effectiveness in addressing complex queries due to imperfect retrieval results. In contrast, our proposed method, CoRAG (Chain-of-Retrieval Augmented Generation), allows the model to dynamically reformulate the query based on the evolving state. To train CoRAG effectively, we utilize rejection sampling to automatically generate intermediate retrieval chains, thereby augmenting existing RAG datasets that only provide the correct final answer. At test time, we propose various decoding strategies to scale the model's test-time compute by controlling the length and number of sampled retrieval chains. Experimental results across multiple benchmarks validate the efficacy of CoRAG, particularly in multi-hop question answering tasks, where we observe more than $10$ points improvement in EM score compared to strong baselines. On the KILT benchmark, CoRAG establishes a new state-of-the-art performance across a diverse range of knowledge-intensive tasks. Furthermore, we offer comprehensive analyses to understand the scaling behavior of CoRAG, laying the groundwork for future research aimed at developing factual and grounded foundation models. Liang Wang 0046, Haonan Chen 0005, Nan Yang 0002, Xiaolong Huang 0002, Zhicheng Dou, Furu Wei |
NeurIPS | 1 |
| 2024 | Learning to Rank in Generative RetrievalabstractGenerative retrieval stands out as a promising new paradigm in text retrieval that aims to generate identifier strings of relevant passages as the retrieval target. This generative paradigm taps into powerful generative language models, distinct from traditional sparse or dense retrieval methods. However, only learning to generate is insufficient for generative retrieval. Generative retrieval learns to generate identifiers of relevant passages as an intermediate goal and then converts predicted identifiers into the final passage rank list. The disconnect between the learning objective of autoregressive models and the desired passage ranking target leads to a learning gap. To bridge this gap, we propose a learning-to-rank framework for generative retrieval, dubbed LTRGR. LTRGR enables generative retrieval to learn to rank passages directly, optimizing the autoregressive model toward the final passage ranking target via a rank loss. This framework only requires an additional learning-to-rank training phase to enhance current generative retrieval systems and does not add any burden to the inference stage. We conducted experiments on three public benchmarks, and the results demonstrate that LTRGR achieves state-of-the-art performance among generative retrieval methods. The code and checkpoints are released at https://github.com/liyongqi67/LTRGR. Yongqi Li 0001, Nan Yang 0002, Liang Wang 0046, Furu Wei, Wenjie Li 0002 |
AAAI | 3 |
| 2024 | Improving Text Embeddings with Large Language ModelsabstractIn this paper, we introduce a novel and simple method for obtaining high-quality text embeddings using only synthetic data and less than 1k training steps.Unlike existing methods that often depend on multi-stage intermediate pretraining with billions of weakly-supervised text pairs, followed by fine-tuning with a few labeled datasets, our method does not require building complex training pipelines or relying on manually collected datasets that are often constrained by task diversity and language coverage.We leverage proprietary LLMs to generate diverse synthetic data for hundreds of thousands of text embedding tasks across 93 languages.We then fine-tune open-source decoder-only LLMs on the synthetic data using standard contrastive loss.Experiments demonstrate that our method achieves strong performance on highly competitive text embedding benchmarks without using any labeled data.Furthermore, when fine-tuned with a mixture of synthetic and labeled data, our model sets new state-of-the-art results on the BEIR and MTEB benchmarks. Liang Wang 0046, Nan Yang 0002, Xiaolong Huang 0002, Linjun Yang, Rangan Majumder, Furu Wei |
ACL (1) | 1 |
| 2024 | Learning to Retrieve In-Context Examples for Large Language ModelsabstractLarge language models (LLMs) have demonstrated their ability to learn in-context, allowing them to perform various tasks based on a few input-output examples.However, the effectiveness of in-context learning is heavily reliant on the quality of the selected examples.In this paper, we propose a novel framework to iteratively train dense retrievers that can identify high-quality in-context examples for LLMs.Our framework initially trains a reward model based on LLM feedback to evaluate the quality of candidate examples, followed by knowledge distillation to train a bi-encoder based dense retriever.Our experiments on a suite of 30 tasks demonstrate that our framework significantly enhances in-context learning performance.Furthermore, we show the generalization ability of our framework to unseen tasks during training.An in-depth analysis reveals that our model improves performance by retrieving examples with similar patterns, and the gains are consistent across LLMs of varying sizes.The code and data are available at https://github.com/microsoft/LMOps/ tree/main/llm_retriever. Liang Wang 0046, Nan Yang 0002, Furu Wei |
EACL (1) | 1 |
| 2024 | LongEmbed: Extending Embedding Models for Long Context RetrievalabstractEmbedding models play a pivotal role in modern NLP applications such as document retrieval.However, existing embedding models are limited to encoding short documents of typically 512 tokens, restrained from application scenarios requiring long inputs.This paper explores context window extension of existing embedding models, pushing their input length to a maximum of 32,768.We begin by evaluating the performance of existing embedding models using our newly constructed LONGEM-BED benchmark, which includes two synthetic and four real-world tasks, featuring documents of varying lengths and dispersed target information.The benchmarking results highlight huge opportunities for enhancement in current models.Via comprehensive experiments, we demonstrate that training-free context window extension strategies can effectively increase the input length of these models by several folds.Moreover, comparison of models using Absolute Position Encoding (APE) and Rotary Position Encoding (RoPE) reveals the superiority of RoPE-based embedding models in context window extension, offering empirical guidance for future models.Our benchmark, code and trained models will be released to advance the research in long context embedding models. Liang Wang 0046, Nan Yang 0002, Yifan Song 0002, Furu Wei, Sujian Li |
EMNLP | 2 |
| 2024 | PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise TrainingabstractLarge Language Models (LLMs) are trained with a pre-defined context length, restricting their use in scenarios requiring long inputs. Previous efforts for adapting LLMs to a longer length usually requires fine-tuning with this target length (Full-length fine-tuning), suffering intensive training cost. To decouple train length from target length for efficient context window extension, we propose Positional Skip-wisE (PoSE) training that smartly simulates long inputs using a fixed context window. This is achieved by first dividing the original context window into several chunks, then designing distinct skipping bias terms to manipulate the position indices of each chunk. These bias terms and the lengths of each chunk are altered for every training example, allowing the model to adapt to all positions within target length. Experimental results show that PoSE greatly reduces memory and time overhead compared with Full-length fine-tuning, with minimal impact on performance. Leveraging this advantage, we have successfully extended the LLaMA model to 128k tokens using a 2k training context window. Furthermore, we empirically confirm that PoSE is compatible with all RoPE-based LLMs and position interpolation strategies. Notably, our method can potentially support infinite length, limited only by memory usage in inference. With ongoing progress for efficient inference, we believe PoSE can further scale the context window beyond 128k. Nan Yang 0002, Liang Wang 0046, Yifan Song 0002, Furu Wei, Sujian Li |
ICLR | 3 |
| 2024 | Fine-Tuning LLaMA for Multi-Stage Text RetrievalabstractWhile large language models (LLMs) have shown impressive NLP capabilities, existing IR applications mainly focus on prompting LLMs to generate query expansions or generating permutations for listwise reranking. In this study, we leverage LLMs directly to serve as components in the widely used multi-stage text ranking pipeline. Specifically, we fine-tune the open-source LLaMA-2 model as a dense retriever (repLLaMA) and a pointwise reranker (rankLLaMA). This is performed for both passage and document retrieval tasks using the MS MARCO training data. Our study shows that finetuned LLM retrieval models outperform smaller models. They are more effective and exhibit greater generalizability, requiring only a straightforward training strategy. Moreover, our pipeline allows for the fine-tuning of LLMs at each stage of a multi-stage retrieval pipeline. This demonstrates the strong potential for optimizing LLMs to enhance a variety of retrieval tasks. Furthermore, as LLMs are naturally pre-trained with longer contexts, they can directly represent longer documents. This eliminates the need for heuristic segmenting and pooling strategies to rank long documents. On the MS MARCO and BEIR datasets, our repLLaMA-rankLLaMA pipeline demonstrates a high level of effectiveness. Xueguang Ma, Liang Wang 0046, Nan Yang 0002, Furu Wei, Jimmy Lin |
SIGIR | 2 |
| 2023 | Multiview Identifiers Enhanced Generative RetrievalabstractInstead of simply matching a query to preexisting passages, generative retrieval generates identifier strings of passages as the retrieval target.At a cost, the identifier must be distinctive enough to represent a passage.Current approaches use either a numeric ID or a text piece (such as a title or substrings) as the identifier.However, these identifiers cannot cover a passage's content well.As such, we are motivated to propose a new type of identifier, synthetic identifiers, that are generated based on the content of a passage and could integrate contextualized information that text pieces lack.Furthermore, we simultaneously consider multiview identifiers, including synthetic identifiers, titles, and substrings.These views of identifiers complement each other and facilitate the holistic ranking of passages from multiple perspectives.We conduct a series of experiments on three public datasets, and the results indicate that our proposed approach performs the best in generative retrieval, demonstrating its effectiveness and robustness.The code is released at https://github.com/liyongqi67/MINDER. Yongqi Li 0001, Nan Yang 0002, Liang Wang 0046, Furu Wei, Wenjie Li 0002 |
ACL (1) | 3 |
| 2023 | SimLM: Pre-training with Representation Bottleneck for Dense Passage RetrievalabstractLiang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Liang Wang 0046, Nan Yang 0002, Xiaolong Huang 0002, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei |
ACL (1) | 1 |
| 2023 | Query2doc: Query Expansion with Large Language ModelsabstractThis paper introduces a simple yet effective query expansion approach, denoted as query2doc, to improve both sparse and dense retrieval systems.The proposed method first generates pseudo-documents by few-shot prompting large language models (LLMs), and then expands the query with generated pseudodocuments.LLMs are trained on web-scale text corpora and are adept at knowledge memorization.The pseudo-documents from LLMs often contain highly relevant information that can aid in query disambiguation and guide the retrievers.Experimental results demonstrate that query2doc boosts the performance of BM25 by 3% to 15% on ad-hoc IR datasets, such as MS-MARCO and TREC DL, without any model fine-tuning.Furthermore, our method also benefits state-of-the-art dense retrievers in terms of both in-domain and out-ofdomain results. Liang Wang 0046, Nan Yang 0002, Furu Wei |
EMNLP | 1 |
| 2023 | Generative retrieval for conversational question answering
Yongqi Li 0001, Nan Yang 0002, Liang Wang 0046, Furu Wei, Wenjie Li 0002 |
Inf. Process. Manag. | 3 |
| 2022 | SimKGC: Simple Contrastive Knowledge Graph Completion with Pre-trained Language ModelsabstractKnowledge graph completion (KGC) aims to reason over known facts and infer the missing links. Text-based methods such as KGBERT (Yao et al., 2019) learn entity representations from natural language descriptions, and have the potential for inductive KGC. However, the performance of text-based methods still largely lag behind graph embedding-based methods like TransE (Bordes et al., 2013) and RotatE (Sun et al., 2019b). In this paper, we identify that the key issue is efficient contrastive learning. To improve the learning efficiency, we introduce three types of negatives: in-batch negatives, pre-batch negatives, and self-negatives which act as a simple form of hard negatives. Combined with InfoNCE loss, our proposed model SimKGC can substantially outperform embedding-based methods on several benchmark datasets. In terms of mean reciprocal rank (MRR), we advance the state-of-the-art by +19% on WN18RR, +6.8% on the Wikidata5M transductive setting, and +22% on the Wikidata5M inductive setting. Thorough analyses are conducted to gain insights into each component. Our code is available at https://github.com/intfloat/SimKGC . Liang Wang 0046, Zhuoyu Wei, Jingming Liu |
ACL (1) | 1 |
| 2021 | Aligning Cross-lingual Sentence Representations with Dual Momentum ContrastabstractIn this paper, we propose to align sentence representations from different languages into a unified embedding space, where semantic similarities (both cross-lingual and monolingual) can be computed with a simple dot product.Pre-trained language models are finetuned with the translation ranking task.Existing work (Feng et al., 2020) uses sentences within the same batch as negatives, which can suffer from the issue of easy negatives.We adapt MoCo (He et al., 2020) to further improve the quality of alignment.As the experimental results show, the sentence representations produced by our model achieve the new state-of-the-art on several tasks, including Tatoeba en-zh similarity search (Artetxe and Schwenk, 2019b), BUCC en-zh bitext mining, and semantic textual similarity on 7 datasets. Liang Wang 0046, Jingming Liu |
EMNLP (1) | 1 |
| 2019 | Denoising based Sequence-to-Sequence Pre-training for Text GenerationabstractLiang Wang, Wei Zhao, Ruoyu Jia, Sujian Li, Jingming Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Liang Wang 0046, Ruoyu Jia, Sujian Li, Jingming Liu |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Multi-Perspective Context Aggregation for Semi-supervised Cloze-style Reading ComprehensionabstractCloze-style reading comprehension has been a popular task for measuring the progress of natural language understanding in recent years. In this paper, we design a novel multi-perspective framework, which can be seen as the joint training of heterogeneous experts and aggregate context information from different perspectives. Each perspective is modeled by a simple aggregation module. The outputs of multiple aggregation modules are fed into a one-timestep pointer network to get the final answer. At the same time, to tackle the problem of insufficient labeled data, we propose an efficient sampling mechanism to automatically generate more training examples by matching the distribution of candidates between labeled and unlabeled data. We conduct our experiments on a recently released cloze-test dataset CLOTH (Xie et al., 2017), which consists of nearly 100k questions designed by professional teachers. Results show that our method achieves new state-of-the-art performance over previous strong baselines. Liang Wang 0046, Sujian Li, Kewei Shen, Ruoyu Jia, Jingming Liu |
COLING | 1 |
| 2017 | Learning to Rank Semantic Coherence for Topic SegmentationabstractTopic segmentation plays an important role for discourse parsing and information retrieval.Due to the absence of training data, previous work mainly adopts unsupervised methods to rank semantic coherence between paragraphs for topic segmentation.In this paper, we present an intuitive and simple idea to automatically create a "quasi" training dataset, which includes a large amount of text pairs from the same or different documents with different semantic coherence.With the training corpus, we design a symmetric CNN neural network to model text pairs and rank the semantic coherence within the learning to rank framework.Experiments show that our algorithm is able to achieve competitive performance over strong baselines on several real-world datasets. Liang Wang 0046, Sujian Li, Yajuan Lü, Houfeng Wang |
EMNLP | 1 |
| 2014 | Text-level Discourse Dependency ParsingabstractPrevious researches on Text-level discourse parsing mainly made use of constituency structure to parse the whole document into one discourse tree. In this paper, we present the limitations of constituency based dis-course parsing and first propose to use de-pendency structure to directly represent the relations between elementary discourse units (EDUs). The state-of-the-art depend-ency parsing techniques, the Eisner algo-rithm and maximum spanning tree (MST) algorithm, are adopted to parse an optimal discourse dependency tree based on the arc-factored model and the large-margin learn-ing techniques. Experiments show that our discourse dependency parsers achieve a competitive performance on text-level dis-course parsing. 1 Sujian Li, Liang Wang 0046, Ziqiang Cao, Wenjie Li 0002 |
ACL (1) | 2 |