EDBT 2026 Demo / reviewers in the wild / expert
Nan Yang 0002
dblp:51/1629-2
· DBLP profile ↗
43ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0001-7379-2609ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal EmbeddingsabstractMultimodal embedding models, built upon causal Vision Language Models (VLMs), have shown promise in various tasks.However, current approaches face three limitations: causal attention in VLM backbones is suboptimal for embedding tasks; scalability issues due to reliance on high-quality labeled paired data for contrastive learning; and limited diversity in training objectives and data.To address these issues, we propose MoCa, a two-stage framework for transforming pre-trained VLMs into bidirectional multimodal embedding models.The first stage, Modality-aware Continual Pre-training, introduces a joint reconstruction objective that simultaneously denoises interleaved texts and images, enhancing bidirectional context-aware reasoning.The second stage, Heterogeneous Contrastive Fine-tuning, leverages diverse, semantically rich multimodal data beyond simple image-caption pairs to enhance generalization and alignment.Our method addresses the stated limitations by introducing bidirectional attention through continual pre-training, scaling effectively with massive unlabeled datasets via joint reconstruction objectives, and utilizing diverse multimodal data for enhanced representation robustness.Experiments demonstrate that MoCa consistently improves performance across MMEB and ViDoRe-v2 benchmarks, achieving new state-of-the-arts, and exhibits strong scalability with both model size and training data on MMEB.We have released the model weights and data on our project page https://haon-chen.github.io/MoCa/. Haonan Chen 0005, Yuping Luo, Liang Wang 0046, Nan Yang 0002, Furu Wei, Zhicheng Dou |
ACL (1) | 5 |
| 2026 | Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM HallucinationsabstractWen Luo, Guangyue Peng, Wei Li, Shaohang Wei, Feifan Song, Liang Wang, Nan Yang, Xingxing Zhang, Jing Jin, Furu Wei, Houfeng Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wen Luo 0001, Guangyue Peng, Wei Li 0101, Shaohang Wei, Feifan Song 0001, Liang Wang 0046, Nan Yang 0002, Xingxing Zhang 0002, Furu Wei, Houfeng Wang |
ACL (1) | 7 |
| 2025 | Generative Representational Instruction TuningabstractAll text-based language problems can be reduced to either generation or embedding. Current models only perform well at one or the other. We introduce generative representational instruction tuning (GRIT) whereby a large language model is trained to handle both generative and embedding tasks by distinguishing between them through instructions. Compared to other open models, our resulting GritLM-7B is among the top models on the Massive Text Embedding Benchmark (MTEB) and outperforms various models up to its size on a range of generative tasks. By scaling up further, GritLM-8x7B achieves even stronger generative performance while still being among the best embedding models. Notably, we find that GRIT matches training on only generative or embedding data, thus we can unify both at no performance loss. Among other benefits, the unification via GRIT speeds up Retrieval-Augmented Generation (RAG) by > 60% for long documents, by no longer requiring separate retrieval and generation models. Models, code, etc. are freely available at https://github.com/ContextualAI/gritlm. Niklas Muennighoff, Hongjin Su, Liang Wang 0046, Nan Yang 0002, Furu Wei, Tao Yu 0009, Amanpreet Singh, Douwe Kiela |
ICLR | 4 |
| 2025 | Little Giants: Synthesizing High-Quality Embedding Data at ScaleabstractHaonan Chen, Liang Wang, Nan Yang, Yutao Zhu, Ziliang Zhao, Furu Wei, Zhicheng Dou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Haonan Chen 0005, Liang Wang 0046, Nan Yang 0002, Yutao Zhu 0001, Ziliang Zhao 0001, Furu Wei, Zhicheng Dou |
NAACL (Long Papers) | 3 |
| 2025 | Chain-of-Retrieval Augmented GenerationabstractThis paper introduces an approach for training o1-like RAG models that retrieve and reason over relevant information step by step before generating the final answer. Conventional RAG methods usually perform a single retrieval step before the generation process, which limits their effectiveness in addressing complex queries due to imperfect retrieval results. In contrast, our proposed method, CoRAG (Chain-of-Retrieval Augmented Generation), allows the model to dynamically reformulate the query based on the evolving state. To train CoRAG effectively, we utilize rejection sampling to automatically generate intermediate retrieval chains, thereby augmenting existing RAG datasets that only provide the correct final answer. At test time, we propose various decoding strategies to scale the model's test-time compute by controlling the length and number of sampled retrieval chains. Experimental results across multiple benchmarks validate the efficacy of CoRAG, particularly in multi-hop question answering tasks, where we observe more than $10$ points improvement in EM score compared to strong baselines. On the KILT benchmark, CoRAG establishes a new state-of-the-art performance across a diverse range of knowledge-intensive tasks. Furthermore, we offer comprehensive analyses to understand the scaling behavior of CoRAG, laying the groundwork for future research aimed at developing factual and grounded foundation models. Liang Wang 0046, Haonan Chen 0005, Nan Yang 0002, Xiaolong Huang 0002, Zhicheng Dou, Furu Wei |
NeurIPS | 3 |
| 2024 | Learning to Rank in Generative RetrievalabstractGenerative retrieval stands out as a promising new paradigm in text retrieval that aims to generate identifier strings of relevant passages as the retrieval target. This generative paradigm taps into powerful generative language models, distinct from traditional sparse or dense retrieval methods. However, only learning to generate is insufficient for generative retrieval. Generative retrieval learns to generate identifiers of relevant passages as an intermediate goal and then converts predicted identifiers into the final passage rank list. The disconnect between the learning objective of autoregressive models and the desired passage ranking target leads to a learning gap. To bridge this gap, we propose a learning-to-rank framework for generative retrieval, dubbed LTRGR. LTRGR enables generative retrieval to learn to rank passages directly, optimizing the autoregressive model toward the final passage ranking target via a rank loss. This framework only requires an additional learning-to-rank training phase to enhance current generative retrieval systems and does not add any burden to the inference stage. We conducted experiments on three public benchmarks, and the results demonstrate that LTRGR achieves state-of-the-art performance among generative retrieval methods. The code and checkpoints are released at https://github.com/liyongqi67/LTRGR. Yongqi Li 0001, Nan Yang 0002, Liang Wang 0046, Furu Wei, Wenjie Li 0002 |
AAAI | 2 |
| 2024 | Improving Text Embeddings with Large Language ModelsabstractIn this paper, we introduce a novel and simple method for obtaining high-quality text embeddings using only synthetic data and less than 1k training steps.Unlike existing methods that often depend on multi-stage intermediate pretraining with billions of weakly-supervised text pairs, followed by fine-tuning with a few labeled datasets, our method does not require building complex training pipelines or relying on manually collected datasets that are often constrained by task diversity and language coverage.We leverage proprietary LLMs to generate diverse synthetic data for hundreds of thousands of text embedding tasks across 93 languages.We then fine-tune open-source decoder-only LLMs on the synthetic data using standard contrastive loss.Experiments demonstrate that our method achieves strong performance on highly competitive text embedding benchmarks without using any labeled data.Furthermore, when fine-tuned with a mixture of synthetic and labeled data, our model sets new state-of-the-art results on the BEIR and MTEB benchmarks. Liang Wang 0046, Nan Yang 0002, Xiaolong Huang 0002, Linjun Yang, Rangan Majumder, Furu Wei |
ACL (1) | 2 |
| 2024 | Learning to Retrieve In-Context Examples for Large Language ModelsabstractLarge language models (LLMs) have demonstrated their ability to learn in-context, allowing them to perform various tasks based on a few input-output examples.However, the effectiveness of in-context learning is heavily reliant on the quality of the selected examples.In this paper, we propose a novel framework to iteratively train dense retrievers that can identify high-quality in-context examples for LLMs.Our framework initially trains a reward model based on LLM feedback to evaluate the quality of candidate examples, followed by knowledge distillation to train a bi-encoder based dense retriever.Our experiments on a suite of 30 tasks demonstrate that our framework significantly enhances in-context learning performance.Furthermore, we show the generalization ability of our framework to unseen tasks during training.An in-depth analysis reveals that our model improves performance by retrieving examples with similar patterns, and the gains are consistent across LLMs of varying sizes.The code and data are available at https://github.com/microsoft/LMOps/ tree/main/llm_retriever. Liang Wang 0046, Nan Yang 0002, Furu Wei |
EACL (1) | 2 |
| 2024 | LongEmbed: Extending Embedding Models for Long Context RetrievalabstractEmbedding models play a pivotal role in modern NLP applications such as document retrieval.However, existing embedding models are limited to encoding short documents of typically 512 tokens, restrained from application scenarios requiring long inputs.This paper explores context window extension of existing embedding models, pushing their input length to a maximum of 32,768.We begin by evaluating the performance of existing embedding models using our newly constructed LONGEM-BED benchmark, which includes two synthetic and four real-world tasks, featuring documents of varying lengths and dispersed target information.The benchmarking results highlight huge opportunities for enhancement in current models.Via comprehensive experiments, we demonstrate that training-free context window extension strategies can effectively increase the input length of these models by several folds.Moreover, comparison of models using Absolute Position Encoding (APE) and Rotary Position Encoding (RoPE) reveals the superiority of RoPE-based embedding models in context window extension, offering empirical guidance for future models.Our benchmark, code and trained models will be released to advance the research in long context embedding models. Liang Wang 0046, Nan Yang 0002, Yifan Song 0002, Furu Wei, Sujian Li |
EMNLP | 3 |
| 2024 | PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise TrainingabstractLarge Language Models (LLMs) are trained with a pre-defined context length, restricting their use in scenarios requiring long inputs. Previous efforts for adapting LLMs to a longer length usually requires fine-tuning with this target length (Full-length fine-tuning), suffering intensive training cost. To decouple train length from target length for efficient context window extension, we propose Positional Skip-wisE (PoSE) training that smartly simulates long inputs using a fixed context window. This is achieved by first dividing the original context window into several chunks, then designing distinct skipping bias terms to manipulate the position indices of each chunk. These bias terms and the lengths of each chunk are altered for every training example, allowing the model to adapt to all positions within target length. Experimental results show that PoSE greatly reduces memory and time overhead compared with Full-length fine-tuning, with minimal impact on performance. Leveraging this advantage, we have successfully extended the LLaMA model to 128k tokens using a 2k training context window. Furthermore, we empirically confirm that PoSE is compatible with all RoPE-based LLMs and position interpolation strategies. Notably, our method can potentially support infinite length, limited only by memory usage in inference. With ongoing progress for efficient inference, we believe PoSE can further scale the context window beyond 128k. Nan Yang 0002, Liang Wang 0046, Yifan Song 0002, Furu Wei, Sujian Li |
ICLR | 2 |
| 2024 | Fine-Tuning LLaMA for Multi-Stage Text RetrievalabstractWhile large language models (LLMs) have shown impressive NLP capabilities, existing IR applications mainly focus on prompting LLMs to generate query expansions or generating permutations for listwise reranking. In this study, we leverage LLMs directly to serve as components in the widely used multi-stage text ranking pipeline. Specifically, we fine-tune the open-source LLaMA-2 model as a dense retriever (repLLaMA) and a pointwise reranker (rankLLaMA). This is performed for both passage and document retrieval tasks using the MS MARCO training data. Our study shows that finetuned LLM retrieval models outperform smaller models. They are more effective and exhibit greater generalizability, requiring only a straightforward training strategy. Moreover, our pipeline allows for the fine-tuning of LLMs at each stage of a multi-stage retrieval pipeline. This demonstrates the strong potential for optimizing LLMs to enhance a variety of retrieval tasks. Furthermore, as LLMs are naturally pre-trained with longer contexts, they can directly represent longer documents. This eliminates the need for heuristic segmenting and pooling strategies to rank long documents. On the MS MARCO and BEIR datasets, our repLLaMA-rankLLaMA pipeline demonstrates a high level of effectiveness. Xueguang Ma, Liang Wang 0046, Nan Yang 0002, Furu Wei, Jimmy Lin |
SIGIR | 3 |
| 2023 | Multiview Identifiers Enhanced Generative RetrievalabstractInstead of simply matching a query to preexisting passages, generative retrieval generates identifier strings of passages as the retrieval target.At a cost, the identifier must be distinctive enough to represent a passage.Current approaches use either a numeric ID or a text piece (such as a title or substrings) as the identifier.However, these identifiers cannot cover a passage's content well.As such, we are motivated to propose a new type of identifier, synthetic identifiers, that are generated based on the content of a passage and could integrate contextualized information that text pieces lack.Furthermore, we simultaneously consider multiview identifiers, including synthetic identifiers, titles, and substrings.These views of identifiers complement each other and facilitate the holistic ranking of passages from multiple perspectives.We conduct a series of experiments on three public datasets, and the results indicate that our proposed approach performs the best in generative retrieval, demonstrating its effectiveness and robustness.The code is released at https://github.com/liyongqi67/MINDER. Yongqi Li 0001, Nan Yang 0002, Liang Wang 0046, Furu Wei, Wenjie Li 0002 |
ACL (1) | 2 |
| 2023 | SimLM: Pre-training with Representation Bottleneck for Dense Passage RetrievalabstractLiang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Liang Wang 0046, Nan Yang 0002, Xiaolong Huang 0002, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei |
ACL (1) | 2 |
| 2023 | Query2doc: Query Expansion with Large Language ModelsabstractThis paper introduces a simple yet effective query expansion approach, denoted as query2doc, to improve both sparse and dense retrieval systems.The proposed method first generates pseudo-documents by few-shot prompting large language models (LLMs), and then expands the query with generated pseudodocuments.LLMs are trained on web-scale text corpora and are adept at knowledge memorization.The pseudo-documents from LLMs often contain highly relevant information that can aid in query disambiguation and guide the retrievers.Experimental results demonstrate that query2doc boosts the performance of BM25 by 3% to 15% on ad-hoc IR datasets, such as MS-MARCO and TREC DL, without any model fine-tuning.Furthermore, our method also benefits state-of-the-art dense retrievers in terms of both in-domain and out-ofdomain results. Liang Wang 0046, Nan Yang 0002, Furu Wei |
EMNLP | 2 |
| 2023 | Generative retrieval for conversational question answering
Yongqi Li 0001, Nan Yang 0002, Liang Wang 0046, Furu Wei, Wenjie Li 0002 |
Inf. Process. Manag. | 2 |
| 2021 | xMoCo: Cross Momentum Contrastive Learning for Open-Domain Question AnsweringabstractNan Yang, Furu Wei, Binxing Jiao, Daxing Jiang, Linjun Yang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Nan Yang 0002, Furu Wei, Binxing Jiao, Daxing Jiang, Linjun Yang |
ACL/IJCNLP (1) | 1 |
| 2021 | InfoXLM: An Information-Theoretic Framework for Cross-Lingual Language Model Pre-TrainingabstractZewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, Heyan Huang, Ming Zhou. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zewen Chi, Li Dong 0004, Furu Wei, Nan Yang 0002, Saksham Singhal, Wenhui Wang 0003, Xianling Mao, Heyan Huang, Ming Zhou 0001 |
NAACL-HLT | 4 |
| 2020 | UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingabstractWe propose to pre-train a unified language model for both autoencoding and partially autoregressive language modeling tasks using a novel training procedure, referred to as a pseudo-masked language model (PMLM). Given an input text with masked tokens, we rely on conventional masks to learn inter-relations between corrupted tokens and context via autoencoding, and pseudo masks to learn intra-relations between masked spans via partially autoregressive modeling. With well-designed position embeddings and self-attention masks, the context encodings are reused to avoid redundant computation. Moreover, conventional masks used for autoencoding provide global masking information, so that all the position embeddings are accessible in partially autoregressive language modeling. In addition, the two tasks pre-train a unified language model as a bidirectional encoder and a sequence-to-sequence decoder, respectively. Our experiments show that the unified language models pre-trained using PMLM achieve new state-of-the-art results on a wide range of language understanding and generation tasks across several widely used benchmarks. The code and pre-trained models are available at https://github.com/microsoft/unilm. Hangbo Bao, Li Dong 0004, Furu Wei, Wenhui Wang 0003, Nan Yang 0002, Xiaodong Liu 0003, Yu Wang 0009, Jianfeng Gao 0001, Ming Zhou 0001, Hsiao-Wuen Hon |
ICML | 5 |
| 2020 | MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersabstractPre-trained language models (e.g., BERT (Devlin et al., 2018) and its variants) have achieved remarkable success in varieties of NLP tasks. However, these models usually consist of hundreds of millions of parameters which brings challenges for fine-tuning and online serving in real-life applications due to latency and capacity constraints. In this work, we present a simple and effective approach to compress large Transformer (Vaswani et al., 2017) based pre-trained models, termed as deep self-attention distillation. The small model (student) is trained by deeply mimicking the self-attention module, which plays a vital role in Transformer networks, of the large model (teacher). Specifically, we propose distilling the self-attention module of the last Transformer layer of the teacher, which is effective and flexible for the student. Furthermore, we introduce the scaled dot-product between values in the self-attention module as the new deep self-attention knowledge, in addition to the attention distributions (i.e., the scaled dot-product of queries and keys) that have been used in existing works. Moreover, we show that introducing a teacher assistant (Mirzadeh et al., 2019) also helps the distillation of large pre-trained Transformer models. Experimental results demonstrate that our monolingual model outperforms state-of-the-art baselines in different parameter size of student models. In particular, it retains more than 99% accuracy on SQuAD 2.0 and several GLUE benchmark tasks using 50% of the Transformer parameters and computations of the teacher model. We also obtain competitive results in applying deep self-attention distillation to multilingual pre-trained models. Wenhui Wang 0003, Furu Wei, Li Dong 0004, Hangbo Bao, Nan Yang 0002, Ming Zhou 0001 |
NeurIPS | 5 |
| 2020 | A Joint Sentence Scoring and Selection Framework for Neural Extractive Document SummarizationabstractExtractive document summarization methods aim to extract important sentences to form a summary. Previous works perform this task by first scoring all sentences in the document then selecting most informative ones; while we propose to jointly learn the two steps with a novel end-to-end neural network framework. Specifically, the sentences in the input document are represented as real-valued vectors through a neural document encoder. Then the method builds the output summary by extracting important sentences one by one. Different from previous works, the proposed joint sentence scoring and selection framework directly predicts the relative sentence importance score according to both sentence content and previously selected sentences. We evaluate the proposed framework with two realizations: a hierarchical recurrent neural network based model; and a pre-training based model that uses BERT as the document encoder. Experiments on two datasets show that the proposed joint framework outperforms the state-of-the-art extractive summarization models which treat sentence scoring and selection as two subtasks. Qingyu Zhou, Nan Yang 0002, Furu Wei, Shaohan Huang, Ming Zhou 0001, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Read + Verify: Machine Reading Comprehension with Unanswerable QuestionsabstractMachine reading comprehension with unanswerable questions aims to abstain from answering when no answer can be inferred. In addition to extract answers, previous works usually predict an additional “no-answer” probability to detect unanswerable cases. However, they fail to validate the answerability of the question by verifying the legitimacy of the predicted answer. To address this problem, we propose a novel read-then-verify system, which not only utilizes a neural reader to extract candidate answers and produce no-answer probabilities, but also leverages an answer verifier to decide whether the predicted answer is entailed by the input snippets. Moreover, we introduce two auxiliary losses to help the reader better handle answer extraction as well as no-answer detection, and investigate three different architectures for the answer verifier. Our experiments on the SQuAD 2.0 dataset show that our system obtains a score of 74.2 F1 on test set, achieving state-of-the-art results at the time of submission (Aug. 28th, 2018). Furu Wei, Yuxing Peng 0001, Zhen Huang 0006, Nan Yang 0002, Dongsheng Li 0001 |
AAAI | 5 |
| 2019 | Unified Language Model Pre-training for Natural Language Understanding and GenerationabstractThis paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks. The model is pre-trained using three types of language modeling tasks: unidirectional, bidirectional, and sequence-to-sequence prediction. The unified modeling is achieved by employing a shared Transformer network and utilizing specific self-attention masks to control what context the prediction conditions on. UniLM compares favorably with BERT on the GLUE benchmark, and the SQuAD 2.0 and CoQA question answering tasks. Moreover, UniLM achieves new state-of-the-art results on five natural language generation datasets, including improving the CNN/DailyMail abstractive summarization ROUGE-L to 40.51 (2.04 absolute improvement), the Gigaword abstractive summarization ROUGE-L to 35.75 (0.86 absolute improvement), the CoQA generative question answering F1 score to 82.5 (37.1 absolute improvement), the SQuAD question generation BLEU-4 to 22.12 (3.75 absolute improvement), and the DSTC7 document-grounded dialog response generation NIST-4 to 2.67 (human performance is 2.65). The code and pre-trained models are available at https://github.com/microsoft/unilm. Li Dong 0004, Nan Yang 0002, Wenhui Wang 0003, Furu Wei, Xiaodong Liu 0003, Yu Wang 0009, Jianfeng Gao 0001, Ming Zhou 0001, Hsiao-Wuen Hon |
NeurIPS | 2 |
| 2018 | S-Net: From Answer Extraction to Answer Synthesis for Machine Reading ComprehensionabstractIn this paper, we present a novel approach to machine reading comprehension for the MS-MARCO dataset. Unlike the SQuAD dataset that aims to answer a question with exact text spans in a passage, the MS-MARCO dataset defines the task as answering a question from multiple passages and the words in the answer are not necessary in the passages. We therefore develop an extraction-then-synthesis framework to synthesize answers from extraction results. Specifically, the answer extraction model is first employed to predict the most important sub-spans from the passage as evidence, and the answer synthesis model takes the evidence as additional features along with the question and passage to further elaborate the final answers. We build the answer extraction model with state-of-the-art neural networks for single passage reading comprehension, and propose an additional task of passage ranking to help answer extraction in multiple passages. The answer synthesis model is based on the sequence-to-sequence neural networks with extracted evidences as features. Experiments show that our extraction-then-synthesis method outperforms state-of-the-art methods. Chuanqi Tan, Furu Wei, Nan Yang 0002, Bowen Du 0001, Weifeng Lv, Ming Zhou 0001 |
AAAI | 3 |
| 2018 | Sequential Copying NetworksabstractCopying mechanism shows effectiveness in sequence-to-sequence based neural network models for text generation tasks, such as abstractive sentence summarization and question generation. However, existing works on modeling copying or pointing mechanism only considers single word copying from the source sentences. In this paper, we propose a novel copying framework, named Sequential Copying Networks (SeqCopyNet), which not only learns to copy single words, but also copies sequences from the input sentence. It leverages the pointer networks to explicitly select a sub-span from the source side to target side, and integrates this sequential copying mechanism to the generation process in the encoder-decoder paradigm. Experiments on abstractive sentence summarization and question generation tasks show that the proposed SeqCopyNet can copy meaningful spans and outperforms the baseline models. Qingyu Zhou, Nan Yang 0002, Furu Wei, Ming Zhou 0001 |
AAAI | 2 |
| 2018 | Neural Document Summarization by Jointly Learning to Score and Select SentencesabstractSentence scoring and sentence selection are two main steps in extractive document summarization systems.However, previous works treat them as two separated subtasks.In this paper, we present a novel end-to-end neural network framework for extractive document summarization by jointly learning to score and select sentences.It first reads the document sentences with a hierarchical encoder to obtain the representation of sentences.Then it builds the output summary by extracting sentences one by one.Different from previous methods, our approach integrates the selection strategy into the scoring model, which directly predicts the relative importance given previously selected sentences.Experiments on the CNN/Daily Mail dataset show that the proposed framework significantly outperforms the state-of-the-art extractive summarization models. Qingyu Zhou, Nan Yang 0002, Furu Wei, Shaohan Huang, Ming Zhou 0001, Tiejun Zhao |
ACL (1) | 2 |
| 2018 | Attention-Guided Answer Distillation for Machine Reading ComprehensionabstractDespite that current reading comprehension systems have achieved significant advancements, their promising performances are often obtained at the cost of making an ensemble of numerous models. Besides, existing approaches are also vulnerable to adversarial attacks. This paper tackles these problems by leveraging knowledge distillation, which aims to transfer knowledge from an ensemble model to a single model. We first demonstrate that vanilla knowledge distillation applied to answer span prediction is effective for reading comprehension systems. We then propose two novel approaches that not only penalize the prediction on confusing answers but also guide the training with alignment information distilled from the ensemble. Experiments show that our best student model has only a slight drop of 0.4% F1 on the SQuAD test set compared to the ensemble teacher, while running 12x faster during inference. It even outperforms the teacher on adversarial SQuAD datasets and NarrativeQA benchmark. Yuxing Peng 0001, Furu Wei, Zhen Huang 0006, Dongsheng Li 0001, Nan Yang 0002, Ming Zhou 0001 |
EMNLP | 6 |
| 2018 | I Know There Is No Answer: Modeling Answer Validation for Machine Reading Comprehension
Chuanqi Tan, Furu Wei, Qingyu Zhou, Nan Yang 0002, Weifeng Lv, Ming Zhou 0001 |
NLPCC (1) | 4 |
| 2018 | Context-Aware Answer Sentence Selection With Hierarchical Gated Recurrent Neural NetworksabstractIn this paper, we study the task of reading comprehension style answer sentence selection that aims to select the best sentence from a given passage to answer a question. Unlike most previous works that match the question and each candidate sentence separately, we observe that the context information among sentences in the same passage plays a vital role in this task. We propose modeling context information with hierarchical gated recurrent neural networks. Specifically, we first apply a word level recurrent neural network to model the context independent matching between the question and each candidate sentence. We then employ a sentence level recurrent neural network to incorporate the context information among all candidate sentences. Moreover, we introduce the gate mechanism to select matching information before feeding into recurrent neural networks at both word and sentence level. Experiments on the WikiQA and SQuAD datasets show that our model outperforms state-of-the-art methods. Chuanqi Tan, Furu Wei, Qingyu Zhou, Nan Yang 0002, Bowen Du 0001, Weifeng Lv, Ming Zhou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Dependency-to-Dependency Neural Machine TranslationabstractRecent research has proven that syntactic knowledge is effective to improve the performance of neural machine translation (NMT). Most previous work focuses on leveraging either source or target syntax in the recurrent neural network (RNN) based encoder–decoder model. In this paper, we simultaneously use both source and target dependency tree to improve the NMT model. First, we propose a simple but effective syntax-aware encoder to incorporate source dependency tree into NMT. The new encoder enriches each source state with dependence relations from the tree. Then, we propose a novel sequence-to-dependence framework. In this framework, the target translation and its corresponding dependence tree are jointly constructed and modeled. During decoding, the tree structure is used as context to facilitate word generations. Finally, we extend the sequence-to-dependence framework with the syntax-aware encoder to build a dependence-NMT model and apply the dependence-based framework to the Transformer. Experimental results on several translation tasks show that both source and target dependence structures can improve the translation quality and their effects can be accumulated. Shuangzhi Wu, Dongdong Zhang 0001, Zhirui Zhang, Nan Yang 0002, Mu Li 0001, Ming Zhou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Gated Self-Matching Networks for Reading Comprehension and Question AnsweringabstractIn this paper, we present the gated selfmatching networks for reading comprehension style question answering, which aims to answer questions from a given passage.We first match the question and passage with gated attention-based recurrent networks to obtain the question-aware passage representation.Then we propose a self-matching attention mechanism to refine the representation by matching the passage against itself, which effectively encodes information from the whole passage.We finally employ the pointer networks to locate the positions of answers from the passages.We conduct extensive experiments on the SQuAD dataset.The single model achieves 71.3% on the evaluation metrics of exact match on the hidden test set, while the ensemble model further boosts the results to 75.9%.At the time of submission of the paper, our model holds the first place on the SQuAD leaderboard for both single and ensemble model. Wenhui Wang 0003, Nan Yang 0002, Furu Wei, Baobao Chang, Ming Zhou 0001 |
ACL (1) | 2 |
| 2017 | Sequence-to-Dependency Neural Machine TranslationabstractNowadays a typical Neural Machine Translation (NMT) model generates translations from left to right as a linear sequence, during which latent syntactic structures of the target sentences are not explicitly concerned.Inspired by the success of using syntactic knowledge of target language for improving statistical machine translation, in this paper we propose a novel Sequence-to-Dependency Neural Machine Translation (SD-NMT) method, in which the target word sequence and its corresponding dependency structure are jointly constructed and modeled, and this structure is used as context to facilitate word generations.Experimental results show that the proposed method significantly outperforms state-of-the-art baselines on Chinese-English and Japanese-English translation tasks. Shuangzhi Wu, Dongdong Zhang 0001, Nan Yang 0002, Mu Li 0001, Ming Zhou 0001 |
ACL (1) | 3 |
| 2017 | Selective Encoding for Abstractive Sentence SummarizationabstractWe propose a selective encoding model to extend the sequence-to-sequence framework for abstractive sentence summarization.It consists of a sentence encoder, a selective gate network, and an attention equipped decoder.The sentence encoder and decoder are built with recurrent neural networks.The selective gate network constructs a second level sentence representation by controlling the information flow from encoder to decoder.The second level representation is tailored for sentence summarization task, which leads to better performance.We evaluate our model on the English Gigaword, DUC 2004 and MSR abstractive sentence summarization datasets.The experimental results show that the proposed selective encoding model outperforms the state-ofthe-art baseline models. Qingyu Zhou, Nan Yang 0002, Furu Wei, Ming Zhou 0001 |
ACL (1) | 2 |
| 2017 | Neural Question Generation from Text: A Preliminary Study
Qingyu Zhou, Nan Yang 0002, Furu Wei, Chuanqi Tan, Hangbo Bao, Ming Zhou 0001 |
NLPCC | 2 |
| 2016 | Jointly Modeling Topics and Intents with Global Order StructureabstractModeling document structure is of great importance for discourse analysis and related applications. The goal of this research is to capture the document intent structure by modeling documents as a mixture of topic words and rhetorical words. While the topics are relatively unchanged through one document, the rhetorical functions of sentences usually change following certain orders in discourse. We propose GMM-LDA, a topic modeling based Bayesian unsupervised model, to analyze the document intent structure cooperated with order information. Our model is flexible that has the ability to combine the annotations and do supervised learning. Additionally, entropic regularization can be introduced to model the significant divergence between topics and intents. We perform experiments in both unsupervised and supervised settings, results show the superiority of our model over several state-of-the-art baselines. Jun Zhu 0001, Nan Yang 0002, Tian Tian 0001, Ming Zhou 0001, Bo Zhang 0010 |
AAAI | 3 |
| 2016 | Improving Attention Modeling with Implicit Distortion and Fertility for Machine TranslationabstractIn neural machine translation, the attention mechanism facilitates the translation process by producing a soft alignment between the source sentence and the target sentence. However, without dedicated distortion and fertility models seen in traditional SMT systems, the learned alignment may not be accurate, which can lead to low translation quality. In this paper, we propose two novel models to improve attention-based neural machine translation. We propose a recurrent attention mechanism as an implicit distortion model, and a fertility conditioned decoder as an implicit fertility model. We conduct experiments on large-scale Chinese–English translation tasks. The results show that our models significantly improve both the alignment and translation quality compared to the original attention mechanism and several other variations. Shujie Liu 0001, Nan Yang 0002, Mu Li 0001, Ming Zhou 0001, Kenny Q. Zhu |
COLING | 3 |
| 2016 | Sentiment Embeddings with Applications to Sentiment AnalysisabstractWe propose learning sentiment-specific word embeddings dubbed sentiment embeddings in this paper. Existing word embedding learning algorithms typically only use the contexts of words but ignore the sentiment of texts. It is problematic for sentiment analysis because the words with similar contexts but opposite sentiment polarity, such asgoodandbad, are mapped to neighboring word vectors. We address this issue by encoding sentiment information of texts (e.g., sentences and words) together with contexts of words in sentiment embeddings. By combining context and sentiment level evidences, the nearest neighbors in sentiment embedding space are semantically similar and it favors words with the same sentiment polarity. In order to learn sentiment embeddings effectively, we develop a number of neural networks with tailoring loss functions, and collect massive texts automatically with sentiment signals like emoticons as the training data. Sentiment embeddings can be naturally used as word features for a variety of sentiment analysis tasks without feature engineering. We apply sentiment embeddings to word-level sentiment analysis, sentence level sentiment classification, and building sentiment lexicons. Experimental results show that sentiment embeddings consistently outperform context-based embeddings on several benchmark datasets of these tasks. This work provides insights on the design of neural networks for learning task-specific word embeddings in other natural language processing tasks. Duyu Tang, Furu Wei, Bing Qin 0001, Nan Yang 0002, Ting Liu 0001, Ming Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2015 | Modeling Mention, Context and Entity with Neural Networks for Entity Disambiguation
Yaming Sun, Lei Lin 0001, Duyu Tang, Nan Yang 0002, Zhenzhou Ji, Xiaolong Wang 0001 |
IJCAI | 4 |
| 2014 | A Recursive Recurrent Neural Network for Statistical Machine TranslationabstractIn this paper, we propose a novel recursive recurrent neural network (R 2 NN) to model the end-to-end decoding process for statistical machine translation.R 2 NN is a combination of recursive neural network and recurrent neural network, and in turn integrates their respective capabilities: (1) new information can be used to generate the next hidden state, like recurrent neural networks, so that language model and translation model can be integrated naturally; (2) a tree structure can be built, as recursive neural networks, so as to generate the translation candidates in a bottom up manner.A semi-supervised training approach is proposed to train the parameters, and the phrase pair embedding is explored to model translation confidence directly.Experiments on a Chinese to English translation task show that our proposed R 2 NN can outperform the stateof-the-art baseline by about 1.5 points in BLEU. Shujie Liu 0001, Nan Yang 0002, Mu Li 0001, Ming Zhou 0001 |
ACL (1) | 2 |
| 2014 | Learning Sentiment-Specific Word Embedding for Twitter Sentiment ClassificationabstractWe present a method that learns word embedding for Twitter sentiment classification in this paper.Most existing algorithms for learning continuous word representations typically only model the syntactic context of words but ignore the sentiment of text.This is problematic for sentiment analysis as they usually map words with similar syntactic context but opposite sentiment polarity, such as good and bad, to neighboring word vectors.We address this issue by learning sentimentspecific word embedding (SSWE), which encodes sentiment information in the continuous representation of words.Specifically, we develop three neural networks to effectively incorporate the supervision from sentiment polarity of text (e.g.sentences or tweets) in their loss functions.To obtain large scale training corpora, we learn the sentiment-specific word embedding from massive distant-supervised tweets collected by positive and negative emoticons.Experiments on applying SS-WE to a benchmark Twitter sentiment classification dataset in SemEval 2013 show that (1) the SSWE feature performs comparably with hand-crafted features in the top-performed system; (2) the performance is further improved by concatenating SSWE with existing feature set. Duyu Tang, Furu Wei, Nan Yang 0002, Ming Zhou 0001, Ting Liu 0001, Bing Qin 0001 |
ACL (1) | 3 |
| 2014 | Radical-Enhanced Chinese Character Embedding
Yaming Sun, Lei Lin 0001, Nan Yang 0002, Zhenzhou Ji, Xiaolong Wang 0001 |
ICONIP (2) | 3 |
| 2013 | Word Alignment Modeling with Context Dependent Deep Neural Network
Nan Yang 0002, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Nenghai Yu |
ACL (1) | 1 |
| 2013 | Punctuation Prediction with Transition-based Parsing
Dongdong Zhang 0001, Shuangzhi Wu, Nan Yang 0002, Mu Li 0001 |
ACL (1) | 3 |
| 2012 | A Ranking-based Approach to Word Reordering for Statistical Machine Translation
Nan Yang 0002, Mu Li 0001, Dongdong Zhang 0001, Nenghai Yu |
ACL (1) | 1 |