EDBT 2026 Demo / reviewers in the wild / expert
Furu Wei
dblp:72/5870
· DBLP profile ↗
268ranked-venue papers
7as first author
153since 2021 · last 2026
0000-0002-7810-5852ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 238 · 2 first-author · 142 since 2021Graphics, computer vision, multimedia, augmented reality and games · 59 · 30 since 2021Databases, data management, data science and information retrieval · 26 · 7 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reasoning with Exploration: An Entropy PerspectiveabstractBalancing exploration and exploitation is a central goal in reinforcement learning (RL). Despite recent advances in enhancing language model (LM) reasoning, most methods lean toward exploitation, and increasingly encounter performance plateaus. In this work, we revisit entropy -- a signal of exploration in RL -- and examine its relationship to exploratory reasoning in LMs. Through empirical analysis, we uncover positive correlations between high-entropy regions and three types of exploratory reasoning actions: (1) pivotal tokens that determine or connect logical steps, (2) reflective actions such as self-verification and correction, and (3) rare behaviors under-explored by the base LMs. Motivated by this, we introduce a minimal modification to standard RL with only one line of code: augmenting the advantage function with an entropy-based term. Unlike traditional maximum-entropy methods which encourage exploration by promoting uncertainty, we encourage exploration by promoting deeper and longer reasoning chains. Notably, our method achieves significant gains on the Pass@K metric -- an upper-bound estimator of LM reasoning capabilities -- even when evaluated with extremely large K values, pushing the boundaries of LM reasoning. Daixuan Cheng, Shaohan Huang, Xuekai Zhu, Bo Dai 0026, Wayne Xin Zhao, Furu Wei |
AAAI | 7 |
| 2026 | MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal EmbeddingsabstractMultimodal embedding models, built upon causal Vision Language Models (VLMs), have shown promise in various tasks.However, current approaches face three limitations: causal attention in VLM backbones is suboptimal for embedding tasks; scalability issues due to reliance on high-quality labeled paired data for contrastive learning; and limited diversity in training objectives and data.To address these issues, we propose MoCa, a two-stage framework for transforming pre-trained VLMs into bidirectional multimodal embedding models.The first stage, Modality-aware Continual Pre-training, introduces a joint reconstruction objective that simultaneously denoises interleaved texts and images, enhancing bidirectional context-aware reasoning.The second stage, Heterogeneous Contrastive Fine-tuning, leverages diverse, semantically rich multimodal data beyond simple image-caption pairs to enhance generalization and alignment.Our method addresses the stated limitations by introducing bidirectional attention through continual pre-training, scaling effectively with massive unlabeled datasets via joint reconstruction objectives, and utilizing diverse multimodal data for enhanced representation robustness.Experiments demonstrate that MoCa consistently improves performance across MMEB and ViDoRe-v2 benchmarks, achieving new state-of-the-arts, and exhibits strong scalability with both model size and training data on MMEB.We have released the model weights and data on our project page https://haon-chen.github.io/MoCa/. Haonan Chen 0005, Yuping Luo, Liang Wang 0046, Nan Yang 0002, Furu Wei, Zhicheng Dou |
ACL (1) | 6 |
| 2026 | VFA: Empowering Multilingual MLLMs via Vision-Free AdaptationabstractYixia Li, Yaqing Shi, Zhiwen Ruan, Dongdong Zhang, Lingjie Jiang, Shaohan Huang, Yun Chen, Guanhua Chen, Furu Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yixia Li, Yaqing Shi, Zhiwen Ruan, Lingjie Jiang, Shaohan Huang, Furu Wei |
ACL (1) | 9 |
| 2026 | Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM HallucinationsabstractWen Luo, Guangyue Peng, Wei Li, Shaohang Wei, Feifan Song, Liang Wang, Nan Yang, Xingxing Zhang, Jing Jin, Furu Wei, Houfeng Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wen Luo 0001, Guangyue Peng, Wei Li 0101, Shaohang Wei, Feifan Song 0001, Liang Wang 0046, Nan Yang 0002, Xingxing Zhang 0002, Furu Wei, Houfeng Wang |
ACL (1) | 10 |
| 2026 | Towards Stable and Effective Reinforcement Learning for Mixture-of-ExpertsabstractDi Zhang, Xun Wu, Shaohan Huang, Lingjie Jiang, Yaru Hao, Li Dong, Zewen Chi, Zhifang Sui, Furu Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaohan Huang, Lingjie Jiang, Yaru Hao, Li Dong 0004, Zewen Chi, Zhifang Sui, Furu Wei |
ACL (1) | 9 |
| 2025 | Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning TasksabstractFangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann, Adrian de Wynter, Xun Wang, Si-Qing Chen, Michael J. Wooldridge, Janet B. Pierrehumbert, Furu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Fangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann, Adrian de Wynter, Xun Wang 0012, Michael J. Wooldridge, Janet B. Pierrehumbert, Furu Wei |
ACL (1) | 10 |
| 2025 | Autoregressive Speech Synthesis without Vector QuantizationabstractLingwei Meng, Long Zhou, Shujie Liu, Sanyuan Chen, Bing Han, Shujie Hu, Yanqing Liu, Jinyu Li, Sheng Zhao, Xixin Wu, Helen M. Meng, Furu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Lingwei Meng, Shujie Liu 0001, Sanyuan Chen, Bing Han 0008, Shujie Hu, Jinyu Li 0001, Sheng Zhao 0002, Xixin Wu, Helen M. Meng, Furu Wei |
ACL (1) | 12 |
| 2025 | Bitnet.cpp: Efficient Edge Inference for Ternary LLMsabstractJinheng Wang, Hansong Zhou, Ting Song, Shijie Cao, Yan Xia, Ting Cao, Jianyu Wei, Shuming Ma, Hongyu Wang, Furu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jinheng Wang, Hansong Zhou, Shijie Cao, Yan Xia 0005, Jianyu Wei, Shuming Ma, Furu Wei |
ACL (1) | 10 |
| 2025 | Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm PerspectiveabstractYiyao Yu, Yuxiang Zhang, Dongdong Zhang, Xiao Liang, Hengyuan Zhang, Xingxing Zhang, Mahmoud Khademi, Hany Hassan Awadalla, Junjie Wang, Yujiu Yang, Furu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yiyao Yu, Dongdong Zhang 0001, Xingxing Zhang 0002, Mahmoud Khademi, Hany Hassan, Junjie Wang 0011, Yujiu Yang 0001, Furu Wei |
ACL (1) | 11 |
| 2025 | ShifCon: Enhancing Non-Dominant Language Capabilities with a Shift-based Multilingual Contrastive FrameworkabstractHengyuan Zhang, Chenming Shang, Sizhe Wang, Dongdong Zhang, Yiyao Yu, Feng Yao, Renliang Sun, Yujiu Yang, Furu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chenming Shang, Dongdong Zhang 0001, Yiyao Yu, Renliang Sun, Yujiu Yang 0001, Furu Wei |
ACL (1) | 9 |
| 2025 | MMLU-CF: A Contamination-free Multi-task Language Understanding BenchmarkabstractMultiple-choice question (MCQ) datasets like Massive Multitask Language Understanding (MMLU) are widely used to evaluate the commonsense, understanding, and problem-solving abilities of large language models (LLMs). However, the open-source nature of these benchmarks and the broad sources of training data for LLMs have inevitably led to benchmark contamination, resulting in unreliable evaluation. To alleviate this issue, we propose the contamination-free MCQ benchmark called MMLU-CF, which reassesses LLMs’ understanding of world knowledge by averting both unintentional and malicious data contamination. To mitigate unintentional data contamination, we source questions from a broader domain of over 200 billion webpages and apply three specifically designed decontamination rules. To prevent malicious data contamination, we divide the benchmark into validation and test sets with similar difficulty and subject distributions. The test set remains closed-source to ensure reliable results, while the validation set is publicly available to promote transparency and facilitate independent evaluation. The performance gap between these two sets of LLMs will indicate the contamination degree on the validation set in the future. We evaluated over 40 mainstream LLMs on the MMLU-CF. Compared to the original MMLU, not only LLMs’ performances significantly dropped but also the performance rankings of them changed considerably. This indicates the effectiveness of our approach in establishing a contamination-free and fairer evaluation standard. Qihao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui 0001, Qinzheng Sun, Shaoguang Mao, Qiufeng Yin, Scarlett Li, Furu Wei |
ACL (1) | 11 |
| 2025 | ALYMPICS: LLM Agents Meet Game TheoryabstractGame theory is a branch of mathematics that studies strategic interactions among rational agents. We propose Alympics (Olympics for Agents), a systematic framework utilizing Large Language Model (LLM) agents for empirical game theory research. Alympics creates a versatile platform for studying complex game theory problems, bridging the gap between theoretical game theory and empirical investigations by providing a controlled environment for simulating human-like strategic interactions with LLM agents. In our pilot case study, the “Water Allocation Challenge”, we explore Alympics through a challenging strategic game focused on the multi-round auction of scarce survival resources. This study demonstrates the framework’s ability to qualitatively and quantitatively analyze game determinants, strategies, and outcomes. Additionally, we conduct a comprehensive human assessment and an in-depth evaluation of LLM agents in rational strategic decision-making scenarios. Our findings highlight LLM agents’ potential to advance game theory knowledge and expand the understanding of their proficiency in emulating human strategic behavior. Shaoguang Mao, Yuzhe Cai, Yan Xia 0005, Wenshan Wu, Xun Wang 0012, Qiang Guan, Tao Ge 0001, Furu Wei |
COLING | 9 |
| 2025 | PEACE: Empowering Geologic Map Holistic Understanding with MLLMsabstractGeologic map, as a fundamental diagram in geology science, provides critical insights into the structure and composition of Earth’s subsurface and surface. These maps are indispensable in various fields, including disaster assessment, resource exploration, and civil engineering. Despite their significance, current Multimodal Large Language Models (MLLMs) often fall short in geologic map understanding. This gap is primarily due to the challenging nature of cartographic generalization, which involves handling high-resolution map, managing multiple associated components, and requiring domain-specific knowledge. To quantify this gap, we construct GeoMap-Bench, the first-ever benchmark for evaluating MLLMs in geologic map understanding, which assesses the full-scale abilities in extracting, referring, grounding, reasoning, and analyzing. To bridge this gap, we introduce GeoMap-Agent, the inaugural agent designed for geologic map understanding, which features three modules: Hierarchical Information Extraction (HIE), Domain Knowledge Injection (DKI), and Prompt-enhanced Question Answering (PEQA). Inspired by the interdisciplinary collaboration among human scientists, an AI expert group acts as consultants, utilizing a diverse tool pool to comprehensively analyze questions. Through comprehensive experiments, GeoMap-Agent achieves an overall score of 0.811 on GeoMap-Bench, significantly outperforming 0.369 of GPT-4o. Our work, emPowering gEologic mAp holistiC undErstanding (PEACE) with MLLMs, paves the way for advanced AI applications in geology, enhancing the efficiency and accuracy of geological investigations. The code and data are available at https://github.com/microsoft/PEACE. Yangyu Huang, Qihao Zhao, Zhipeng Gui, Tengchao Lv, Lei Cui 0001, Scarlett Li, Furu Wei |
CVPR | 11 |
| 2025 | NL2Lean: Translating Natural Language into Lean 4 through Multi-Aspect Reinforcement LearningabstractYue Fang, Shaohan Huang, Xin Yu, Haizhen Huang, Zihan Zhang, Weiwei Deng, Furu Wei, Feng Sun, Qi Zhang, Zhi Jin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yue Fang 0001, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066, Zhi Jin 0001 |
EMNLP | 7 |
| 2025 | Textual Aesthetics in Large Language ModelsabstractImage aesthetics is a crucial metric in the field of image generation.However, textual aesthetics has not been sufficiently explored.With the widespread application of large language models (LLMs), previous work has primarily focused on the correctness of content and the helpfulness of responses.Nonetheless, providing responses with textual aesthetics is also an important factor for LLMs, which can offer a cleaner layout and ensure greater consistency and coherence in content.In this work, we introduce a pipeline for aesthetics polishing and help construct a textual aesthetics dataset named TEXAES.We propose a textual aesthetics-powered fine-tuning method based on direct preference optimization, termed TAPO, which leverages textual aesthetics without compromising content correctness.Additionally, we develop two evaluation methods for textual aesthetics based on text and image analysis, respectively.Our experiments demonstrate that using textual aesthetics data and employing the TAPO fine-tuning method not only improves aesthetic scores but also enhances performance on general evaluation datasets such as AlpacalEval and Arena-hard.Our code and data are available at https://github.com/ JackLingjie/ Lingjie Jiang, Shaohan Huang, Furu Wei |
EMNLP | 4 |
| 2025 | Examining False Positives under Inference Scaling for Mathematical ReasoningabstractRecent advancements in language models have led to significant improvements in mathematical reasoning across various benchmarks.However, most of these benchmarks rely on automatic evaluation methods that only compare final answers using heuristics, without verifying the underlying reasoning steps.This limitation results in false positive solutions, where models may produce correct final answers but with flawed deduction paths.In this paper, we systematically examine the prevalence of false positive solutions in mathematical problem solving for language models.We analyze the characteristics and extent of this issue across different open-source models, datasets of varying difficulty levels, and decoding strategies.Specifically, we explore how "false positives" influence the inference time scaling behavior of language models.Our experimental results reveal that: (1) false positive solutions persist across different models, datasets, and decoding methods, (2) sampling-based inference time scaling methods do not alleviate the problem, and (3) the pass@N evaluation metric is more susceptible to "false positives", suggesting a significantly lower scaling ceiling than what automatic evaluations indicate.Additionally, we analyze specific instances of "false positives" and discuss potential limitations in self-improvement techniques and synthetic data generation under such conditions.Our data and code are publicly available at https://github.com/Wloner0809/False- Positives-in-Math. Yu Wang 0089, Furu Wei, Fuli Feng |
EMNLP | 4 |
| 2025 | Boosting Large Language Model for Speech Synthesis: An Empirical StudyabstractLarge language models (LLMs) have made significant advancements in natural language processing and are concurrently extending the language ability to other modalities, such as speech and vision. Nevertheless, most of the previous work focuses on prompting LLMs with perception abilities like auditory comprehension, and the effective approach for augmenting LLMs with speech synthesis capabilities remains ambiguous. In this paper, we conduct a comprehensive empirical exploration of boosting LLMs with the ability to generate speech, by combining pre-trained LLM LLaMA/OPT and text-to-speech synthesis model VALL-E. We compare three integration methods between LLMs and speech synthesis models, including directly fine-tuned LLMs, superposed layers of LLMs and VALL-E, and coupled LLMs and VALL-E using LLMs as a powerful text encoder. Experimental results show that, using LoRA method to fine-tune LLMs directly to boost the speech synthesis capability does not work well, and superposed LLMs and VALL-E can improve the quality of generated speech both in speaker similarity and word error rate (WER). Among these three methods, coupled methods leveraging LLMs as the text encoder can achieve the best performance, making it outperform original speech synthesis models with a consistently better speaker similarity and a significant (10.9%) WER reduction. Hongkun Hao, Shujie Liu 0001, Jinyu Li 0001, Shujie Hu, Rui Wang 0015, Furu Wei |
ICASSP | 7 |
| 2025 | Rethinking DPO-Style Diffusion Aligning Frameworks
Shaohan Huang, Lingjie Jiang, Furu Wei |
ICCV | 4 |
| 2025 | Semi-Parametric Retrieval via Binary Bag-of-Tokens IndexabstractInformation retrieval has transitioned from standalone systems into essential components across broader applications, with indexing efficiency, cost-effectiveness, and freshness becoming increasingly critical yet often overlooked. In this paper, we introduce SemI-parametric Disentangled Retrieval (SiDR), a bi-encoder retrieval framework that decouples retrieval index from neural parameters to enable efficient, low-cost, and parameter-agnostic indexing for emerging use cases. Specifically, in addition to using embeddings as indexes like existing neural retrieval methods, SiDR supports a non-parametric tokenization index for search, achieving BM25-like indexing complexity with significantly better effectiveness. Our comprehensive evaluation across 16 retrieval benchmarks demonstrates that SiDR outperforms both neural and term-based retrieval baselines under the same indexing workload: (i) When using an embedding-based index, SiDR exceeds the performance of conventional neural retrievers while maintaining similar training complexity; (ii) When using a tokenization-based index, SiDR drastically reduces indexing cost and time, matching the complexity of traditional term-based retrieval, while consistently outperforming BM25 on all in-domain datasets; (iii) Additionally, we introduce a late parametric mechanism that matches BM25 index preparation time while outperforming other neural retrieval baselines in effectiveness. Jiawei Zhou 0003, Li Dong 0004, Furu Wei, Lei Chen 0002 |
ICLR | 3 |
| 2025 | Scaling Optimal LR Across Token HorizonsabstractState-of-the-art LLMs are powered by scaling -- scaling model size, training tokens, and cluster size. It is economically infeasible to extensively tune hyperparameters for the largest runs. Instead, approximately optimal hyperparameters must be inferred or transferred from smaller experiments. Hyperparameter transfer across model sizes has been studied in muP. However, hyperparameter transfer across training tokens -- or token horizon -- has not been studied yet. To remedy this we conduct a large-scale empirical study on how optimal learning rate (LR) depends on the token horizon in LLM training. We first demonstrate that the optimal LR changes significantly with token horizon -- longer training necessitates smaller LR. Secondly, we demonstrate that the optimal LR follows a scaling law and that the optimal LR for longer horizons can be accurately estimated from shorter horizons via such scaling laws. We also provide a rule-of-thumb for transferring LR across token horizons with zero overhead over current practices. Lastly, we provide evidence that LLama-1 used too high LR, and thus argue that hyperparameter transfer across data size is an overlooked component of LLM training. Johan Bjorck, Alon Benhaim, Vishrav Chaudhary, Furu Wei |
ICLR | 4 |
| 2025 | Self-Boosting Large Language Models with Synthetic Preference DataabstractThrough alignment with human preferences, Large Language Models (LLMs) have advanced significantly in generating honest, harmless, and helpful responses. However, collecting high-quality preference data is a resource-intensive and creativity-demanding process, especially for the continual improvement of LLMs. We introduce SynPO, a self-boosting paradigm that leverages synthetic preference data for model alignment. SynPO employs an iterative mechanism wherein a self-prompt generator creates diverse prompts, and a response improver refines model responses progressively. This approach trains LLMs to autonomously learn the generative rewards for their own outputs and eliminates the need for large-scale annotation of prompts and human preferences. After four SynPO iterations, Llama3-8B and Mistral-7B show significant enhancements in instruction-following abilities, achieving over 22.1% win rate improvements on AlpacaEval 2.0 and ArenaHard. Simultaneously, SynPO improves the general performance of LLMs on various tasks, validated by a 3.2 to 5.0 average score increase on the well-recognized Open LLM leaderboard. Qingxiu Dong, Li Dong 0004, Xingxing Zhang 0002, Zhifang Sui, Furu Wei |
ICLR | 5 |
| 2025 | Data Selection via Optimal Control for Language ModelsabstractThis work investigates the selection of high-quality pre-training data from massive corpora to enhance LMs' capabilities for downstream usage.
We formulate data selection as a generalized Optimal Control problem, which can be solved theoretically by Pontryagin's Maximum Principle (PMP), yielding a set of necessary conditions that characterize the relationship between optimal data selection and LM training dynamics.
Based on these theoretical results, we introduce **P**MP-based **D**ata **S**election (**PDS**), a framework that approximates optimal data selection by solving the PMP conditions.
In our experiments, we adopt PDS to select data from CommmonCrawl and show that the PDS-selected corpus accelerates the learning of LMs and constantly boosts their performance on a wide range of downstream tasks across various model sizes.
Moreover, the benefits of PDS extend to ~400B models trained on ~10T tokens, as evidenced by the extrapolation of the test loss curves according to the Scaling Laws.
PDS also improves data utilization when the pre-training data is limited, by reducing the data demand by 1.8 times, which helps mitigate the quick exhaustion of available web-crawled corpora. Our code, model, and data can be found at https://github.com/microsoft/LMOps/tree/main/data_selection. Yuxian Gu, Li Dong 0004, Hongning Wang, Yaru Hao, Qingxiu Dong, Furu Wei, Minlie Huang |
ICLR | 6 |
| 2025 | Preference Optimization for Reasoning with Pseudo FeedbackabstractPreference optimization techniques, such as Direct Preference Optimization (DPO), are frequently employed to enhance the reasoning capabilities of large language models (LLMs) in domains like mathematical reasoning and coding, typically following supervised fine-tuning. These methods rely on high-quality labels for reasoning tasks to generate preference pairs; however, the availability of reasoning datasets with human-verified labels is limited.
In this study, we introduce a novel approach to generate pseudo feedback for reasoning tasks by framing the labeling of solutions to reason problems as an evaluation against associated \emph{test cases}.
We explore two forms of pseudo feedback based on test cases: one generated by frontier LLMs and the other by extending self-consistency to multi-test-case.
We conduct experiments on both mathematical reasoning and coding tasks using pseudo feedback for preference optimization, and observe improvements across both tasks. Specifically, using Mathstral-7B as our base model, we improve MATH results from 58.3 to 68.6, surpassing both NuminaMath-72B and GPT-4-Turbo-1106-preview. In GSM8K and College Math, our scores increase from 85.6 to 90.3 and from 34.3 to 42.3, respectively. Building on Deepseek-coder-7B-v1.5, we achieve a score of 24.3 on LiveCodeBench (from 21.1), surpassing Claude-3-Haiku. Fangkai Jiao, Geyang Guo, Nancy F. Chen, Shafiq R. Joty, Furu Wei |
ICLR | 6 |
| 2025 | ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video GenerationabstractText-to-video (T2V) models have recently undergone rapid and substantial advancements. Nevertheless, due to limitations in data and computational resources, achieving efficient generation of long videos with rich motion dynamics remains a significant challenge.
To generate high-quality, dynamic, and temporally consistent long videos, this paper presents ARLON, a novel framework that boosts diffusion Transformers with autoregressive (\textbf{AR}) models for long (\textbf{LON}) video generation, by integrating the coarse spatial and long-range temporal information provided by the AR model to guide the DiT model effectively.
Specifically, ARLON incorporates several key innovations:
1) A latent Vector Quantized Variational Autoencoder (VQ-VAE) compresses the input latent space of the DiT model into compact and highly quantized visual tokens, bridging the AR and DiT models and balancing the learning complexity and information density;
2) An adaptive norm-based semantic injection module integrates the coarse discrete visual units from the AR model into the DiT model, ensuring effective guidance during video generation;
3) To enhance the tolerance capability of noise introduced from the AR inference, the DiT model is trained with coarser visual latent tokens incorporated with an uncertainty sampling module.
Experimental results demonstrate that ARLON significantly outperforms the baseline OpenSora-V1.2 on eight out of eleven metrics selected from VBench, with notable improvements in dynamic degree and aesthetic quality, while delivering competitive results on the remaining three and simultaneously accelerating the generation process. In addition, ARLON achieves state-of-the-art performance in long video generation, outperforming other open-source models in this domain.
Detailed analyses of the improvements in inference efficiency are presented, alongside a practical application that demonstrates the generation of long videos using progressive text prompts. Project page: \url{http://aka.ms/arlon}. Zongyi Li, Shujie Hu, Shujie Liu 0001, Jeongsoo Choi, Lingwei Meng, Jinyu Li 0001, Furu Wei |
ICLR | 10 |
| 2025 | Generative Representational Instruction TuningabstractAll text-based language problems can be reduced to either generation or embedding. Current models only perform well at one or the other. We introduce generative representational instruction tuning (GRIT) whereby a large language model is trained to handle both generative and embedding tasks by distinguishing between them through instructions. Compared to other open models, our resulting GritLM-7B is among the top models on the Massive Text Embedding Benchmark (MTEB) and outperforms various models up to its size on a range of generative tasks. By scaling up further, GritLM-8x7B achieves even stronger generative performance while still being among the best embedding models. Notably, we find that GRIT matches training on only generative or embedding data, thus we can unify both at no performance loss. Among other benefits, the unification via GRIT speeds up Retrieval-Augmented Generation (RAG) by > 60% for long documents, by no longer requiring separate retrieval and generation models. Models, code, etc. are freely available at https://github.com/ContextualAI/gritlm. Niklas Muennighoff, Hongjin Su, Liang Wang 0046, Nan Yang 0002, Furu Wei, Tao Yu 0009, Amanpreet Singh, Douwe Kiela |
ICLR | 5 |
| 2025 | Differential TransformerabstractTransformer tends to overallocate attention to irrelevant context. In this work, we introduce Diff Transformer, which amplifies attention to the relevant context while canceling noise. Specifically, the differential attention mechanism calculates attention scores as the difference between two separate softmax attention maps. The subtraction cancels noise, promoting the emergence of sparse attention patterns. Experimental results on language modeling show that Diff Transformer outperforms Transformer in various settings of scaling up model size and training tokens. More intriguingly, it offers notable advantages in practical applications, such as long-context modeling, key information retrieval, hallucination mitigation, in-context learning, and reduction of activation outliers. By being less distracted by irrelevant context, Diff Transformer can mitigate hallucination in question answering and text summarization. For in-context learning, Diff Transformer not only enhances accuracy but is also more robust to order permutation, which was considered as a chronic robustness issue. The results position Diff Transformer as a highly effective and promising architecture for large language models. Tianzhu Ye, Li Dong 0004, Yuqing Xia, Yutao Sun, Gao Huang 0001, Furu Wei |
ICLR | 7 |
| 2025 | Imagine While Reasoning in Space: Multimodal Visualization-of-ThoughtabstractChain-of-Thought (CoT) prompting has proven highly effective for enhancing complex reasoning in Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). Yet, it struggles in complex spatial reasoning tasks. Nonetheless, human cognition extends beyond language alone, enabling the remarkable capability to think in both words and images. Inspired by this mechanism, we propose a new reasoning paradigm, Multimodal Visualization-of-Thought (MVoT). It enables visual thinking in MLLMs by generating image visualizations of their reasoning traces. To ensure high-quality visualization, we introduce token discrepancy loss into autoregressive MLLMs. This innovation significantly improves both visual coherence and fidelity. We validate this approach through several dynamic spatial reasoning tasks. Experimental results reveal that MVoT demonstrates competitive performance across tasks. Moreover, it exhibits robust and reliable improvements in the most challenging scenarios where CoT fails. Ultimately, MVoT establishes new possibilities for complex reasoning tasks where visual thinking can effectively complement verbal reasoning. Chengzu Li, Wenshan Wu, Huanyu Zhang 0002, Yan Xia 0005, Shaoguang Mao, Li Dong 0004, Ivan Vulic, Furu Wei |
ICML | 8 |
| 2025 | Little Giants: Synthesizing High-Quality Embedding Data at ScaleabstractHaonan Chen, Liang Wang, Nan Yang, Yutao Zhu, Ziliang Zhao, Furu Wei, Zhicheng Dou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Haonan Chen 0005, Liang Wang 0046, Nan Yang 0002, Yutao Zhu 0001, Ziliang Zhao 0001, Furu Wei, Zhicheng Dou |
NAACL (Long Papers) | 6 |
| 2025 | K-Level Reasoning: Establishing Higher Order Beliefs in Large Language Models for Strategic ReasoningabstractYadong Zhang, Shaoguang Mao, Tao Ge, Xun Wang, Yan Xia, Man Lan, Furu Wei. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shaoguang Mao, Tao Ge 0001, Xun Wang 0012, Yan Xia 0005, Man Lan, Furu Wei |
NAACL (Long Papers) | 7 |
| 2025 | Reward Reasoning ModelsabstractReward models play a critical role in guiding large language models toward outputs that align with human expectations. However, an open challenge remains in effectively utilizing test-time compute to enhance reward model performance. In this work, we introduce Reward Reasoning Models (RRMs), which are specifically designed to execute a deliberate reasoning process before generating final rewards. Through chain-of-thought reasoning, RRMs leverage additional test-time compute for complex queries where appropriate rewards are not immediately apparent. To develop RRMs, we implement a reinforcement learning framework that fosters self-evolved reward reasoning capabilities without requiring explicit reasoning traces as training data. Experimental results demonstrate that RRMs achieve superior performance on reward modeling benchmarks across diverse domains. Notably, we show that RRMs can adaptively exploit test-time compute to further improve reward accuracy. The pretrained models are available at https://huggingface.co/Reward-Reasoning. Zewen Chi, Li Dong 0004, Qingxiu Dong, Shaohan Huang, Furu Wei |
NeurIPS | 7 |
| 2025 | Think Only When You Need with Large Hybrid-Reasoning ModelsabstractRecent Large Reasoning Models (LRMs) have shown substantially improved reasoning capabilities over traditional Large Language Models (LLMs) by incorporating extended thinking processes prior to producing final responses. However, excessively lengthy thinking introduces substantial overhead in terms of token consumption and latency, which is unnecessary for simple queries. In this work, we introduce Large Hybrid-Reasoning Models (LHRMs), the first kind of model capable of adaptively determining whether to perform reasoning based on the contextual information of user queries. To achieve this, we propose a two-stage training pipeline comprising Hybrid Fine-Tuning (HFT) as a cold start, followed by online reinforcement learning with the proposed Hybrid Group Policy Optimization (HGPO) to implicitly learn to select the appropriate reasoning mode. Furthermore, we introduce a metric called Hybrid Accuracy to quantitatively assess the model’s capability for hybrid reasoning. Extensive experimental results show that LHRMs can adaptively perform hybrid reasoning on queries of varying difficulty and type. It outperforms existing LRMs and LLMs in reasoning and general capabilities while significantly improving efficiency. Together, our work advocates for a reconsideration of the appropriate use of extended reasoning processes and provides a solid starting point for building hybrid reasoning systems. Lingjie Jiang, Shaohan Huang, Qingxiu Dong, Zewen Chi, Li Dong 0004, Xingxing Zhang 0002, Tengchao Lv, Lei Cui 0001, Furu Wei |
NeurIPS | 10 |
| 2025 | Chain-of-Retrieval Augmented GenerationabstractThis paper introduces an approach for training o1-like RAG models that retrieve and reason over relevant information step by step before generating the final answer. Conventional RAG methods usually perform a single retrieval step before the generation process, which limits their effectiveness in addressing complex queries due to imperfect retrieval results. In contrast, our proposed method, CoRAG (Chain-of-Retrieval Augmented Generation), allows the model to dynamically reformulate the query based on the evolving state. To train CoRAG effectively, we utilize rejection sampling to automatically generate intermediate retrieval chains, thereby augmenting existing RAG datasets that only provide the correct final answer. At test time, we propose various decoding strategies to scale the model's test-time compute by controlling the length and number of sampled retrieval chains. Experimental results across multiple benchmarks validate the efficacy of CoRAG, particularly in multi-hop question answering tasks, where we observe more than $10$ points improvement in EM score compared to strong baselines. On the KILT benchmark, CoRAG establishes a new state-of-the-art performance across a diverse range of knowledge-intensive tasks. Furthermore, we offer comprehensive analyses to understand the scaling behavior of CoRAG, laying the groundwork for future research aimed at developing factual and grounded foundation models. Liang Wang 0046, Haonan Chen 0005, Nan Yang 0002, Xiaolong Huang 0002, Zhicheng Dou, Furu Wei |
NeurIPS | 6 |
| 2025 | Towards Thinking-Optimal Scaling of Test-Time Compute for LLM ReasoningabstractRecent studies have shown that making a model spend more time thinking through longer Chain of Thoughts (CoTs) enables it to gain significant improvements in complex reasoning tasks. While current researches continue to explore the benefits of increasing test-time compute by extending the CoT lengths of Large Language Models (LLMs), we are concerned about a potential issue hidden behind the current pursuit of test-time scaling: Would excessively scaling the CoT length actually bring adverse effects to a model's reasoning performance? Our explorations on mathematical reasoning tasks reveal an unexpected finding that scaling with longer CoTs can indeed impair the reasoning performance of LLMs in certain domains. Moreover, we discover that there exists an optimal scaled length distribution that differs across different domains. Based on these insights, we propose a Thinking-Optimal Scaling strategy. Our method first uses a small set of seed data with varying response length distributions to teach the model to adopt different reasoning efforts for deep thinking. Then, the model selects its shortest correct response under different reasoning efforts on additional problems for self-improvement. Our self-improved models built upon Qwen2.5-32B-Instruct outperform other distillation-based 32B o1-like models across various math benchmarks, and achieve performance on par with the teacher model QwQ-32B-Preview that produces the seed data. Wenkai Yang, Shuming Ma, Yankai Lin 0001, Furu Wei |
NeurIPS | 4 |
| 2025 | Text to Image Generation with Bidirectional Multiway TransformersabstractIn this study, we explore the potential of Multiway Transformers for text-to-image generation to achieve performance improvements through a concise and efficient decoupled model design and the inference efficiency provided by bidirectional encoding. We propose a method for improving the image tokenizer using pretrained Vision Transformers. Next, we employ bidirectional Multiway Transformers to restore the masked visual tokens combined with the unmasked text tokens. On the MS-COCO benchmark, our Multiway Transformers outperform vanilla Transformers, achieving superior FID scores and confirming the efficacy of the modality-specific parameter computation design. Ablation studies reveal that the fusion of visual and text tokens in bidirectional encoding contributes to improved model performance. Additionally, our proposed tokenizer outperforms VQGAN in image reconstruction quality and enhances the text-to-image generation results. By incorporating the additional CC-3M dataset for intermediate finetuning on our model with 688M parameters, we achieve competitive results with a finetuned FID score of 4.98 on MS-COCO. Hangbo Bao, Li Dong 0004, Furu Wei |
Comput. Vis. Media | 4 |
| 2025 | BitNet: 1-bit Pre-training for Large Language ModelsabstractThe increasing size of large language models (LLMs) has posed challenges for deployment and raised concerns about environmental impact due to high energy consumption. Previous research typically applies quantization after pre-training. While these methods avoid the need for model retraining, they often cause notable accuracy loss at extremely low bit-widths. In this work, we explore the feasibility and scalability of 1-bit pre-training. We introduce BitNet b1 and BitNet b1.58, the scalable and stable 1-bit Transformer architecture designed for LLMs. Specifically, we introduce BitLinear as a drop-in replacement of the nn.Linear layer in order to train 1-bit weights from scratch. Experimental results show that BitNet b1 achieves competitive performance, compared to state-of-the-art 8-bit quantization methods and FP16 Transformer baselines. With the ternary weight, BitNet b1.58 matches the half-precision Transformer LLM with the same model size and training tokens in terms of both perplexity and end-task performance, while being significantly more cost-effective in terms of latency, memory, throughput, and energy consumption. More profoundly, BitNet defines a new scaling law and recipe for training new generations of LLMs that are both high-performance and cost-effective. It enables a new computation paradigm and opens the door for designing specific hardware optimized for 1-bit LLMs. Hongyu Wang 0009, Shuming Ma, Lingxiao Ma, Lei Wang 0222, Wenhui Wang 0003, Li Dong 0004, Shaohan Huang, Huaijie Wang, Jilong Xue, Yi Wu 0013, Furu Wei |
J. Mach. Learn. Res. | 12 |
| 2024 | Learning to Rank in Generative RetrievalabstractGenerative retrieval stands out as a promising new paradigm in text retrieval that aims to generate identifier strings of relevant passages as the retrieval target. This generative paradigm taps into powerful generative language models, distinct from traditional sparse or dense retrieval methods. However, only learning to generate is insufficient for generative retrieval. Generative retrieval learns to generate identifiers of relevant passages as an intermediate goal and then converts predicted identifiers into the final passage rank list. The disconnect between the learning objective of autoregressive models and the desired passage ranking target leads to a learning gap. To bridge this gap, we propose a learning-to-rank framework for generative retrieval, dubbed LTRGR. LTRGR enables generative retrieval to learn to rank passages directly, optimizing the autoregressive model toward the final passage ranking target via a rank loss. This framework only requires an additional learning-to-rank training phase to enhance current generative retrieval systems and does not add any burden to the inference stage. We conducted experiments on three public benchmarks, and the results demonstrate that LTRGR achieves state-of-the-art performance among generative retrieval methods. The code and checkpoints are released at https://github.com/liyongqi67/LTRGR. Yongqi Li 0001, Nan Yang 0002, Liang Wang 0046, Furu Wei, Wenjie Li 0002 |
AAAI | 4 |
| 2024 | Text Diffusion with Reinforced ConditioningabstractDiffusion models have demonstrated exceptional capability in generating high-quality images, videos, and audio. Due to their adaptiveness in iterative refinement, they provide a strong potential for achieving better non-autoregressive sequence generation. However, existing text diffusion models still fall short in their performance due to a challenge in handling the discreteness of language. This paper thoroughly analyzes text diffusion models and uncovers two significant limitations: degradation of self-conditioning during training and misalignment between training and sampling. Motivated by our findings, we propose a novel Text Diffusion model called TReC, which mitigates the degradation with Reinforced Conditioning and the misalignment by Time-Aware Variance Scaling. Our extensive experiments demonstrate the competitiveness of TReC against autoregressive, non-autoregressive, and diffusion baselines. Moreover, qualitative analysis shows its advanced ability to fully utilize the diffusion process in refining samples. Yuxuan Liu 0011, Tianchi Yang, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066 |
AAAI | 6 |
| 2024 | HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria DecompositionabstractYuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yuxuan Liu 0011, Tianchi Yang, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066 |
ACL (1) | 6 |
| 2024 | Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language ModelsabstractTianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, Ji-Rong Wen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Wenyang Luo, Haoyang Huang, Dongdong Zhang 0001, Xiaolei Wang 0005, Wayne Xin Zhao, Furu Wei, Ji-Rong Wen |
ACL (1) | 7 |
| 2024 | Improving Text Embeddings with Large Language ModelsabstractIn this paper, we introduce a novel and simple method for obtaining high-quality text embeddings using only synthetic data and less than 1k training steps.Unlike existing methods that often depend on multi-stage intermediate pretraining with billions of weakly-supervised text pairs, followed by fine-tuning with a few labeled datasets, our method does not require building complex training pipelines or relying on manually collected datasets that are often constrained by task diversity and language coverage.We leverage proprietary LLMs to generate diverse synthetic data for hundreds of thousands of text embedding tasks across 93 languages.We then fine-tune open-source decoder-only LLMs on the synthetic data using standard contrastive loss.Experiments demonstrate that our method achieves strong performance on highly competitive text embedding benchmarks without using any labeled data.Furthermore, when fine-tuned with a mixture of synthetic and labeled data, our model sets new state-of-the-art results on the BEIR and MTEB benchmarks. Liang Wang 0046, Nan Yang 0002, Xiaolong Huang 0002, Linjun Yang, Rangan Majumder, Furu Wei |
ACL (1) | 6 |
| 2024 | Respond in my Language: Mitigating Language Inconsistency in Response Generation based on Large Language ModelsabstractLarge Language Models (LLMs) show strong instruction understanding ability across multiple languages.However, they are easily biased towards English in instruction tuning, and generate English responses even given non-English instructions.In this paper, we investigate the language inconsistent generation problem in monolingual instruction tuning.We find that instruction tuning in English increases the models' preference for English responses.It attaches higher probabilities to English responses than to responses in the same language as the instruction.Based on the findings, we alleviate the language inconsistent generation problem by counteracting the model preference for English responses in both the training and inference stages.Specifically, we propose Pseudo-Inconsistent Penalization (PIP) which prevents the model from generating English responses when given non-English language prompts during training, and Prior Enhanced Decoding (PED) which improves the language-consistent prior by leveraging the untuned base language model.Experimental results show that our two methods significantly improve the language consistency of the model without requiring any multilingual data 1 . Qin Jin, Haoyang Huang, Furu Wei |
ACL (1) | 5 |
| 2024 | Calibrating LLM-Based EvaluatorabstractRecent advancements in large language models (LLMs) and their emergent capabilities make LLM a promising reference-free evaluator on the quality of natural language generation, and a competent alternative to human evaluation. However, hindered by the closed-source or high computational demand to host and tune, there is a lack of practice to further calibrate an off-the-shelf LLM-based evaluator towards better human alignment. In this work, we propose AutoCalibrate, a multi-stage, gradient-free approach to automatically calibrate and align an LLM-based evaluator toward human preference. Instead of explicitly modeling human preferences, we first implicitly encompass them within a set of human labels. Then, an initial set of scoring criteria is drafted by the language model itself, leveraging in-context learning on different few-shot examples. To further calibrate this set of criteria, we select the best performers and re-draft them with self-refinement. Our experiments on multiple text quality evaluation datasets illustrate a significant improvement in correlation with expert evaluation through calibration. Our comprehensive qualitative analysis conveys insightful intuitions and observations on the essence of effective scoring criteria. Yuxuan Liu 0011, Tianchi Yang, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066 |
LREC/COLING | 6 |
| 2024 | Learning to Retrieve In-Context Examples for Large Language ModelsabstractLarge language models (LLMs) have demonstrated their ability to learn in-context, allowing them to perform various tasks based on a few input-output examples.However, the effectiveness of in-context learning is heavily reliant on the quality of the selected examples.In this paper, we propose a novel framework to iteratively train dense retrievers that can identify high-quality in-context examples for LLMs.Our framework initially trains a reward model based on LLM feedback to evaluate the quality of candidate examples, followed by knowledge distillation to train a bi-encoder based dense retriever.Our experiments on a suite of 30 tasks demonstrate that our framework significantly enhances in-context learning performance.Furthermore, we show the generalization ability of our framework to unseen tasks during training.An in-depth analysis reveals that our model improves performance by retrieving examples with similar patterns, and the gains are consistent across LLMs of varying sizes.The code and data are available at https://github.com/microsoft/LMOps/ tree/main/llm_retriever. Liang Wang 0046, Nan Yang 0002, Furu Wei |
EACL (1) | 3 |
| 2024 | Language Models as Inductive ReasonersabstractZonglin Yang, Li Dong, Xinya Du, Hao Cheng, Erik Cambria, Xiaodong Liu, Jianfeng Gao, Furu Wei. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zonglin Yang 0001, Li Dong 0004, Xinya Du, Hao Cheng 0002, Erik Cambria, Xiaodong Liu 0003, Jianfeng Gao 0001, Furu Wei |
EACL (1) | 8 |
| 2024 | TextDiffuser-2: Unleashing the Power of Language Models for Text Rendering
Jingye Chen, Yupan Huang, Tengchao Lv, Lei Cui 0001, Qifeng Chen 0001, Furu Wei |
ECCV (5) | 6 |
| 2024 | Instruction Pre-Training: Language Models are Supervised Multitask LearnersabstractUnsupervised multitask pre-training has been the critical method behind the recent success of language models (LMs).However, supervised multitask learning still holds significant promise, as scaling it in the post-training stage trends towards better generalization.In this paper, we explore supervised multitask pretraining by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response pairs to pre-train LMs.The instruction-response pairs are generated by an efficient instruction synthesizer built on open-source models.In our experiments, we synthesize 200M instruction-response pairs covering 40+ task categories to verify the effectiveness of Instruction Pre-Training.In pre-training from scratch, Instruction Pre-Training not only consistently enhances pre-trained base models but also benefits more from further instruction tuning.In continual pre-training, Instruction Pre-Training enables Llama3-8B to be comparable to or even outperform Llama3-70B.Our model, code, and data are available at https://github.com/microsoft/LMOps.Ins: When is the finale of season 7? Let's think step by step. Daixuan Cheng, Yuxian Gu, Shaohan Huang, Junyu Bi, Minlie Huang, Furu Wei |
EMNLP | 6 |
| 2024 | Chain-of-Dictionary Prompting Elicits Translation in Large Language ModelsabstractLarge language models (LLMs) have shown surprisingly good performance in multilingual neural machine translation (MNMT) even if not being trained explicitly for translation.Yet, they still struggle with translating low-resource languages.As supported by our experiments, a bilingual dictionary between the source and the target language could help.Motivated by the fact that multilingual training effectively improves cross-lingual performance, we show that a chained multilingual dictionary with words expressed in more languages can provide more information to better enhance the LLM translation.To this end, we present a novel framework, COD, Chain-of-Dictionary Prompting, which augments LLMs with prior knowledge with the chains of multilingual dictionaries for a subset of input words to elicit translation abilities for LLMs.Experiments indicate that ChatGPT and InstructGPT still have room for improvement in translating many language pairs.And COD elicits large gains by up to 13x chrF++ points for MNMT (3.08 to 42.63 for English to Serbian written in Cyrillic script) on FLORES-200 full devtest set.We demonstrate the importance of chaining the multilingual dictionaries, as well as the superiority of COD to few-shot in-context learning for low-resource languages.Using COD helps ChatGPT to obviously surpass the SOTA translator NLLB 3.3B. Hongyuan Lu, Haoyang Huang, Wai Lam, Furu Wei |
EMNLP | 6 |
| 2024 | LongEmbed: Extending Embedding Models for Long Context RetrievalabstractEmbedding models play a pivotal role in modern NLP applications such as document retrieval.However, existing embedding models are limited to encoding short documents of typically 512 tokens, restrained from application scenarios requiring long inputs.This paper explores context window extension of existing embedding models, pushing their input length to a maximum of 32,768.We begin by evaluating the performance of existing embedding models using our newly constructed LONGEM-BED benchmark, which includes two synthetic and four real-world tasks, featuring documents of varying lengths and dispersed target information.The benchmarking results highlight huge opportunities for enhancement in current models.Via comprehensive experiments, we demonstrate that training-free context window extension strategies can effectively increase the input length of these models by several folds.Moreover, comparison of models using Absolute Position Encoding (APE) and Rotary Position Encoding (RoPE) reveals the superiority of RoPE-based embedding models in context window extension, offering empirical guidance for future models.Our benchmark, code and trained models will be released to advance the research in long context embedding models. Liang Wang 0046, Nan Yang 0002, Yifan Song 0002, Furu Wei, Sujian Li |
EMNLP | 6 |
| 2024 | In-context Autoencoder for Context Compression in a Large Language ModelabstractWe propose the In-context Autoencoder (ICAE), leveraging the power of a large language model (LLM) to compress a long context into short compact memory slots that can be directly conditioned on by the LLM for various purposes. ICAE is first pretrained using both autoencoding and language modeling objectives on massive text data, enabling it to generate memory slots that accurately and comprehensively represent the original context. Then, it is fine-tuned on instruction data for producing desirable responses to various prompts. Experiments demonstrate that our lightweight ICAE, introducing about 1% additional parameters, effectively achieves $4\times$ context compression based on Llama, offering advantages in both improved latency and GPU memory cost during inference, and showing an interesting insight in memorization as well as potential for scalability. These promising results imply a novel perspective on the connection between working memory in cognitive science and representation learning in LLMs, revealing ICAE's significant implications in addressing the long context problem and suggesting further research in LLM context management. Our data, code and models are available at https://github.com/getao/icae. Tao Ge 0001, Jing Hu 0001, Xun Wang 0012, Furu Wei |
ICLR | 6 |
| 2024 | Adapting Large Language Models via Reading ComprehensionabstractWe explore how continued pre-training on domain-specific corpora influences large language models, revealing that training on the raw corpora endows the model with domain knowledge, but drastically hurts its prompting ability for question answering. Taken inspiration from human learning via reading comprehension--practice after reading improves the ability to answer questions based on the learned knowledge--we propose a simple method for transforming raw corpora into reading comprehension texts. Each raw text is enriched with a series of tasks related to its content. Our method, highly scalable and applicable to any pre-training corpora, consistently enhances performance across various tasks in three different domains: biomedicine, finance, and law. Notably, our 7B language model achieves competitive performance with domain-specific models of much larger scales, such as BloombergGPT-50B. Furthermore, we demonstrate that domain-specific reading comprehension texts can improve the model's performance even on general benchmarks, showing the potential to develop a general model across even more domains. Our model, code, and data are available at https://github.com/microsoft/LMOps. Daixuan Cheng, Shaohan Huang, Furu Wei |
ICLR | 3 |
| 2024 | MiniLLM: Knowledge Distillation of Large Language ModelsabstractKnowledge Distillation (KD) is a promising technique for reducing the high computational demand of large language models (LLMs). However, previous KD methods are primarily applied to white-box classification models or training small models to imitate black-box model APIs like ChatGPT. How to effectively distill the knowledge of white-box LLMs into small models is still under-explored, which becomes more important with the prosperity of open-source LLMs. In this work, we propose a KD approach that distills LLMs into smaller language models. We first replace the forward Kullback-Leibler divergence (KLD) objective in the standard KD approaches with reverse KLD, which is more suitable for KD on generative language models, to prevent the student model from overestimating the low-probability regions of the teacher distribution. Then, we derive an effective optimization approach to learn this objective. The student models are named MiniLLM. Extensive experiments in the instruction-following setting show that MiniLLM generates more precise responses with higher overall quality, lower exposure bias, better calibration, and higher long-text generation performance than the baselines. Our method is scalable for different model families
with 120M to 13B parameters. Our code, data, and model checkpoints can be found in https://github.com/microsoft/LMOps/tree/main/minillm. Yuxian Gu, Li Dong 0004, Furu Wei, Minlie Huang |
ICLR | 3 |
| 2024 | Kosmos-G: Generating Images in Context with Multimodal Large Language ModelsabstractRecent advancements in subject-driven image generation have made significant strides. However, current methods still fall short in diverse application scenarios, as they require test-time tuning and cannot accept interleaved multi-image and text input. These limitations keep them far from the ultimate goal of "image as a foreign language in image generation." This paper presents Kosmos-G, a model that leverages the advanced multimodal perception capabilities of Multimodal Large Language Models (MLLMs) to tackle the aforementioned challenge. Our approach aligns the output space of MLLM with CLIP using the textual modality as an anchor and performs compositional instruction tuning on curated data. Kosmos-G demonstrates an impressive capability of zero-shot subject-driven generation with interleaved multi-image and text input. Notably, the score distillation instruction tuning requires no modifications to the image decoder. This allows for a seamless substitution of CLIP and effortless integration with a myriad of U-Net techniques ranging from fine-grained controls to personalized image decoder variants. We posit Kosmos-G as an initial attempt towards the goal of "image as a foreign language in image generation." Xichen Pan, Li Dong 0004, Shaohan Huang, Zhiliang Peng, Wenhu Chen, Furu Wei |
ICLR | 6 |
| 2024 | Grounding Multimodal Large Language Models to the WorldabstractWe introduce Kosmos-2, a Multimodal Large Language Model (MLLM), enabling new capabilities of perceiving object descriptions (e.g., bounding boxes) and grounding text to the visual world. Specifically, we represent text spans (i.e., referring expressions and noun phrases) as links in Markdown, i.e., [text span](bounding boxes), where object descriptions are sequences of location tokens. To train the model, we construct a large-scale dataset about grounded image-text pairs (GrIT) together with multimodal corpora. In addition to the existing capabilities of MLLMs (e.g., perceiving general modalities, following instructions, and performing in-context learning), Kosmos-2 integrates the grounding capability to downstream applications, while maintaining the conventional capabilities of MLLMs (e.g., perceiving general modalities, following instructions, and performing in-context learning). Kosmos-2 is evaluated on a wide range of tasks, including (i) multimodal grounding, such as referring expression comprehension and phrase grounding, (ii) multimodal referring, such as referring expression generation, (iii) perception-language tasks, and (iv) language understanding and generation. This study sheds a light on the big convergence of language, multimodal perception, and world modeling, which is a key step toward artificial general intelligence. Code can be found in [https://aka.ms/kosmos-2](https://aka.ms/kosmos-2). Zhiliang Peng, Wenhui Wang 0003, Li Dong 0004, Yaru Hao, Shaohan Huang, Shuming Ma, Qixiang Ye, Furu Wei |
ICLR | 8 |
| 2024 | Mixture of LoRA ExpertsabstractLoRA has gained widespread acceptance in the fine-tuning of large pre-trained models to cater to a diverse array of downstream tasks, showcasing notable effectiveness and efficiency, thereby solidifying its position as one of the most prevalent fine-tuning techniques. Due to the modular nature of LoRA's plug-and-play plugins, researchers have delved into the amalgamation of multiple LoRAs to empower models to excel across various downstream tasks. Nonetheless, extant approaches for LoRA fusion grapple with inherent challenges. Direct arithmetic merging may result in the loss of the original pre-trained model's generative capabilities or the distinct identity of LoRAs, thereby yielding suboptimal outcomes. On the other hand, Reference tuning-based fusion exhibits limitations concerning the requisite flexibility for the effective combination of multiple LoRAs. In response to these challenges, this paper introduces the Mixture of LoRA Experts (MoLE) approach, which harnesses hierarchical control and unfettered branch selection. The MoLE approach not only achieves superior LoRA fusion performance in comparison to direct arithmetic merging but also retains the crucial flexibility for combining LoRAs effectively. Extensive experimental evaluations conducted in both the Natural Language Processing (NLP) and Vision \& Language (V\&L) domains substantiate the efficacy of MoLE. Shaohan Huang, Furu Wei |
ICLR | 3 |
| 2024 | PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise TrainingabstractLarge Language Models (LLMs) are trained with a pre-defined context length, restricting their use in scenarios requiring long inputs. Previous efforts for adapting LLMs to a longer length usually requires fine-tuning with this target length (Full-length fine-tuning), suffering intensive training cost. To decouple train length from target length for efficient context window extension, we propose Positional Skip-wisE (PoSE) training that smartly simulates long inputs using a fixed context window. This is achieved by first dividing the original context window into several chunks, then designing distinct skipping bias terms to manipulate the position indices of each chunk. These bias terms and the lengths of each chunk are altered for every training example, allowing the model to adapt to all positions within target length. Experimental results show that PoSE greatly reduces memory and time overhead compared with Full-length fine-tuning, with minimal impact on performance. Leveraging this advantage, we have successfully extended the LLaMA model to 128k tokens using a 2k training context window. Furthermore, we empirically confirm that PoSE is compatible with all RoPE-based LLMs and position interpolation strategies. Notably, our method can potentially support infinite length, limited only by memory usage in inference. With ongoing progress for efficient inference, we believe PoSE can further scale the context window beyond 128k. Nan Yang 0002, Liang Wang 0046, Yifan Song 0002, Furu Wei, Sujian Li |
ICLR | 6 |
| 2024 | MathScale: Scaling Instruction Tuning for Mathematical ReasoningabstractLarge language models (LLMs) have demonstrated remarkable capabilities in problem-solving. However, their proficiency in solving mathematical problems remains inadequate. We propose MathScale, a simple and scalable method to create high-quality mathematical reasoning data using frontier LLMs (e.g., GPT-3.5). Inspired by the cognitive mechanism in human mathematical learning, it first extracts topics and knowledge points from seed math questions and then build a concept graph, which is subsequently used to generate new math questions. MathScale exhibits effective scalability along the size axis of the math dataset that we generate. As a result, we create a mathematical reasoning dataset (MathScaleQA) containing two million math question-answer pairs. To evaluate mathematical reasoning abilities of LLMs comprehensively, we construct MWPBench, a benchmark of Math Word Problems, which is a collection of 9 datasets (including GSM8K and MATH) covering K-12, college, and competition level math problems. We apply MathScaleQA to fine-tune open-source LLMs (e.g., LLaMA-2 and Mistral), resulting in significantly improved capabilities in mathematical reasoning. Evaluated on MWPBench, MathScale-7B achieves state-of-the-art performance across all datasets, surpassing its best peers of equivalent size by 42.8% in micro average accuracy and 43.6% in macro average accuracy, respectively. Zhengyang Tang, Xingxing Zhang 0002, Benyou Wang, Furu Wei |
ICML | 4 |
| 2024 | KOSMOS-E : Learning to Follow Instruction for Robotic GraspingabstractTuning on instruction-following data has been shown to enhance the capabilities and controllability of language models, but the idea is less explored in the robotic field. In this work, we introduce KOSMOS-E, a Multimodal Large Language Model (MLLM) that leverages instruction-following robotic grasping data to enhance capabilities for precise and intricate robotic grasping maneuvers. To achieve this, we craft a large-scale instruction-following robotic grasping dataset, termed INSTRUCT-GRASP, primarily comprising two aspects: (i) grasp a single object following varying levels of granularity descriptions, e.g., different angles and aspects, and (ii) grasp a specific object within a multi-object environment following specific attributes, e.g., color and shape. Extensive experiments show the effectiveness of KOSMOS-E on robotic grasping tasks across a variety of environments. Shaohan Huang, Li Dong 0004, Wenhui Wang 0003, Shuming Ma, Furu Wei |
IROS | 7 |
| 2024 | Not All Metrics Are Guilty: Improving NLG Evaluation by Diversifying ReferencesabstractTianyi Tang, Hongyuan Lu, Yuchen Jiang, Haoyang Huang, Dongdong Zhang, Xin Zhao, Tom Kocmi, Furu Wei. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Hongyuan Lu, Haoyang Huang, Dongdong Zhang 0001, Wayne Xin Zhao, Tom Kocmi, Furu Wei |
NAACL-HLT | 8 |
| 2024 | Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-CollaborationabstractZhenhailong Wang, Shaoguang Mao, Wenshan Wu, Tao Ge, Furu Wei, Heng Ji. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhenhailong Wang, Shaoguang Mao, Wenshan Wu, Tao Ge 0001, Furu Wei, Heng Ji 0001 |
NAACL-HLT | 5 |
| 2024 | xRAG: Extreme Context Compression for Retrieval-augmented Generation with One TokenabstractThis paper introduces xRAG, an innovative context compression method tailored for retrieval-augmented generation. xRAG reinterprets document embeddings in dense retrieval--traditionally used solely for retrieval--as features from the retrieval modality. By employing a modality fusion methodology, xRAG seamlessly integrates these embeddings into the language model representation space, effectively eliminating the need for their textual counterparts and achieving an extreme compression rate.
In xRAG, the only trainable component is the modality bridge, while both the retriever and the language model remain frozen. This design choice allows for the reuse of offline-constructed document embeddings and preserves the plug-and-play nature of retrieval augmentation.
Experimental results demonstrate that xRAG achieves an average improvement of over 10% across six knowledge-intensive tasks, adaptable to various language model backbones, ranging from a dense 7B model to an 8x7B Mixture of Experts configuration. xRAG not only significantly outperforms previous context compression methods but also matches the performance of uncompressed models on several datasets, while reducing overall FLOPs by a factor of 3.53. Our work pioneers new directions in retrieval-augmented generation from the perspective of multimodality fusion, and we hope it lays the foundation for future efficient and scalable retrieval-augmented systems. Xin Cheng 0002, Xun Wang 0012, Xingxing Zhang 0002, Tao Ge 0001, Furu Wei, Huishuai Zhang, Dongyan Zhao 0001 |
NeurIPS | 6 |
| 2024 | You Only Cache Once: Decoder-Decoder Architectures for Language ModelsabstractWe introduce a decoder-decoder architecture, YOCO, for large language models, which only caches key-value pairs once. It consists of two components, i.e., a cross-decoder stacked upon a self-decoder. The self-decoder efficiently encodes global key-value (KV) caches that are reused by the cross-decoder via cross-attention. The overall model behaves like a decoder-only Transformer, although YOCO only caches once. The design substantially reduces GPU memory demands, yet retains global attention capability. Additionally, the computation flow enables prefilling to early exit without changing the final output, thereby significantly speeding up the prefill stage. Experimental results demonstrate that YOCO achieves favorable performance compared to Transformer in various settings of scaling up model size and number of training tokens. We also extend YOCO to 1M context length with near-perfect needle retrieval accuracy. The profiling results show that YOCO improves inference memory, prefill latency, and throughput by orders of magnitude across context lengths and model sizes. Yutao Sun, Li Dong 0004, Shaohan Huang, Wenhui Wang 0003, Shuming Ma, Quanlu Zhang, Jianyong Wang 0001, Furu Wei |
NeurIPS | 9 |
| 2024 | Multi-Head Mixture-of-ExpertsabstractSparse Mixtures of Experts (SMoE) scales model capacity without significant increases in computational costs. However, it exhibits the low expert activation issue, i.e., only a small subset of experts are activated for optimization, leading to suboptimal performance and limiting its effectiveness in learning a larger number of experts in complex tasks. In this paper, we propose Multi-Head Mixture-of-Experts (MH-MoE). MH-MoE split each input token into multiple sub-tokens, then these sub-tokens are assigned to and processed by a diverse set of experts in parallel, and seamlessly reintegrated into the original token form. The above operations enables MH-MoE to significantly enhance expert activation while collectively attend to information from various representation spaces within different experts to deepen context understanding. Besides, it's worth noting that our MH-MoE is straightforward to implement and decouples from other SMoE frameworks, making it easy to integrate with these frameworks for enhanced performance. Extensive experimental results across different parameter scales (300M to 7B) and three pre-training tasks—English-focused language modeling, multi-lingual language modeling and masked multi-modality modeling—along with multiple downstream validation tasks, demonstrate the effectiveness of MH-MoE. Shaohan Huang, Wenhui Wang 0003, Shuming Ma, Li Dong 0004, Furu Wei |
NeurIPS | 6 |
| 2024 | Multimodal Large Language Models Make Text-to-Image Generative Models Align BetterabstractRecent studies have demonstrated the exceptional potentials of leveraging human preference datasets to refine text-to-image generative models, enhancing the alignment between generated images and textual prompts. Despite these advances, current human preference datasets are either prohibitively expensive to construct or suffer from a lack of diversity in preference dimensions, resulting in limited applicability for instruction tuning in open-source text-to-image generative models and hinder further exploration. To address these challenges and promote the alignment of generative models through instruction tuning, we leverage multimodal large language models to create VisionPrefer, a high-quality and fine-grained preference dataset that captures multiple preference aspects. We aggregate feedback from AI annotators across four aspects: prompt-following, aesthetic, fidelity, and harmlessness to construct VisionPrefer. To validate the effectiveness of VisionPrefer, we train a reward model VP-Score over VisionPrefer to guide the training of text-to-image generative models and the preference prediction accuracy of VP-Score is comparable to human annotators. Furthermore, we use two reinforcement learning methods to supervised fine-tune generative models to evaluate the performance of VisionPrefer, and extensive experimental results demonstrate that VisionPrefer significantly improves text-image alignment in compositional image generation across diverse aspects, e.g., aesthetic, and generalizes better than previous human-preference metrics across various image distributions. Moreover, VisionPrefer indicates that the integration of AI-generated synthetic data as a supervisory signal is a promising avenue for achieving improved alignment with human preferences in vision generative models. Shaohan Huang, Guolong Wang 0001, Furu Wei |
NeurIPS | 5 |
| 2024 | Boosting Text-to-Video Generative Model with MLLMs FeedbackabstractRecent advancements in text-to-video generative models, such as Sora, have showcased impressive capabilities. These models have attracted significant interest for their potential applications. However, they often rely on extensive datasets of variable quality, which can result in generated videos that lack aesthetic appeal and do not accurately reflect the input text prompts. A promising approach to mitigate these issues is to leverage Reinforcement Learning from Human Feedback (RLHF), which aims to align the outputs of text-to-video generative with human preferences. However, the considerable costs associated with manual annotation have led to a scarcity of comprehensive preference datasets. In response to this challenge, our study begins by investigating the efficacy of Multimodal Large Language Models (MLLMs) generated annotations in capturing video preferences, discovering a high degree of concordance with human judgments. Building upon this finding, we utilize MLLMs to perform fine-grained video preference annotations across two dimensions, resulting in the creation of VideoPrefer, which includes 135,000 preference annotations. Utilizing this dataset, we introduce VideoRM, the first general-purpose reward model tailored for video preference in the text-to-video domain. Our comprehensive experiments confirm the effectiveness of both VideoPrefer and VideoRM, representing a significant step forward in the field. Shaohan Huang, Guolong Wang 0001, Furu Wei |
NeurIPS | 5 |
| 2024 | Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language ModelsabstractLarge language models (LLMs) have exhibited impressive performance in language comprehension and various reasoning tasks. However, their abilities in spatial reasoning, a crucial aspect of human cognition, remain relatively unexplored. Human possess a remarkable ability to create mental images of unseen objects and actions through a process known as the Mind's Eye, enabling the imagination of the unseen world. Inspired by this cognitive capacity, we propose Visualization-of-Thought (VoT) prompting. VoT aims to elicit spatial reasoning of LLMs by visualizing their reasoning traces, thereby guiding subsequent reasoning steps. We employed VoT for multi-hop spatial reasoning tasks, including natural language navigation, visual navigation, and visual tiling in 2D grid worlds. Experimental results demonstrated that VoT significantly enhances the spatial reasoning abilities of LLMs. Notably, VoT outperformed existing multimodal large language models (MLLMs) in these tasks. While VoT works surprisingly well on LLMs, the ability to generate mental images to facilitate spatial reasoning resembles the mind's eye process, suggesting its potential viability in MLLMs. Please find the dataset and codes in our [project page](https://microsoft.github.io/visualization-of-thought). Wenshan Wu, Shaoguang Mao, Yan Xia 0005, Li Dong 0004, Lei Cui 0001, Furu Wei |
NeurIPS | 7 |
| 2024 | Fine-Tuning LLaMA for Multi-Stage Text RetrievalabstractWhile large language models (LLMs) have shown impressive NLP capabilities, existing IR applications mainly focus on prompting LLMs to generate query expansions or generating permutations for listwise reranking. In this study, we leverage LLMs directly to serve as components in the widely used multi-stage text ranking pipeline. Specifically, we fine-tune the open-source LLaMA-2 model as a dense retriever (repLLaMA) and a pointwise reranker (rankLLaMA). This is performed for both passage and document retrieval tasks using the MS MARCO training data. Our study shows that finetuned LLM retrieval models outperform smaller models. They are more effective and exhibit greater generalizability, requiring only a straightforward training strategy. Moreover, our pipeline allows for the fine-tuning of LLMs at each stage of a multi-stage retrieval pipeline. This demonstrates the strong potential for optimizing LLMs to enhance a variety of retrieval tasks. Furthermore, as LLMs are naturally pre-trained with longer contexts, they can directly represent longer documents. This eliminates the need for heuristic segmenting and pooling strategies to rank long documents. On the MS MARCO and BEIR datasets, our repLLaMA-rankLLaMA pipeline demonstrates a high level of effectiveness. Xueguang Ma, Liang Wang 0046, Nan Yang 0002, Furu Wei, Jimmy Lin |
SIGIR | 4 |
| 2024 | DeepNet: Scaling Transformers to 1,000 LayersabstractIn this paper, we propose a simple yet effective method to stabilize extremely deep Transformers. Specifically, we introduce a new normalization function (DeepNorm) to modify the residual connection in Transformer, accompanying with theoretically derived initialization. In-depth theoretical analysis shows that model updates can be bounded in a stable way. The proposed method combines the best of two worlds, i.e., good performance of Post-LN and stable training of Pre-LN, makingDeepNorma preferred alternative. We successfully scale Transformers up to 1,000 layers (i.e., 2,500 attention and feed-forward network sublayers) without difficulty, which is one order of magnitude deeper than previous deep Transformers. Extensive experiments demonstrate thatDeepNethas superior performance across various benchmarks, including machine translation, language modeling (i.e., BERT, GPT) and vision pre-training (i.e., BEiT). Remarkably, on a multilingual benchmark with 7,482 translation directions, our 200-layer model with 3.2B parameters significantly outperforms the 48-layer state-of-the-art model with 12B parameters by 5 BLEU points, which indicates a promising scaling direction. Our code is available athttps://aka.ms/torchscale. Hongyu Wang 0009, Shuming Ma, Li Dong 0004, Shaohan Huang, Dongdong Zhang 0001, Furu Wei |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | VioLA: Conditional Language Models for Speech Recognition, Synthesis, and TranslationabstractRecent research shows a big convergence in model architecture, training objectives, and inference methods across various tasks for different modalities. In this paper, we proposeVioLA, a single auto-regressive Transformer decoder-only network that unifies various cross-modal tasks involving speech and text, such as speech-to-text, text-to-text, text-to-speech, and speech-to-speech tasks, as a conditional language model task via multi-task learning framework. To accomplish this, we first convert the speech utterances to discrete tokens (similar to the textual data) using an offline neural codec encoder. In such a way, all these tasks are converted to token-based sequence prediction problems, which can be naturally handled with one conditional language model. We further integrate task IDs (TID), language IDs (LID), and LSTM-based acoustic embedding into the proposed model to enhance the modeling capability of handling different languages and tasks. Experimental results demonstrate that the proposedVioLAmodel can support both single-modal and cross-modal tasks well, and the decoder-only model achieves a comparable and even better performance than the strong baselines. Tianrui Wang, Yu Wu 0012, Shujie Liu 0001, Yashesh Gaur, Zhuo Chen 0006, Jinyu Li 0001, Furu Wei |
IEEE ACM Trans. Audio Speech Lang. Process. | 9 |
| 2024 | SpeechLM: Enhanced Speech Pre-Training With Unpaired Textual DataabstractHow to boost speech pre-training with textual data is an unsolved problem due to the fact that speech and text are very different modalities with distinct characteristics. In this paper, we propose a cross-modalSpeechandLanguageModel (SpeechLM) to explicitly align speech and text pre-training with a pre-defined unified discrete representation. Specifically, we introduce two alternative discrete tokenizers to bridge the speech and text modalities, including phoneme-unit and hidden-unit tokenizers, which can be trained using unpaired speech or a small amount of paired speech-text data. Based on the trained tokenizers, we convert the unlabeled speech and text data into tokens of phoneme units or hidden units. The pre-training objective is designed to unify the speech and the text into the same discrete semantic space with a unified Transformer network. We evaluate SpeechLM on various spoken language processing tasks including speech recognition, speech translation, and universal representation evaluation framework SUPERB, demonstrating significant improvements on content-related tasks. Code and models are available athttps://aka.ms/SpeechLM. Sanyuan Chen, Yu Wu 0012, Shuo Ren 0002, Shujie Liu 0001, Zhuoyuan Yao, Xun Gong 0005, Li-Rong Dai 0001, Jinyu Li 0001, Furu Wei |
IEEE ACM Trans. Audio Speech Lang. Process. | 11 |
| 2024 | Generic-to-Specific Distillation of Masked AutoencodersabstractTo transfer the representation capacity of large pre-trained models to lightweight models, knowledge distillation has been widely explored. However, conventional single-stage distillation methods are prone to getting stuck in the transfer of task-specific knowledge, making it difficult to retain task-agnostic knowledge which is crucial for model generalization. In this study, we propose generic-to-specific distillation (G2SD), to boost lightweight models under the assistance of large models pre-trained by masked image modeling. In generic distillation, the decoder of a small model is encouraged to align feature predictions with that of a large model, so that task-agnostic knowledge can be transferred. In specific distillation, predictions of the small model are encouraged to be consistent with those of the large model, to guarantee task performance. G2SD is also applicable for heterogeneous settings(i.e., distilling from ViT to CNN). With G2SD, the ViT-Small model respectively achieves 98.9%, 98.4%, 99.3% and 98.9% accuracies when compared with its teachers (ViT-Base) for image classification, object detection, semantic segmentation and video recognition tasks. The lightweight ResNet models are improved to a new height on image classification task. The code is available at github.com/pengzhiliang/G2SD. Zhiliang Peng, Li Dong 0004, Furu Wei, Qixiang Ye, Jianbin Jiao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | VatLM: Visual-Audio-Text Pre-Training With Unified Masked Prediction for Speech Representation LearningabstractAlthough speech is a simple and effective way for humans to communicate with the outside world, a more realistic speech interaction contains multimodal information, e.g., vision, text. How to design a unified framework to integrate different modal information and leverage different resources (e.g., visual-audio pairs, audio-text pairs, unlabeled speech, and unlabeled text) to facilitate speech representation learning was not well explored. In this paper, we propose a unified cross-modal representation learning frameworkVatLM(Visual-Audio-Text Language Model). The proposedVatLMemploys a unified backbone network to model the modality-independent information and utilizes three simple modality-dependent modules to preprocess visual, speech, and text inputs. In order to integrate these three modalities into one shared semantic space,VatLMis optimized with a masked prediction task of unified tokens, given by our proposed unified tokenizer. We evaluate the pre-trainedVatLMon audio-visual related downstream tasks, including audio-visual speech recognition (AVSR), and visual speech recognition (VSR) tasks. Results show that the proposedVatLMoutperforms previous state-of-the-art models, such as the audio-visual pre-trained AV-HuBERT model, and analysis also demonstrates thatVatLMis capable of aligning different modalities into the same space. To facilitate future research, we release the code and pre-trained models athttps://aka.ms/vatlm. Qiushi Zhu, Shujie Liu 0001, Binxing Jiao, Jie Zhang 0042, Li-Rong Dai 0001, Daxin Jiang, Jinyu Li 0001, Furu Wei |
IEEE Trans. Multim. | 10 |
| 2023 | TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsabstractText recognition is a long-standing research problem for document digitalization. Existing approaches are usually built based on CNN for image understanding and RNN for char-level text generation. In addition, another language model is usually needed to improve the overall accuracy as a post-processing step. In this paper, we propose an end-to-end text recognition approach with pre-trained image Transformer and text Transformer models, namely TrOCR, which leverages the Transformer architecture for both image understanding and wordpiece-level text generation. The TrOCR model is simple but effective, and can be pre-trained with large-scale synthetic data and fine-tuned with human-labeled datasets. Experiments show that the TrOCR model outperforms the current state-of-the-art models on the printed, handwritten and scene text recognition tasks. The TrOCR models and code are publicly available at https://aka.ms/trocr. Minghao Li 0004, Tengchao Lv, Jingye Chen, Lei Cui 0001, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Zhoujun Li 0001, Furu Wei |
AAAI | 9 |
| 2023 | MoEC: Mixture of Expert ClustersabstractSparsely Mixture of Experts (MoE) has received great interest due to its promising scaling capability with affordable computational overhead. MoE models convert dense layers into sparse experts, and utilize a gated routing network to make experts conditionally activated. However, as the number of experts grows, MoE with outrageous parameters suffers from overfitting and sparse data allocation. Such problems are especially severe on tasks with limited data, thus hindering the progress towards improving performance by scaling up. We verify that there exists a performance upper bound of scaling up sparse MoE. In this work, we propose Mixture of Expert Clusters — a general approach to enable expert layers to learn more diverse and appropriate knowledge by imposing variance-based constraints on the routing stage. Given this, we could further propose a cluster-level expert dropout strategy specifically designed for the expert cluster structure. Our experiments reveal that MoEC could improve performance on machine translation and natural language understanding tasks. MoEC plays a positive role in mitigating overfitting and sparse data allocation problems, thus fully releasing the potential of large-scale sparse models. Shaohan Huang, Furu Wei |
AAAI | 4 |
| 2023 | Multiview Identifiers Enhanced Generative RetrievalabstractInstead of simply matching a query to preexisting passages, generative retrieval generates identifier strings of passages as the retrieval target.At a cost, the identifier must be distinctive enough to represent a passage.Current approaches use either a numeric ID or a text piece (such as a title or substrings) as the identifier.However, these identifiers cannot cover a passage's content well.As such, we are motivated to propose a new type of identifier, synthetic identifiers, that are generated based on the content of a passage and could integrate contextualized information that text pieces lack.Furthermore, we simultaneously consider multiview identifiers, including synthetic identifiers, titles, and substrings.These views of identifiers complement each other and facilitate the holistic ranking of passages from multiple perspectives.We conduct a series of experiments on three public datasets, and the results indicate that our proposed approach performs the best in generative retrieval, demonstrating its effectiveness and robustness.The code is released at https://github.com/liyongqi67/MINDER. Yongqi Li 0001, Nan Yang 0002, Liang Wang 0046, Furu Wei, Wenjie Li 0002 |
ACL (1) | 4 |
| 2023 | Pre-Training to Learn in ContextabstractIn-context learning, where pre-trained language models learn to perform tasks from task examples and instructions in their contexts, has attracted much attention in the NLP community.However, the ability of in-context learning is not fully exploited because language models are not explicitly trained to learn in context.To this end, we propose PICL (Pretraining for In-Context Learning), a framework to enhance the language models' in-context learning ability by pre-training the model on a large collection of "intrinsic tasks" in the general plain-text corpus using the simple language modeling objective.PICL encourages the model to infer and perform tasks by conditioning on the contexts while maintaining task generalization of pre-trained models.We evaluate the in-context learning performance of the model trained with PICL on seven widelyused text classification datasets and the SUPER-NATURALINSTRCTIONS benchmark, which contains 100+ NLP tasks formulated to text generation.Our experiments show that PICL is more effective and task-generalizable than a range of baselines, outperforming larger language models with nearly 4x parameters.The code is publicly available at https://github. com/thu-coai/PICL. Yuxian Gu, Li Dong 0004, Furu Wei, Minlie Huang |
ACL (1) | 3 |
| 2023 | Dual-Alignment Pre-training for Cross-lingual Sentence EmbeddingabstractZiheng Li, Shaohan Huang, Zihan Zhang, Zhi-Hong Deng, Qiang Lou, Haizhen Huang, Jian Jiao, Furu Wei, Weiwei Deng, Qi Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ziheng Li 0003, Shaohan Huang, Zhi-Hong Deng 0001, Qiang Lou, Haizhen Huang, Jian Jiao 0007, Furu Wei, Qi Zhang 0066 |
ACL (1) | 8 |
| 2023 | Beyond English-Centric Bitexts for Better Multilingual Language Representation LearningabstractBarun Patra, Saksham Singhal, Shaohan Huang, Zewen Chi, Li Dong, Furu Wei, Vishrav Chaudhary, Xia Song. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Barun Patra, Saksham Singhal, Shaohan Huang, Zewen Chi, Li Dong 0004, Furu Wei, Vishrav Chaudhary |
ACL (1) | 6 |
| 2023 | A Length-Extrapolatable TransformerabstractYutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, Furu Wei. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yutao Sun, Li Dong 0004, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Furu Wei |
ACL (1) | 9 |
| 2023 | SimLM: Pre-training with Representation Bottleneck for Dense Passage RetrievalabstractLiang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Liang Wang 0046, Nan Yang 0002, Xiaolong Huang 0002, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei |
ACL (1) | 8 |
| 2023 | GanLM: Encoder-Decoder Pre-training with an Auxiliary DiscriminatorabstractJian Yang, Shuming Ma, Li Dong, Shaohan Huang, Haoyang Huang, Yuwei Yin, Dongdong Zhang, Liqun Yang, Furu Wei, Zhoujun Li. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jian Yang 0030, Shuming Ma, Li Dong 0004, Shaohan Huang, Haoyang Huang, Yuwei Yin, Dongdong Zhang 0001, Liqun Yang, Furu Wei, Zhoujun Li 0001 |
ACL (1) | 9 |
| 2023 | Generic-to-Specific Distillation of Masked AutoencodersabstractLarge vision Transformers (ViTs) driven by self-supervised pre-training mechanisms achieved unprecedented progress. Lightweight ViT models limited by the model capacity, however, benefit little from those pre-training mechanisms. Knowledge distillation defines a paradigm to transfer representations from large (teacher) models to small (student) ones. However, the conventional single-stage distillation easily gets stuck on task-specific transfer, failing to retain the task-agnostic knowledge crucial for model generalization. In this study, we propose generic-to-specific distillation (G2SD), to tap the potential of small ViT models under the supervision of large models pretrained by masked autoencoders. In generic distillation, decoder of the small model is encouraged to align feature predictions with hidden representations of the large model, so that task-agnostic knowledge can be transferred. In specific distillation, predictions of the small model are constrained to be consistent with those of the large model, to transfer task-specific features which guarantee task performance. With G2SD, the vanilla ViT-Small model respectively achieves 98.7%, 98.1% and 99.3% the performance of its teacher (ViT-Base) for image classification, object detection, and semantic segmentation, setting a solid baseline for two-stage vision distillation. Code will be available at https://github.com/pengzhiliang/G2SD Zhiliang Peng, Li Dong 0004, Furu Wei, Jianbin Jiao, Qixiang Ye |
CVPR | 4 |
| 2023 | Image as a Foreign Language: BEIT Pretraining for Vision and Vision-Language TasksabstractA big convergence of language, vision, and multimodal pretraining is emerging. In this work, we introduce a general-purpose multimodal foundation model BEIT-3, which achieves excellent transfer performance on both vision and vision-language tasks. Specifically, we advance the big convergence from three aspects: backbone architecture, pretraining task, and model scaling up. We use Multiway Transformers for general-purpose modeling, where the modular architecture enables both deep fusion and modality-specific encoding. Based on the shared backbone, we perform masked “language” modeling on images (Imglish), texts (English), and image-text pairs (“parallel sentences”) in a unified manner. Experimental results show that BEIT-3 obtains remarkable performance on object detection (COCO), semantic segmentation (ADE20K), image classification (ImageNet), visual reasoning (NLVR2), visual question answering (VQAv2), image captioning (COCO), and cross-modal retrieval (Flickr30K, COCO). Wenhui Wang 0003, Hangbo Bao, Li Dong 0004, Johan Bjorck, Zhiliang Peng, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, Furu Wei |
CVPR | 11 |
| 2023 | Non-Contrastive Learning Meets Language-Image Pre-TrainingabstractContrastive language-image pre-training (CLIP) serves as a de-facto standard to align images and texts. Nonetheless, the loose correlation between images and texts of webcrawled data renders the contrastive objective data inefficient and craving for a large training batch size. In this work, we explore the validity of non-contrastive language-image pre-training (nCLIP), and study whether nice properties exhibited in visual self-supervised models can emerge. We empirically observe that the non-contrastive objective benefits representation learning while sufficiently underperforming under zero-shot recognition. Based on the above study, we further introduce xCLIP, a multi-tasking framework combining CLIP and nCLIP, and show that nCLIP aids CLIP in enhancing feature semantics. The synergy between two objectives lets xCLIP enjoy the best of both worlds: superior performance in both zero-shot transfer and representation learning. Systematic evaluation is conducted spanning a wide variety of downstream tasks including zero-shot classification, out-of-domain classification, retrieval, visual representation learning, and textual representation learning, showcasing a consistent performance gain and validating the effectiveness of xCLIP. The code and pre-trained models will be publicly available at https://github.com/shallowtoil/xclip. Jinghao Zhou, Li Dong 0004, Zhe Gan, Furu Wei |
CVPR | 5 |
| 2023 | HanoiT: Enhancing Context-aware Translation via Selective Context
Jian Yang 0030, Yuwei Yin, Shuming Ma, Liqun Yang, Hongcheng Guo, Haoyang Huang, Dongdong Zhang 0001, Yutao Zeng, Zhoujun Li 0001, Furu Wei |
DASFAA (3) | 10 |
| 2023 | UPRISE: Universal Prompt Retrieval for Improving Zero-Shot EvaluationabstractDaixuan Cheng, Shaohan Huang, Junyu Bi, Yuefeng Zhan, Jianfeng Liu, Yujing Wang, Hao Sun, Furu Wei, Weiwei Deng, Qi Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Daixuan Cheng, Shaohan Huang, Junyu Bi, Yuefeng Zhan, Yujing Wang 0002, Hao Sun 0015, Furu Wei, Qi Zhang 0066 |
EMNLP | 8 |
| 2023 | Syllogistic Reasoning for Legal Judgment AnalysisabstractLegal judgment assistants are developing fast due to impressive progress of large language models (LLMs).However, people can hardly trust the results generated by a model without reliable analysis of legal judgement.For legal practitioners, it is common practice to utilize syllogistic reasoning to select and evaluate the arguments of the parties as part of the legal decision-making process.But the development of syllogistic reasoning for legal judgment analysis is hindered by the lack of resources: (1) there is no large-scale syllogistic reasoning dataset for legal judgment analysis, and (2) there is no set of established benchmarks for legal judgment analysis.In this paper, we construct and manually correct a syllogistic reasoning dataset for legal judgment analysis.The dataset contains 11,239 criminal cases which cover 4 criminal elements, 80 charges and 124 articles.We also select a set of large language models as benchmarks, and conduct a in-depth analysis of the capacity of their legal judgment analysis. Wentao Deng, Jiahuan Pei, Keyi Kong, Furu Wei, Zhaochun Ren, Zhumin Chen, Pengjie Ren |
EMNLP | 5 |
| 2023 | Democratizing Reasoning Ability: Tailored Learning from Large Language ModelabstractZhaoyang Wang, Shaohan Huang, Yuxuan Liu, Jiahai Wang, Minghui Song, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Shaohan Huang, Yuxuan Liu 0011, Jiahai Wang, Minghui Song, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066 |
EMNLP | 8 |
| 2023 | Query2doc: Query Expansion with Large Language ModelsabstractThis paper introduces a simple yet effective query expansion approach, denoted as query2doc, to improve both sparse and dense retrieval systems.The proposed method first generates pseudo-documents by few-shot prompting large language models (LLMs), and then expands the query with generated pseudodocuments.LLMs are trained on web-scale text corpora and are adept at knowledge memorization.The pseudo-documents from LLMs often contain highly relevant information that can aid in query disambiguation and guide the retrievers.Experimental results demonstrate that query2doc boosts the performance of BM25 by 3% to 15% on ad-hoc IR datasets, such as MS-MARCO and TREC DL, without any model fine-tuning.Furthermore, our method also benefits state-of-the-art dense retrievers in terms of both in-domain and out-ofdomain results. Liang Wang 0046, Nan Yang 0002, Furu Wei |
EMNLP | 3 |
| 2023 | Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech TranslationabstractDirect speech-to-speech translation (S2ST) is an attractive research topic with many advantages compared to cascaded S2ST. However, direct S2ST suffers from the data scarcity problem because the corpora from the speech of the source language to the speech of the target language are very rare. To address this issue, we propose in this paper a Speech2S model, which is jointly pre-trained with unpaired speech and bilingual text data for direct speech-to-speech translation tasks. By effectively leveraging the paired text data, Speech2S is capable of modeling the cross-lingual speech conversion from source to target language. We verify the performance of the proposed Speech2S on Europarl-ST and VoxPopuli datasets. Experimental results demonstrate that Speech2S gets an improvement of about 5 BLEU scores compared to encoder-only pre-training models, and achieves a competitive or even better performance than existing state-of-the-art models1. Shujie Liu 0001, Lei He 0005, Jinyu Li 0001, Furu Wei |
ICASSP | 8 |
| 2023 | Corrupted Image Modeling for Self-Supervised Visual Pre-Training
Li Dong 0004, Hangbo Bao, Xinggang Wang, Furu Wei |
ICLR | 5 |
| 2023 | Prototypical Calibration for Few-shot Learning of Language Models
Zhixiong Han, Yaru Hao, Li Dong 0004, Yutao Sun, Furu Wei |
ICLR | 5 |
| 2023 | Visually-Augmented Language Modeling
Weizhi Wang, Li Dong 0004, Hao Cheng 0002, Haoyu Song 0002, Xiaodong Liu 0003, Xifeng Yan, Jianfeng Gao 0001, Furu Wei |
ICLR | 8 |
| 2023 | Are More Layers Beneficial to Graph Transformers?
Haiteng Zhao, Shuming Ma, Dongdong Zhang 0001, Zhi-Hong Deng 0001, Furu Wei |
ICLR | 5 |
| 2023 | BEATs: Audio Pre-Training with Acoustic TokenizersabstractWe introduce a self-supervised learning (SSL) framework BEATs for general audio representation pre-training, where we optimize an acoustic tokenizer and an audio SSL model by iterations. Unlike the previous audio SSL models that employ reconstruction loss for pre-training, our audio SSL model is trained with the discrete label prediction task, where the labels are generated by a semantic-rich acoustic tokenizer. We propose an iterative pipeline to jointly optimize the tokenizer and the pre-trained model, aiming to abstract high-level semantics and discard the redundant details for audio. The experimental results demonstrate our acoustic tokenizers can generate discrete labels with rich audio semantics and our audio SSL models achieve state-of-the-art (SOTA) results across various audio classification benchmarks, even outperforming previous models that use more training data and model parameters significantly. Specifically, we set a new SOTA mAP 50.6% on AudioSet-2M without using any external data, and 98.1% accuracy on ESC-50. The code and pre-trained models are available at https://aka.ms/beats. Sanyuan Chen, Yu Wu 0012, Chengyi Wang 0002, Shujie Liu 0001, Daniel Tompkins, Zhuo Chen 0006, Wanxiang Che, Xiangzhan Yu, Furu Wei |
ICML | 9 |
| 2023 | Magneto: A Foundation TransformerabstractA big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name ”Transformers”, the above areas use different implementations for better performance, e.g., Post-LayerNorm for BERT, and Pre-LayerNorm for GPT and vision Transformers. We call for the development of Foundation Transformer for true general-purpose modeling, which serves as a go-to architecture for various tasks and modalities with guaranteed training stability. In this work, we introduce a Transformer variant, named Magneto, to fulfill the goal. Specifically, we propose Sub-LayerNorm for good expressivity, and the initialization strategy theoretically derived from DeepNet for stable scaling up. Extensive experiments demonstrate its superior performance and better stability than the de facto Transformer variants designed for various applications, including language modeling (i.e., BERT, and GPT), machine translation, vision pretraining (i.e., BEiT), speech recognition, and multimodal pretraining (i.e., BEiT-3). Hongyu Wang 0009, Shuming Ma, Shaohan Huang, Li Dong 0004, Wenhui Wang 0003, Zhiliang Peng, Payal Bajaj, Saksham Singhal, Alon Benhaim, Barun Patra, Zhun Liu, Vishrav Chaudhary, Furu Wei |
ICML | 15 |
| 2023 | Extensible Prompts for Language Models on Zero-shot Language Style CustomizationabstractWe propose eXtensible Prompt (X-Prompt) for prompting a large language model (LLM) beyond natural language (NL). X-Prompt instructs an LLM with not only NL but also an extensible vocabulary of imaginary words. Registering new imaginary words allows us to instruct the LLM to comprehend concepts that are difficult to describe with NL words, thereby making a prompt more descriptive. Also, these imaginary words are designed to be out-of-distribution (OOD) robust so that they can be (re)used like NL words in various prompts, distinguishing X-Prompt from soft prompt that is for fitting in-distribution data. We propose context-augmented learning (CAL) to learn imaginary words for general usability, enabling them to work properly in OOD (unseen) prompts. We experiment X-Prompt for zero-shot language style customization as a case study. The promising results of X-Prompt demonstrate its potential to facilitate advanced interaction beyond the natural language interface, bridging the communication gap between humans and LLMs. Tao Ge 0001, Jing Hu 0001, Li Dong 0004, Shaoguang Mao, Yan Xia 0005, Xun Wang 0012, Furu Wei |
NeurIPS | 8 |
| 2023 | TextDiffuser: Diffusion Models as Text PaintersabstractDiffusion models have gained increasing attention for their impressive generation abilities but currently struggle with rendering accurate and coherent text. To address this issue, we introduce TextDiffuser, focusing on generating images with visually appealing text that is coherent with backgrounds. TextDiffuser consists of two stages: first, a Transformer model generates the layout of keywords extracted from text prompts, and then diffusion models generate images conditioned on the text prompt and the generated layout. Additionally, we contribute the first large-scale text images dataset with OCR annotations, MARIO-10M, containing 10 million image-text pairs with text recognition, detection, and character-level segmentation annotations. We further collect the MARIO-Eval benchmark to serve as a comprehensive tool for evaluating text rendering quality. Through experiments and user studies, we demonstrate that TextDiffuser is flexible and controllable to create high-quality text images using text prompts alone or together with text template images, and conduct text inpainting to reconstruct incomplete images with text. We will make the code, model and dataset publicly available. Jingye Chen, Yupan Huang, Tengchao Lv, Lei Cui 0001, Qifeng Chen 0001, Furu Wei |
NeurIPS | 6 |
| 2023 | On the Pareto Front of Multilingual Neural Machine TranslationabstractIn this work, we study how the performance of a given direction changes with its sampling ratio in Multilingual Neural Machine Translation (MNMT). By training over 200 multilingual models with various model sizes, data sizes, and language directions, we find it interesting that the performance of certain translation direction does not always improve with the increase of its weight in the multi-task optimization objective. Accordingly, scalarization method leads to a multitask trade-off front that deviates from the traditional Pareto front when there exists data imbalance in the training corpus, which poses a great challenge to improve the overall performance of all directions. Based on our observations, we propose the Double Power Law to predict the unique performance trade-off front in MNMT, which is robust across various languages, data adequacy, and the number of tasks. Finally, we formulate the sample ratio selection problem in MNMT as an optimization problem based on the Double Power Law. Extensive experiments show that it achieves better performance than temperature searching and gradient manipulation methods with only 1/5 to 1/2 of the total training budget. We release the code at https://github.com/pkunlp-icler/ParetoMNMT for reproduction. Liang Chen 0024, Shuming Ma, Dongdong Zhang 0001, Furu Wei, Baobao Chang |
NeurIPS | 4 |
| 2023 | Optimizing Prompts for Text-to-Image GenerationabstractWell-designed prompts can guide text-to-image models to generate amazing images. However, the performant prompts are often model-specific and misaligned with user input. Instead of laborious human engineering, we propose prompt adaptation, a general framework that automatically adapts original user input to model-preferred prompts. Specifically, we first perform supervised fine-tuning with a pretrained language model on a small collection of manually engineered prompts. Then we use reinforcement learning to explore better prompts. We define a reward function that encourages the policy to generate more aesthetically pleasing images while preserving the original user intentions. Experimental results on Stable Diffusion show that our method outperforms manual prompt engineering in terms of both automatic metrics and human preference ratings. Moreover, reinforcement learning further boosts performance, especially on out-of-domain prompts. Yaru Hao, Zewen Chi, Li Dong 0004, Furu Wei |
NeurIPS | 4 |
| 2023 | Language Is Not All You Need: Aligning Perception with Language ModelsabstractA big convergence of language, multimodal perception, action, and world modeling is a key step toward artificial general intelligence. In this work, we introduce KOSMOS-1, a Multimodal Large Language Model (MLLM) that can perceive general modalities, learn in context (i.e., few-shot), and follow instructions (i.e., zero-shot). Specifically, we train KOSMOS-1 from scratch on web-scale multi-modal corpora, including arbitrarily interleaved text and images, image-caption pairs, and text data. We evaluate various settings, including zero-shot, few-shot, and multimodal chain-of-thought prompting, on a wide range of tasks without any gradient updates or finetuning. Experimental results show that KOSMOS-1 achieves impressive performance on (i) language understanding, generation, and even OCR-free NLP (directly fed with document images), (ii) perception-language tasks, including multimodal dialogue, image captioning, visual question answering, and (iii) vision tasks, such as image recognition with descriptions (specifying classification via text instructions). We also show that MLLMs can benefit from cross-modal transfer, i.e., transfer knowledge from language to multimodal, and from multimodal to language. In addition, we introduce a dataset of Raven IQ test, which diagnoses the nonverbal reasoning capability of MLLMs. Shaohan Huang, Li Dong 0004, Wenhui Wang 0003, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui 0001, Owais Khan Mohammed, Barun Patra, Kriti Aggarwal, Zewen Chi, Johan Bjorck, Vishrav Chaudhary, Subhojit Som, Furu Wei |
NeurIPS | 18 |
| 2023 | Augmenting Language Models with Long-Term MemoryabstractExisting large language models (LLMs) can only afford fix-sized inputs due to the input length limit, preventing them from utilizing rich long-context information from past inputs. To address this, we propose a framework, Language Models Augmented with Long-Term Memory (LongMem), which enables LLMs to memorize long history. We design a novel decoupled network architecture with the original backbone LLM frozen as a memory encoder and an adaptive residual side-network as a memory retriever and reader. Such a decoupled memory design can easily cache and update long-term past contexts for memory retrieval without suffering from memory staleness. Enhanced with memory-augmented adaptation training, LongMem can thus memorize long past context and use long-term memory for language modeling. The proposed memory retrieval module can handle unlimited-length context in its memory bank to benefit various downstream tasks. Typically, LongMem can enlarge the long-form memory to 65k tokens and thus cache many-shot extra demonstration examples as long-form memory for in-context learning. Experiments show that our method outperforms strong long-context models on ChapterBreak, a challenging long-context modeling benchmark, and achieves remarkable improvements on memory-augmented in-context learning over LLMs. The results demonstrate that the proposed method is effective in helping language models to memorize and utilize long-form contents. Weizhi Wang, Li Dong 0004, Hao Cheng 0002, Xiaodong Liu 0003, Xifeng Yan, Jianfeng Gao 0001, Furu Wei |
NeurIPS | 7 |
| 2023 | Generative retrieval for conversational question answering
Yongqi Li 0001, Nan Yang 0002, Liang Wang 0046, Furu Wei, Wenjie Li 0002 |
Inf. Process. Manag. | 4 |
| 2023 | GTrans: Grouping and Fusing Transformer Layers for Neural Machine TranslationabstractTransformer structure, stacked by a sequence of encoder and decoder network layers, achieves significant development in neural machine translation. However, vanilla Transformer mainly exploits the top-layer representation, assuming the lower layers provide trivial or redundant information and thus ignoring the bottom-layer feature that is potentially valuable. In this work, we propose theGroup-Transformer model (GTrans) that flexibly divides multi-layer representations of both encoder and decoder into different groups and then fuses these group features to generate target words. To corroborate the effectiveness of the proposed method, extensive experiments and analytic experiments are conducted on three bilingual translation benchmarks and three multilingual translation tasks, including the IWLST-14, IWLST-17, LDC, WMT-14, WMT-21 and OPUS-100 benchmark. Experimental and analytical results demonstrate that our model outperforms its Transformer counterparts by a consistent gain. Furthermore, it can be successfully scaled up to 60 encoder layers and 36 decoder layers. Jian Yang 0030, Yuwei Yin, Liqun Yang, Shuming Ma, Haoyang Huang, Dongdong Zhang 0001, Furu Wei, Zhoujun Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2022 | Sequence Level Contrastive Learning for Text SummarizationabstractContrastive learning models have achieved great success in unsupervised visual representation learning, which maximize the similarities between feature representations of different views of the same image, while minimize the similarities between feature representations of views of different images. In text summarization, the output summary is a shorter form of the input document and they have similar meanings. In this paper, we propose a contrastive learning model for supervised abstractive text summarization, where we view a document, its gold summary and its model generated summaries as different views of the same mean representation and maximize the similarities between them during training. We improve over a strong sequence-to-sequence text generation model (i.e., BART) on three different summarization datasets. Human evaluation also shows that our model achieves better faithfulness ratings compared to its counterpart without contrastive objectives. We release our code at https://github.com/xssstory/SeqCo. Shusheng Xu, Xingxing Zhang 0002, Yi Wu 0013, Furu Wei |
AAAI | 4 |
| 2022 | CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual EntailmentabstractCLIP has shown a remarkable zero-shot capability on a wide range of vision tasks.Previously, CLIP is only regarded as a powerful visual encoder.However, after being pretrained by language supervision from a large amount of image-caption pairs, CLIP itself should also have acquired some few-shot abilities for vision-language tasks.In this work, we empirically show that CLIP can be a strong vision-language few-shot learner by leveraging the power of language.We first evaluate CLIP's zero-shot performance on a typical visual question answering task and demonstrate a zero-shot cross-modality transfer capability of CLIP on the visual entailment task.Then we propose a parameter-efficient fine-tuning strategy to boost the few-shot performance on the vqa task.We achieve competitive zero/fewshot results on the visual question answering and visual entailment tasks without introducing any additional pre-training procedure. Haoyu Song 0002, Li Dong 0004, Weinan Zhang 0003, Ting Liu 0001, Furu Wei |
ACL (1) | 5 |
| 2022 | SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language ProcessingabstractJunyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Junyi Ao, Rui Wang 0073, Chengyi Wang 0002, Shuo Ren 0002, Yu Wu 0012, Shujie Liu 0001, Tom Ko, Qing Li 0001, Yu Zhang 0006, Zhihua Wei 0001, Yao Qian, Jinyu Li 0001, Furu Wei |
ACL (1) | 14 |
| 2022 | Towards Making the Most of Cross-Lingual Transfer for Zero-Shot Neural Machine TranslationabstractGuanhua Chen, Shuming Ma, Yun Chen, Dongdong Zhang, Jia Pan, Wenping Wang, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Guanhua Chen 0001, Shuming Ma, Yun Chen 0007, Dongdong Zhang 0001, Jia Pan 0001, Wenping Wang 0001, Furu Wei |
ACL (1) | 7 |
| 2022 | XLM-E: Cross-lingual Language Model Pre-training via ELECTRAabstractZewen Chi, Shaohan Huang, Li Dong, Shuming Ma, Bo Zheng, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Zewen Chi, Shaohan Huang, Li Dong 0004, Shuming Ma, Bo Zheng 0010, Saksham Singhal, Payal Bajaj, Xianling Mao, Heyan Huang, Furu Wei |
ACL (1) | 11 |
| 2022 | StableMoE: Stable Routing Strategy for Mixture of ExpertsabstractThe Mixture-of-Experts (MoE) technique can scale up the model size of Transformers with an affordable computational overhead.We point out that existing learning-to-route MoE methods suffer from the routing fluctuation issue, i.e., the target expert of the same input may change along with training, but only one expert will be activated for the input during inference.The routing fluctuation tends to harm sample efficiency because the same input updates different experts but only one is finally used.In this paper, we propose STABLEMOE with two training stages to address the routing fluctuation problem.In the first training stage, we learn a balanced and cohesive routing strategy and distill it into a lightweight router decoupled from the backbone model.In the second training stage, we utilize the distilled router to determine the token-to-expert assignment and freeze it for a stable routing strategy.We validate our method on language modeling and multilingual machine translation.The results show that STABLEMOE outperforms existing MoE methods in terms of both convergence speed and performance. Damai Dai, Li Dong 0004, Shuming Ma, Bo Zheng 0010, Zhifang Sui, Baobao Chang, Furu Wei |
ACL (1) | 7 |
| 2022 | Knowledge Neurons in Pretrained TransformersabstractLarge-scale pretrained language models are surprisingly good at recalling factual knowledge presented in the training corpus (Petroni et al., 2019; Jiang et al., 2020b).In this paper, we present preliminary studies on how factual knowledge is stored in pretrained Transformers by introducing the concept of knowledge neurons.Specifically, we examine the fill-in-the-blank cloze task for BERT.Given a relational fact, we propose a knowledge attribution method to identify the neurons that express the fact.We find that the activation of such knowledge neurons is positively correlated to the expression of their corresponding facts.In our case studies, we attempt to leverage knowledge neurons to edit (such as update, and erase) specific factual knowledge without fine-tuning.Our results shed light on understanding the storage of knowledge within pretrained Transformers.The code is available at https://github.com/ Hunter-DDM/knowledge-neurons. Damai Dai, Li Dong 0004, Yaru Hao, Zhifang Sui, Baobao Chang, Furu Wei |
ACL (1) | 6 |
| 2022 | Neural Label Search for Zero-Shot Multi-Lingual Extractive SummarizationabstractIn zero-shot multilingual extractive text summarization, a model is typically trained on English summarization dataset and then applied on summarization datasets of other languages.Given English gold summaries and documents, sentence-level labels for extractive summarization are usually generated using heuristics.However, these monolingual labels created on English datasets may not be optimal on datasets of other languages, for that there is the syntactic or semantic discrepancy between different languages.In this way, it is possible to translate the English dataset to other languages and obtain different sets of labels again using heuristics.To fully leverage the information of these different sets of labels, we propose NLSSum (Neural Label Search for Summarization), which jointly learns hierarchical weights for these different sets of labels together with our summarization model.We conduct multilingual zero-shot summarization experiments on MLSUM and WikiLingua datasets, and we achieve state-of-the-art results using both human and automatic evaluations across these two datasets. Ruipeng Jia, Xingxing Zhang 0002, Yanan Cao 0001, Zheng Lin 0001, Shi Wang 0002, Furu Wei |
ACL (1) | 6 |
| 2022 | MarkupLM: Pre-training of Text and Markup Language for Visually Rich Document UnderstandingabstractMultimodal pre-training with text, layout, and image has made significant progress for Visually Rich Document Understanding (VRDU), especially the fixed-layout documents such as scanned document images.While, there are still a large number of digital documents where the layout information is not fixed and needs to be interactively and dynamically rendered for visualization, making existing layout-based pre-training approaches not easy to apply.In this paper, we propose MarkupLM for document understanding tasks with markup languages as the backbone, such as HTML/XMLbased documents, where text and markup information is jointly pre-trained.Experiment results show that the pre-trained MarkupLM significantly outperforms the existing strong baseline models on several document understanding tasks.The pre-trained model and code will be publicly available at https:// aka.ms/markuplm. Yiheng Xu, Lei Cui 0001, Furu Wei |
ACL (1) | 4 |
| 2022 | Attention Temperature Matters in Abstractive Summarization DistillationabstractRecent progress of abstractive text summarization largely relies on large pre-trained sequence-to-sequence Transformer models, which are computationally expensive.This paper aims to distill these large models into smaller ones for faster inference and with minimal performance loss.Pseudo-labeling based methods are popular in sequence-tosequence model distillation.In this paper, we find simply manipulating attention temperatures in Transformers can make pseudo labels easier to learn for student models.Our experiments on three summarization datasets show our proposed method consistently improves vanilla pseudo-labeling based methods.Further empirical analysis shows that both pseudo labels and summaries produced by our students are shorter and more abstractive.Our code is available at https://github. com/Shengqiang-Zhang/plate. Shengqiang Zhang, Xingxing Zhang 0002, Hangbo Bao, Furu Wei |
ACL (1) | 4 |
| 2022 | Swin Transformer V2: Scaling Up Capacity and ResolutionabstractWe present techniques for scaling Swin Transformer [35] up to 3 billion parameters and making it capable of training with images of up to 1,536x1,536 resolution. By scaling up capacity and resolution, Swin Transformer sets new records on four representative vision benchmarks: 84.0% top-1 accuracy on ImageNet- V2 image classification, 63.1 / 54.4 box / mask mAP on COCO object detection, 59.9 mIoU on ADE20K semantic segmentation, and 86.8% top-1 accuracy on Kinetics-400 video action classification. We tackle issues of training instability, and study how to effectively transfer models pre-trained at low resolutions to higher resolution ones. To this aim, several novel technologies are proposed: 1) a residual post normalization technique and a scaled cosine attention approach to improve the stability of large vision models; 2) a log-spaced continuous position bias technique to effectively transfer models pre-trained at low-resolution images and windows to their higher-resolution counterparts. In addition, we share our crucial implementation details that lead to significant savings of GPU memory consumption and thus make it feasi-ble to train large vision models with regular GPUs. Using these techniques and self-supervised pre-training, we suc-cessfully train a strong 3 billion Swin Transformer model and effectively transfer it to various vision tasks involving high-resolution images or windows, achieving the state-of-the-art accuracy on a variety of benchmarks. Code is avail-able at https://github.com/microsoft/Swin-Transformer. Han Hu 0001, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Yue Cao 0001, Zheng Zhang 0022, Li Dong 0004, Furu Wei, Baining Guo |
CVPR | 11 |
| 2022 | EdgeFormer: A Parameter-Efficient Transformer for On-Device Seq2seq GenerationabstractWe introduce EDGEFORMER -a parameterefficient Transformer for on-device seq2seq generation under the strict computation and memory constraints.Compared with the previous parameter-efficient Transformers, EDGE-FORMER applies two novel principles for costeffective parameterization, allowing it to perform better given the same parameter budget; moreover, EDGEFORMER is further enhanced by layer adaptation innovation that is proposed for improving the network with shared layers.Extensive experiments show EDGEFORMER can effectively outperform previous parameterefficient Transformer baselines and achieve competitive results under both the computation and memory constraints.Given the promising results, we release EDGELM 1 -the pretrained version of EDGEFORMER, which is the first publicly available pretrained on-device seq2seq model that can be easily fine-tuned for seq2seq tasks with strong results, facilitating on-device seq2seq generation in practice. Tao Ge 0001, Furu Wei |
EMNLP | 3 |
| 2022 | Zero-shot Cross-lingual Transfer of Prompt-based Tuning with a Unified Multilingual PromptabstractPrompt-based tuning has been proven effective for pretrained language models (PLMs).While most of the existing work focuses on the monolingual prompts, we study the multilingual prompts for multilingual PLMs, especially in the zero-shot cross-lingual setting.To alleviate the effort of designing different prompts for multiple languages, we propose a novel model that uses a unified prompt for all languages, called UniPrompt.Different from the discrete prompts and soft prompts, the unified prompt is model-based and languageagnostic.Specifically, the unified prompt is initialized by a multilingual PLM to produce language-independent representation, after which is fused with the text input.During inference, the prompts can be pre-computed so that no extra computation cost is needed.To collocate with the unified prompt, we propose a new initialization method for the target label word to further improve the model's transferability across languages.Extensive experiments show that our proposed methods can significantly outperform the strong baselines across different languages.We release data and code to facilitate future research 1 . Lianzhe Huang, Shuming Ma, Dongdong Zhang 0001, Furu Wei, Houfeng Wang |
EMNLP | 4 |
| 2022 | PromptBERT: Improving BERT Sentence Embeddings with PromptsabstractTing Jiang, Jian Jiao, Shaohan Huang, Zihan Zhang, Deqing Wang, Fuzhen Zhuang, Furu Wei, Haizhen Huang, Denvy Deng, Qi Zhang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Jian Jiao 0007, Shaohan Huang, Fuzhen Zhuang, Furu Wei, Haizhen Huang, Denvy Deng, Qi Zhang 0066 |
EMNLP | 7 |
| 2022 | Distilled Dual-Encoder Model for Vision-Language UnderstandingabstractOn vision-language understanding (VLU) tasks, fusion-encoder vision-language models achieve superior results but sacrifice efficiency because of the simultaneous encoding of images and text.On the contrary, the dual-encoder model that separately encodes images and text has the advantage in efficiency, while failing on VLU tasks due to the lack of deep cross-modal interactions.To get the best of both worlds, we propose DIDE 1 , a framework that distills the knowledge of the fusion-encoder teacher model into the dual-encoder student model.Since the cross-modal interaction is the key to the superior performance of teacher model but is absent in the student model, we encourage the student not only to mimic the predictions of teacher, but also to calculate the cross-modal attention distributions and align with the teacher.Experimental results demonstrate that DIDE is competitive with the fusion-encoder teacher model in performance (only a 1% drop) while enjoying 4× faster inference.Further analyses reveal that the proposed cross-modal attention distillation is crucial to the success of our framework. Zekun Wang 0001, Wenhui Wang 0003, Ming Liu 0004, Bing Qin 0001, Furu Wei |
EMNLP | 6 |
| 2022 | SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-trainingabstractThe rapid development of single-modal pretraining has prompted researchers to pay more attention to cross-modal pre-training methods.In this paper, we propose a unified-modal speech-unit-text pre-training model, SpeechUT, to connect the representations of a speech encoder and a text decoder with a shared unit encoder.Leveraging hidden-unit as an interface to align speech and text, we can decompose the speech-to-text model into a speech-to-unit model and a unit-to-text model, which can be jointly pre-trained with unpaired speech and text data respectively.Our proposed SpeechUT is fine-tuned and evaluated on automatic speech recognition (ASR) and speech translation (ST) tasks.Experimental results show that SpeechUT gets substantial improvements over strong baselines, and achieves state-of-the-art performance on both the Lib-riSpeech ASR and MuST-C ST tasks.To better understand the proposed SpeechUT, detailed analyses are conducted.The code and pretrained models are available at https://aka. ms/SpeechUT. Junyi Ao, Shujie Liu 0001, Li-Rong Dai 0001, Jinyu Li 0001, Furu Wei |
EMNLP | 7 |
| 2022 | Unispeech-Sat: Universal Speech Representation Learning With Speaker Aware Pre-TrainingabstractSelf-supervised learning (SSL) is a long-standing goal for speech processing, since it utilizes large-scale unlabeled data and avoids extensive human labeling. Recent years have witnessed great successes in applying self-supervised learning in speech recognition, while limited exploration was attempted in applying SSL for modeling speaker characteristics. In this paper, we aim to improve the existing SSL framework for speaker representation learning. Two methods are introduced for enhancing the unsupervised speaker information extraction. First, we apply multi-task learning to the current SSL framework, where we integrate utterance-wise contrastive loss with the SSL objective function. Second, for better speaker discrimination, we propose an utterance mixing strategy for data augmentation, where additional overlapped utterances are created unsupervisely and incorporated during training. We integrate the proposed methods into the HuBERT framework. Experiment results on the SUPERB benchmark show that the proposed system achieves state-of-the-art performance in universal representation learning, especially for speaker identification oriented tasks. An ablation study is performed verifying the efficacy of each proposed method. Finally, we scale up the training dataset to 94 thousand hours of public audio data and achieve further performance improvement in all SUPERB tasks. Sanyuan Chen, Yu Wu 0012, Chengyi Wang 0002, Zhengyang Chen, Zhuo Chen 0006, Shujie Liu 0001, Jian Wu 0027, Yao Qian, Furu Wei, Jinyu Li 0001, Xiangzhan Yu |
ICASSP | 9 |
| 2022 | BEiT: BERT Pre-Training of Image Transformers
Hangbo Bao, Li Dong 0004, Furu Wei |
ICLR | 4 |
| 2022 | A Unified Strategy for Multilingual Grammatical Error Correction with Pre-trained Cross-Lingual Language ModelabstractSynthetic data construction of Grammatical Error Correction (GEC) for non-English languages relies heavily on human-designed and language-specific rules, which produce limited error-corrected patterns. In this paper, we propose a generic and language-independent strategy for multilingual GEC, which can train a GEC system effectively for a new non-English language with only two easy-to-access resources: 1) a pre-trained cross-lingual language model (PXLM) and 2) parallel translation data between English and the language. Our approach creates diverse parallel GEC data without any language-specific operations by taking the non-autoregressive translation generated by PXLM and the gold translation as error-corrected sentence pairs. Then, we reuse PXLM to initialize the GEC model and pre-train it with the synthetic data generated by itself, which yields further improvement. We evaluate our approach on three public benchmarks of GEC in different languages. It achieves the state-of-the-art results on the NLPCC 2018 Task 2 dataset (Chinese) and obtains competitive performance on Falko-Merlin (German) and RULEC-GEC (Russian). Further analysis demonstrates that our data construction method is complementary to rule-based approaches. Xin Sun 0013, Tao Ge 0001, Shuming Ma, Jingjing Li 0007, Furu Wei, Houfeng Wang |
IJCAI | 5 |
| 2022 | High-resource Language-specific Training for Multilingual Neural Machine TranslationabstractMultilingual neural machine translation (MNMT) trained in multiple language pairs has attracted considerable attention due to fewer model parameters and lower training costs by sharing knowledge among multiple languages. Nonetheless, multilingual training is plagued by language interference degeneration in shared parameters because of the negative interference among different translation directions, especially on high-resource languages. In this paper, we propose the multilingual translation model with the high-resource language-specific training (HLT-MT) to alleviate the negative interference, which adopts the two-stage training with the language-specific selection mechanism. Specifically, we first train the multilingual model only with the high-resource pairs and select the language-specific modules at the top of the decoder to enhance the translation quality of high-resource directions. Next, the model is further trained on all available corpora to transfer knowledge from high-resource languages (HRLs) to low-resource languages (LRLs). Experimental results show that HLT-MT outperforms various strong baselines on WMT-10 and OPUS-100 benchmarks. Furthermore, the analytic experiments validate the effectiveness of our method in mitigating the negative interference in multilingual training. Jian Yang 0030, Yuwei Yin, Shuming Ma, Dongdong Zhang 0001, Zhoujun Li 0001, Furu Wei |
IJCAI | 6 |
| 2022 | UM4: Unified Multilingual Multiple Teacher-Student Model for Zero-Resource Neural Machine TranslationabstractMost translation tasks among languages belong to the zero-resource translation problem where parallel corpora are unavailable. Multilingual neural machine translation (MNMT) enables one-pass translation using shared semantic space for all languages compared to the two-pass pivot translation but often underperforms the pivot-based method. In this paper, we propose a novel method, named as Unified Multilingual Multiple teacher-student Model for NMT (UM4). Our method unifies source-teacher, target-teacher, and pivot-teacher models to guide the student model for the zero-resource translation. The source teacher and target teacher force the student to learn the direct source-target translation by the distilled knowledge on both source and target sides. The monolingual corpus is further leveraged by the pivot-teacher model to enhance the student model. Experimental results demonstrate that our model of 72 directions significantly outperforms previous methods on the WMT benchmark. Jian Yang 0030, Yuwei Yin, Shuming Ma, Dongdong Zhang 0001, Shuangzhi Wu, Hongcheng Guo, Zhoujun Li 0001, Furu Wei |
IJCAI | 8 |
| 2022 | Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Speech DataabstractThis paper studies a novel pre-training technique with unpaired speech data, Speech2C, for encoder-decoder based automatic speech recognition (ASR).Within a multi-task learning framework, we introduce two pre-training tasks for the encoderdecoder network using acoustic units, i.e., pseudo codes, derived from an offline clustering model.One is to predict the pseudo codes via masked language modeling in encoder output, like HuBERT model, while the other lets the decoder learn to reconstruct pseudo codes autoregressively instead of generating textual scripts.In this way, the decoder learns to reconstruct original speech information with codes before learning to generate correct text.Comprehensive experiments on the LibriSpeech corpus show that the proposed Speech2C can relatively reduce the word error rate (WER) by 19.2% over the method without decoder pre-training, and also outperforms significantly the state-of-the-art wav2vec 2.0 and Hu-BERT on fine-tuning subsets of 10h and 100h.We release our code and model at https://github.com/microsoft/SpeechT5/tree/main/Speech2C. Junyi Ao, Shujie Liu 0001, Haizhou Li 0001, Tom Ko, Li-Rong Dai 0001, Jinyu Li 0001, Yao Qian, Furu Wei |
INTERSPEECH | 10 |
| 2022 | Why does Self-Supervised Learning for Speech Recognition Benefit Speaker Recognition?abstractRecently, self-supervised learning (SSL) has demonstrated strong performance in speaker recognition, even if the pretraining objective is designed for speech recognition.In this paper, we study which factor leads to the success of selfsupervised learning on speaker-related tasks, e.g.speaker verification (SV), through a series of carefully designed experiments.Our empirical results on the Voxceleb-1 dataset suggest that the benefit of SSL to SV task is from a combination of mask speech prediction loss, data scale, and model size, while the SSL quantizer has a minor impact.We further employ the integrated gradients attribution method and loss landscape visualization to understand the effectiveness of self-supervised learning for speaker recognition performance. Sanyuan Chen, Yu Wu 0012, Chengyi Wang 0002, Shujie Liu 0001, Zhuo Chen 0006, Gang Liu 0001, Jinyu Li 0001, Jian Wu 0027, Xiangzhan Yu, Furu Wei |
INTERSPEECH | 11 |
| 2022 | Speech Pre-training with Acoustic PieceabstractPrevious speech pre-training methods, such as wav2vec2.0 and HuBERT, pre-train a Transformer encoder to learn deep representations from audio data, with objectives predicting either elements from latent vector quantized space or pre-generated labels (known as target codes) with offline clustering. However, those training signals (quantized elements or codes) are independent across different tokens without considering their relations. According to our observation and analysis, the target codes share obvious patterns aligned with phonemized text data. Based on that, we propose to leverage those patterns to better pre-train the model considering the relations among the codes. The patterns we extracted, called "acoustic piece"s, are from the sentence piece result of HuBERT codes. With the acoustic piece as the training signal, we can implicitly bridge the input audio and natural language, which benefits audio-to-text tasks, such as automatic speech recognition (ASR). Simple but effective, our method "HuBERT-AP" significantly outperforms strong baselines on the LibriSpeech ASR task. Shuo Ren 0002, Shujie Liu 0001, Yu Wu 0012, Furu Wei |
INTERSPEECH | 5 |
| 2022 | Supervision-Guided Codebooks for Masked Prediction in Speech Pre-trainingabstractRecently, masked prediction pre-training has seen remarkable progress in self-supervised learning (SSL) for speech recognition.It usually requires a codebook obtained in an unsupervised way, making it less accurate and difficult to interpret.We propose two supervision-guided codebook generation approaches to improve automatic speech recognition (ASR) performance and also the pre-training efficiency, either through decoding with a hybrid ASR system to generate phoneme-level alignments (named PBERT), or performing clustering on the supervised speech features extracted from an end-to-end CTC model (named CTC clustering).Both the hybrid and CTC models are trained on the same small amount of labeled speech as used in fine-tuning.Experiments demonstrate significant superiority of our methods to various SSL and self-training baselines, with up to 17.0% relative WER reduction.Our pre-trained models also show good transferability in a non-ASR speech task. Chengyi Wang 0002, Yu Wu 0012, Sanyuan Chen, Jinyu Li 0001, Shujie Liu 0001, Furu Wei |
INTERSPEECH | 7 |
| 2022 | Separating Long-Form Speech with Group-wise Permutation Invariant TrainingabstractMulti-talker conversational speech processing has drawn many interests for various applications such as meeting transcription.Speech separation is often required to handle overlapped speech that is commonly observed in conversation.Although the original utterancelevel permutation invariant training-based continuous speech separation approach has proven to be effective in various conditions, it lacks the ability to leverage the long-span relationship of utterances and is computationally inefficient due to the highly overlapped sliding windows.To overcome these drawbacks, we propose a novel training scheme named Group-PIT, which allows direct training of the speech separation models on the long-form speech with a low computational cost for label assignment.Two different speech separation approaches with Group-PIT are explored, including direct long-span speech separation and short-span speech separation with long-span tracking.The experiments on the simulated meeting-style data demonstrate the effectiveness of our proposed approaches, especially in dealing with a very long speech input. Wangyou Zhang, Zhuo Chen 0006, Naoyuki Kanda, Shujie Liu 0001, Jinyu Li 0001, Sefik Emre Eskimez, Takuya Yoshioka, Zhong Meng, Yanmin Qian, Furu Wei |
INTERSPEECH | 11 |
| 2022 | LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingabstractSelf-supervised pre-training techniques have achieved remarkable progress in Document AI. Most multimodal pre-trained models use a masked language modeling objective to learn bidirectional representations on the text modality, but they differ in pre-training objectives for the image modality. This discrepancy adds difficulty to multimodal representation learning. In this paper, we propose LayoutLMv3 to pre-train multimodal Transformers for Document AI with unified text and image masking. Additionally, LayoutLMv3 is pre-trained with a word-patch alignment objective to learn cross-modal alignment by predicting whether the corresponding image patch of a text word is masked. The simple unified architecture and training objectives make LayoutLMv3 a general-purpose pre-trained model for both text-centric and image-centric Document AI tasks. Experimental results show that LayoutLMv3 achieves state-of-the-art performance not only in text-centric tasks, including form understanding, receipt understanding, and document visual question answering, but also in image-centric tasks such as document image classification and document layout analysis. The code and models are publicly available at https://aka.ms/layoutlmv3. Yupan Huang, Tengchao Lv, Lei Cui 0001, Yutong Lu, Furu Wei |
ACM Multimedia | 5 |
| 2022 | DiT: Self-supervised Pre-training for Document Image TransformerabstractImage Transformer has recently achieved significant progress for natural image understanding, either using supervised (ViT, DeiT, etc.) or self-supervised (BEiT, MAE, etc.) pre-training techniques. In this paper, we propose DiT, a self-supervised pre-trained Document Image Transformer model using large-scale unlabeled text images for Document AI tasks, which is essential since no supervised counterparts ever exist due to the lack of human-labeled document images. We leverage DiT as the backbone network in a variety of vision-based Document AI tasks, including document image classification, document layout analysis, table detection as well as text detection for OCR. Experiment results have illustrated that the self-supervised pre-trained DiT model achieves new state-of-the-art results on these downstream tasks, e.g. document image classification (91.11 - 92.69), document layout analysis (91.0 - 94.9), table detection (94.23 - 96.55) and text detection for OCR (93.07 - 94.29). The code and pre-trained models are publicly available at https://aka.ms/msdit. Yiheng Xu, Tengchao Lv, Lei Cui 0001, Cha Zhang, Furu Wei |
ACM Multimedia | 6 |
| 2022 | VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-ExpertsabstractWe present a unified Vision-Language pretrained Model (VLMo) that jointly learns a dual encoder and a fusion encoder with a modular Transformer network. Specifically, we introduce Multiway Transformer, where each block contains a pool of modality-specific experts and a shared self-attention layer. Because of the modeling flexibility of Multiway Transformer, pretrained VLMo can be fine-tuned as a fusion encoder for vision-language classification tasks, or used as a dual encoder for efficient image-text retrieval. Moreover, we propose a stagewise pre-training strategy, which effectively leverages large-scale image-only and text-only data besides image-text pairs. Experimental results show that VLMo achieves state-of-the-art results on various vision-language tasks, including VQA, NLVR2 and image-text retrieval. Hangbo Bao, Wenhui Wang 0003, Li Dong 0004, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, Furu Wei |
NeurIPS | 9 |
| 2022 | On the Representation Collapse of Sparse Mixture of ExpertsabstractSparse mixture of experts provides larger model capacity while requiring a constant computational overhead. It employs the routing mechanism to distribute input tokens to the best-matched experts according to their hidden representations. However, learning such a routing mechanism encourages token clustering around expert centroids, implying a trend toward representation collapse. In this work, we propose to estimate the routing scores between tokens and experts on a low-dimensional hypersphere. We conduct extensive experiments on cross-lingual language model pre-training and fine-tuning on downstream tasks. Experimental results across seven multilingual benchmarks show that our method achieves consistent gains. We also present a comprehensive analysis on the representation and routing behaviors of our models. Our method alleviates the representation collapse issue and achieves more consistent routing than the baseline mixture-of-experts methods. Zewen Chi, Li Dong 0004, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xianling Mao, Heyan Huang, Furu Wei |
NeurIPS | 12 |
| 2022 | Kformer: Knowledge Injection in Transformer Feed-Forward Layers
Yunzhi Yao, Shaohan Huang, Li Dong 0004, Furu Wei, Huajun Chen, Ningyu Zhang 0001 |
NLPCC (1) | 4 |
| 2022 | Transforming Wikipedia Into Augmented Data for Query-Focused SummarizationabstractThe limited size of existing query-focused summarization datasets renders training data-driven summarization models challenging. Meanwhile, the manual construction of a query-focused summarization corpus is costly and time-consuming. In this paper, we use Wikipedia to automatically collect a large query-focused summarization dataset (named WikiRef) of more than 280,000 examples, which can serve as a means of data augmentation. We also develop a BERT-based query-focused summarization model (Q-BERT) to extract sentences from the documents as summaries. To better adapt a huge model containing millions of parameters to tiny benchmarks, we identify and fine-tune only a sparse subnetwork, which corresponds to a small fraction of the whole model parameters. Experimental results on three DUC benchmarks show that the model pre-trained on WikiRefhas already achieved reasonable performance. After fine-tuning on the specific benchmark datasets, the model with data augmentation outperforms strong comparison systems. Moreover, both our proposed Q-BERT model and subnetwork fine-tuning further improve the model performance. Li Dong 0004, Furu Wei, Bing Qin 0001, Ting Liu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Self-Attention Attribution: Interpreting Information Interactions Inside TransformerabstractThe great success of Transformer-based models benefits from the powerful multi-head self-attention mechanism, which learns token dependencies and encodes contextual information from the input. Prior work strives to attribute model decisions to individual input features with different saliency measures, but they fail to explain how these input features interact with each other to reach predictions. In this paper, we propose a self-attention attribution method to interpret the information interactions inside Transformer. We take BERT as an example to conduct extensive studies. Firstly, we apply self-attention attribution to identify the important attention heads, while others can be pruned with marginal performance degradation. Furthermore, we extract the most salient dependencies in each layer to construct an attribution tree, which reveals the hierarchical interactions inside Transformer. Finally, we show that the attribution results can be used as adversarial patterns to implement non-targeted attacks towards BERT. Yaru Hao, Li Dong 0004, Furu Wei, Ke Xu 0001 |
AAAI | 3 |
| 2021 | Improving Pretrained Cross-Lingual Language Models via Self-Labeled Word AlignmentabstractZewen Chi, Li Dong, Bo Zheng, Shaohan Huang, Xian-Ling Mao, Heyan Huang, Furu Wei. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zewen Chi, Li Dong 0004, Bo Zheng 0010, Shaohan Huang, Xianling Mao, Heyan Huang, Furu Wei |
ACL/IJCNLP (1) | 7 |
| 2021 | SemFace: Pre-training Encoder and Decoder with a Semantic Interface for Neural Machine TranslationabstractShuo Ren, Long Zhou, Shujie Liu, Furu Wei, Ming Zhou, Shuai Ma. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Shuo Ren 0002, Shujie Liu 0001, Furu Wei, Ming Zhou 0001, Shuai Ma 0001 |
ACL/IJCNLP (1) | 4 |
| 2021 | Instantaneous Grammatical Error Correction with Shallow Aggressive DecodingabstractXin Sun, Tao Ge, Furu Wei, Houfeng Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xin Sun 0013, Tao Ge 0001, Furu Wei, Houfeng Wang |
ACL/IJCNLP (1) | 3 |
| 2021 | LayoutLMv2: Multi-modal Pre-training for Visually-rich Document UnderstandingabstractYang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, Lidong Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yang Xu 0049, Yiheng Xu, Tengchao Lv, Lei Cui 0001, Furu Wei, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Wanxiang Che, Min Zhang 0005, Lidong Zhou |
ACL/IJCNLP (1) | 5 |
| 2021 | xMoCo: Cross Momentum Contrastive Learning for Open-Domain Question AnsweringabstractNan Yang, Furu Wei, Binxing Jiao, Daxing Jiang, Linjun Yang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Nan Yang 0002, Furu Wei, Binxing Jiao, Daxing Jiang, Linjun Yang |
ACL/IJCNLP (1) | 2 |
| 2021 | Consistency Regularization for Cross-Lingual Fine-TuningabstractBo Zheng, Li Dong, Shaohan Huang, Wenhui Wang, Zewen Chi, Saksham Singhal, Wanxiang Che, Ting Liu, Xia Song, Furu Wei. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Bo Zheng 0010, Li Dong 0004, Shaohan Huang, Wenhui Wang 0003, Zewen Chi, Saksham Singhal, Wanxiang Che, Ting Liu 0001, Furu Wei |
ACL/IJCNLP (1) | 10 |
| 2021 | Jointly Learning to Repair Code and Generate Commit MessageabstractWe propose a novel task of jointly repairing program codes and generating commit messages.Code repair and commit message generation are two essential and related tasks for software development.However, existing work usually performs the two tasks independently.We construct a multilingual triple dataset including buggy code, fixed code, and commit messages for this novel task.We provide the cascaded models as baseline, which are enhanced with different training approaches, including the teacher-student method, the multi-task method, and the backtranslation method.To deal with the error propagation problem of the cascaded method, the joint model is proposed that can both repair the code and generate the commit message in a unified framework.Experimental results show that the enhanced cascaded model with teacher-student method and multitask-learning method achieves the best score on different metrics of automated code repair, and the joint model behaves better than the cascaded model on commit message generation. Jiaqi Bai 0001, Ambrosio Blanco, Shujie Liu 0001, Furu Wei, Ming Zhou 0001, Zhoujun Li 0001 |
EMNLP (1) | 5 |
| 2021 | Zero-Shot Cross-Lingual Transfer of Neural Machine Translation with Multilingual Pretrained EncodersabstractPrevious work mainly focuses on improving cross-lingual transfer for NLU tasks with a multilingual pretrained encoder (MPE), or improving the performance on supervised machine translation with BERT.However, it is under-explored that whether the MPE can help to facilitate the cross-lingual transferability of NMT model.In this paper, we focus on a zero-shot cross-lingual transfer task in NMT.In this task, the NMT model is trained with parallel dataset of only one language pair and an off-the-shelf MPE, then it is directly tested on zero-shot language pairs.We propose SixT, a simple yet effective model for this task.SixT leverages the MPE with a two-stage training schedule and gets further improvement with a position disentangled encoder and a capacity-enhanced decoder.Using this method, SixT significantly outperforms mBART, a pretrained multilingual encoderdecoder model explicitly designed for NMT, with an average improvement of 7.1 BLEU on zero-shot any-to-English test sets across 14 source languages.Furthermore, with much less training computation cost and training data, our model achieves better performance on 15 any-to-English test sets than CRISS and m2m-100, two strong multilingual NMT baselines. Guanhua Chen 0001, Shuming Ma, Yun Chen 0007, Li Dong 0004, Dongdong Zhang 0001, Jia Pan 0001, Wenping Wang 0001, Furu Wei |
EMNLP (1) | 8 |
| 2021 | mT6: Multilingual Pretrained Text-to-Text Transformer with Translation PairsabstractZewen Chi, Li Dong, Shuming Ma, Shaohan Huang, Saksham Singhal, Xian-Ling Mao, Heyan Huang, Xia Song, Furu Wei. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zewen Chi, Li Dong 0004, Shuming Ma, Shaohan Huang, Saksham Singhal, Xianling Mao, Heyan Huang, Furu Wei |
EMNLP (1) | 9 |
| 2021 | LayoutReader: Pre-training of Text and Layout for Reading Order DetectionabstractReading order detection is the cornerstone to understanding visually-rich documents (e.g., receipts and forms).Unfortunately, no existing work took advantage of advanced deep learning models because it is too laborious to annotate a large enough dataset.We observe that the reading order of WORD documents is embedded in their XML metadata; meanwhile, it is easy to convert WORD documents to PDFs or images.Therefore, in an automated manner, we construct ReadingBank, a benchmark dataset that contains reading order, text, and layout information for 500,000 document images covering a wide spectrum of document types.This first-ever large-scale dataset unleashes the power of deep neural networks for reading order detection.Specifically, our proposed LayoutReader captures the text and layout information for reading order prediction using the seq2seq model.It performs almost perfectly in reading order detection and significantly improves both open-source and commercial OCR engines in ordering text lines in their results in our experiments.The dataset and models are publicly available at https: //aka.ms/layoutreader. Zilong Wang 0002, Yiheng Xu, Lei Cui 0001, Jingbo Shang, Furu Wei |
EMNLP (1) | 5 |
| 2021 | Beyond Preserved Accuracy: Evaluating Loyalty and Robustness of BERT CompressionabstractRecent studies on compression of pretrained language models (e.g., BERT) usually use preserved accuracy as the metric for evaluation.In this paper, we propose two new metrics, label loyalty and probability loyalty that measure how closely a compressed model (i.e., student) mimics the original model (i.e., teacher).We also explore the effect of compression with regard to robustness under adversarial attacks.We benchmark quantization, pruning, knowledge distillation and progressive module replacing with loyalty and robustness.By combining multiple compression techniques, we provide a practical strategy to achieve better accuracy, loyalty and robustness. 1 Canwen Xu, Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Julian J. McAuley, Furu Wei |
EMNLP (1) | 6 |
| 2021 | Allocating Large Vocabulary Capacity for Cross-Lingual Language Model Pre-TrainingabstractCompared to monolingual models, crosslingual models usually require a more expressive vocabulary to represent all languages adequately.We find that many languages are under-represented in recent cross-lingual language models due to the limited vocabulary capacity.To this end, we propose an algorithm VOCAP to determine the desired vocabulary capacity of each language.However, increasing the vocabulary size significantly slows down the pre-training speed.In order to address the issues, we propose k-NN-based target sampling to accelerate the expensive softmax.Our experiments show that the multilingual vocabulary learned with VOCAP benefits cross-lingual language model pre-training.Moreover, k-NN-based target sampling mitigates the side-effects of increasing the vocabulary size while achieving comparable performance and faster pre-training speed.The code and the pretrained multilingual vocabularies are available at https://github. com/bozheng-hit/VoCapXLM. Bo Zheng 0010, Li Dong 0004, Shaohan Huang, Saksham Singhal, Wanxiang Che, Ting Liu 0001, Furu Wei |
EMNLP (1) | 8 |
| 2021 | Improving Sequence-to-Sequence Pre-training via Sequence Span RewritingabstractIn this paper, we propose Sequence Span Rewriting (SSR), a self-supervised task for sequence-to-sequence (Seq2Seq) pre-training.SSR learns to refine the machine-generated imperfect text spans into ground truth text.SSR provides more fine-grained and informative supervision in addition to the original textinfilling objective.Compared to the prevalent text infilling objectives for Seq2Seq pretraining, SSR is naturally more consistent with many downstream generation tasks that require sentence rewriting (e.g., text summarization, question generation, grammatical error correction, and paraphrase generation).We conduct extensive experiments by using SSR to improve the typical Seq2Seq pre-trained model T5 in a continual pre-training setting and show substantial improvements over T5 on various natural language generation tasks. 1 Wangchunshu Zhou, Tao Ge 0001, Canwen Xu, Ke Xu 0001, Furu Wei |
EMNLP (1) | 5 |
| 2021 | UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled DataabstractIn this paper, we propose a unified pre-training approach called UniSpeech to learn speech representations with both labeled and unlabeled data, in which supervised phonetic CTC learning and phonetically-aware contrastive self-supervised learning are conducted in a multi-task learning manner. The resultant representations can capture information more correlated with phonetic structures and improve the generalization across languages and domains. We evaluate the effectiveness of UniSpeech for cross-lingual representation learning on public CommonVoice corpus. The results show that UniSpeech outperforms self-supervised pretraining and supervised transfer learning for speech recognition by a maximum of 13.4% and 26.9% relative phone error rate reductions respectively (averaged over all testing languages). The transferability of UniSpeech is also verified on a domain-shift speech recognition task, i.e., a relative word error rate reduction of 6% against the previous approach. Chengyi Wang 0002, Yu Wu 0012, Yao Qian, Ken'ichi Kumatani, Shujie Liu 0001, Furu Wei, Michael Zeng 0001, Xuedong Huang 0001 |
ICML | 6 |
| 2021 | InfoXLM: An Information-Theoretic Framework for Cross-Lingual Language Model Pre-TrainingabstractZewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, Heyan Huang, Ming Zhou. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zewen Chi, Li Dong 0004, Furu Wei, Nan Yang 0002, Saksham Singhal, Wenhui Wang 0003, Xianling Mao, Heyan Huang, Ming Zhou 0001 |
NAACL-HLT | 3 |
| 2021 | Blow the Dog Whistle: A Chinese Dataset for Cant Understanding with Common Sense and World KnowledgeabstractCanwen Xu, Wangchunshu Zhou, Tao Ge, Ke Xu, Julian McAuley, Furu Wei. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Canwen Xu, Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Julian J. McAuley, Furu Wei |
NAACL-HLT | 6 |
| 2021 | Learning to Select Relevant Knowledge for Neural Machine Translation
Jian Yang 0030, Juncheng Wan, Shuming Ma, Haoyang Huang, Dongdong Zhang 0001, Yong Yu 0001, Zhoujun Li 0001, Furu Wei |
NLPCC (1) | 8 |
| 2020 | Cross-Lingual Natural Language Generation via Pre-TrainingabstractIn this work we focus on transferring supervision signals of natural language generation (NLG) tasks between multiple languages. We propose to pretrain the encoder and the decoder of a sequence-to-sequence model under both monolingual and cross-lingual settings. The pre-training objective encourages the model to represent different languages in the shared space, so that we can conduct zero-shot cross-lingual transfer. After the pre-training procedure, we use monolingual data to fine-tune the pre-trained model on downstream NLG tasks. Then the sequence-to-sequence model trained in a single language can be directly evaluated beyond that language (i.e., accepting multi-lingual input and producing multi-lingual output). Experimental results on question generation and abstractive summarization show that our model outperforms the machine-translation-based pipeline methods for zero-shot cross-lingual generation. Moreover, cross-lingual transfer improves NLG performance of low-resource languages by leveraging rich-resource language data. Our implementation and data are available at https://github.com/CZWin32768/xnlg. Zewen Chi, Li Dong 0004, Furu Wei, Wenhui Wang 0003, Xianling Mao, Heyan Huang |
AAAI | 3 |
| 2020 | Fact-Aware Sentence Split and Rephrase with Permutation Invariant TrainingabstractSentence Split and Rephrase aims to break down a complex sentence into several simple sentences with its meaning preserved. Previous studies tend to address the issue by seq2seq learning from parallel sentence pairs, which takes a complex sentence as input and sequentially generates a series of simple sentences. However, the conventional seq2seq learning has two limitations for this task: (1) it does not take into account the facts stated in the long sentence; As a result, the generated simple sentences may miss or inaccurately state the facts in the original sentence. (2) The order variance of the simple sentences to be generated may confuse the seq2seq model during training because the simple sentences derived from the long source sentence could be in any order.To overcome the challenges, we first propose the Fact-aware Sentence Encoding, which enables the model to learn facts from the long sentence and thus improves the precision of sentence split; then we introduce Permutation Invariant Training to alleviate the effects of order variance in seq2seq learning for this task. Experiments on the WebSplit-v1.0 benchmark dataset show that our approaches can largely improve the performance over the previous seq2seq learning approaches. Moreover, an extrinsic evaluation on oie-benchmark verifies the effectiveness of our approaches by an observation that splitting long sentences with our state-of-the-art model as preprocessing is helpful for improving OpenIE performance. Yinuo Guo, Tao Ge 0001, Furu Wei |
AAAI | 3 |
| 2020 | Harvesting and Refining Question-Answer Pairs for Unsupervised QAabstractQuestion Answering (QA) has shown great success thanks to the availability of largescale datasets and the effectiveness of neural models.Recent research works have attempted to extend these successes to the settings with few or no labeled data available.In this work, we introduce two approaches to improve unsupervised QA.First, we harvest lexically and syntactically divergent questions from Wikipedia to automatically construct a corpus of question-answer pairs (named as REFQA).Second, we take advantage of the QA model to extract more appropriate answers, which iteratively refines data over RE-FQA.We conduct experiments 1 on SQuAD 1.1, and NewsQA by fine-tuning BERT without access to manually annotated data.Our approach outperforms previous unsupervised approaches by a large margin and is competitive with early supervised models.We also show the effectiveness of our approach in the fewshot learning setting. Zhongli Li, Wenhui Wang 0003, Li Dong 0004, Furu Wei, Ke Xu 0001 |
ACL | 4 |
| 2020 | Unsupervised Fine-tuning for Text ClusteringabstractFine-tuning with pre-trained language models (e.g.BERT) has achieved great success in many language understanding tasks in supervised settings (e.g.text classification).However, relatively little work has been focused on applying pre-trained models in unsupervised settings, such as text clustering.In this paper, we propose a novel method to fine-tune pre-trained models unsupervisedly for text clustering, which simultaneously learns text representations and cluster assignments using a clustering oriented loss.Experiments on three text clustering datasets (namely TREC-6, Yelp, and DBpedia) show that our model outperforms the baseline methods and achieves stateof-the-art results. Shaohan Huang, Furu Wei, Lei Cui 0001, Xingxing Zhang 0002, Ming Zhou 0001 |
COLING | 2 |
| 2020 | DocBank: A Benchmark Dataset for Document Layout AnalysisabstractDocument layout analysis usually relies on computer vision models to understand documents while ignoring textual information that is vital to capture.Meanwhile, high quality labeled datasets with both visual and textual information are still insufficient.In this paper, we present DocBank, a benchmark dataset that contains 500K document pages with fine-grained tokenlevel annotations for document layout analysis.DocBank is constructed using a simple yet effective way with weak supervision from the L A T E X documents available on the arXiv.com.With DocBank, models from different modalities can be compared fairly and multi-modal approaches will be further investigated and boost the performance of document layout analysis.We build several strong baselines and manually split train/dev/test sets for evaluation.Experiment results show that models trained on DocBank accurately recognize the layout information for a variety of documents.The DocBank dataset is publicly available at https: //github.com/doc-analysis/DocBank. Minghao Li 0004, Yiheng Xu, Lei Cui 0001, Shaohan Huang, Furu Wei, Zhoujun Li 0001, Ming Zhou 0001 |
COLING | 5 |
| 2020 | At Which Level Should We Extract? An Empirical Analysis on Extractive Document SummarizationabstractExtractive methods have been proven effective in automatic document summarization.Previous works perform this task by identifying informative contents at sentence level.However, it is unclear whether performing extraction at sentence level is the best solution.In this work, we show that unnecessity and redundancy issues exist when extracting full sentences, and extracting sub-sentential units is a promising alternative.Specifically, we propose extracting sub-sentential units based on the constituency parsing tree.A neural extractive model which leverages the subsentential information and extracts them is presented.Extensive experiments and analyses show that extracting sub-sentential units performs competitively comparing to full sentence extraction under the evaluation of both automatic and human evaluations.Hopefully, our work could provide some inspiration of the basic extraction units in extractive summarization for future research. Qingyu Zhou, Furu Wei, Ming Zhou 0001 |
COLING | 2 |
| 2020 | Multimodal Matching Transformer for Live CommentingabstractAutomatic live commenting aims to provide real-time comments on videos for viewers. It encourages users engagement on online video sites, and is also a good benchmark for video-to-text generation. Recent work on this task adopts encoder-decoder models to generate comments. However, these methods do not model the interaction between videos and comments explicitly, so they tend to generate popular comments that are often irrelevant to the videos. In this work, we aim to improve the relevance between live comments and videos by modeling the cross-modal interactions among different modalities. To this end, we propose a multimodal matching transformer to capture the relationships among comments, vision, and audio. The proposed model is based on the transformer framework and can iteratively learn the attention-aware representations for each modality. We evaluate the model on a publicly available live commenting dataset. Experiments show that the multimodal matching transformer model outperforms the state-of-the-art methods. Chaoqun Duan, Lei Cui 0001, Shuming Ma, Furu Wei, Conghui Zhu, Tiejun Zhao |
ECAI | 4 |
| 2020 | Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Xiujun Li, Xi Yin 0006, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu 0006, Lei Zhang 0001, Houdong Hu, Li Dong 0004, Furu Wei, Yejin Choi 0001, Jianfeng Gao 0001 |
ECCV (30) | 10 |
| 2020 | Improving the Efficiency of Grammatical Error Correction with Erroneous Span Detection and CorrectionabstractWe propose a novel language-independent approach to improve the efficiency for Grammatical Error Correction (GEC) by dividing the task into two subtasks: Erroneous Span Detection (ESD) and Erroneous Span Correction (ESC).ESD identifies grammatically incorrect text spans with an efficient sequence tagging model.Then, ESC leverages a seq2seq model to take the sentence with annotated erroneous spans as input and only outputs the corrected text for these spans.Experiments show our approach performs comparably to conventional seq2seq approaches in both English and Chinese GEC benchmarks with less than 50% time cost for inference. Mengyun Chen, Tao Ge 0001, Xingxing Zhang 0002, Furu Wei, Ming Zhou 0001 |
EMNLP (1) | 4 |
| 2020 | Language Generation with Multi-Hop Reasoning on Commonsense Knowledge GraphabstractDespite the success of generative pre-trained language models on a series of text generation tasks, they still suffer in cases where reasoning over underlying commonsense knowledge is required during generation.Existing approaches that integrate commonsense knowledge into generative pre-trained language models simply transfer relational knowledge by post-training on individual knowledge triples while ignoring rich connections within the knowledge graph.We argue that exploiting both the structural and semantic information of the knowledge graph facilitates commonsenseaware text generation.In this paper, we propose Generation with Multi-Hop Reasoning Flow (GRF) that enables pre-trained models with dynamic multi-hop reasoning on multirelational paths extracted from the external commonsense knowledge graph.We empirically show that our model outperforms existing baselines on three text generation tasks that require reasoning over commonsense knowledge.We also demonstrate the effectiveness of the dynamic multi-hop reasoning module with reasoning paths inferred by the model that provide rationale to the generation. 1 Haozhe Ji, Pei Ke, Shaohan Huang, Furu Wei, Xiaoyan Zhu 0001, Minlie Huang |
EMNLP (1) | 4 |
| 2020 | BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingabstractIn this paper, we propose a novel model compression approach to effectively compress BERT by progressive module replacing.Our approach first divides the original BERT into several modules and builds their compact substitutes.Then, we randomly replace the original modules with their substitutes to train the compact modules to mimic the behavior of the original modules.We progressively increase the probability of replacement through the training.In this way, our approach brings a deeper level of interaction between the original and compact models.Compared to the previous knowledge distillation approaches for BERT compression, our approach does not introduce any additional loss function.Our approach outperforms existing knowledge distillation approaches on GLUE benchmark, showing a new perspective of model compression.1 Canwen Xu, Wangchunshu Zhou, Tao Ge 0001, Furu Wei, Ming Zhou 0001 |
EMNLP (1) | 4 |
| 2020 | Pre-training for Abstractive Document Summarization by Reinstating Source TextabstractAbstractive document summarization is usually modeled as a sequence-to-sequence (SEQ2SEQ) learning problem.Unfortunately, training large SEQ2SEQ based summarization models on limited supervised summarization data is challenging.This paper presents three sequence-to-sequence pre-training (in shorthand, STEP) objectives which allow us to pre-train a SEQ2SEQ based abstractive summarization model on unlabeled text.The main idea is that, given an input text artificially constructed from a document, a model is pre-trained to reinstate the original document.These objectives include sentence reordering, next sentence generation and masked document generation, which have close relations with the abstractive document summarization task.Experiments on two benchmark summarization datasets (i.e., CNN/DailyMail and New York Times) show that all three objectives can improve performance upon baselines.Compared to models pre-trained on large-scale data (≥160GB), our method, with only 19GB text for pre-training, achieves comparable results, which demonstrates its effectiveness.Code and models are public available at https://github.com/ zoezou2015/abs_pretraining. Xingxing Zhang 0002, Wei Lu 0011, Furu Wei, Ming Zhou 0001 |
EMNLP (1) | 4 |
| 2020 | VL-BERT: Pre-training of Generic Visual-Linguistic Representations
Weijie Su 0002, Xizhou Zhu, Bin Li 0025, Lewei Lu, Furu Wei, Jifeng Dai |
ICLR | 6 |
| 2020 | Self-Adversarial Learning with Comparative Discrimination for Text Generation
Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Furu Wei, Ming Zhou 0001 |
ICLR | 4 |
| 2020 | UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingabstractWe propose to pre-train a unified language model for both autoencoding and partially autoregressive language modeling tasks using a novel training procedure, referred to as a pseudo-masked language model (PMLM). Given an input text with masked tokens, we rely on conventional masks to learn inter-relations between corrupted tokens and context via autoencoding, and pseudo masks to learn intra-relations between masked spans via partially autoregressive modeling. With well-designed position embeddings and self-attention masks, the context encodings are reused to avoid redundant computation. Moreover, conventional masks used for autoencoding provide global masking information, so that all the position embeddings are accessible in partially autoregressive language modeling. In addition, the two tasks pre-train a unified language model as a bidirectional encoder and a sequence-to-sequence decoder, respectively. Our experiments show that the unified language models pre-trained using PMLM achieve new state-of-the-art results on a wide range of language understanding and generation tasks across several widely used benchmarks. The code and pre-trained models are available at https://github.com/microsoft/unilm. Hangbo Bao, Li Dong 0004, Furu Wei, Wenhui Wang 0003, Nan Yang 0002, Xiaodong Liu 0003, Yu Wang 0009, Jianfeng Gao 0001, Ming Zhou 0001, Hsiao-Wuen Hon |
ICML | 3 |
| 2020 | LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingabstractPre-training techniques have been verified successfully in a variety of NLP tasks in recent years. Despite the widespread use of pre-training models for NLP applications, they almost exclusively focus on text-level manipulation, while neglecting layout and style information that is vital for document image understanding. In this paper, we propose the LayoutLM to jointly model interactions between text and layout information across scanned document images, which is beneficial for a great number of real-world document image understanding tasks such as information extraction from scanned documents. Furthermore, we also leverage image features to incorporate words' visual information into LayoutLM. To the best of our knowledge, this is the first time that text and layout are jointly learned in a single framework for document-level pre-training. It achieves new state-of-the-art results in several downstream tasks, including form understanding (from 70.72 to 79.27), receipt understanding (from 94.02 to 95.24) and document image classification (from 93.07 to 94.42). The code and pre-trained LayoutLM models are publicly available at https://aka.ms/layoutlm. Yiheng Xu, Minghao Li 0004, Lei Cui 0001, Shaohan Huang, Furu Wei, Ming Zhou 0001 |
KDD | 5 |
| 2020 | TableBank: Table Benchmark for Image-based Table Detection and RecognitionabstractWe present TableBank, a new image-based table detection and recognition dataset built with novel weak supervision from Word and Latex documents on the internet. Existing research for image-based table detection and recognition usually fine-tunes pre-trained models on out-of-domain data with a few thousand human-labeled examples, which is difficult to generalize on real-world applications. With TableBank that contains 417K high quality labeled tables, we build several strong baselines using state-of-the-art models with deep neural networks. We make TableBank publicly available and hope it will empower more deep learning approaches in the table detection and recognition task. The dataset and models can be downloaded from https://github.com/doc-analysis/TableBank. Minghao Li 0004, Lei Cui 0001, Shaohan Huang, Furu Wei, Ming Zhou 0001, Zhoujun Li 0001 |
LREC | 4 |
| 2020 | MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersabstractPre-trained language models (e.g., BERT (Devlin et al., 2018) and its variants) have achieved remarkable success in varieties of NLP tasks. However, these models usually consist of hundreds of millions of parameters which brings challenges for fine-tuning and online serving in real-life applications due to latency and capacity constraints. In this work, we present a simple and effective approach to compress large Transformer (Vaswani et al., 2017) based pre-trained models, termed as deep self-attention distillation. The small model (student) is trained by deeply mimicking the self-attention module, which plays a vital role in Transformer networks, of the large model (teacher). Specifically, we propose distilling the self-attention module of the last Transformer layer of the teacher, which is effective and flexible for the student. Furthermore, we introduce the scaled dot-product between values in the self-attention module as the new deep self-attention knowledge, in addition to the attention distributions (i.e., the scaled dot-product of queries and keys) that have been used in existing works. Moreover, we show that introducing a teacher assistant (Mirzadeh et al., 2019) also helps the distillation of large pre-trained Transformer models. Experimental results demonstrate that our monolingual model outperforms state-of-the-art baselines in different parameter size of student models. In particular, it retains more than 99% accuracy on SQuAD 2.0 and several GLUE benchmark tasks using 50% of the Transformer parameters and computations of the teacher model. We also obtain competitive results in applying deep self-attention distillation to multilingual pre-trained models. Wenhui Wang 0003, Furu Wei, Li Dong 0004, Hangbo Bao, Nan Yang 0002, Ming Zhou 0001 |
NeurIPS | 2 |
| 2020 | BERT Loses Patience: Fast and Robust Inference with Early ExitabstractIn this paper, we propose Patience-based Early Exit, a straightforward yet effective inference method that can be used as a plug-and-play technique to simultaneously improve the efficiency and robustness of a pretrained language model (PLM). To achieve this, our approach couples an internal-classifier with each layer of a PLM and dynamically stops inference when the intermediate predictions of the internal classifiers do not change for a pre-defined number of steps. Our approach improves inference efficiency as it allows the model to make a prediction with fewer layers. Meanwhile, experimental results with an ALBERT model show that our method can improve the accuracy and robustness of the model by preventing it from overthinking and exploiting multiple classifiers for prediction, yielding a better accuracy-speed trade-off compared to existing early exit methods. Wangchunshu Zhou, Canwen Xu, Tao Ge 0001, Julian J. McAuley, Ke Xu 0001, Furu Wei |
NeurIPS | 6 |
| 2020 | A Joint Sentence Scoring and Selection Framework for Neural Extractive Document SummarizationabstractExtractive document summarization methods aim to extract important sentences to form a summary. Previous works perform this task by first scoring all sentences in the document then selecting most informative ones; while we propose to jointly learn the two steps with a novel end-to-end neural network framework. Specifically, the sentences in the input document are represented as real-valued vectors through a neural document encoder. Then the method builds the output summary by extracting important sentences one by one. Different from previous works, the proposed joint sentence scoring and selection framework directly predicts the relative sentence importance score according to both sentence content and previously selected sentences. We evaluate the proposed framework with two realizations: a hierarchical recurrent neural network based model; and a pre-training based model that uses BERT as the document encoder. Experiments on two datasets show that the proposed joint framework outperforms the state-of-the-art extractive summarization models which treat sentence scoring and selection as two subtasks. Qingyu Zhou, Nan Yang 0002, Furu Wei, Shaohan Huang, Ming Zhou 0001, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Response Generation by Context-Aware Prototype EditingabstractOpen domain response generation has achieved remarkable progress in recent years, but sometimes yields short and uninformative responses. We propose a new paradigm, prototypethen-edit for response generation, that first retrieves a prototype response from a pre-defined index and then edits the prototype response according to the differences between the prototype context and current context. Our motivation is that the retrieved prototype provides a good start-point for generation because it is grammatical and informative, and the post-editing process further improves the relevance and coherence of the prototype. In practice, we design a contextaware editing model that is built upon an encoder-decoder framework augmented with an editing vector. We first generate an edit vector by considering lexical differences between a prototype context and current context. After that, the edit vector and the prototype response representation are fed to a decoder to generate a new response. Experiment results on a large scale dataset demonstrate that our new paradigm significantly increases the relevance, diversity and originality of generation results, compared to traditional generative models. Furthermore, our model outperforms retrieval-based methods in terms of relevance and originality. Yu Wu 0012, Furu Wei, Shaohan Huang, Yunli Wang, Zhoujun Li 0001, Ming Zhou 0001 |
AAAI | 2 |
| 2019 | Read + Verify: Machine Reading Comprehension with Unanswerable QuestionsabstractMachine reading comprehension with unanswerable questions aims to abstain from answering when no answer can be inferred. In addition to extract answers, previous works usually predict an additional “no-answer” probability to detect unanswerable cases. However, they fail to validate the answerability of the question by verifying the legitimacy of the predicted answer. To address this problem, we propose a novel read-then-verify system, which not only utilizes a neural reader to extract candidate answers and produce no-answer probabilities, but also leverages an answer verifier to decide whether the predicted answer is entailed by the input snippets. Moreover, we introduce two auxiliary losses to help the reader better handle answer extraction as well as no-answer detection, and investigate three different architectures for the answer verifier. Our experiments on the SQuAD 2.0 dataset show that our system obtains a score of 74.2 F1 on test set, achieving state-of-the-art results at the time of submission (Aug. 28th, 2018). Furu Wei, Yuxing Peng 0001, Zhen Huang 0006, Nan Yang 0002, Dongsheng Li 0001 |
AAAI | 2 |
| 2019 | Dictionary-Guided Editing Networks for Paraphrase GenerationabstractAn intuitive way for a human to write paraphrase sentences is to replace words or phrases in the original sentence with their corresponding synonyms and make necessary changes to ensure the new sentences are fluent and grammatically correct. We propose a novel approach to modeling the process with dictionary-guided editing networks which effectively conduct rewriting on the source sentence to generate paraphrase sentences. It jointly learns the selection of the appropriate word level and phrase level paraphrase pairs in the context of the original sentence from an off-the-shelf dictionary as well as the generation of fluent natural language sentences. Specifically, the system retrieves a set of word level and phrase level paraphrase pairs derived from the Paraphrase Database (PPDB) for the original sentence, which is used to guide the decision of which the words might be deleted or inserted with the soft attention mechanism under the sequence-to-sequence framework. We conduct experiments on two benchmark datasets for paraphrase generation, namely the MSCOCO and Quora dataset. The automatic evaluation results demonstrate that our dictionary-guided editing networks outperforms the baseline methods. On human evaluation, results indicate that the generated paraphrases are grammatically correct and relevant to the input sentence. Shaohan Huang, Yu Wu 0012, Furu Wei, Zhongzhi Luan |
AAAI | 3 |
| 2019 | LiveBot: Generating Live Video Comments Based on Visual and Textual ContextsabstractWe introduce the task of automatic live commenting. Live commenting, which is also called “video barrage”, is an emerging feature on online video sites that allows real-time comments from viewers to fly across the screen like bullets or roll at the right side of the screen. The live comments are a mixture of opinions for the video and the chit chats with other comments. Automatic live commenting requires AI agents to comprehend the videos and interact with human viewers who also make the comments, so it is a good testbed of an AI agent’s ability to deal with both dynamic vision and language. In this work, we construct a large-scale live comment dataset with 2,361 videos and 895,929 live comments. Then, we introduce two neural models to generate live comments based on the visual and textual contexts, which achieve better performance than previous neural baselines such as the sequence-to-sequence model. Finally, we provide a retrieval-based evaluation protocol for automatic live commenting where the model is asked to sort a set of candidate comments based on the log-likelihood score, and evaluated on metrics such as mean-reciprocal-rank. Putting it all together, we demonstrate the first “LiveBot”. The datasets and the codes can be found at https://github.com/lancopku/livebot. Shuming Ma, Lei Cui 0001, Damai Dai, Furu Wei, Xu Sun 0001 |
AAAI | 4 |
| 2019 | Automatic Grammatical Error Correction for Sequence-to-sequence Text Generation: An Empirical StudyabstractSequence-to-sequence (seq2seq) models have achieved tremendous success in text generation tasks.However, there is no guarantee that they can always generate sentences without grammatical errors.In this paper, we present a preliminary empirical study on whether and how much automatic grammatical error correction can help improve seq2seq text generation.We conduct experiments across various seq2seq text generation tasks including machine translation, formality style transfer, sentence compression and simplification.Experiments show the state-of-the-art grammatical error correction system can improve the grammaticality of generated text and can bring taskoriented improvements in the tasks where target sentences are in a formal style. Tao Ge 0001, Xingxing Zhang 0002, Furu Wei, Ming Zhou 0001 |
ACL (1) | 3 |
| 2019 | HIBERT: Document Level Pre-training of Hierarchical Bidirectional Transformers for Document SummarizationabstractNeural extractive summarization models usually employ a hierarchical encoder for document encoding and they are trained using sentence-level labels, which are created heuristically using rule-based methods.Training the hierarchical encoder with these inaccurate labels is challenging.Inspired by the recent work on pre-training transformer sentence encoders (Devlin et al., 2018), we propose HIBERT (as shorthand for HIerachical Bidirectional Encoder Representations from Transformers) for document encoding and a method to pre-train it using unlabeled data.We apply the pre-trained HIBERT to our summarization model and it outperforms its randomly initialized counterpart by 1.25 ROUGE on the CNN/Dailymail dataset and by 2.0 ROUGE on a version of New York Times dataset.We also achieve the state-of-the-art performance on these two datasets. Xingxing Zhang 0002, Furu Wei, Ming Zhou 0001 |
ACL (1) | 2 |
| 2019 | BERT-based Lexical SubstitutionabstractPrevious studies on lexical substitution tend to obtain substitute candidates by finding the target word's synonyms from lexical resources (e.g., WordNet) and then rank the candidates based on its contexts.These approaches have two limitations: (1) They are likely to overlook good substitute candidates that are not the synonyms of the target words in the lexical resources;(2) They fail to take into account the substitution's influence on the global context of the sentence.To address these issues, we propose an end-toend BERT-based lexical substitution approach which can propose and validate substitute candidates without using any annotated data or manually curated resources.Our approach first applies dropout to the target word's embedding for partially masking the word, allowing BERT to take balanced consideration of the target word's semantics and contexts for proposing substitute candidates, and then validates the candidates based on their substitution's influence on the global contextualized representation of the sentence.Experiments show our approach performs well in both proposing and ranking substitute candidates, achieving the state-of-the-art results in both LS07 and LS14 benchmarks. Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Furu Wei, Ming Zhou 0001 |
ACL (1) | 4 |
| 2019 | Retrieval-Enhanced Adversarial Training for Neural Response GenerationabstractDialogue systems are usually built on either generation-based or retrieval-based approaches, yet they do not benefit from the advantages of different models.In this paper, we propose a Retrieval-Enhanced Adversarial Training (REAT) method for neural response generation.Distinct from existing approaches, the REAT method leverages an encoder-decoder framework in terms of an adversarial training paradigm, while taking advantage of N-best response candidates from a retrieval-based system to construct the discriminator.An empirical study on a large scale public available benchmark dataset shows that the REAT method significantly outperforms the vanilla Seq2Seq model as well as the conventional adversarial training approach. Qingfu Zhu, Lei Cui 0001, Weinan Zhang 0003, Furu Wei, Ting Liu 0001 |
ACL (1) | 4 |
| 2019 | Learning to Ask Unanswerable Questions for Machine Reading ComprehensionabstractMachine reading comprehension with unanswerable questions is a challenging task.In this work, we propose a data augmentation technique by automatically generating relevant unanswerable questions according to an answerable question paired with its corresponding paragraph that contains the answer.We introduce a pair-to-sequence model for unanswerable question generation, which effectively captures the interactions between the question and the paragraph.We also present a way to construct training data for our question generation models by leveraging the existing reading comprehension dataset.Experimental results show that the pair-to-sequence model performs consistently better compared with the sequence-to-sequence baseline.We further use the automatically generated unanswerable questions as a means of data augmentation on the SQuAD 2.0 dataset, yielding 1.9 absolute F1 improvement with BERT-base model and 1.7 absolute F1 improvement with BERT-large model. Li Dong 0004, Furu Wei, Wenhui Wang 0003, Bing Qin 0001, Ting Liu 0001 |
ACL (1) | 3 |
| 2019 | Visualizing and Understanding the Effectiveness of BERTabstractYaru Hao, Li Dong, Furu Wei, Ke Xu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yaru Hao, Li Dong 0004, Furu Wei, Ke Xu 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Video Dialog via Progressive Inference and Cross-TransformerabstractWeike Jin, Zhou Zhao, Mao Gu, Jun Xiao, Furu Wei, Yueting Zhuang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Weike Jin, Zhou Zhao 0001, Mao Gu, Jun Xiao 0001, Furu Wei, Yueting Zhuang |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Unified Language Model Pre-training for Natural Language Understanding and GenerationabstractThis paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks. The model is pre-trained using three types of language modeling tasks: unidirectional, bidirectional, and sequence-to-sequence prediction. The unified modeling is achieved by employing a shared Transformer network and utilizing specific self-attention masks to control what context the prediction conditions on. UniLM compares favorably with BERT on the GLUE benchmark, and the SQuAD 2.0 and CoQA question answering tasks. Moreover, UniLM achieves new state-of-the-art results on five natural language generation datasets, including improving the CNN/DailyMail abstractive summarization ROUGE-L to 40.51 (2.04 absolute improvement), the Gigaword abstractive summarization ROUGE-L to 35.75 (0.86 absolute improvement), the CoQA generative question answering F1 score to 82.5 (37.1 absolute improvement), the SQuAD question generation BLEU-4 to 22.12 (3.75 absolute improvement), and the DSTC7 document-grounded dialog response generation NIST-4 to 2.67 (human performance is 2.65). The code and pre-trained models are available at https://github.com/microsoft/unilm. Li Dong 0004, Nan Yang 0002, Wenhui Wang 0003, Furu Wei, Xiaodong Liu 0003, Yu Wang 0009, Jianfeng Gao 0001, Ming Zhou 0001, Hsiao-Wuen Hon |
NeurIPS | 4 |
| 2019 | Neural Melody Composition from Lyrics
Hangbo Bao, Shaohan Huang, Furu Wei, Lei Cui 0001, Yu Wu 0012, Chuanqi Tan, Ming Zhou 0001 |
NLPCC (1) | 3 |
| 2019 | Document-Based Question Answering Improves Query-Focused Multi-document Summarization
Weikang Li, Xingxing Zhang 0002, Yunfang Wu, Furu Wei, Ming Zhou 0001 |
NLPCC (2) | 4 |
| 2018 | Faithful to the Original: Fact Aware Neural Abstractive SummarizationabstractUnlike extractive summarization, abstractive summarization has to fuse different parts of the source text, which inclines to create fake facts. Our preliminary study reveals nearly 30% of the outputs from a state-of-the-art neural summarization system suffer from this problem. While previous abstractive summarization approaches usually focus on the improvement of informativeness, we argue that faithfulness is also a vital prerequisite for a practical abstractive summarization system. To avoid generating fake facts in a summary, we leverage open information extraction and dependency parse technologies to extract actual fact descriptions from the source text. The dual-attention sequence-to-sequence framework is then proposed to force the generation conditioned on both the source text and the extracted fact descriptions. Experiments on the Gigaword benchmark dataset demonstrate that our model can greatly reduce fake summaries by 80%. Notably, the fact descriptions also bring significant improvement on informativeness since they often condense the meaning of the source text. Ziqiang Cao, Furu Wei, Wenjie Li 0002, Sujian Li |
AAAI | 2 |
| 2018 | S-Net: From Answer Extraction to Answer Synthesis for Machine Reading ComprehensionabstractIn this paper, we present a novel approach to machine reading comprehension for the MS-MARCO dataset. Unlike the SQuAD dataset that aims to answer a question with exact text spans in a passage, the MS-MARCO dataset defines the task as answering a question from multiple passages and the words in the answer are not necessary in the passages. We therefore develop an extraction-then-synthesis framework to synthesize answers from extraction results. Specifically, the answer extraction model is first employed to predict the most important sub-spans from the passage as evidence, and the answer synthesis model takes the evidence as additional features along with the question and passage to further elaborate the final answers. We build the answer extraction model with state-of-the-art neural networks for single passage reading comprehension, and propose an additional task of passage ranking to help answer extraction in multiple passages. The answer synthesis model is based on the sequence-to-sequence neural networks with extracted evidences as features. Experiments show that our extraction-then-synthesis method outperforms state-of-the-art methods. Chuanqi Tan, Furu Wei, Nan Yang 0002, Bowen Du 0001, Weifeng Lv, Ming Zhou 0001 |
AAAI | 2 |
| 2018 | Sequential Copying NetworksabstractCopying mechanism shows effectiveness in sequence-to-sequence based neural network models for text generation tasks, such as abstractive sentence summarization and question generation. However, existing works on modeling copying or pointing mechanism only considers single word copying from the source sentences. In this paper, we propose a novel copying framework, named Sequential Copying Networks (SeqCopyNet), which not only learns to copy single words, but also copies sequences from the input sentence. It leverages the pointer networks to explicitly select a sub-span from the source side to target side, and integrates this sequential copying mechanism to the generation process in the encoder-decoder paradigm. Experiments on abstractive sentence summarization and question generation tasks show that the proposed SeqCopyNet can copy meaningful spans and outperforms the baseline models. Qingyu Zhou, Nan Yang 0002, Furu Wei, Ming Zhou 0001 |
AAAI | 3 |
| 2018 | Hierarchical Attention Flow for Multiple-Choice Reading ComprehensionabstractIn this paper, we focus on multiple-choice reading comprehension which aims to answer a question given a passage and multiple candidate options. We present the hierarchical attention flow to adequately leverage candidate options to model the interactions among passages, questions and candidate options. We observe that leveraging candidate options to boost evidence gathering from the passages play a vital role in this task, which is ignored in previous works. In addition, we explicitly model the option correlations with attention mechanism to obtain better option representations, which are further fed into a bilinear layer to obtain the ranking score for each option. On a large-scale multiple-choice reading comprehension dataset (i.e. the RACE dataset), the proposed model outperforms two previous neural network baselines on both RACE-M and RACE-H subsets and yields the state-of-the-art overall results. Furu Wei, Bing Qin 0001, Ting Liu 0001 |
AAAI | 2 |
| 2018 | Retrieve, Rerank and Rewrite: Soft Template Based Neural SummarizationabstractMost previous seq2seq summarization systems purely depend on the source text to generate summaries, which tends to work unstably.Inspired by the traditional template-based summarization approaches, this paper proposes to use existing summaries as soft templates to guide the seq2seq model.To this end, we use a popular IR platform to Retrieve proper summaries as candidate templates.Then, we extend the seq2seq framework to jointly conduct template Reranking and templateaware summary generation (Rewriting).Experiments show that, in terms of informativeness, our model significantly outperforms the state-of-the-art methods, and even soft templates themselves demonstrate high competitiveness.In addition, the import of high-quality external summaries improves the stability and readability of generated summaries. Ziqiang Cao, Wenjie Li 0002, Sujian Li, Furu Wei |
ACL (1) | 4 |
| 2018 | Neural Document Summarization by Jointly Learning to Score and Select SentencesabstractSentence scoring and sentence selection are two main steps in extractive document summarization systems.However, previous works treat them as two separated subtasks.In this paper, we present a novel end-to-end neural network framework for extractive document summarization by jointly learning to score and select sentences.It first reads the document sentences with a hierarchical encoder to obtain the representation of sentences.Then it builds the output summary by extracting sentences one by one.Different from previous methods, our approach integrates the selection strategy into the scoring model, which directly predicts the relative importance given previously selected sentences.Experiments on the CNN/Daily Mail dataset show that the proposed framework significantly outperforms the state-of-the-art extractive summarization models. Qingyu Zhou, Nan Yang 0002, Furu Wei, Shaohan Huang, Ming Zhou 0001, Tiejun Zhao |
ACL (1) | 3 |
| 2018 | Fluency Boost Learning and Inference for Neural Grammatical Error CorrectionabstractMost of the neural sequence-to-sequence (seq2seq) models for grammatical error correction (GEC) have two limitations: (1) a seq2seq model may not be well generalized with only limited error-corrected data; (2) a seq2seq model may fail to completely correct a sentence with multiple errors through normal seq2seq inference.We attempt to address these limitations by proposing a fluency boost learning and inference mechanism.Fluency boosting learning generates fluency-boost sentence pairs during training, enabling the error correction model to learn how to improve a sentence's fluency from more instances, while fluency boosting inference allows the model to correct a sentence incrementally through multi-round seq2seq inference until the sentence's fluency stops increasing.Experiments show our approaches improve the performance of seq2seq models for GEC, achieving state-of-the-art results on both CoNLL-2014 and JFLEG benchmark datasets. Tao Ge 0001, Furu Wei, Ming Zhou 0001 |
ACL (1) | 2 |
| 2018 | Fine-grained Coordinated Cross-lingual Text Stream Alignment for Endless Language Knowledge AcquisitionabstractThis paper proposes to study fine-grained coordinated cross-lingual text stream alignment through a novel information network decipherment paradigm.We use Burst Information Networks as media to represent text streams and present a simple yet effective network decipherment algorithm with diverse clues to decipher the networks for accurate text stream alignment.Experiments on Chinese-English news streams show our approach not only outperforms previous approaches on bilingual lexicon extraction from coordinated text streams but also can harvest high-quality alignments from large amounts of streaming data for endless language knowledge mining, which makes it promising to be a new paradigm for automatic language knowledge acquisition. Tao Ge 0001, Qing Dou, Heng Ji 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001 |
EMNLP | 7 |
| 2018 | Attention-Guided Answer Distillation for Machine Reading ComprehensionabstractDespite that current reading comprehension systems have achieved significant advancements, their promising performances are often obtained at the cost of making an ensemble of numerous models. Besides, existing approaches are also vulnerable to adversarial attacks. This paper tackles these problems by leveraging knowledge distillation, which aims to transfer knowledge from an ensemble model to a single model. We first demonstrate that vanilla knowledge distillation applied to answer span prediction is effective for reading comprehension systems. We then propose two novel approaches that not only penalize the prediction on confusing answers but also guide the training with alignment information distilled from the ensemble. Experiments show that our best student model has only a slight drop of 0.4% F1 on the SQuAD test set compared to the ensemble teacher, while running 12x faster during inference. It even outperforms the teacher on adversarial SQuAD datasets and NarrativeQA benchmark. Yuxing Peng 0001, Furu Wei, Zhen Huang 0006, Dongsheng Li 0001, Nan Yang 0002, Ming Zhou 0001 |
EMNLP | 3 |
| 2018 | Neural Latent Extractive Document SummarizationabstractExtractive summarization models require sentence-level labels, which are usually created heuristically (e.g., with rule-based methods) given that most summarization datasets only have document-summary pairs.Since these labels might be suboptimal, we propose a latent variable extractive model where sentences are viewed as latent variables and sentences with activated variables are used to infer gold summaries.During training the loss comes directly from gold summaries.Experiments on the CNN/Dailymail dataset show that our model improves over a strong extractive baseline trained on heuristically approximated labels and also performs competitively to several recent models. Xingxing Zhang 0002, Mirella Lapata, Furu Wei, Ming Zhou 0001 |
EMNLP | 3 |
| 2018 | Attention-Fused Deep Matching Network for Natural Language InferenceabstractNatural language inference aims to predict whether a premise sentence can infer another hypothesis sentence. Recent progress on this task only relies on a shallow interaction between sentence pairs, which is insufficient for modeling complex relations. In this paper, we present an attention-fused deep matching network (AF-DMN) for natural language inference. Unlike existing models, AF-DMN takes two sentences as input and iteratively learns the attention-aware representations for each side by multi-level interactions. Moreover, we add a self-attention mechanism to fully exploit local context information within each sentence. Experiment results show that AF-DMN achieves state-of-the-art performance and outperforms strong baselines on Stanford natural language inference (SNLI), multi-genre natural language inference (MultiNLI), and Quora duplicate questions datasets. Chaoqun Duan, Lei Cui 0001, Xinchi Chen, Furu Wei, Conghui Zhu, Tiejun Zhao |
IJCAI | 4 |
| 2018 | Reinforced Mnemonic Reader for Machine Reading ComprehensionabstractIn this paper, we introduce the Reinforced Mnemonic Reader for machine reading comprehension tasks, which enhances previous attentive readers in two aspects. First, a reattention mechanism is proposed to refine current attentions by directly accessing to past attentions that are temporally memorized in a multi-round alignment architecture, so as to avoid the problems of attention redundancy and attention deficiency. Second, a new optimization approach, called dynamic-critical reinforcement learning, is introduced to extend the standard supervised method. It always encourages to predict a more acceptable answer so as to address the convergence suppression problem occurred in traditional reinforcement learning algorithms. Extensive experiments on the Stanford Question Answering Dataset (SQuAD) show that our model achieves state-of-the-art results. Meanwhile, our model outperforms previous systems by over 6% in terms of both Exact Match and F1 metrics on two adversarial SQuAD datasets. Yuxing Peng 0001, Zhen Huang 0006, Xipeng Qiu, Furu Wei, Ming Zhou 0001 |
IJCAI | 5 |
| 2018 | Multiway Attention Networks for Modeling Sentence PairsabstractModeling sentence pairs plays the vital role for judging the relationship between two sentences, such as paraphrase identification, natural language inference, and answer sentence selection. Previous work achieves very promising results using neural networks with attention mechanism. In this paper, we propose the multiway attention networks which employ multiple attention functions to match sentence pairs under the matching-aggregation framework. Specifically, we design four attention functions to match words in corresponding sentences. Then, we aggregate the matching information from each function, and combine the information from all functions to obtain the final representation. Experimental results demonstrate that the proposed multiway attention networks improve the result on the Quora Question Pairs, SNLI, MultiNLI, and answer sentence selection task on the SQuAD dataset. Chuanqi Tan, Furu Wei, Wenhui Wang 0003, Weifeng Lv, Ming Zhou 0001 |
IJCAI | 2 |
| 2018 | XiaoIce Band: A Melody and Arrangement Generation Framework for Pop MusicabstractWith the development of knowledge of music composition and the recent increase in demand, an increasing number of companies and research institutes have begun to study the automatic generation of music. However, previous models have limitations when applying to song generation, which requires both the melody and arrangement. Besides, many critical factors related to the quality of a song such as chord progression and rhythm patterns are not well addressed. In particular, the problem of how to ensure the harmony of multi-track music is still underexplored. To this end, we present a focused study on pop music generation, in which we take both chord and rhythm influence of melody generation and the harmony of music arrangement into consideration. We propose an end-to-end melody and arrangement generation framework, called XiaoIce Band, which generates a melody track with several accompany tracks played by several types of instruments. Specifically, we devise a Chord based Rhythm and Melody Cross-Generation Model (CRMCG) to generate melody with chord progressions. Then, we propose a Multi-Instrument Co-Arrangement Model (MICA) using multi-task learning for multi-track music arrangement. Finally, we conduct extensive experiments on a real-world dataset, where the results demonstrate the effectiveness of XiaoIce Band. Hongyuan Zhu 0001, Qi Liu 0003, Nicholas Jing Yuan, Chuan Qin 0002, Kun Zhang 0015, Guang Zhou, Furu Wei, Yuanchun Xu, Enhong Chen |
KDD | 8 |
| 2018 | EventWiki: A Knowledge Base of Major Events
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001 |
LREC | 5 |
| 2018 | SeRI: A Dataset for Sub-event Relation Inference from an Encyclopedia
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001 |
NLPCC (2) | 5 |
| 2018 | I Know There Is No Answer: Modeling Answer Validation for Machine Reading Comprehension
Chuanqi Tan, Furu Wei, Qingyu Zhou, Nan Yang 0002, Weifeng Lv, Ming Zhou 0001 |
NLPCC (1) | 2 |
| 2018 | Context-Aware Answer Sentence Selection With Hierarchical Gated Recurrent Neural NetworksabstractIn this paper, we study the task of reading comprehension style answer sentence selection that aims to select the best sentence from a given passage to answer a question. Unlike most previous works that match the question and each candidate sentence separately, we observe that the context information among sentences in the same passage plays a vital role in this task. We propose modeling context information with hierarchical gated recurrent neural networks. Specifically, we first apply a word level recurrent neural network to model the context independent matching between the question and each candidate sentence. We then employ a sentence level recurrent neural network to incorporate the context information among all candidate sentences. Moreover, we introduce the gate mechanism to select matching information before feeding into recurrent neural networks at both word and sentence level. Experiments on the WikiQA and SQuAD datasets show that our model outperforms state-of-the-art methods. Chuanqi Tan, Furu Wei, Qingyu Zhou, Nan Yang 0002, Bowen Du 0001, Weifeng Lv, Ming Zhou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Sentence Relations for Extractive Summarization with Deep Neural NetworksabstractSentence regression is a type of extractive summarization that achieves state-of-the-art performance and is commonly used in practical systems. The most challenging task within the sentence regression framework is to identify discriminative features to represent each sentence. In this article, we study the use of sentence relations, e.g., Contextual Sentence Relations (CSR), Title Sentence Relations (TSR), and Query Sentence Relations (QSR), so as to improve the performance of sentence regression. CSR, TSR, and QSR refer to the relations between a main body sentence and its local context, its document title, and a given query, respectively. We propose a deep neural network model, Sentence Relation-based Summarization (SRSum), that consists of five sub-models, PriorSum, CSRSum, TSRSum, QSRSum, and SFSum. PriorSum encodes the latent semantic meaning of a sentence using a bi-gram convolutional neural network. SFSum encodes the surface information of a sentence, e.g., sentence length, sentence position, and so on. CSRSum, TSRSum, and QSRSum are three sentence relation sub-models corresponding to CSR, TSR, and QSR, respectively. CSRSum evaluates the ability of each sentence to summarize its local contexts. Specifically, CSRSum applies a CSR-based word-level and sentence-level attention mechanism to simulate the context-aware reading of a human reader, where words and sentences that have anaphoric relations or local summarization abilities are easily remembered and paid attention to. TSRSum evaluates the semantic closeness of each sentence with respect to its title, which usually reflects the main ideas of a document. TSRSum applies a TSR-based attention mechanism to simulate people’s reading ability with the main idea (title) in mind. QSRSum evaluates the relevance of each sentence with given queries for the query-focused summarization. QSRSum applies a QSR-based attention mechanism to simulate the attentive reading of a human reader with some queries in mind. The mechanism can recognize which parts of the given queries are more likely answered by a sentence under consideration. Finally as a whole, SRSum automatically learns useful latent features by jointly learning representations of query sentences, content sentences, and title sentences as well as their relations. We conduct extensive experiments on six benchmark datasets, including generic multi-document summarization and query-focused multi-document summarization. On both tasks, SRSum achieves comparable or superior performance compared with state-of-the-art approaches in terms of multiple ROUGE metrics. Pengjie Ren, Zhumin Chen, Zhaochun Ren, Furu Wei, Liqiang Nie, Jun Ma 0001, Maarten de Rijke |
ACM Trans. Inf. Syst. | 4 |
| 2017 | Improving Multi-Document Summarization via Text ClassificationabstractDeveloped so far, multi-document summarization has reached its bottleneck due to the lack of sufficient training data and diverse categories of documents. Text classification just makes up for these deficiencies. In this paper, we propose a novel summarization system called TCSum, which leverages plentiful text classification data to improve the performance of multi-document summarization. TCSum projects documents onto distributed representations which act as a bridge between text classification and summarization. It also utilizes the classification results to produce summaries of different styles. Extensive experiments on DUC generic multi-document summarization datasets show that, TCSum can achieve the state-of-the-art performance without using any hand-crafted features and has the capability to catch the variations of summary styles with respect to different text categories. Ziqiang Cao, Wenjie Li 0002, Sujian Li, Furu Wei |
AAAI | 4 |
| 2017 | Gated Self-Matching Networks for Reading Comprehension and Question AnsweringabstractIn this paper, we present the gated selfmatching networks for reading comprehension style question answering, which aims to answer questions from a given passage.We first match the question and passage with gated attention-based recurrent networks to obtain the question-aware passage representation.Then we propose a self-matching attention mechanism to refine the representation by matching the passage against itself, which effectively encodes information from the whole passage.We finally employ the pointer networks to locate the positions of answers from the passages.We conduct extensive experiments on the SQuAD dataset.The single model achieves 71.3% on the evaluation metrics of exact match on the hidden test set, while the ensemble model further boosts the results to 75.9%.At the time of submission of the paper, our model holds the first place on the SQuAD leaderboard for both single and ensemble model. Wenhui Wang 0003, Nan Yang 0002, Furu Wei, Baobao Chang, Ming Zhou 0001 |
ACL (1) | 3 |
| 2017 | Selective Encoding for Abstractive Sentence SummarizationabstractWe propose a selective encoding model to extend the sequence-to-sequence framework for abstractive sentence summarization.It consists of a sentence encoder, a selective gate network, and an attention equipped decoder.The sentence encoder and decoder are built with recurrent neural networks.The selective gate network constructs a second level sentence representation by controlling the information flow from encoder to decoder.The second level representation is tailored for sentence summarization task, which leads to better performance.We evaluate our model on the English Gigaword, DUC 2004 and MSR abstractive sentence summarization datasets.The experimental results show that the proposed selective encoding model outperforms the state-ofthe-art baseline models. Qingyu Zhou, Nan Yang 0002, Furu Wei, Ming Zhou 0001 |
ACL (1) | 3 |
| 2017 | Learning to Generate Product Reviews from AttributesabstractLi Dong, Shaohan Huang, Furu Wei, Mirella Lapata, Ming Zhou, Ke Xu. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Li Dong 0004, Shaohan Huang, Furu Wei, Mirella Lapata, Ming Zhou 0001, Ke Xu 0001 |
EACL (1) | 3 |
| 2017 | Entity Linking for Queries by Searching Wikipedia SentencesabstractWe present a simple yet effective approach for linking entities in queries.The key idea is to search sentences similar to a query from Wikipedia articles and directly use the human-annotated entities in the similar sentences as candidate entities for the query.Then, we employ a rich set of features, such as link-probability, contextmatching, word embeddings, and relatedness among candidate entities as well as their related entities, to rank the candidates under a regression based framework.The advantages of our approach lie in two aspects, which contribute to the ranking process and final linking result.First, it can greatly reduce the number of candidate entities by filtering out irrelevant entities with the words in the query.Second, we can obtain the query sensitive prior probability in addition to the static linkprobability derived from all Wikipedia articles.We conduct experiments on two benchmark datasets on entity linking for queries, namely the ERD14 dataset and the GERDAQ dataset.Experimental results show that our method outperforms state-of-the-art systems and yields 75.0% in F1 on the ERD14 dataset and 56.9% on the GERDAQ dataset. Chuanqi Tan, Furu Wei, Pengjie Ren, Weifeng Lv, Ming Zhou 0001 |
EMNLP | 2 |
| 2017 | Neural Question Generation from Text: A Preliminary Study
Qingyu Zhou, Nan Yang 0002, Furu Wei, Chuanqi Tan, Hangbo Bao, Ming Zhou 0001 |
NLPCC | 3 |
| 2017 | Leveraging Contextual Sentence Relations for Extractive Summarization Using a Neural Attention ModelabstractAs a framework for extractive summarization, sentence regression has achieved state-of-the-art performance in several widely-used practical systems. The most challenging task within the sentence regression framework is to identify discriminative features to encode a sentence into a feature vector. So far, sentence regression approaches have neglected to use features that capture contextual relations among sentences. Pengjie Ren, Zhumin Chen, Zhaochun Ren, Furu Wei, Jun Ma 0001, Maarten de Rijke |
SIGIR | 4 |
| 2016 | TGSum: Build Tweet Guided Multi-Document Summarization DatasetabstractThe development of summarization research has been significantly hampered by the costly acquisition of reference summaries. This paper proposes an effective way to automatically collect large scales of news-related multi-document summaries with reference to social media's reactions. We utilize two types of social labels in tweets, i.e., hashtags and hyper-links. Hashtags are used to cluster documents into different topic sets. Also, a tweet with a hyper-link often highlights certain key points of the corresponding document. We synthesize a linked document cluster to form a reference summary which can cover most key points. To this aim, we adopt the ROUGE metrics to measure the coverage ratio, and develop an Integer Linear Programming solution to discover the sentence set reaching the upper bound of ROUGE. Since we allow summary sentences to be selected from both documents and high-quality tweets, the generated reference summaries could be abstractive. Both informativeness and readability of the collected summaries are verified by manual judgment. In addition, we train a Support Vector Regression summarizer on DUC generic multi-document summarization benchmarks. With the collected data as extra training resource, the performance of the summarizer improves a lot on all the test sets. We release this dataset for further research. Ziqiang Cao, Chengyao Chen, Wenjie Li 0002, Sujian Li, Furu Wei, Ming Zhou 0001 |
AAAI | 5 |
| 2016 | AttSum: Joint Learning of Focusing and Summarization with Neural AttentionabstractQuery relevance ranking and sentence saliency ranking are the two main tasks in extractive query-focused summarization. Previous supervised summarization systems often perform the two tasks in isolation. However, since reference summaries are the trade-off between relevance and saliency, using them as supervision, neither of the two rankers could be trained well. This paper proposes a novel summarization system called AttSum, which tackles the two tasks jointly. It automatically learns distributed representations for sentences as well as the document cluster. Meanwhile, it applies the attention mechanism to simulate the attentive reading of human behavior when a query is given. Extensive experiments are conducted on DUC query-focused summarization benchmark datasets. Without using any hand-crafted features, AttSum achieves competitive performance. We also observe that the sentences recognized to focus on the query indeed meet the query need. Ziqiang Cao, Wenjie Li 0002, Sujian Li, Furu Wei, Yanran Li |
COLING | 4 |
| 2016 | A Redundancy-Aware Sentence Regression Framework for Extractive SummarizationabstractExisting sentence regression methods for extractive summarization usually model sentence importance and redundancy in two separate processes. They first evaluate the importance f(s) of each sentence s and then select sentences to generate a summary based on both the importance scores and redundancy among sentences. In this paper, we propose to model importance and redundancy simultaneously by directly evaluating the relative importance f(s|S) of a sentence s given a set of selected sentences S. Specifically, we present a new framework to conduct regression with respect to the relative gain of s given S calculated by the ROUGE metric. Besides the single sentence features, additional features derived from the sentence relations are incorporated. Experiments on the DUC 2001, 2002 and 2004 multi-document summarization datasets show that the proposed method outperforms state-of-the-art extractive summarization approaches. Pengjie Ren, Furu Wei, Zhumin Chen, Jun Ma 0001, Ming Zhou 0001 |
COLING | 2 |
| 2016 | Solving and Generating Chinese Character RiddlesabstractChinese character riddle is a riddle game in which the riddle solution is a single Chinese character.It is closely connected with the shape, pronunciation or meaning of Chinese characters.The riddle description (sentence) is usually composed of phrases with rich linguistic phenomena (such as pun, simile, and metaphor), which are associated to different parts (namely radicals) of the solution character.In this paper, we propose a statistical framework to solve and generate Chinese character riddles.Specifically, we learn the alignments and rules to identify the metaphors between phrases in riddles and radicals in characters.Then, in the solving phase, we utilize a dynamic programming method to combine the identified metaphors to obtain candidate solutions.In the riddle generation phase, we use a template-based method and a replacement-based method to obtain candidate riddle descriptions.We then use Ranking SVM to rerank the candidates both in the solving and generation process.Experimental results in the solving task show that the proposed method outperforms baseline methods.We also get very promising results in the generation task according to human judges. Chuanqi Tan, Furu Wei, Li Dong 0004, Weifeng Lv, Ming Zhou 0001 |
EMNLP | 2 |
| 2016 | Unsupervised Word and Dependency Path Embeddings for Aspect Term Extraction
Yichun Yin, Furu Wei, Li Dong 0004, Kaimeng Xu, Ming Zhang 0004, Ming Zhou 0001 |
IJCAI | 2 |
| 2016 | Adaptive Multi-Compositionality for Recursive Neural Network ModelsabstractRecursive neural network models have achieved promising results in many natural language processing tasks. The main difference among these models lies in the composition function, i.e., how to obtain the vector representation for a phrase or sentence using the representations of words it contains. This paper introduces a novel Adaptive Multi-Compositionality (AdaMC) layer to recursive neural network models. The basic idea is to use more than one composition function and adaptively select them depending on input vectors. We develop a general framework to model the semantic composition as a distribution of these composition functions. The composition functions and parameters used for adaptive selection are jointly learnt from the supervision of specific tasks. We integrate AdaMC into existing recursive neural network models and conduct extensive experiments on the Stanford Sentiment Treebank and semantic relation classification task. The experimental results demonstrate that AdaMC improves the performance of recursive neural network models and outperforms the baseline methods. Li Dong 0004, Furu Wei, Ke Xu 0001, Shixia Liu, Ming Zhou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Relation Classification Via Modeling Augmented Dependency PathsabstractPrevious research on relation classification has verified the effectiveness of using dependency shortest paths or dependency subtrees. How to efficiently unify these two kinds of dependency information in relation classification is still an open problem. In this paper, we propose a novel structure, termed augmented dependency path (ADP), which is composed of the shortest dependency path between two entities and the subtrees attached to the shortest path. To exploit the semantic representation behind the ADP structure, we develop the dependency-based neural networks (DepNN) model which combines the advantages of the recursive neural network (RNN) and the convolutional neural network (CNN). In DepNN, RNN is designed to model the dependency subtrees since it is good at capturing the hierarchical structures. Then, the semantic representation in subtrees is passed to the nodes on the shortest path and CNN is used to get the most important features on the ADP. Experiments on the SemEval-2010 dataset show that the ADP structure including both the shortest dependency path and the attached subtrees is helpful to classify the semantic relations between two entities and our proposed method can achieve the state-of-the-art performance. Yang Liu 0124, Sujian Li, Furu Wei, Heng Ji 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Sentiment Embeddings with Applications to Sentiment AnalysisabstractWe propose learning sentiment-specific word embeddings dubbed sentiment embeddings in this paper. Existing word embedding learning algorithms typically only use the contexts of words but ignore the sentiment of texts. It is problematic for sentiment analysis because the words with similar contexts but opposite sentiment polarity, such asgoodandbad, are mapped to neighboring word vectors. We address this issue by encoding sentiment information of texts (e.g., sentences and words) together with contexts of words in sentiment embeddings. By combining context and sentiment level evidences, the nearest neighbors in sentiment embedding space are semantically similar and it favors words with the same sentiment polarity. In order to learn sentiment embeddings effectively, we develop a number of neural networks with tailoring loss functions, and collect massive texts automatically with sentiment signals like emoticons as the training data. Sentiment embeddings can be naturally used as word features for a variety of sentiment analysis tasks without feature engineering. We apply sentiment embeddings to word-level sentiment analysis, sentence level sentiment classification, and building sentiment lexicons. Experimental results show that sentiment embeddings consistently outperform context-based embeddings on several benchmark datasets of these tasks. This work provides insights on the design of neural networks for learning task-specific word embeddings in other natural language processing tasks. Duyu Tang, Furu Wei, Bing Qin 0001, Nan Yang 0002, Ting Liu 0001, Ming Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2016 | An Uncertainty-Aware Approach for Exploratory Microblog RetrievalabstractAlthough there has been a great deal of interest in analyzing customer opinions and breaking news in microblogs, progress has been hampered by the lack of an effective mechanism to discover and retrieve data of interest from microblogs. To address this problem, we have developed an uncertainty-aware visual analytics approach to retrieve salient posts, users, and hashtags. We extend an existing ranking technique to compute a multifaceted retrieval result: the mutual reinforcement rank of a graph node, the uncertainty of each rank, and the propagation of uncertainty among different graph nodes. To illustrate the three facets, we have also designed a composite visualization with three visual components: a graph visualization, an uncertainty glyph, and a flow map. The graph visualization with glyphs, the flow map, and the uncertainty analysis together enable analysts to effectively find the most uncertain results and interactively refine them. We have applied our approach to several Twitter datasets. Qualitative evaluation and two real-world case studies demonstrate the promise of our approach for retrieving high-quality microblog data. Mengchen Liu, Shixia Liu, Xizhou Zhu, Qinying Liao, Furu Wei, Shimei Pan |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2015 | Ranking with Recursive Neural Networks and Its Application to Multi-Document SummarizationabstractWe develop a Ranking framework upon Recursive Neural Networks (R2N2) to rank sentences for multi-document summarization. It formulates the sentence ranking task as a hierarchical regression process, which simultaneously measures the salience of a sentence and its constituents (e.g., phrases) in the parsing tree. This enables us to draw on word-level to sentence-level supervisions derived from reference summaries.In addition, recursive neural networks are used to automatically learn ranking features over the tree, with hand-crafted feature vectors of words as inputs. Hierarchical regressions are then conducted with learned features concatenating raw features.Ranking scores of sentences and words are utilized to effectively select informative and non-redundant sentences to generate summaries.Experiments on the DUC 2001, 2002 and 2004 multi-document summarization datasets show that R2N2 outperforms state-of-the-art extractive summarization approaches. Ziqiang Cao, Furu Wei, Li Dong 0004, Sujian Li, Ming Zhou 0001 |
AAAI | 2 |
| 2015 | Question Answering over Freebase with Multi-Column Convolutional Neural NetworksabstractLi Dong, Furu Wei, Ming Zhou, Ke Xu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Li Dong 0004, Furu Wei, Ming Zhou 0001, Ke Xu 0001 |
ACL (1) | 2 |
| 2015 | Cold-Start Expert Finding in Community Question Answering via Graph Regularization
Zhou Zhao 0001, Furu Wei, Ming Zhou 0001, Wilfred Ng |
DASFAA (1) | 2 |
| 2015 | Crowd-Selection Query Processing in Crowdsourcing Databases: A Task-Driven ApproachabstractCrowd-selection is essential to crowdsourcing applications, since choosing the right workers with particular expertise to carry out specific crowdsourced tasks is extremely important. The central problem is simple but tricky: given a crowdsourced task, who is the right worker to ask? Currently, most existing work has mainly studied the problem of crowd-selection for simple crowdsourced tasks such as decision making and sentiment analysis. Their crowd-selection procedures are based on the trustworthiness of workers. However, for some complex tasks such as document review and question answering, selecting workers based on the latent category of tasks is a better solution. In this paper, we formulate a new problem of task-driven crowd-selection for complex tasks. We first develop a Bayesian generative model to exploit "who knows what" for the workers in the crowdsourcing environment. The model provides a principle and natural framework for capturing the latent skills of workers as well as the latent categories of crowdsourced tasks. The inference of the latent skills of workers is based on past resolved crowdsourced tasks with feedback scores. We assume that the feedback scores can illustrate the performance of the workers for the tasks. We then devise a variational algorithm that transforms the latent skill inference with the proposed model into a standard optimization problem, which can be solved efficiently. We verify the performance of our method through extensive experiments on the data collected from three well-known crowdsourcing platforms for question answering tasks such as Quora, Yahoo! Answer and Stack Overflow. Zhou Zhao 0001, Furu Wei, Ming Zhou 0001, Weikeng Chen, Wilfred Ng |
EDBT | 2 |
| 2015 | A Hybrid Neural Model for Type Classification of Entity Mentions
Li Dong 0004, Furu Wei, Ming Zhou 0001, Ke Xu 0001 |
IJCAI | 2 |
| 2015 | A Statistical Parsing Framework for Sentiment ClassificationabstractWe present a statistical parsing framework for sentence-level sentiment classification in this article. Unlike previous works that use syntactic parsing results for sentiment analysis, we develop a statistical parser to directly analyze the sentiment structure of a sentence. We show that complicated phenomena in sentiment analysis (e.g., negation, intensification, and contrast) can be handled the same way as simple and straightforward sentiment expressions in a unified and probabilistic way. We formulate the sentiment grammar upon Context-Free Grammars (CFGs), and provide a formal description of the sentiment parsing framework. We develop the parsing model to obtain possible sentiment parse trees for a sentence, from which the polarity model is proposed to derive the sentiment strength and polarity, and the ranking model is dedicated to selecting the best sentiment tree. We train the parser directly from examples of sentences annotated only with sentiment polarity labels but without any syntactic annotations or polarity annotations of constituents within sentences. Therefore we can obtain training data easily. In particular, we train a sentiment parser, s.parser, from a large amount of review sentences with users' ratings as rough sentiment polarity labels. Extensive experiments on existing benchmark data sets show significant improvements over baseline sentiment classification approaches. Li Dong 0004, Furu Wei, Shujie Liu 0001, Ming Zhou 0001, Ke Xu 0001 |
Comput. Linguistics | 2 |
| 2015 | Cross-lingual Sentiment Lexicon Learning With Bilingual Word Graph Label PropagationabstractIn this article we address the task of cross-lingual sentiment lexicon learning, which aims to automatically generate sentiment lexicons for the target languages with available English sentiment lexicons. We formalize the task as a learning problem on a bilingual word graph, in which the intra-language relations among the words in the same language and the inter-language relations among the words between different languages are properly represented. With the words in the English sentiment lexicon as seeds, we propose a bilingual word graph label propagation approach to induce sentiment polarities of the unlabeled words in the target language. Particularly, we show that both synonym and antonym word relations can be used to build the intra-language relation, and that the word alignment information derived from bilingual parallel sentences can be effectively leveraged to build the inter-language relation. The evaluation of Chinese sentiment lexicon learning shows that the proposed approach outperforms existing approaches in both precision and recall. Experiments conducted on the NTCIR data set further demonstrate the effectiveness of the learned sentiment lexicon in sentence-level sentiment classification. Dehong Gao, Furu Wei, Wenjie Li 0002, Ming Zhou 0001 |
Comput. Linguistics | 2 |
| 2015 | A Joint Segmentation and Classification Framework for Sentence Level Sentiment ClassificationabstractIn this paper, we propose a joint segmentation and classification framework for sentence-level sentiment classification. It is widely recognized that phrasal information is crucial for sentiment classification. However, existing sentiment classification algorithms typically split a sentence as a word sequence, which does not effectively handle the inconsistent sentiment polarity between a phrase and the words it contains, such as {“not bad,” “bad”} and {“a great deal of,” “great”}. We address this issue by developing a joint framework for sentence-level sentiment classification. It simultaneously generates useful segmentations and predicts sentence-level polarity based on the segmentation results. Specifically, we develop a candidate generation model to produce segmentation candidates of a sentence; a segmentation ranking model to score the usefulness of a segmentation candidate for sentiment classification; and a classification model for predicting the sentiment polarity of a segmentation. We train the joint framework directly from sentences annotated with only sentiment polarity, without using any syntactic or sentiment annotations in segmentation level. We conduct experiments for sentiment classification on two benchmark datasets: a tweet dataset and a review dataset. Experimental results show that: 1) our method performs comparably with state-of-the-art methods on both datasets; 2) joint modeling segmentation and classification outperforms pipelined baseline methods in various experimental settings. Duyu Tang, Bing Qin 0001, Furu Wei, Li Dong 0004, Ting Liu 0001, Ming Zhou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Adaptive Multi-Compositionality for Recursive Neural Models with Applications to Sentiment AnalysisabstractRecursive neural models have achieved promising results in many natural language processing tasks. The main difference among these models lies in the composition function, i.e., how to obtain the vector representation for a phrase or sentence using the representations of words it contains. This paper introduces a novel Adaptive Multi-Compositionality (AdaMC) layer to recursive neural models. The basic idea is to use more than one composition functions and adaptively select them depending on the input vectors. We present a general framework to model each semantic composition as a distribution over these composition functions. The composition functions and parameters used for adaptive selection are learned jointly from data. We integrate AdaMC into existing recursive neural models and conduct extensive experiments on the Stanford Sentiment Treebank. The results illustrate that AdaMC significantly outperforms state-of-the-art sentiment classification methods. It helps push the best accuracy of sentence-level negative/positive classification from 85.4% up to 88.5%. Li Dong 0004, Furu Wei, Ming Zhou 0001, Ke Xu 0001 |
AAAI | 2 |
| 2014 | Learning Sentiment-Specific Word Embedding for Twitter Sentiment ClassificationabstractWe present a method that learns word embedding for Twitter sentiment classification in this paper.Most existing algorithms for learning continuous word representations typically only model the syntactic context of words but ignore the sentiment of text.This is problematic for sentiment analysis as they usually map words with similar syntactic context but opposite sentiment polarity, such as good and bad, to neighboring word vectors.We address this issue by learning sentimentspecific word embedding (SSWE), which encodes sentiment information in the continuous representation of words.Specifically, we develop three neural networks to effectively incorporate the supervision from sentiment polarity of text (e.g.sentences or tweets) in their loss functions.To obtain large scale training corpora, we learn the sentiment-specific word embedding from massive distant-supervised tweets collected by positive and negative emoticons.Experiments on applying SS-WE to a benchmark Twitter sentiment classification dataset in SemEval 2013 show that (1) the SSWE feature performs comparably with hand-crafted features in the top-performed system; (2) the performance is further improved by concatenating SSWE with existing feature set. Duyu Tang, Furu Wei, Nan Yang 0002, Ming Zhou 0001, Ting Liu 0001, Bing Qin 0001 |
ACL (1) | 2 |
| 2014 | SocialTransfer: Transferring Social Knowledge for Cold-Start CowdsourcingabstractAn essential component of building a successful crowdsourcing market is effective task matching, which matches a given task to the right crowdworkers. In order to provide high- quality task matching, crowdsourcing systems rely on past task-solving activities of crowdworkers. However, the average number of past activities of crowdworkers in most crowd- sourcing systems is very small. We call the workers who have only solved a small number of tasks cold-start crowdworkers. We observe that most of the workers in crowdsourcing systems are cold-start crowdworkers, and crowdsourcing systems actually enjoy great benefits from cold-start crowd-workers. However, the problem of task matching with the presence of many cold-start crowdworkers has not been well studied. We propose a new approach to address this issue. Our main idea, motivated by the prevalence of online social networks, is to transfer the knowledge about crowdworkers in their social networks to crowdsourcing systems for task matching. We propose a SocialTransfer model for cold-start crowdsourcing, which not only infers the expertise of warm- start crowdworkers from their past activities, but also transfers the expertise knowledge to cold-start crowdworkers via social connections. We evaluate the SocialTransfer model on the well-known crowdsourcing system Quora, using knowledge from the popular social network Twitter. Experimental results show that, by transferring social knowledge, our method achieves significant improvements over the state-of-the-art methods. Zhou Zhao 0001, James Cheng, Furu Wei, Ming Zhou 0001, Wilfred Ng, Yingjun Wu |
CIKM | 3 |
| 2014 | Building Large-Scale Twitter-Specific Sentiment Lexicon : A Representation Learning Approach
Duyu Tang, Furu Wei, Bing Qin 0001, Ming Zhou 0001, Ting Liu 0001 |
COLING | 2 |
| 2014 | A Joint Segmentation and Classification Framework for Sentiment AnalysisabstractIn this paper, we propose a joint segmentation and classification framework for sentiment analysis.Existing sentiment classification algorithms typically split a sentence as a word sequence, which does not effectively handle the inconsistent sentiment polarity between a phrase and the words it contains, such as "not bad" and "a great deal of ".We address this issue by developing a joint segmentation and classification framework (JSC), which simultaneously conducts sentence segmentation and sentence-level sentiment classification.Specifically, we use a log-linear model to score each segmentation candidate, and exploit the phrasal information of top-ranked segmentations as features to build the sentiment classifier.A marginal log-likelihood objective function is devised for the segmentation model, which is optimized for enhancing the sentiment classification performance.The joint model is trained only based on the annotated sentiment polarity of sentences, without any segmentation annotations.Experiments on a benchmark Twitter sentiment classification dataset in SemEval 2013 show that, our joint model performs comparably with the state-of-the-art methods. Duyu Tang, Furu Wei, Bing Qin 0001, Li Dong 0004, Ting Liu 0001, Ming Zhou 0001 |
EMNLP | 2 |
| 2014 | Answer Extraction with Multiple Extraction Engines for Web-Based Question Answering
Furu Wei, Ming Zhou 0001 |
NLPCC | 2 |
| 2013 | The Automated Acquisition of Suggestions from TweetsabstractThis paper targets at automatically detecting and classifying user's suggestions from tweets. The short and informal nature of tweets, along with the imbalanced characteristics of suggestion tweets, makes the task extremely challenging. To this end, we develop a classification framework on Factorization Machines, which is effective and efficient especially in classification tasks with feature sparsity settings. Moreover, we tackle the imbalance problem by introducing cost-sensitive learning techniques in Factorization Machines. Extensively experimental studies on a manually annotated real-life data set show that the proposed approach significantly improves the baseline approach, and yields the precision of 71.06% and recall of 67.86%. We also investigate the reason why Factorization Machines perform better. Finally, we introduce the first manually annotated dataset for suggestion classification. Li Dong 0004, Furu Wei, Yajuan Duan, Ming Zhou 0001, Ke Xu 0001 |
AAAI | 2 |
| 2013 | Entity Linking for Tweets
Haocheng Wu, Ming Zhou 0001, Furu Wei |
ACL (1) | 5 |
| 2013 | Exploring hypergraph-based semi-supervised ranking for query-oriented summarization
Wei Wang 0013, Sujian Li, Wenjie Li 0002, Furu Wei |
Inf. Sci. | 5 |
| 2013 | Named entity recognition for tweetsabstractTwo main challenges of Named Entity Recognition (NER) for tweets are the insufficient information in a tweet and the lack of training data. We propose a novel method consisting of three core elements: (1) normalization of tweets; (2) combination of a K-Nearest Neighbors (KNN) classifier with a linear Conditional Random Fields (CRF) model; and (3) semisupervised learning framework. The tweet normalization preprocessing corrects common ill-formed words using a global linear model. The KNN-based classifier conducts prelabeling to collect global coarse evidence across tweets while the CRF model conducts sequential labeling to capture fine-grained information encoded in a tweet. The semisupervised learning plus the gazetteers alleviate the lack of training data. Extensive experiments show the advantages of our method over the baselines as well as the effectiveness of normalization, KNN, and semisupervised learning. Furu Wei, Shaodian Zhang, Ming Zhou 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2013 | Constrained Text Coclustering with Supervised and Unsupervised ConstraintsabstractIn this paper, we propose a novel constrained coclustering method to achieve two goals. First, we combine information-theoretic coclustering and constrained clustering to improve clustering performance. Second, we adopt both supervised and unsupervised constraints to demonstrate the effectiveness of our algorithm. The unsupervised constraints are automatically derived from existing knowledge sources, thus saving the effort and cost of using manually labeled constraints. To achieve our first goal, we develop a two-sided hidden Markov random field (HMRF) model to represent both document and word constraints. We then use an alternating expectation maximization (EM) algorithm to optimize the model. We also propose two novel methods to automatically construct and incorporate document and word constraints to support unsupervised constrained clustering: 1) automatically construct document constraints based on overlapping named entities (NE) extracted by an NE extractor; 2) automatically construct word constraints based on their semantic distance inferred from WordNet. The results of our evaluation over two benchmark data sets demonstrate the superiority of our approaches against a number of existing approaches. Yangqiu Song, Shimei Pan, Shixia Liu, Furu Wei, Michelle X. Zhou, Weihong Qian |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2012 | Collective Nominal Semantic Role Labeling for TweetsabstractTweets have become an increasingly popular source of fresh information. We investigate the task of Nominal Semantic Role Labeling (NSRL) for tweets, which aims to identify predicate-argument structures defined by nominals in tweets. Studies of this task can help fine-grained information extraction and retrieval from tweets. There are two main challenges in this task: 1) The lack of information in a single tweet, rooted in the short and noisy nature of tweets; and 2) recovery of implicit arguments. We propose jointly conducting NSRL on multiple similar tweets using a graphical model, leveraging the redundancy in tweets to tackle these challenges. Extensive evaluations on a human annotated data set demonstrate that our method outperforms two baselines with an absolute gain of 2.7% in F1. Zhongyang Fu, Furu Wei, Ming Zhou 0001 |
AAAI | 3 |
| 2012 | Exacting Social Events for Tweets Using a Factor GraphabstractSocial events are events that occur between people where at least one person is aware of the other and of the event taking place. Extracting social events can play an important role in a wide range of applications, such as the construction of social network. In this paper, we introduce the task of social event extraction for tweets, an important source of fresh events. One main challenge is the lack of information in a single tweet, which is rooted in the short and noise-prone nature of tweets. We propose to collectively extract social events from multiple similar tweets using a novel factor graph, to harvest the redundance in tweets, i.e., the repeated occurrences of a social event in several tweets. We evaluate our method on a human annotated data set, and show that it outperforms all baselines, with an absolute gain of 21% in F1. Xiangyang Zhou, Zhongyang Fu, Furu Wei, Ming Zhou 0001 |
AAAI | 4 |
| 2012 | Joint Inference of Named Entity Recognition and Normalization for Tweets
Ming Zhou 0001, Xiangyang Zhou, Zhongyang Fu, Furu Wei |
ACL (1) | 5 |
| 2012 | Cross-Lingual Mixture Model for Sentiment Classification
Xinfan Meng, Furu Wei, Ming Zhou 0001, Houfeng Wang |
ACL (1) | 2 |
| 2012 | Breaking news on twitterabstractAfter the news of Osama Bin Laden's death leaked through Twitter, many people wondered if Twitter would fundamentally change the way we produce, spread, and consume news. In this paper we provide an in-depth analysis of how the news broke and spread on Twitter. We confirm the claim that Twitter broke the news first, and find evidence that Twitter had convinced a large number of its audience before mainstream media confirmed the news. We also discover that attention on Twitter was highly concentrated on a small number of "opinion leaders" and identify three groups of opinion leaders who played key roles in spreading the news: individuals affiliated with media played a large part in breaking the news, mass media brought the news to a wider audience and provided eager Twitter users with content on external sites, and celebrities helped to spread the news and stimulate conversation. Our findings suggest Twitter has great potential as a news medium. Mengdie Hu, Shixia Liu, Furu Wei, Yingcai Wu, John T. Stasko, Kwan-Liu Ma |
CHI | 3 |
| 2012 | Graph-based collective classification for tweetsabstractIn this paper, we address the problem of classifying tweets into topical categories. Because of the short, noisy and ambiguous nature of tweets, we propose to collectively conduct the classification by exploiting the context information (i.e. related tweets) other than individually as in conventional text classification methods. In particular, we augment the content-based representation of text with tweets sharing same #hashtag or URL, which results in a tweet graph. We then formulate the tweet classification task under a graph optimization framework. We investigate three popular approaches, namely, Loopy Belief Propagation (LBP), Relaxation Labeling (RL), and Iterative Classification Algorithm (ICA). Extensive experiment results show that the graph-based tweet classification approach remarkably improves the performance, while the ICA model with relationship of sharing the same #hashtag gives the best result on separate tweet graph. Yajuan Duan, Furu Wei, Ming Zhou 0001, Harry Shum |
CIKM | 2 |
| 2012 | Twitter Topic Summarization by Ranking Tweets using Social Influence and Content Quality
Yajuan Duan, Furu Wei, Ming Zhou 0001, Harry Shum |
COLING | 3 |
| 2012 | Graph-Based Multi-Tweet Summarization using Social Signals
Furu Wei, Ming Zhou 0001 |
COLING | 3 |
| 2012 | Entity-centric topic-oriented opinion summarization in twitterabstractMicroblogging services, such as Twitter, have become popular channels for people to express their opinions towards a broad range of topics. Twitter generates a huge volume of instant messages (i.e. tweets) carrying users' sentiments and attitudes every minute, which both necessitates automatic opinion summarization and poses great challenges to the summarization system. In this paper, we study the problem of opinion summarization for entities, such as celebrities and brands, in Twitter. We propose an entity-centric topic-based opinion summarization framework, which aims to produce opinion summaries in accordance with topics and remarkably emphasizing the insight behind the opinions. To this end, we first mine topics from #hashtags, the human-annotated semantic tags in tweets. We integrate the #hashtags as weakly supervised information into topic modeling algorithms to obtain better interpretation and representation for calculating the similarity among them, and adopt Affinity Propagation algorithm to group #hashtags into coherent topics. Subsequently, we use templates generalized from paraphrasing to identify tweets with deep insights, which reveal reasons, express demands or reflect viewpoints. Afterwards, we develop a target (i.e. entity) dependent sentiment classification approach to identifying the opinion towards a given target (i.e. entity) of tweets. Finally, the opinion summary is generated through integrating information from dimensions of topic, opinion and insight, as well as other factors (e.g. topic relevancy, redundancy and language styles) in an unified optimization framework. We conduct extensive experiments on a real-life data set to evaluate the performance of individual opinion summarization modules as well as the quality of the produced summary. The promising experiment results show the effectiveness of the proposed framework and algorithms. Xinfan Meng, Furu Wei, Ming Zhou 0001, Sujian Li, Houfeng Wang |
KDD | 2 |
| 2011 | Recognizing Named Entities in Tweets
Shaodian Zhang, Furu Wei, Ming Zhou 0001 |
ACL | 3 |
| 2011 | Topic sentiment analysis in twitter: a graph-based hashtag sentiment classification approachabstractTwitter is one of the biggest platforms where massive instant messages (i.e. tweets) are published every day. Users tend to express their real feelings freely in Twitter, which makes it an ideal source for capturing the opinions towards various interesting topics, such as brands, products or celebrities, etc. Naturally, people may anticipate an approach to receiving the common sentiment tendency towards these topics directly rather than through reading the huge amount of tweets about them. On the other side, Hashtags, starting with a symbol "#" ahead of keywords or phrases, are widely used in tweets as coarse-grained topics. In this paper, instead of presenting the sentiment polarity of each tweet relevant to the topic, we focus our study on hashtag-level sentiment classification. This task aims to automatically generate the overall sentiment polarity for a given hashtag in a certain time period, which markedly differs from the conventional sentence-level and document-level sentiment analysis. Our investigation illustrates that three types of information is useful to address the task, including (1) sentiment polarity of tweets containing the hashtag; (2) hashtags co-occurrence relationship and (3) the literal meaning of hashtags. Consequently, in order to incorporate the first two types of information into a classification framework where hashtags can be classified collectively, we propose a novel graph model and investigate three approximate collective classification algorithms for inference. Going one step further, we show that the performance can be remarkably improved using an enhanced boosting classification setting in which we employ the literal meaning of hashtags as semi-supervised information. Experimental results on a real-life data set consisting of 29,195 tweets and 2,181 hashtags show the effectiveness of the proposed model and algorithms. Furu Wei, Ming Zhou 0001, Ming Zhang 0004 |
CIKM | 2 |
| 2011 | QuickView: advanced search of tweetsabstractTweets have become a comprehensive repository for real-time information. However, it is often hard for users to quickly get information they are interested in from tweets, owing to the sheer volume of tweets as well as their noisy and informal nature. We present QuickView, an NLP-based tweet search platform to tackle this issue. Specifically, it exploits a series of natural language processing technologies, such as tweet normalization, named entity recognition, semantic role labeling, sentiment analysis, tweet classification, to extract useful information, i.e., named entities, events, opinions, etc., from a large volume of tweets. Then, non-noisy tweets, together with the mined information, are indexed, on top of which two brand new scenarios are enabled, i.e., categorized browsing and advanced search, allowing users to effectively access either the tweets or fine-grained information they are interested in. Long Jiang, Furu Wei, Ming Zhou 0001 |
SIGIR | 3 |
| 2011 | Semantic-Preserving Word Clouds by Seam CarvingabstractAbstract Word clouds are proliferating on the Internet and have received much attention in visual analytics. Although word clouds can help users understand the major content of a document collection quickly, their ability to visually compare documents is limited. This paper introduces a new method to create semantic‐preserving word clouds by leveraging tailored seam carving, a well‐established content‐aware image resizing operator. The method can optimize a word cloud layout by removing a left‐to‐right or top‐to‐bottom seam iteratively and gracefully from the layout. Each seam is a connected path of low energy regions determined by a Gaussian‐based energy function. With seam carving, we can pack the word cloud compactly and effectively, while preserving its overall semantic structure. Furthermore, we design a set of interactive visualization techniques for the created word clouds to facilitate visual text analysis and comparison. Case studies are conducted to demonstrate the effectiveness and usefulness of our techniques. Yingcai Wu, Thomas Provan, Furu Wei, Shixia Liu, Kwan-Liu Ma |
Comput. Graph. Forum | 3 |
| 2010 | Constrained Coclustering for Textual DocumentsabstractIn this paper, we present a constrained co-clustering approach for clustering textual documents. Our approach combines the benefits of information-theoretic co-clustering and constrained clustering. We use a two-sided hidden Markov random field (HMRF) to model both the document and word constraints. We also develop an alternating expectation maximization (EM) algorithm to optimize the constrained co-clustering model. We have conducted two sets of experiments on a benchmark data set: (1) using human-provided category labels to derive document and word constraints for semi-supervised document clustering, and (2) using automatically extracted named entities to derive document constraints for unsupervised document clustering. Compared to several representative constrained clustering and co-clustering approaches, our approach is shown to be more effective for high-dimensional, sparse text data. Yangqiu Song, Shimei Pan, Shixia Liu, Furu Wei, Michelle X. Zhou, Weihong Qian |
AAAI | 4 |
| 2010 | Context preserving dynamic word cloud visualizationabstractIn this paper, we introduce a visualization method that couples a trend chart with word clouds to illustrate temporal content evolutions in a set of documents. Specifically, we use a trend chart to encode the overall semantic evolution of document content over time. In our work, semantic evolution of a document collection is modeled by varied significance of document content, represented by a set of representative keywords, at different time points. At each time point, we also use a word cloud to depict the representative keywords. Since the words in a word cloud may vary one from another over time (e.g., words with increased importance), we use geometry meshes and an adaptive force-directed model to lay out word clouds to highlight the word differences between any two subsequent word clouds. Our method also ensures semantic coherence and spatial stability of word clouds over time. Our work is embodied in an interactive visual analysis system that helps users to perform text analysis and derive insights from a large collection of documents. Our preliminary evaluation demonstrates the usefulness and usability of our work. Weiwei Cui 0001, Yingcai Wu, Shixia Liu, Furu Wei, Michelle X. Zhou, Huamin Qu |
PacificVis | 4 |
| 2010 | TIARA: a visual exploratory text analytic systemabstractIn this paper, we present a novel exploratory visual analytic system called TIARA (Text Insight via Automated Responsive Analytics), which combines text analytics and interactive visualization to help users explore and analyze large collections of text. Given a collection of documents, TIARA first uses topic analysis techniques to summarize the documents into a set of topics, each of which is represented by a set of keywords. In addition to extracting topics, TIARA derives time-sensitive keywords to depict the content evolution of each topic over time. To help users understand the topic-based summarization results, TIARA employs several interactive text visualization techniques to explain the summarization results and seamlessly link such results to the original text. We have applied TIARA to several real-world applications, including email summarization and patient record analysis. To measure the effectiveness of TIARA, we have conducted several experiments. Our experimental results and initial user feedback suggest that TIARA is effective in aiding users in their exploratory text analytic tasks. Furu Wei, Shixia Liu, Yangqiu Song, Shimei Pan, Michelle X. Zhou, Weihong Qian, Lei Shi 0002 |
KDD | 1 |
| 2010 | iRANK: A rank-learn-combine framework for unsupervised ensemble rankingabstractAbstract The authors address the problem of unsupervised ensemble ranking. Traditional approaches either combine multiple ranking criteria into a unified representation to obtain an overall ranking score or to utilize certain rank fusion or aggregation techniques to combine the ranking results. Beyond the aforementioned “combine‐then‐rank” and “rank‐then‐combine” approaches, the authors propose a novel “rank‐learn‐combine” ranking framework, called Interactive Ranking (iRANK), which allows two base rankers to “teach” each other before combination during the ranking process by providing their own ranking results as feedback to the others to boost the ranking performance. This mutual ranking refinement process continues until the two base rankers cannot learn from each other any more. The overall performance is improved by the enhancement of the base rankers through the mutual learning mechanism. The authors further design two ranking refinement strategies to efficiently and effectively use the feedback based on reasonable assumptions and rational analysis. Although iRANK is applicable to many applications, as a case study, they apply this framework to the sentence ranking problem in query‐focused summarization and evaluate its effectiveness on the DUC 2005 and 2006 data sets. The results are encouraging with consistent and promising improvements. Furu Wei, Wenjie Li 0002, Shixia Liu |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2010 | A document-sensitive graph model for multi-document summarization
Furu Wei, Wenjie Li 0002, Qin Lu 0001, Yanxiang He |
Knowl. Inf. Syst. | 1 |
| 2010 | OpinionSeer: Interactive Visualization of Hotel Customer FeedbackabstractThe rapid development of Web technology has resulted in an increasing number of hotel customers sharing their opinions on the hotel services. Effective visual analysis of online customer opinions is needed, as it has a significant impact on building a successful business. In this paper, we present OpinionSeer, an interactive visualization system that could visually analyze a large collection of online hotel customer reviews. The system is built on a new visualization-centric opinion mining technique that considers uncertainty for faithfully modeling and analyzing customer opinions. A new visual representation is developed to convey customer opinions by augmenting well-established scatterplots and radial visualization. To provide multiple-level exploration, we introduce subjective logic to handle and organize subjective opinions with degrees of uncertainty. Several case studies illustrate the effectiveness and usefulness of OpinionSeer on analyzing relationships among multiple data dimensions and comparing opinions of different groups. Aside from data on hotel customer feedback, OpinionSeer could also be applied to visually analyze customer opinions on other products or services. Yingcai Wu, Furu Wei, Shixia Liu, Norman Au, Weiwei Cui 0001, Hong Zhou 0004, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2009 | HyperSum: hypergraph based semi-supervised sentence ranking for query-oriented summarizationabstractGraph based sentence ranking algorithms such as PageRank and HITS have been successfully used in query-oriented summarization. With these algorithms, the documents to be summarized are often modeled as a text graph where nodes represent sentences and edges represent pairwise similarity relationships between two sentences. A deficiency of conventional graph modeling is its incapability of naturally and effectively representing complex group relationships shared among multiple objects. Simply squeezing complex relationships into pairwise ones will inevitably lead to loss of information which can be useful for ranking and learning. In this paper, we propose to take advantage of hypergraph, i.e. a generalization of graph, to remedy this defect. In a text hypergraph, nodes still represent sentences, yet hyperedges are allowed to connect more than two sentences. With a text hypergraph, we are thus able to integrate both group relationships formulated among multiple sentences and pairwise relationships formulated between two sentences in a unified framework. As essential work, it is first addressed in the paper that how a text hypergraph can be built for summarization by applying clustering techniques. Then, a hypergraph based semi-supervised sentence ranking algorithm is developed for query-oriented extractive summarization, where the influence of query is propagated to sentences through the structure of the constructed text hypergraph. When evaluated on DUC data sets, performance of the proposed approach is remarkable. Wei Wang 0013, Furu Wei, Wenjie Li 0002, Sujian Li |
CIKM | 2 |
| 2009 | iRANK: an interactive ranking framework and its application in query-focused summarizationabstractWe address the problem of unsupervised ensemble ranking in this paper. Traditional approaches either combine multiple ranking criteria into a unified representation to obtain an overall ranking score or to utilize certain rank fusion or aggregation techniques to combine the ranking results. Beyond the aforementioned combine-then-rank and rank-then-combine approaches, we propose a novel rank-learn-combine ranking framework, called Interactive Ranking (iRANK), which allows two base rankers to "teach" each other before combination during the ranking process by providing their own ranking results as feedback to the others so as to boost the ranking performance. This mutual ranking refinement process continues until the two base rankers cannot learn from each other any more. The overall performance is improved by the enhancement of the base rankers through the mutual learning mechanism. We apply this framework to the sentence ranking problem in query-focused summarization and evaluate its effectiveness on the DUC 2005 data set. The results are encouraging with consistent and promising improvements. Furu Wei, Wenjie Li 0002, Wei Wang 0013, Yanxiang He |
CIKM | 1 |
| 2009 | Applying two-level reinforcement ranking in query-oriented multidocument summarizationabstractAbstract Sentence ranking is the issue of most concern in document summarization today. While traditional feature‐based approaches evaluate sentence significance and rank the sentences relying on the features that are particularly designed to characterize the different aspects of the individual sentences, the newly emerging graph‐based ranking algorithms (such as the PageRank‐like algorithms) recursively compute sentence significance using the global information in a text graph that links sentences together. In general, the existing PageRank‐like algorithms can model well the phenomena that a sentence is important if it is linked by many other important sentences. Or they are capable of modeling the mutual reinforcement among the sentences in the text graph. However, when dealing with multidocument summarization these algorithms often assemble a set of documents into one large file. The document dimension is totally ignored. In this article we present a framework to model the two‐level mutual reinforcement among sentences as well as documents. Under this framework we design and develop a novel ranking algorithm such that the document reinforcement is taken into account in the process of sentence ranking. The convergence issue is examined. We also explore an interesting and important property of the proposed algorithm. When evaluated on the DUC 2005 and 2006 query‐oriented multidocument summarization datasets, significant results are achieved. Furu Wei, Wenjie Li 0002, Qin Lu 0001, Yanxiang He |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2008 | PNR2: Ranking Sentences with Positive and Negative Reinforcement for Query-Oriented Update Summarization
Wenjie Li 0002, Furu Wei, Qin Lu 0001, Yanxiang He |
COLING | 2 |
| 2008 | A Cluster-Sensitive Graph Model for Query-Oriented Multi-document Summarization
Furu Wei, Wenjie Li 0002, Qin Lu 0001, Yanxiang He |
ECIR | 1 |
| 2008 | Exploiting the Role of Position Feature in Chinese Relation Extraction
Peng Zhang 0002, Wenjie Li 0002, Furu Wei, Qin Lu 0001, Yuexian Hou |
LREC | 3 |
| 2008 | Exploiting the Role of Named Entities in Query-Oriented Document Summarization
Wenjie Li 0002, Furu Wei, Ouyang You, Qin Lu 0001, Yanxiang He |
PRICAI | 2 |
| 2008 | Query-sensitive mutual reinforcement chain and its application in query-oriented multi-document summarizationabstractSentence ranking is the issue of most concern in document summarization. Early researchers have presented the mutual reinforcement principle (MR) between sentence and term for simultaneous key phrase and salient sentence extraction in generic single-document summarization. In this work, we extend the MR to the mutual reinforcement chain (MRC) of three different text granularities, i.e., document, sentence and terms. The aim is to provide a general reinforcement framework and a formal mathematical modeling for the MRC. Going one step further, we incorporate the query influence into the MRC to cope with the need for query-oriented multi-document summarization. While the previous summarization approaches often calculate the similarity regardless of the query, we develop a query-sensitive similarity to measure the affinity between the pair of texts. When evaluated on the DUC 2005 dataset, the experimental results suggest that the proposed query-sensitive MRC (Qs-MRC) is a promising approach for summarization. Furu Wei, Wenjie Li 0002, Qin Lu 0001, Yanxiang He |
SIGIR | 1 |