VLDB 2026 Research / reviewers in the wild / expert
Li Dong 0004
dblp:85/5090-4
· DBLP profile ↗
82ranked-venue papers
12as first author
52since 2021 · last 2026
0000-0003-3083-7170ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 78 · 12 first-author · 50 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Stable and Effective Reinforcement Learning for Mixture-of-ExpertsabstractDi Zhang, Xun Wu, Shaohan Huang, Lingjie Jiang, Yaru Hao, Li Dong, Zewen Chi, Zhifang Sui, Furu Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaohan Huang, Lingjie Jiang, Yaru Hao, Li Dong 0004, Zewen Chi, Zhifang Sui, Furu Wei |
ACL (1) | 6 |
| 2025 | Semi-Parametric Retrieval via Binary Bag-of-Tokens IndexabstractInformation retrieval has transitioned from standalone systems into essential components across broader applications, with indexing efficiency, cost-effectiveness, and freshness becoming increasingly critical yet often overlooked. In this paper, we introduce SemI-parametric Disentangled Retrieval (SiDR), a bi-encoder retrieval framework that decouples retrieval index from neural parameters to enable efficient, low-cost, and parameter-agnostic indexing for emerging use cases. Specifically, in addition to using embeddings as indexes like existing neural retrieval methods, SiDR supports a non-parametric tokenization index for search, achieving BM25-like indexing complexity with significantly better effectiveness. Our comprehensive evaluation across 16 retrieval benchmarks demonstrates that SiDR outperforms both neural and term-based retrieval baselines under the same indexing workload: (i) When using an embedding-based index, SiDR exceeds the performance of conventional neural retrievers while maintaining similar training complexity; (ii) When using a tokenization-based index, SiDR drastically reduces indexing cost and time, matching the complexity of traditional term-based retrieval, while consistently outperforming BM25 on all in-domain datasets; (iii) Additionally, we introduce a late parametric mechanism that matches BM25 index preparation time while outperforming other neural retrieval baselines in effectiveness. Jiawei Zhou 0003, Li Dong 0004, Furu Wei, Lei Chen 0002 |
ICLR | 2 |
| 2025 | Self-Boosting Large Language Models with Synthetic Preference DataabstractThrough alignment with human preferences, Large Language Models (LLMs) have advanced significantly in generating honest, harmless, and helpful responses. However, collecting high-quality preference data is a resource-intensive and creativity-demanding process, especially for the continual improvement of LLMs. We introduce SynPO, a self-boosting paradigm that leverages synthetic preference data for model alignment. SynPO employs an iterative mechanism wherein a self-prompt generator creates diverse prompts, and a response improver refines model responses progressively. This approach trains LLMs to autonomously learn the generative rewards for their own outputs and eliminates the need for large-scale annotation of prompts and human preferences. After four SynPO iterations, Llama3-8B and Mistral-7B show significant enhancements in instruction-following abilities, achieving over 22.1% win rate improvements on AlpacaEval 2.0 and ArenaHard. Simultaneously, SynPO improves the general performance of LLMs on various tasks, validated by a 3.2 to 5.0 average score increase on the well-recognized Open LLM leaderboard. Qingxiu Dong, Li Dong 0004, Xingxing Zhang 0002, Zhifang Sui, Furu Wei |
ICLR | 2 |
| 2025 | Data Selection via Optimal Control for Language ModelsabstractThis work investigates the selection of high-quality pre-training data from massive corpora to enhance LMs' capabilities for downstream usage.
We formulate data selection as a generalized Optimal Control problem, which can be solved theoretically by Pontryagin's Maximum Principle (PMP), yielding a set of necessary conditions that characterize the relationship between optimal data selection and LM training dynamics.
Based on these theoretical results, we introduce **P**MP-based **D**ata **S**election (**PDS**), a framework that approximates optimal data selection by solving the PMP conditions.
In our experiments, we adopt PDS to select data from CommmonCrawl and show that the PDS-selected corpus accelerates the learning of LMs and constantly boosts their performance on a wide range of downstream tasks across various model sizes.
Moreover, the benefits of PDS extend to ~400B models trained on ~10T tokens, as evidenced by the extrapolation of the test loss curves according to the Scaling Laws.
PDS also improves data utilization when the pre-training data is limited, by reducing the data demand by 1.8 times, which helps mitigate the quick exhaustion of available web-crawled corpora. Our code, model, and data can be found at https://github.com/microsoft/LMOps/tree/main/data_selection. Yuxian Gu, Li Dong 0004, Hongning Wang, Yaru Hao, Qingxiu Dong, Furu Wei, Minlie Huang |
ICLR | 2 |
| 2025 | Differential TransformerabstractTransformer tends to overallocate attention to irrelevant context. In this work, we introduce Diff Transformer, which amplifies attention to the relevant context while canceling noise. Specifically, the differential attention mechanism calculates attention scores as the difference between two separate softmax attention maps. The subtraction cancels noise, promoting the emergence of sparse attention patterns. Experimental results on language modeling show that Diff Transformer outperforms Transformer in various settings of scaling up model size and training tokens. More intriguingly, it offers notable advantages in practical applications, such as long-context modeling, key information retrieval, hallucination mitigation, in-context learning, and reduction of activation outliers. By being less distracted by irrelevant context, Diff Transformer can mitigate hallucination in question answering and text summarization. For in-context learning, Diff Transformer not only enhances accuracy but is also more robust to order permutation, which was considered as a chronic robustness issue. The results position Diff Transformer as a highly effective and promising architecture for large language models. Tianzhu Ye, Li Dong 0004, Yuqing Xia, Yutao Sun, Gao Huang 0001, Furu Wei |
ICLR | 2 |
| 2025 | Imagine While Reasoning in Space: Multimodal Visualization-of-ThoughtabstractChain-of-Thought (CoT) prompting has proven highly effective for enhancing complex reasoning in Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). Yet, it struggles in complex spatial reasoning tasks. Nonetheless, human cognition extends beyond language alone, enabling the remarkable capability to think in both words and images. Inspired by this mechanism, we propose a new reasoning paradigm, Multimodal Visualization-of-Thought (MVoT). It enables visual thinking in MLLMs by generating image visualizations of their reasoning traces. To ensure high-quality visualization, we introduce token discrepancy loss into autoregressive MLLMs. This innovation significantly improves both visual coherence and fidelity. We validate this approach through several dynamic spatial reasoning tasks. Experimental results reveal that MVoT demonstrates competitive performance across tasks. Moreover, it exhibits robust and reliable improvements in the most challenging scenarios where CoT fails. Ultimately, MVoT establishes new possibilities for complex reasoning tasks where visual thinking can effectively complement verbal reasoning. Chengzu Li, Wenshan Wu, Huanyu Zhang 0002, Yan Xia 0005, Shaoguang Mao, Li Dong 0004, Ivan Vulic, Furu Wei |
ICML | 6 |
| 2025 | Reward Reasoning ModelsabstractReward models play a critical role in guiding large language models toward outputs that align with human expectations. However, an open challenge remains in effectively utilizing test-time compute to enhance reward model performance. In this work, we introduce Reward Reasoning Models (RRMs), which are specifically designed to execute a deliberate reasoning process before generating final rewards. Through chain-of-thought reasoning, RRMs leverage additional test-time compute for complex queries where appropriate rewards are not immediately apparent. To develop RRMs, we implement a reinforcement learning framework that fosters self-evolved reward reasoning capabilities without requiring explicit reasoning traces as training data. Experimental results demonstrate that RRMs achieve superior performance on reward modeling benchmarks across diverse domains. Notably, we show that RRMs can adaptively exploit test-time compute to further improve reward accuracy. The pretrained models are available at https://huggingface.co/Reward-Reasoning. Zewen Chi, Li Dong 0004, Qingxiu Dong, Shaohan Huang, Furu Wei |
NeurIPS | 3 |
| 2025 | Think Only When You Need with Large Hybrid-Reasoning ModelsabstractRecent Large Reasoning Models (LRMs) have shown substantially improved reasoning capabilities over traditional Large Language Models (LLMs) by incorporating extended thinking processes prior to producing final responses. However, excessively lengthy thinking introduces substantial overhead in terms of token consumption and latency, which is unnecessary for simple queries. In this work, we introduce Large Hybrid-Reasoning Models (LHRMs), the first kind of model capable of adaptively determining whether to perform reasoning based on the contextual information of user queries. To achieve this, we propose a two-stage training pipeline comprising Hybrid Fine-Tuning (HFT) as a cold start, followed by online reinforcement learning with the proposed Hybrid Group Policy Optimization (HGPO) to implicitly learn to select the appropriate reasoning mode. Furthermore, we introduce a metric called Hybrid Accuracy to quantitatively assess the model’s capability for hybrid reasoning. Extensive experimental results show that LHRMs can adaptively perform hybrid reasoning on queries of varying difficulty and type. It outperforms existing LRMs and LLMs in reasoning and general capabilities while significantly improving efficiency. Together, our work advocates for a reconsideration of the appropriate use of extended reasoning processes and provides a solid starting point for building hybrid reasoning systems. Lingjie Jiang, Shaohan Huang, Qingxiu Dong, Zewen Chi, Li Dong 0004, Xingxing Zhang 0002, Tengchao Lv, Lei Cui 0001, Furu Wei |
NeurIPS | 6 |
| 2025 | Text to Image Generation with Bidirectional Multiway TransformersabstractIn this study, we explore the potential of Multiway Transformers for text-to-image generation to achieve performance improvements through a concise and efficient decoupled model design and the inference efficiency provided by bidirectional encoding. We propose a method for improving the image tokenizer using pretrained Vision Transformers. Next, we employ bidirectional Multiway Transformers to restore the masked visual tokens combined with the unmasked text tokens. On the MS-COCO benchmark, our Multiway Transformers outperform vanilla Transformers, achieving superior FID scores and confirming the efficacy of the modality-specific parameter computation design. Ablation studies reveal that the fusion of visual and text tokens in bidirectional encoding contributes to improved model performance. Additionally, our proposed tokenizer outperforms VQGAN in image reconstruction quality and enhances the text-to-image generation results. By incorporating the additional CC-3M dataset for intermediate finetuning on our model with 688M parameters, we achieve competitive results with a finetuned FID score of 4.98 on MS-COCO. Hangbo Bao, Li Dong 0004, Furu Wei |
Comput. Vis. Media | 2 |
| 2025 | BitNet: 1-bit Pre-training for Large Language ModelsabstractThe increasing size of large language models (LLMs) has posed challenges for deployment and raised concerns about environmental impact due to high energy consumption. Previous research typically applies quantization after pre-training. While these methods avoid the need for model retraining, they often cause notable accuracy loss at extremely low bit-widths. In this work, we explore the feasibility and scalability of 1-bit pre-training. We introduce BitNet b1 and BitNet b1.58, the scalable and stable 1-bit Transformer architecture designed for LLMs. Specifically, we introduce BitLinear as a drop-in replacement of the nn.Linear layer in order to train 1-bit weights from scratch. Experimental results show that BitNet b1 achieves competitive performance, compared to state-of-the-art 8-bit quantization methods and FP16 Transformer baselines. With the ternary weight, BitNet b1.58 matches the half-precision Transformer LLM with the same model size and training tokens in terms of both perplexity and end-task performance, while being significantly more cost-effective in terms of latency, memory, throughput, and energy consumption. More profoundly, BitNet defines a new scaling law and recipe for training new generations of LLMs that are both high-performance and cost-effective. It enables a new computation paradigm and opens the door for designing specific hardware optimized for 1-bit LLMs. Hongyu Wang 0009, Shuming Ma, Lingxiao Ma, Lei Wang 0222, Wenhui Wang 0003, Li Dong 0004, Shaohan Huang, Huaijie Wang, Jilong Xue, Yi Wu 0013, Furu Wei |
J. Mach. Learn. Res. | 6 |
| 2024 | Language Models as Inductive ReasonersabstractZonglin Yang, Li Dong, Xinya Du, Hao Cheng, Erik Cambria, Xiaodong Liu, Jianfeng Gao, Furu Wei. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zonglin Yang 0001, Li Dong 0004, Xinya Du, Hao Cheng 0002, Erik Cambria, Xiaodong Liu 0003, Jianfeng Gao 0001, Furu Wei |
EACL (1) | 2 |
| 2024 | MiniLLM: Knowledge Distillation of Large Language ModelsabstractKnowledge Distillation (KD) is a promising technique for reducing the high computational demand of large language models (LLMs). However, previous KD methods are primarily applied to white-box classification models or training small models to imitate black-box model APIs like ChatGPT. How to effectively distill the knowledge of white-box LLMs into small models is still under-explored, which becomes more important with the prosperity of open-source LLMs. In this work, we propose a KD approach that distills LLMs into smaller language models. We first replace the forward Kullback-Leibler divergence (KLD) objective in the standard KD approaches with reverse KLD, which is more suitable for KD on generative language models, to prevent the student model from overestimating the low-probability regions of the teacher distribution. Then, we derive an effective optimization approach to learn this objective. The student models are named MiniLLM. Extensive experiments in the instruction-following setting show that MiniLLM generates more precise responses with higher overall quality, lower exposure bias, better calibration, and higher long-text generation performance than the baselines. Our method is scalable for different model families
with 120M to 13B parameters. Our code, data, and model checkpoints can be found in https://github.com/microsoft/LMOps/tree/main/minillm. Yuxian Gu, Li Dong 0004, Furu Wei, Minlie Huang |
ICLR | 2 |
| 2024 | Kosmos-G: Generating Images in Context with Multimodal Large Language ModelsabstractRecent advancements in subject-driven image generation have made significant strides. However, current methods still fall short in diverse application scenarios, as they require test-time tuning and cannot accept interleaved multi-image and text input. These limitations keep them far from the ultimate goal of "image as a foreign language in image generation." This paper presents Kosmos-G, a model that leverages the advanced multimodal perception capabilities of Multimodal Large Language Models (MLLMs) to tackle the aforementioned challenge. Our approach aligns the output space of MLLM with CLIP using the textual modality as an anchor and performs compositional instruction tuning on curated data. Kosmos-G demonstrates an impressive capability of zero-shot subject-driven generation with interleaved multi-image and text input. Notably, the score distillation instruction tuning requires no modifications to the image decoder. This allows for a seamless substitution of CLIP and effortless integration with a myriad of U-Net techniques ranging from fine-grained controls to personalized image decoder variants. We posit Kosmos-G as an initial attempt towards the goal of "image as a foreign language in image generation." Xichen Pan, Li Dong 0004, Shaohan Huang, Zhiliang Peng, Wenhu Chen, Furu Wei |
ICLR | 2 |
| 2024 | Grounding Multimodal Large Language Models to the WorldabstractWe introduce Kosmos-2, a Multimodal Large Language Model (MLLM), enabling new capabilities of perceiving object descriptions (e.g., bounding boxes) and grounding text to the visual world. Specifically, we represent text spans (i.e., referring expressions and noun phrases) as links in Markdown, i.e., [text span](bounding boxes), where object descriptions are sequences of location tokens. To train the model, we construct a large-scale dataset about grounded image-text pairs (GrIT) together with multimodal corpora. In addition to the existing capabilities of MLLMs (e.g., perceiving general modalities, following instructions, and performing in-context learning), Kosmos-2 integrates the grounding capability to downstream applications, while maintaining the conventional capabilities of MLLMs (e.g., perceiving general modalities, following instructions, and performing in-context learning). Kosmos-2 is evaluated on a wide range of tasks, including (i) multimodal grounding, such as referring expression comprehension and phrase grounding, (ii) multimodal referring, such as referring expression generation, (iii) perception-language tasks, and (iv) language understanding and generation. This study sheds a light on the big convergence of language, multimodal perception, and world modeling, which is a key step toward artificial general intelligence. Code can be found in [https://aka.ms/kosmos-2](https://aka.ms/kosmos-2). Zhiliang Peng, Wenhui Wang 0003, Li Dong 0004, Yaru Hao, Shaohan Huang, Shuming Ma, Qixiang Ye, Furu Wei |
ICLR | 3 |
| 2024 | KOSMOS-E : Learning to Follow Instruction for Robotic GraspingabstractTuning on instruction-following data has been shown to enhance the capabilities and controllability of language models, but the idea is less explored in the robotic field. In this work, we introduce KOSMOS-E, a Multimodal Large Language Model (MLLM) that leverages instruction-following robotic grasping data to enhance capabilities for precise and intricate robotic grasping maneuvers. To achieve this, we craft a large-scale instruction-following robotic grasping dataset, termed INSTRUCT-GRASP, primarily comprising two aspects: (i) grasp a single object following varying levels of granularity descriptions, e.g., different angles and aspects, and (ii) grasp a specific object within a multi-object environment following specific attributes, e.g., color and shape. Extensive experiments show the effectiveness of KOSMOS-E on robotic grasping tasks across a variety of environments. Shaohan Huang, Li Dong 0004, Wenhui Wang 0003, Shuming Ma, Furu Wei |
IROS | 4 |
| 2024 | You Only Cache Once: Decoder-Decoder Architectures for Language ModelsabstractWe introduce a decoder-decoder architecture, YOCO, for large language models, which only caches key-value pairs once. It consists of two components, i.e., a cross-decoder stacked upon a self-decoder. The self-decoder efficiently encodes global key-value (KV) caches that are reused by the cross-decoder via cross-attention. The overall model behaves like a decoder-only Transformer, although YOCO only caches once. The design substantially reduces GPU memory demands, yet retains global attention capability. Additionally, the computation flow enables prefilling to early exit without changing the final output, thereby significantly speeding up the prefill stage. Experimental results demonstrate that YOCO achieves favorable performance compared to Transformer in various settings of scaling up model size and number of training tokens. We also extend YOCO to 1M context length with near-perfect needle retrieval accuracy. The profiling results show that YOCO improves inference memory, prefill latency, and throughput by orders of magnitude across context lengths and model sizes. Yutao Sun, Li Dong 0004, Shaohan Huang, Wenhui Wang 0003, Shuming Ma, Quanlu Zhang, Jianyong Wang 0001, Furu Wei |
NeurIPS | 2 |
| 2024 | Multi-Head Mixture-of-ExpertsabstractSparse Mixtures of Experts (SMoE) scales model capacity without significant increases in computational costs. However, it exhibits the low expert activation issue, i.e., only a small subset of experts are activated for optimization, leading to suboptimal performance and limiting its effectiveness in learning a larger number of experts in complex tasks. In this paper, we propose Multi-Head Mixture-of-Experts (MH-MoE). MH-MoE split each input token into multiple sub-tokens, then these sub-tokens are assigned to and processed by a diverse set of experts in parallel, and seamlessly reintegrated into the original token form. The above operations enables MH-MoE to significantly enhance expert activation while collectively attend to information from various representation spaces within different experts to deepen context understanding. Besides, it's worth noting that our MH-MoE is straightforward to implement and decouples from other SMoE frameworks, making it easy to integrate with these frameworks for enhanced performance. Extensive experimental results across different parameter scales (300M to 7B) and three pre-training tasks—English-focused language modeling, multi-lingual language modeling and masked multi-modality modeling—along with multiple downstream validation tasks, demonstrate the effectiveness of MH-MoE. Shaohan Huang, Wenhui Wang 0003, Shuming Ma, Li Dong 0004, Furu Wei |
NeurIPS | 5 |
| 2024 | Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language ModelsabstractLarge language models (LLMs) have exhibited impressive performance in language comprehension and various reasoning tasks. However, their abilities in spatial reasoning, a crucial aspect of human cognition, remain relatively unexplored. Human possess a remarkable ability to create mental images of unseen objects and actions through a process known as the Mind's Eye, enabling the imagination of the unseen world. Inspired by this cognitive capacity, we propose Visualization-of-Thought (VoT) prompting. VoT aims to elicit spatial reasoning of LLMs by visualizing their reasoning traces, thereby guiding subsequent reasoning steps. We employed VoT for multi-hop spatial reasoning tasks, including natural language navigation, visual navigation, and visual tiling in 2D grid worlds. Experimental results demonstrated that VoT significantly enhances the spatial reasoning abilities of LLMs. Notably, VoT outperformed existing multimodal large language models (MLLMs) in these tasks. While VoT works surprisingly well on LLMs, the ability to generate mental images to facilitate spatial reasoning resembles the mind's eye process, suggesting its potential viability in MLLMs. Please find the dataset and codes in our [project page](https://microsoft.github.io/visualization-of-thought). Wenshan Wu, Shaoguang Mao, Yan Xia 0005, Li Dong 0004, Lei Cui 0001, Furu Wei |
NeurIPS | 5 |
| 2024 | DeepNet: Scaling Transformers to 1,000 LayersabstractIn this paper, we propose a simple yet effective method to stabilize extremely deep Transformers. Specifically, we introduce a new normalization function (DeepNorm) to modify the residual connection in Transformer, accompanying with theoretically derived initialization. In-depth theoretical analysis shows that model updates can be bounded in a stable way. The proposed method combines the best of two worlds, i.e., good performance of Post-LN and stable training of Pre-LN, makingDeepNorma preferred alternative. We successfully scale Transformers up to 1,000 layers (i.e., 2,500 attention and feed-forward network sublayers) without difficulty, which is one order of magnitude deeper than previous deep Transformers. Extensive experiments demonstrate thatDeepNethas superior performance across various benchmarks, including machine translation, language modeling (i.e., BERT, GPT) and vision pre-training (i.e., BEiT). Remarkably, on a multilingual benchmark with 7,482 translation directions, our 200-layer model with 3.2B parameters significantly outperforms the 48-layer state-of-the-art model with 12B parameters by 5 BLEU points, which indicates a promising scaling direction. Our code is available athttps://aka.ms/torchscale. Hongyu Wang 0009, Shuming Ma, Li Dong 0004, Shaohan Huang, Dongdong Zhang 0001, Furu Wei |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Generic-to-Specific Distillation of Masked AutoencodersabstractTo transfer the representation capacity of large pre-trained models to lightweight models, knowledge distillation has been widely explored. However, conventional single-stage distillation methods are prone to getting stuck in the transfer of task-specific knowledge, making it difficult to retain task-agnostic knowledge which is crucial for model generalization. In this study, we propose generic-to-specific distillation (G2SD), to boost lightweight models under the assistance of large models pre-trained by masked image modeling. In generic distillation, the decoder of a small model is encouraged to align feature predictions with that of a large model, so that task-agnostic knowledge can be transferred. In specific distillation, predictions of the small model are encouraged to be consistent with those of the large model, to guarantee task performance. G2SD is also applicable for heterogeneous settings(i.e., distilling from ViT to CNN). With G2SD, the ViT-Small model respectively achieves 98.9%, 98.4%, 99.3% and 98.9% accuracies when compared with its teachers (ViT-Base) for image classification, object detection, semantic segmentation and video recognition tasks. The lightweight ResNet models are improved to a new height on image classification task. The code is available at github.com/pengzhiliang/G2SD. Zhiliang Peng, Li Dong 0004, Furu Wei, Qixiang Ye, Jianbin Jiao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Pre-Training to Learn in ContextabstractIn-context learning, where pre-trained language models learn to perform tasks from task examples and instructions in their contexts, has attracted much attention in the NLP community.However, the ability of in-context learning is not fully exploited because language models are not explicitly trained to learn in context.To this end, we propose PICL (Pretraining for In-Context Learning), a framework to enhance the language models' in-context learning ability by pre-training the model on a large collection of "intrinsic tasks" in the general plain-text corpus using the simple language modeling objective.PICL encourages the model to infer and perform tasks by conditioning on the contexts while maintaining task generalization of pre-trained models.We evaluate the in-context learning performance of the model trained with PICL on seven widelyused text classification datasets and the SUPER-NATURALINSTRCTIONS benchmark, which contains 100+ NLP tasks formulated to text generation.Our experiments show that PICL is more effective and task-generalizable than a range of baselines, outperforming larger language models with nearly 4x parameters.The code is publicly available at https://github. com/thu-coai/PICL. Yuxian Gu, Li Dong 0004, Furu Wei, Minlie Huang |
ACL (1) | 2 |
| 2023 | Beyond English-Centric Bitexts for Better Multilingual Language Representation LearningabstractBarun Patra, Saksham Singhal, Shaohan Huang, Zewen Chi, Li Dong, Furu Wei, Vishrav Chaudhary, Xia Song. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Barun Patra, Saksham Singhal, Shaohan Huang, Zewen Chi, Li Dong 0004, Furu Wei, Vishrav Chaudhary |
ACL (1) | 5 |
| 2023 | A Length-Extrapolatable TransformerabstractYutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, Furu Wei. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yutao Sun, Li Dong 0004, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Furu Wei |
ACL (1) | 2 |
| 2023 | GanLM: Encoder-Decoder Pre-training with an Auxiliary DiscriminatorabstractJian Yang, Shuming Ma, Li Dong, Shaohan Huang, Haoyang Huang, Yuwei Yin, Dongdong Zhang, Liqun Yang, Furu Wei, Zhoujun Li. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jian Yang 0030, Shuming Ma, Li Dong 0004, Shaohan Huang, Haoyang Huang, Yuwei Yin, Dongdong Zhang 0001, Liqun Yang, Furu Wei, Zhoujun Li 0001 |
ACL (1) | 3 |
| 2023 | Generic-to-Specific Distillation of Masked AutoencodersabstractLarge vision Transformers (ViTs) driven by self-supervised pre-training mechanisms achieved unprecedented progress. Lightweight ViT models limited by the model capacity, however, benefit little from those pre-training mechanisms. Knowledge distillation defines a paradigm to transfer representations from large (teacher) models to small (student) ones. However, the conventional single-stage distillation easily gets stuck on task-specific transfer, failing to retain the task-agnostic knowledge crucial for model generalization. In this study, we propose generic-to-specific distillation (G2SD), to tap the potential of small ViT models under the supervision of large models pretrained by masked autoencoders. In generic distillation, decoder of the small model is encouraged to align feature predictions with hidden representations of the large model, so that task-agnostic knowledge can be transferred. In specific distillation, predictions of the small model are constrained to be consistent with those of the large model, to transfer task-specific features which guarantee task performance. With G2SD, the vanilla ViT-Small model respectively achieves 98.7%, 98.1% and 99.3% the performance of its teacher (ViT-Base) for image classification, object detection, and semantic segmentation, setting a solid baseline for two-stage vision distillation. Code will be available at https://github.com/pengzhiliang/G2SD Zhiliang Peng, Li Dong 0004, Furu Wei, Jianbin Jiao, Qixiang Ye |
CVPR | 3 |
| 2023 | Image as a Foreign Language: BEIT Pretraining for Vision and Vision-Language TasksabstractA big convergence of language, vision, and multimodal pretraining is emerging. In this work, we introduce a general-purpose multimodal foundation model BEIT-3, which achieves excellent transfer performance on both vision and vision-language tasks. Specifically, we advance the big convergence from three aspects: backbone architecture, pretraining task, and model scaling up. We use Multiway Transformers for general-purpose modeling, where the modular architecture enables both deep fusion and modality-specific encoding. Based on the shared backbone, we perform masked “language” modeling on images (Imglish), texts (English), and image-text pairs (“parallel sentences”) in a unified manner. Experimental results show that BEIT-3 obtains remarkable performance on object detection (COCO), semantic segmentation (ADE20K), image classification (ImageNet), visual reasoning (NLVR2), visual question answering (VQAv2), image captioning (COCO), and cross-modal retrieval (Flickr30K, COCO). Wenhui Wang 0003, Hangbo Bao, Li Dong 0004, Johan Bjorck, Zhiliang Peng, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, Furu Wei |
CVPR | 3 |
| 2023 | Non-Contrastive Learning Meets Language-Image Pre-TrainingabstractContrastive language-image pre-training (CLIP) serves as a de-facto standard to align images and texts. Nonetheless, the loose correlation between images and texts of webcrawled data renders the contrastive objective data inefficient and craving for a large training batch size. In this work, we explore the validity of non-contrastive language-image pre-training (nCLIP), and study whether nice properties exhibited in visual self-supervised models can emerge. We empirically observe that the non-contrastive objective benefits representation learning while sufficiently underperforming under zero-shot recognition. Based on the above study, we further introduce xCLIP, a multi-tasking framework combining CLIP and nCLIP, and show that nCLIP aids CLIP in enhancing feature semantics. The synergy between two objectives lets xCLIP enjoy the best of both worlds: superior performance in both zero-shot transfer and representation learning. Systematic evaluation is conducted spanning a wide variety of downstream tasks including zero-shot classification, out-of-domain classification, retrieval, visual representation learning, and textual representation learning, showcasing a consistent performance gain and validating the effectiveness of xCLIP. The code and pre-trained models will be publicly available at https://github.com/shallowtoil/xclip. Jinghao Zhou, Li Dong 0004, Zhe Gan, Furu Wei |
CVPR | 2 |
| 2023 | Corrupted Image Modeling for Self-Supervised Visual Pre-Training
Li Dong 0004, Hangbo Bao, Xinggang Wang, Furu Wei |
ICLR | 2 |
| 2023 | Prototypical Calibration for Few-shot Learning of Language Models
Zhixiong Han, Yaru Hao, Li Dong 0004, Yutao Sun, Furu Wei |
ICLR | 3 |
| 2023 | Visually-Augmented Language Modeling
Weizhi Wang, Li Dong 0004, Hao Cheng 0002, Haoyu Song 0002, Xiaodong Liu 0003, Xifeng Yan, Jianfeng Gao 0001, Furu Wei |
ICLR | 2 |
| 2023 | Magneto: A Foundation TransformerabstractA big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name ”Transformers”, the above areas use different implementations for better performance, e.g., Post-LayerNorm for BERT, and Pre-LayerNorm for GPT and vision Transformers. We call for the development of Foundation Transformer for true general-purpose modeling, which serves as a go-to architecture for various tasks and modalities with guaranteed training stability. In this work, we introduce a Transformer variant, named Magneto, to fulfill the goal. Specifically, we propose Sub-LayerNorm for good expressivity, and the initialization strategy theoretically derived from DeepNet for stable scaling up. Extensive experiments demonstrate its superior performance and better stability than the de facto Transformer variants designed for various applications, including language modeling (i.e., BERT, and GPT), machine translation, vision pretraining (i.e., BEiT), speech recognition, and multimodal pretraining (i.e., BEiT-3). Hongyu Wang 0009, Shuming Ma, Shaohan Huang, Li Dong 0004, Wenhui Wang 0003, Zhiliang Peng, Payal Bajaj, Saksham Singhal, Alon Benhaim, Barun Patra, Zhun Liu, Vishrav Chaudhary, Furu Wei |
ICML | 4 |
| 2023 | Extensible Prompts for Language Models on Zero-shot Language Style CustomizationabstractWe propose eXtensible Prompt (X-Prompt) for prompting a large language model (LLM) beyond natural language (NL). X-Prompt instructs an LLM with not only NL but also an extensible vocabulary of imaginary words. Registering new imaginary words allows us to instruct the LLM to comprehend concepts that are difficult to describe with NL words, thereby making a prompt more descriptive. Also, these imaginary words are designed to be out-of-distribution (OOD) robust so that they can be (re)used like NL words in various prompts, distinguishing X-Prompt from soft prompt that is for fitting in-distribution data. We propose context-augmented learning (CAL) to learn imaginary words for general usability, enabling them to work properly in OOD (unseen) prompts. We experiment X-Prompt for zero-shot language style customization as a case study. The promising results of X-Prompt demonstrate its potential to facilitate advanced interaction beyond the natural language interface, bridging the communication gap between humans and LLMs. Tao Ge 0001, Jing Hu 0001, Li Dong 0004, Shaoguang Mao, Yan Xia 0005, Xun Wang 0012, Furu Wei |
NeurIPS | 3 |
| 2023 | Optimizing Prompts for Text-to-Image GenerationabstractWell-designed prompts can guide text-to-image models to generate amazing images. However, the performant prompts are often model-specific and misaligned with user input. Instead of laborious human engineering, we propose prompt adaptation, a general framework that automatically adapts original user input to model-preferred prompts. Specifically, we first perform supervised fine-tuning with a pretrained language model on a small collection of manually engineered prompts. Then we use reinforcement learning to explore better prompts. We define a reward function that encourages the policy to generate more aesthetically pleasing images while preserving the original user intentions. Experimental results on Stable Diffusion show that our method outperforms manual prompt engineering in terms of both automatic metrics and human preference ratings. Moreover, reinforcement learning further boosts performance, especially on out-of-domain prompts. Yaru Hao, Zewen Chi, Li Dong 0004, Furu Wei |
NeurIPS | 3 |
| 2023 | Language Is Not All You Need: Aligning Perception with Language ModelsabstractA big convergence of language, multimodal perception, action, and world modeling is a key step toward artificial general intelligence. In this work, we introduce KOSMOS-1, a Multimodal Large Language Model (MLLM) that can perceive general modalities, learn in context (i.e., few-shot), and follow instructions (i.e., zero-shot). Specifically, we train KOSMOS-1 from scratch on web-scale multi-modal corpora, including arbitrarily interleaved text and images, image-caption pairs, and text data. We evaluate various settings, including zero-shot, few-shot, and multimodal chain-of-thought prompting, on a wide range of tasks without any gradient updates or finetuning. Experimental results show that KOSMOS-1 achieves impressive performance on (i) language understanding, generation, and even OCR-free NLP (directly fed with document images), (ii) perception-language tasks, including multimodal dialogue, image captioning, visual question answering, and (iii) vision tasks, such as image recognition with descriptions (specifying classification via text instructions). We also show that MLLMs can benefit from cross-modal transfer, i.e., transfer knowledge from language to multimodal, and from multimodal to language. In addition, we introduce a dataset of Raven IQ test, which diagnoses the nonverbal reasoning capability of MLLMs. Shaohan Huang, Li Dong 0004, Wenhui Wang 0003, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui 0001, Owais Khan Mohammed, Barun Patra, Kriti Aggarwal, Zewen Chi, Johan Bjorck, Vishrav Chaudhary, Subhojit Som, Furu Wei |
NeurIPS | 2 |
| 2023 | Augmenting Language Models with Long-Term MemoryabstractExisting large language models (LLMs) can only afford fix-sized inputs due to the input length limit, preventing them from utilizing rich long-context information from past inputs. To address this, we propose a framework, Language Models Augmented with Long-Term Memory (LongMem), which enables LLMs to memorize long history. We design a novel decoupled network architecture with the original backbone LLM frozen as a memory encoder and an adaptive residual side-network as a memory retriever and reader. Such a decoupled memory design can easily cache and update long-term past contexts for memory retrieval without suffering from memory staleness. Enhanced with memory-augmented adaptation training, LongMem can thus memorize long past context and use long-term memory for language modeling. The proposed memory retrieval module can handle unlimited-length context in its memory bank to benefit various downstream tasks. Typically, LongMem can enlarge the long-form memory to 65k tokens and thus cache many-shot extra demonstration examples as long-form memory for in-context learning. Experiments show that our method outperforms strong long-context models on ChapterBreak, a challenging long-context modeling benchmark, and achieves remarkable improvements on memory-augmented in-context learning over LLMs. The results demonstrate that the proposed method is effective in helping language models to memorize and utilize long-form contents. Weizhi Wang, Li Dong 0004, Hao Cheng 0002, Xiaodong Liu 0003, Xifeng Yan, Jianfeng Gao 0001, Furu Wei |
NeurIPS | 2 |
| 2022 | CLIP Models are Few-Shot Learners: Empirical Studies on VQA and Visual EntailmentabstractCLIP has shown a remarkable zero-shot capability on a wide range of vision tasks.Previously, CLIP is only regarded as a powerful visual encoder.However, after being pretrained by language supervision from a large amount of image-caption pairs, CLIP itself should also have acquired some few-shot abilities for vision-language tasks.In this work, we empirically show that CLIP can be a strong vision-language few-shot learner by leveraging the power of language.We first evaluate CLIP's zero-shot performance on a typical visual question answering task and demonstrate a zero-shot cross-modality transfer capability of CLIP on the visual entailment task.Then we propose a parameter-efficient fine-tuning strategy to boost the few-shot performance on the vqa task.We achieve competitive zero/fewshot results on the visual question answering and visual entailment tasks without introducing any additional pre-training procedure. Haoyu Song 0002, Li Dong 0004, Weinan Zhang 0003, Ting Liu 0001, Furu Wei |
ACL (1) | 2 |
| 2022 | XLM-E: Cross-lingual Language Model Pre-training via ELECTRAabstractZewen Chi, Shaohan Huang, Li Dong, Shuming Ma, Bo Zheng, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Zewen Chi, Shaohan Huang, Li Dong 0004, Shuming Ma, Bo Zheng 0010, Saksham Singhal, Payal Bajaj, Xianling Mao, Heyan Huang, Furu Wei |
ACL (1) | 3 |
| 2022 | StableMoE: Stable Routing Strategy for Mixture of ExpertsabstractThe Mixture-of-Experts (MoE) technique can scale up the model size of Transformers with an affordable computational overhead.We point out that existing learning-to-route MoE methods suffer from the routing fluctuation issue, i.e., the target expert of the same input may change along with training, but only one expert will be activated for the input during inference.The routing fluctuation tends to harm sample efficiency because the same input updates different experts but only one is finally used.In this paper, we propose STABLEMOE with two training stages to address the routing fluctuation problem.In the first training stage, we learn a balanced and cohesive routing strategy and distill it into a lightweight router decoupled from the backbone model.In the second training stage, we utilize the distilled router to determine the token-to-expert assignment and freeze it for a stable routing strategy.We validate our method on language modeling and multilingual machine translation.The results show that STABLEMOE outperforms existing MoE methods in terms of both convergence speed and performance. Damai Dai, Li Dong 0004, Shuming Ma, Bo Zheng 0010, Zhifang Sui, Baobao Chang, Furu Wei |
ACL (1) | 2 |
| 2022 | Knowledge Neurons in Pretrained TransformersabstractLarge-scale pretrained language models are surprisingly good at recalling factual knowledge presented in the training corpus (Petroni et al., 2019; Jiang et al., 2020b).In this paper, we present preliminary studies on how factual knowledge is stored in pretrained Transformers by introducing the concept of knowledge neurons.Specifically, we examine the fill-in-the-blank cloze task for BERT.Given a relational fact, we propose a knowledge attribution method to identify the neurons that express the fact.We find that the activation of such knowledge neurons is positively correlated to the expression of their corresponding facts.In our case studies, we attempt to leverage knowledge neurons to edit (such as update, and erase) specific factual knowledge without fine-tuning.Our results shed light on understanding the storage of knowledge within pretrained Transformers.The code is available at https://github.com/ Hunter-DDM/knowledge-neurons. Damai Dai, Li Dong 0004, Yaru Hao, Zhifang Sui, Baobao Chang, Furu Wei |
ACL (1) | 2 |
| 2022 | Swin Transformer V2: Scaling Up Capacity and ResolutionabstractWe present techniques for scaling Swin Transformer [35] up to 3 billion parameters and making it capable of training with images of up to 1,536x1,536 resolution. By scaling up capacity and resolution, Swin Transformer sets new records on four representative vision benchmarks: 84.0% top-1 accuracy on ImageNet- V2 image classification, 63.1 / 54.4 box / mask mAP on COCO object detection, 59.9 mIoU on ADE20K semantic segmentation, and 86.8% top-1 accuracy on Kinetics-400 video action classification. We tackle issues of training instability, and study how to effectively transfer models pre-trained at low resolutions to higher resolution ones. To this aim, several novel technologies are proposed: 1) a residual post normalization technique and a scaled cosine attention approach to improve the stability of large vision models; 2) a log-spaced continuous position bias technique to effectively transfer models pre-trained at low-resolution images and windows to their higher-resolution counterparts. In addition, we share our crucial implementation details that lead to significant savings of GPU memory consumption and thus make it feasi-ble to train large vision models with regular GPUs. Using these techniques and self-supervised pre-training, we suc-cessfully train a strong 3 billion Swin Transformer model and effectively transfer it to various vision tasks involving high-resolution images or windows, achieving the state-of-the-art accuracy on a variety of benchmarks. Code is avail-able at https://github.com/microsoft/Swin-Transformer. Han Hu 0001, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Yue Cao 0001, Zheng Zhang 0022, Li Dong 0004, Furu Wei, Baining Guo |
CVPR | 10 |
| 2022 | BEiT: BERT Pre-Training of Image Transformers
Hangbo Bao, Li Dong 0004, Furu Wei |
ICLR | 2 |
| 2022 | VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-ExpertsabstractWe present a unified Vision-Language pretrained Model (VLMo) that jointly learns a dual encoder and a fusion encoder with a modular Transformer network. Specifically, we introduce Multiway Transformer, where each block contains a pool of modality-specific experts and a shared self-attention layer. Because of the modeling flexibility of Multiway Transformer, pretrained VLMo can be fine-tuned as a fusion encoder for vision-language classification tasks, or used as a dual encoder for efficient image-text retrieval. Moreover, we propose a stagewise pre-training strategy, which effectively leverages large-scale image-only and text-only data besides image-text pairs. Experimental results show that VLMo achieves state-of-the-art results on various vision-language tasks, including VQA, NLVR2 and image-text retrieval. Hangbo Bao, Wenhui Wang 0003, Li Dong 0004, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, Furu Wei |
NeurIPS | 3 |
| 2022 | On the Representation Collapse of Sparse Mixture of ExpertsabstractSparse mixture of experts provides larger model capacity while requiring a constant computational overhead. It employs the routing mechanism to distribute input tokens to the best-matched experts according to their hidden representations. However, learning such a routing mechanism encourages token clustering around expert centroids, implying a trend toward representation collapse. In this work, we propose to estimate the routing scores between tokens and experts on a low-dimensional hypersphere. We conduct extensive experiments on cross-lingual language model pre-training and fine-tuning on downstream tasks. Experimental results across seven multilingual benchmarks show that our method achieves consistent gains. We also present a comprehensive analysis on the representation and routing behaviors of our models. Our method alleviates the representation collapse issue and achieves more consistent routing than the baseline mixture-of-experts methods. Zewen Chi, Li Dong 0004, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xianling Mao, Heyan Huang, Furu Wei |
NeurIPS | 2 |
| 2022 | Kformer: Knowledge Injection in Transformer Feed-Forward Layers
Yunzhi Yao, Shaohan Huang, Li Dong 0004, Furu Wei, Huajun Chen, Ningyu Zhang 0001 |
NLPCC (1) | 3 |
| 2022 | Transforming Wikipedia Into Augmented Data for Query-Focused SummarizationabstractThe limited size of existing query-focused summarization datasets renders training data-driven summarization models challenging. Meanwhile, the manual construction of a query-focused summarization corpus is costly and time-consuming. In this paper, we use Wikipedia to automatically collect a large query-focused summarization dataset (named WikiRef) of more than 280,000 examples, which can serve as a means of data augmentation. We also develop a BERT-based query-focused summarization model (Q-BERT) to extract sentences from the documents as summaries. To better adapt a huge model containing millions of parameters to tiny benchmarks, we identify and fine-tune only a sparse subnetwork, which corresponds to a small fraction of the whole model parameters. Experimental results on three DUC benchmarks show that the model pre-trained on WikiRefhas already achieved reasonable performance. After fine-tuning on the specific benchmark datasets, the model with data augmentation outperforms strong comparison systems. Moreover, both our proposed Q-BERT model and subnetwork fine-tuning further improve the model performance. Li Dong 0004, Furu Wei, Bing Qin 0001, Ting Liu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Self-Attention Attribution: Interpreting Information Interactions Inside TransformerabstractThe great success of Transformer-based models benefits from the powerful multi-head self-attention mechanism, which learns token dependencies and encodes contextual information from the input. Prior work strives to attribute model decisions to individual input features with different saliency measures, but they fail to explain how these input features interact with each other to reach predictions. In this paper, we propose a self-attention attribution method to interpret the information interactions inside Transformer. We take BERT as an example to conduct extensive studies. Firstly, we apply self-attention attribution to identify the important attention heads, while others can be pruned with marginal performance degradation. Furthermore, we extract the most salient dependencies in each layer to construct an attribution tree, which reveals the hierarchical interactions inside Transformer. Finally, we show that the attribution results can be used as adversarial patterns to implement non-targeted attacks towards BERT. Yaru Hao, Li Dong 0004, Furu Wei, Ke Xu 0001 |
AAAI | 2 |
| 2021 | Improving Pretrained Cross-Lingual Language Models via Self-Labeled Word AlignmentabstractZewen Chi, Li Dong, Bo Zheng, Shaohan Huang, Xian-Ling Mao, Heyan Huang, Furu Wei. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zewen Chi, Li Dong 0004, Bo Zheng 0010, Shaohan Huang, Xianling Mao, Heyan Huang, Furu Wei |
ACL/IJCNLP (1) | 2 |
| 2021 | Consistency Regularization for Cross-Lingual Fine-TuningabstractBo Zheng, Li Dong, Shaohan Huang, Wenhui Wang, Zewen Chi, Saksham Singhal, Wanxiang Che, Ting Liu, Xia Song, Furu Wei. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Bo Zheng 0010, Li Dong 0004, Shaohan Huang, Wenhui Wang 0003, Zewen Chi, Saksham Singhal, Wanxiang Che, Ting Liu 0001, Furu Wei |
ACL/IJCNLP (1) | 2 |
| 2021 | Zero-Shot Cross-Lingual Transfer of Neural Machine Translation with Multilingual Pretrained EncodersabstractPrevious work mainly focuses on improving cross-lingual transfer for NLU tasks with a multilingual pretrained encoder (MPE), or improving the performance on supervised machine translation with BERT.However, it is under-explored that whether the MPE can help to facilitate the cross-lingual transferability of NMT model.In this paper, we focus on a zero-shot cross-lingual transfer task in NMT.In this task, the NMT model is trained with parallel dataset of only one language pair and an off-the-shelf MPE, then it is directly tested on zero-shot language pairs.We propose SixT, a simple yet effective model for this task.SixT leverages the MPE with a two-stage training schedule and gets further improvement with a position disentangled encoder and a capacity-enhanced decoder.Using this method, SixT significantly outperforms mBART, a pretrained multilingual encoderdecoder model explicitly designed for NMT, with an average improvement of 7.1 BLEU on zero-shot any-to-English test sets across 14 source languages.Furthermore, with much less training computation cost and training data, our model achieves better performance on 15 any-to-English test sets than CRISS and m2m-100, two strong multilingual NMT baselines. Guanhua Chen 0001, Shuming Ma, Yun Chen 0007, Li Dong 0004, Dongdong Zhang 0001, Jia Pan 0001, Wenping Wang 0001, Furu Wei |
EMNLP (1) | 4 |
| 2021 | mT6: Multilingual Pretrained Text-to-Text Transformer with Translation PairsabstractZewen Chi, Li Dong, Shuming Ma, Shaohan Huang, Saksham Singhal, Xian-Ling Mao, Heyan Huang, Xia Song, Furu Wei. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zewen Chi, Li Dong 0004, Shuming Ma, Shaohan Huang, Saksham Singhal, Xianling Mao, Heyan Huang, Furu Wei |
EMNLP (1) | 2 |
| 2021 | Allocating Large Vocabulary Capacity for Cross-Lingual Language Model Pre-TrainingabstractCompared to monolingual models, crosslingual models usually require a more expressive vocabulary to represent all languages adequately.We find that many languages are under-represented in recent cross-lingual language models due to the limited vocabulary capacity.To this end, we propose an algorithm VOCAP to determine the desired vocabulary capacity of each language.However, increasing the vocabulary size significantly slows down the pre-training speed.In order to address the issues, we propose k-NN-based target sampling to accelerate the expensive softmax.Our experiments show that the multilingual vocabulary learned with VOCAP benefits cross-lingual language model pre-training.Moreover, k-NN-based target sampling mitigates the side-effects of increasing the vocabulary size while achieving comparable performance and faster pre-training speed.The code and the pretrained multilingual vocabularies are available at https://github. com/bozheng-hit/VoCapXLM. Bo Zheng 0010, Li Dong 0004, Shaohan Huang, Saksham Singhal, Wanxiang Che, Ting Liu 0001, Furu Wei |
EMNLP (1) | 2 |
| 2021 | InfoXLM: An Information-Theoretic Framework for Cross-Lingual Language Model Pre-TrainingabstractZewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, Heyan Huang, Ming Zhou. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zewen Chi, Li Dong 0004, Furu Wei, Nan Yang 0002, Saksham Singhal, Wenhui Wang 0003, Xianling Mao, Heyan Huang, Ming Zhou 0001 |
NAACL-HLT | 2 |
| 2020 | Cross-Lingual Natural Language Generation via Pre-TrainingabstractIn this work we focus on transferring supervision signals of natural language generation (NLG) tasks between multiple languages. We propose to pretrain the encoder and the decoder of a sequence-to-sequence model under both monolingual and cross-lingual settings. The pre-training objective encourages the model to represent different languages in the shared space, so that we can conduct zero-shot cross-lingual transfer. After the pre-training procedure, we use monolingual data to fine-tune the pre-trained model on downstream NLG tasks. Then the sequence-to-sequence model trained in a single language can be directly evaluated beyond that language (i.e., accepting multi-lingual input and producing multi-lingual output). Experimental results on question generation and abstractive summarization show that our model outperforms the machine-translation-based pipeline methods for zero-shot cross-lingual generation. Moreover, cross-lingual transfer improves NLG performance of low-resource languages by leveraging rich-resource language data. Our implementation and data are available at https://github.com/CZWin32768/xnlg. Zewen Chi, Li Dong 0004, Furu Wei, Wenhui Wang 0003, Xianling Mao, Heyan Huang |
AAAI | 2 |
| 2020 | Harvesting and Refining Question-Answer Pairs for Unsupervised QAabstractQuestion Answering (QA) has shown great success thanks to the availability of largescale datasets and the effectiveness of neural models.Recent research works have attempted to extend these successes to the settings with few or no labeled data available.In this work, we introduce two approaches to improve unsupervised QA.First, we harvest lexically and syntactically divergent questions from Wikipedia to automatically construct a corpus of question-answer pairs (named as REFQA).Second, we take advantage of the QA model to extract more appropriate answers, which iteratively refines data over RE-FQA.We conduct experiments 1 on SQuAD 1.1, and NewsQA by fine-tuning BERT without access to manually annotated data.Our approach outperforms previous unsupervised approaches by a large margin and is competitive with early supervised models.We also show the effectiveness of our approach in the fewshot learning setting. Zhongli Li, Wenhui Wang 0003, Li Dong 0004, Furu Wei, Ke Xu 0001 |
ACL | 3 |
| 2020 | Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Xiujun Li, Xi Yin 0006, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu 0006, Lei Zhang 0001, Houdong Hu, Li Dong 0004, Furu Wei, Yejin Choi 0001, Jianfeng Gao 0001 |
ECCV (30) | 9 |
| 2020 | UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingabstractWe propose to pre-train a unified language model for both autoencoding and partially autoregressive language modeling tasks using a novel training procedure, referred to as a pseudo-masked language model (PMLM). Given an input text with masked tokens, we rely on conventional masks to learn inter-relations between corrupted tokens and context via autoencoding, and pseudo masks to learn intra-relations between masked spans via partially autoregressive modeling. With well-designed position embeddings and self-attention masks, the context encodings are reused to avoid redundant computation. Moreover, conventional masks used for autoencoding provide global masking information, so that all the position embeddings are accessible in partially autoregressive language modeling. In addition, the two tasks pre-train a unified language model as a bidirectional encoder and a sequence-to-sequence decoder, respectively. Our experiments show that the unified language models pre-trained using PMLM achieve new state-of-the-art results on a wide range of language understanding and generation tasks across several widely used benchmarks. The code and pre-trained models are available at https://github.com/microsoft/unilm. Hangbo Bao, Li Dong 0004, Furu Wei, Wenhui Wang 0003, Nan Yang 0002, Xiaodong Liu 0003, Yu Wang 0009, Jianfeng Gao 0001, Ming Zhou 0001, Hsiao-Wuen Hon |
ICML | 2 |
| 2020 | MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersabstractPre-trained language models (e.g., BERT (Devlin et al., 2018) and its variants) have achieved remarkable success in varieties of NLP tasks. However, these models usually consist of hundreds of millions of parameters which brings challenges for fine-tuning and online serving in real-life applications due to latency and capacity constraints. In this work, we present a simple and effective approach to compress large Transformer (Vaswani et al., 2017) based pre-trained models, termed as deep self-attention distillation. The small model (student) is trained by deeply mimicking the self-attention module, which plays a vital role in Transformer networks, of the large model (teacher). Specifically, we propose distilling the self-attention module of the last Transformer layer of the teacher, which is effective and flexible for the student. Furthermore, we introduce the scaled dot-product between values in the self-attention module as the new deep self-attention knowledge, in addition to the attention distributions (i.e., the scaled dot-product of queries and keys) that have been used in existing works. Moreover, we show that introducing a teacher assistant (Mirzadeh et al., 2019) also helps the distillation of large pre-trained Transformer models. Experimental results demonstrate that our monolingual model outperforms state-of-the-art baselines in different parameter size of student models. In particular, it retains more than 99% accuracy on SQuAD 2.0 and several GLUE benchmark tasks using 50% of the Transformer parameters and computations of the teacher model. We also obtain competitive results in applying deep self-attention distillation to multilingual pre-trained models. Wenhui Wang 0003, Furu Wei, Li Dong 0004, Hangbo Bao, Nan Yang 0002, Ming Zhou 0001 |
NeurIPS | 3 |
| 2019 | Data-to-Text Generation with Content Selection and PlanningabstractRecent advances in data-to-text generation have led to the use of large-scale datasets and neural network models which are trained end-to-end, without explicitly modeling what to say and in what order. In this work, we present a neural network architecture which incorporates content selection and planning without sacrificing end-to-end training. We decompose the generation task into two stages. Given a corpus of data records (paired with descriptive documents), we first generate a content plan highlighting which information should be mentioned and in which order and then generate the document while taking the content plan into account. Automatic and human-based evaluation experiments show that our model1 outperforms strong baselines improving the state-of-the-art on the recently released RotoWIRE dataset. Ratish Puduppully, Li Dong 0004, Mirella Lapata |
AAAI | 2 |
| 2019 | Data-to-text Generation with Entity ModelingabstractRecent approaches to data-to-text generation have shown great promise thanks to the use of large-scale datasets and the application of neural network architectures which are trained end-to-end.These models rely on representation learning to select content appropriately, structure it coherently, and verbalize it grammatically, treating entities as nothing more than vocabulary tokens.In this work we propose an entity-centric neural architecture for data-to-text generation.Our model creates entity-specific representations which are dynamically updated.Text is generated conditioned on the data input and entity memory representations using hierarchical attention at each time step.We present experiments on the ROTOWIRE benchmark and a (five times larger) new dataset on the baseball domain which we create.Our results show that the proposed model outperforms competitive baselines in automatic and human evaluation. 1TEAM Inn1 Inn2 Inn3 Inn4 . . .R H E . . .Orioles 1 0 0 0 . . . 2 4 0 . . .Royals 1 0 0 3 . . .9 14 1 . . . Ratish Puduppully, Li Dong 0004, Mirella Lapata |
ACL (1) | 2 |
| 2019 | Learning to Ask Unanswerable Questions for Machine Reading ComprehensionabstractMachine reading comprehension with unanswerable questions is a challenging task.In this work, we propose a data augmentation technique by automatically generating relevant unanswerable questions according to an answerable question paired with its corresponding paragraph that contains the answer.We introduce a pair-to-sequence model for unanswerable question generation, which effectively captures the interactions between the question and the paragraph.We also present a way to construct training data for our question generation models by leveraging the existing reading comprehension dataset.Experimental results show that the pair-to-sequence model performs consistently better compared with the sequence-to-sequence baseline.We further use the automatically generated unanswerable questions as a means of data augmentation on the SQuAD 2.0 dataset, yielding 1.9 absolute F1 improvement with BERT-base model and 1.7 absolute F1 improvement with BERT-large model. Li Dong 0004, Furu Wei, Wenhui Wang 0003, Bing Qin 0001, Ting Liu 0001 |
ACL (1) | 2 |
| 2019 | Visualizing and Understanding the Effectiveness of BERTabstractYaru Hao, Li Dong, Furu Wei, Ke Xu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yaru Hao, Li Dong 0004, Furu Wei, Ke Xu 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Unified Language Model Pre-training for Natural Language Understanding and GenerationabstractThis paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks. The model is pre-trained using three types of language modeling tasks: unidirectional, bidirectional, and sequence-to-sequence prediction. The unified modeling is achieved by employing a shared Transformer network and utilizing specific self-attention masks to control what context the prediction conditions on. UniLM compares favorably with BERT on the GLUE benchmark, and the SQuAD 2.0 and CoQA question answering tasks. Moreover, UniLM achieves new state-of-the-art results on five natural language generation datasets, including improving the CNN/DailyMail abstractive summarization ROUGE-L to 40.51 (2.04 absolute improvement), the Gigaword abstractive summarization ROUGE-L to 35.75 (0.86 absolute improvement), the CoQA generative question answering F1 score to 82.5 (37.1 absolute improvement), the SQuAD question generation BLEU-4 to 22.12 (3.75 absolute improvement), and the DSTC7 document-grounded dialog response generation NIST-4 to 2.67 (human performance is 2.65). The code and pre-trained models are available at https://github.com/microsoft/unilm. Li Dong 0004, Nan Yang 0002, Wenhui Wang 0003, Furu Wei, Xiaodong Liu 0003, Yu Wang 0009, Jianfeng Gao 0001, Ming Zhou 0001, Hsiao-Wuen Hon |
NeurIPS | 1 |
| 2019 | Multitask learning for biomedical named entity recognition with cross-sharing structureabstractBACKGROUND: Biomedical named entity recognition (BioNER) is a fundamental and essential task for biomedical literature mining, which affects the performance of downstream tasks. Most BioNER models rely on domain-specific features or hand-crafted rules, but extracting features from massive data requires much time and human efforts. To solve this, neural network models are used to automatically learn features. Recently, multi-task learning has been applied successfully to neural network models of biomedical literature mining. For BioNER models, using multi-task learning makes use of features from multiple datasets and improves the performance of models. RESULTS: In experiments, we compared our proposed model with other multi-task models and found our model outperformed the others on datasets of gene, protein, disease categories. We also tested the performance of different dataset pairs to find out the best partners of datasets. Besides, we explored and analyzed the influence of different entity types by using sub-datasets. When dataset size was reduced, our model still produced positive results. CONCLUSION: We propose a novel multi-task model for BioNER with the cross-sharing structure to improve the performance of multi-task models. The cross-sharing structure in our model makes use of features from both datasets in the training procedure. Detailed analysis about best partners of datasets and influence between entity categories can provide guidance of choosing proper dataset pairs for multi-task training. Our implementation is available at https://github.com/JogleLew/bioner-cross-sharing . Jiagao Lyu, Li Dong 0004, Ke Xu 0001 |
BMC Bioinform. | 3 |
| 2018 | Coarse-to-Fine Decoding for Neural Semantic ParsingabstractSemantic parsing aims at mapping natural language utterances into structured meaning representations.In this work, we propose a structure-aware neural architecture which decomposes the semantic parsing process into two stages.Given an input utterance, we first generate a rough sketch of its meaning, where low-level information (such as variable names and arguments) is glossed over.Then, we fill in missing details by taking into account the natural language input and the sketch itself.Experimental results on four datasets characteristic of different domains and meaning representations show that our approach consistently improves performance, achieving competitive results despite the use of relatively simple decoders. Li Dong 0004, Mirella Lapata |
ACL (1) | 1 |
| 2018 | Confidence Modeling for Neural Semantic ParsingabstractIn this work we focus on confidence modeling for neural semantic parsers which are built upon sequence-to-sequence models.We outline three major causes of uncertainty, and design various metrics to quantify these factors.These metrics are then used to estimate confidence scores that indicate whether model predictions are likely to be correct.Beyond confidence estimation, we identify which parts of the input contribute to uncertain predictions allowing users to interpret their model, and verify or refine its input.Experimental results show that our confidence model significantly outperforms a widely used method that relies on posterior probability, and improves the quality of interpretation compared to simply relying on attention scores. Li Dong 0004, Chris Quirk, Mirella Lapata |
ACL (1) | 1 |
| 2018 | Proactive Resource Management for LTE in Unlicensed Spectrum: A Deep Learning PerspectiveabstractPerforming cellular long term evolution (LTE) communications in unlicensed spectrum using licensed assisted access LTE (LTE-LAA) is a promising approach to overcome wireless spectrum scarcity. However, to reap the benefits of LTE-LAA, a fair coexistence mechanism with other incumbent WiFi deployments is required. In this paper, a novel deep learning approach is proposed for modeling the resource allocation problem of LTE-LAA small base stations (SBSs). The proposed approach enables multiple SBSs to proactively perform dynamic channel selection, carrier aggregation, and fractional spectrum access while guaranteeing fairness with existing WiFi networks and other LTE-LAA operators. Adopting a proactive coexistence mechanism enables future delay-tolerant LTE-LAA data demands to be served within a given prediction window ahead of their actual arrival time thus avoiding the underutilization of the unlicensed spectrum during off-peak hours while maximizing the total served LTE-LAA traffic load. To this end, a noncooperative game model is formulated in which SBSs are modeled as homo egualis agents that aim at predicting a sequence of future actions and thus achieving long-term equal weighted fairness with wireless local area network and other LTE-LAA operators over a given time horizon. The proposed deep learning algorithm is then shown to reach a mixed-strategy Nash equilibrium, when it converges. Simulation results using real data traces show that the proposed scheme can yield up to 28% and 11% gains over a conventional reactive approach and a proportional fair coexistence mechanism, respectively. The results also show that the proposed framework prevents WiFi performance degradation for a densely deployed LTE-LAA network. Ursula Challita, Li Dong 0004, Walid Saad 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2017 | Learning to Generate Product Reviews from AttributesabstractLi Dong, Shaohan Huang, Furu Wei, Mirella Lapata, Ming Zhou, Ke Xu. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Li Dong 0004, Shaohan Huang, Furu Wei, Mirella Lapata, Ming Zhou 0001, Ke Xu 0001 |
EACL (1) | 1 |
| 2017 | Learning to Paraphrase for Question AnsweringabstractQuestion answering (QA) systems are sensitive to the many different ways natural language expresses the same information need.In this paper we turn to paraphrases as a means of capturing this knowledge and present a general framework which learns felicitous paraphrases for various QA tasks.Our method is trained end-toend using question-answer pairs as a supervision signal.A question and its paraphrases serve as input to a neural scoring model which assigns higher weights to linguistic expressions most likely to yield correct answers.We evaluate our approach on QA over Freebase and answer sentence selection.Experimental results on three datasets show that our framework consistently improves performance, achieving competitive results despite the use of simple QA models. Li Dong 0004, Jonathan Mallinson, Siva Reddy, Mirella Lapata |
EMNLP | 1 |
| 2016 | Language to Logical Form with Neural AttentionabstractSemantic parsing aims at mapping natural language to machine interpretable meaning representations.Traditional approaches rely on high-quality lexicons, manually-built templates, and linguistic features which are either domainor representation-specific.In this paper we present a general method based on an attention-enhanced encoder-decoder model.We encode input utterances into vector representations, and generate their logical forms by conditioning the output sequences or trees on the encoding vectors.Experimental results on four datasets show that our approach performs competitively without using hand-engineered features and is easy to adapt across domains and meaning representations. Li Dong 0004, Mirella Lapata |
ACL (1) | 1 |
| 2016 | Long Short-Term Memory-Networks for Machine ReadingabstractIn this paper we address the question of how to render sequence-level networks better at handling structured input. We propose a machine reading simulator which processes text incrementally from left to right and performs shallow reasoning with memory and attention. The reader extends the Long Short-Term Memory architecture with a memory network in place of a single memory cell. This enables adaptive memory usage during recurrence with neural attention, offering a way to weakly induce relations among tokens. The system is initially designed to process a single sequence but we also demonstrate how to integrate it with an encoder-decoder architecture. Experiments on language modeling, sentiment analysis, and natural language inference show that our model matches or outperforms the state of the art. Jianpeng Cheng 0001, Li Dong 0004, Mirella Lapata |
EMNLP | 2 |
| 2016 | Solving and Generating Chinese Character RiddlesabstractChinese character riddle is a riddle game in which the riddle solution is a single Chinese character.It is closely connected with the shape, pronunciation or meaning of Chinese characters.The riddle description (sentence) is usually composed of phrases with rich linguistic phenomena (such as pun, simile, and metaphor), which are associated to different parts (namely radicals) of the solution character.In this paper, we propose a statistical framework to solve and generate Chinese character riddles.Specifically, we learn the alignments and rules to identify the metaphors between phrases in riddles and radicals in characters.Then, in the solving phase, we utilize a dynamic programming method to combine the identified metaphors to obtain candidate solutions.In the riddle generation phase, we use a template-based method and a replacement-based method to obtain candidate riddle descriptions.We then use Ranking SVM to rerank the candidates both in the solving and generation process.Experimental results in the solving task show that the proposed method outperforms baseline methods.We also get very promising results in the generation task according to human judges. Chuanqi Tan, Furu Wei, Li Dong 0004, Weifeng Lv, Ming Zhou 0001 |
EMNLP | 3 |
| 2016 | Unsupervised Word and Dependency Path Embeddings for Aspect Term Extraction
Yichun Yin, Furu Wei, Li Dong 0004, Kaimeng Xu, Ming Zhang 0004, Ming Zhou 0001 |
IJCAI | 3 |
| 2016 | Adaptive Multi-Compositionality for Recursive Neural Network ModelsabstractRecursive neural network models have achieved promising results in many natural language processing tasks. The main difference among these models lies in the composition function, i.e., how to obtain the vector representation for a phrase or sentence using the representations of words it contains. This paper introduces a novel Adaptive Multi-Compositionality (AdaMC) layer to recursive neural network models. The basic idea is to use more than one composition function and adaptively select them depending on input vectors. We develop a general framework to model the semantic composition as a distribution of these composition functions. The composition functions and parameters used for adaptive selection are jointly learnt from the supervision of specific tasks. We integrate AdaMC into existing recursive neural network models and conduct extensive experiments on the Stanford Sentiment Treebank and semantic relation classification task. The experimental results demonstrate that AdaMC improves the performance of recursive neural network models and outperforms the baseline methods. Li Dong 0004, Furu Wei, Ke Xu 0001, Shixia Liu, Ming Zhou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | Ranking with Recursive Neural Networks and Its Application to Multi-Document SummarizationabstractWe develop a Ranking framework upon Recursive Neural Networks (R2N2) to rank sentences for multi-document summarization. It formulates the sentence ranking task as a hierarchical regression process, which simultaneously measures the salience of a sentence and its constituents (e.g., phrases) in the parsing tree. This enables us to draw on word-level to sentence-level supervisions derived from reference summaries.In addition, recursive neural networks are used to automatically learn ranking features over the tree, with hand-crafted feature vectors of words as inputs. Hierarchical regressions are then conducted with learned features concatenating raw features.Ranking scores of sentences and words are utilized to effectively select informative and non-redundant sentences to generate summaries.Experiments on the DUC 2001, 2002 and 2004 multi-document summarization datasets show that R2N2 outperforms state-of-the-art extractive summarization approaches. Ziqiang Cao, Furu Wei, Li Dong 0004, Sujian Li, Ming Zhou 0001 |
AAAI | 3 |
| 2015 | Question Answering over Freebase with Multi-Column Convolutional Neural NetworksabstractLi Dong, Furu Wei, Ming Zhou, Ke Xu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Li Dong 0004, Furu Wei, Ming Zhou 0001, Ke Xu 0001 |
ACL (1) | 1 |
| 2015 | A Hybrid Neural Model for Type Classification of Entity Mentions
Li Dong 0004, Furu Wei, Ming Zhou 0001, Ke Xu 0001 |
IJCAI | 1 |
| 2015 | A Statistical Parsing Framework for Sentiment ClassificationabstractWe present a statistical parsing framework for sentence-level sentiment classification in this article. Unlike previous works that use syntactic parsing results for sentiment analysis, we develop a statistical parser to directly analyze the sentiment structure of a sentence. We show that complicated phenomena in sentiment analysis (e.g., negation, intensification, and contrast) can be handled the same way as simple and straightforward sentiment expressions in a unified and probabilistic way. We formulate the sentiment grammar upon Context-Free Grammars (CFGs), and provide a formal description of the sentiment parsing framework. We develop the parsing model to obtain possible sentiment parse trees for a sentence, from which the polarity model is proposed to derive the sentiment strength and polarity, and the ranking model is dedicated to selecting the best sentiment tree. We train the parser directly from examples of sentences annotated only with sentiment polarity labels but without any syntactic annotations or polarity annotations of constituents within sentences. Therefore we can obtain training data easily. In particular, we train a sentiment parser, s.parser, from a large amount of review sentences with users' ratings as rough sentiment polarity labels. Extensive experiments on existing benchmark data sets show significant improvements over baseline sentiment classification approaches. Li Dong 0004, Furu Wei, Shujie Liu 0001, Ming Zhou 0001, Ke Xu 0001 |
Comput. Linguistics | 1 |
| 2015 | A Joint Segmentation and Classification Framework for Sentence Level Sentiment ClassificationabstractIn this paper, we propose a joint segmentation and classification framework for sentence-level sentiment classification. It is widely recognized that phrasal information is crucial for sentiment classification. However, existing sentiment classification algorithms typically split a sentence as a word sequence, which does not effectively handle the inconsistent sentiment polarity between a phrase and the words it contains, such as {“not bad,” “bad”} and {“a great deal of,” “great”}. We address this issue by developing a joint framework for sentence-level sentiment classification. It simultaneously generates useful segmentations and predicts sentence-level polarity based on the segmentation results. Specifically, we develop a candidate generation model to produce segmentation candidates of a sentence; a segmentation ranking model to score the usefulness of a segmentation candidate for sentiment classification; and a classification model for predicting the sentiment polarity of a segmentation. We train the joint framework directly from sentences annotated with only sentiment polarity, without using any syntactic or sentiment annotations in segmentation level. We conduct experiments for sentiment classification on two benchmark datasets: a tweet dataset and a review dataset. Experimental results show that: 1) our method performs comparably with state-of-the-art methods on both datasets; 2) joint modeling segmentation and classification outperforms pipelined baseline methods in various experimental settings. Duyu Tang, Bing Qin 0001, Furu Wei, Li Dong 0004, Ting Liu 0001, Ming Zhou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2014 | Adaptive Multi-Compositionality for Recursive Neural Models with Applications to Sentiment AnalysisabstractRecursive neural models have achieved promising results in many natural language processing tasks. The main difference among these models lies in the composition function, i.e., how to obtain the vector representation for a phrase or sentence using the representations of words it contains. This paper introduces a novel Adaptive Multi-Compositionality (AdaMC) layer to recursive neural models. The basic idea is to use more than one composition functions and adaptively select them depending on the input vectors. We present a general framework to model each semantic composition as a distribution over these composition functions. The composition functions and parameters used for adaptive selection are learned jointly from data. We integrate AdaMC into existing recursive neural models and conduct extensive experiments on the Stanford Sentiment Treebank. The results illustrate that AdaMC significantly outperforms state-of-the-art sentiment classification methods. It helps push the best accuracy of sentence-level negative/positive classification from 85.4% up to 88.5%. Li Dong 0004, Furu Wei, Ming Zhou 0001, Ke Xu 0001 |
AAAI | 1 |
| 2014 | A Joint Segmentation and Classification Framework for Sentiment AnalysisabstractIn this paper, we propose a joint segmentation and classification framework for sentiment analysis.Existing sentiment classification algorithms typically split a sentence as a word sequence, which does not effectively handle the inconsistent sentiment polarity between a phrase and the words it contains, such as "not bad" and "a great deal of ".We address this issue by developing a joint segmentation and classification framework (JSC), which simultaneously conducts sentence segmentation and sentence-level sentiment classification.Specifically, we use a log-linear model to score each segmentation candidate, and exploit the phrasal information of top-ranked segmentations as features to build the sentiment classifier.A marginal log-likelihood objective function is devised for the segmentation model, which is optimized for enhancing the sentiment classification performance.The joint model is trained only based on the annotated sentiment polarity of sentences, without any segmentation annotations.Experiments on a benchmark Twitter sentiment classification dataset in SemEval 2013 show that, our joint model performs comparably with the state-of-the-art methods. Duyu Tang, Furu Wei, Bing Qin 0001, Li Dong 0004, Ting Liu 0001, Ming Zhou 0001 |
EMNLP | 4 |
| 2013 | The Automated Acquisition of Suggestions from TweetsabstractThis paper targets at automatically detecting and classifying user's suggestions from tweets. The short and informal nature of tweets, along with the imbalanced characteristics of suggestion tweets, makes the task extremely challenging. To this end, we develop a classification framework on Factorization Machines, which is effective and efficient especially in classification tasks with feature sparsity settings. Moreover, we tackle the imbalance problem by introducing cost-sensitive learning techniques in Factorization Machines. Extensively experimental studies on a manually annotated real-life data set show that the proposed approach significantly improves the baseline approach, and yields the precision of 71.06% and recall of 67.86%. We also investigate the reason why Factorization Machines perform better. Finally, we introduce the first manually annotated dataset for suggestion classification. Li Dong 0004, Furu Wei, Yajuan Duan, Ming Zhou 0001, Ke Xu 0001 |
AAAI | 1 |
| 2012 | MoodLens: an emoticon-based sentiment analysis system for chinese tweetsabstractRecent years have witnessed the explosive growth of online social media. Weibo, a Twitter-like online social network in China, has attracted more than 300 million users in less than three years, with more than 1000 tweets generated in every second. These tweets not only convey the factual information, but also reflect the emotional states of the authors, which are very important for understanding user behaviors. However, a tweet in Weibo is extremely short and the words it contains evolve extraordinarily fast. Moreover, the Chinese corpus of sentiments is still very small, which prevents the conventional keyword-based methods from being used. In light of this, we build a system called MoodLens, which to our best knowledge is the first system for sentiment analysis of Chinese tweets in Weibo. In MoodLens, 95 emoticons are mapped into four categories of sentiments, i.e. angry, disgusting, joyful, and sad, which serve as the class labels of tweets. We then collect over 3.5 million labeled tweets as the corpus and train a fast Naive Bayes classifier, with an empirical precision of 64.3%. MoodLens also implements an incremental learning method to tackle the problem of the sentiment shift and the generation of new words. Using MoodLens for real-time tweets obtained from Weibo, several interesting temporal and spatial patterns are observed. Also, sentiment variations are well captured by MoodLens to effectively detect abnormal events in China. Finally, by using the highly efficient Naive Bayes classifier, MoodLens is capable of online real-time sentiment monitoring. The demo of MoodLens can be found at http://goo.gl/8DQ65. Jichang Zhao, Li Dong 0004, Junjie Wu 0002, Ke Xu 0001 |
KDD | 2 |