Weizhu Chen

dblp:79/2536 · DBLP profile ↗
← Back
100ranked-venue papers
4as first author
66since 2021 · last 2026
0000-0003-4196-7229ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 95 · 4 first-author · 66 since 2021Databases, data management, data science and information retrieval · 22 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 Gold-Medal-Level Olympiad Geometry Solving with Efficient Heuristic Auxiliary Constructions
abstract
Boyan Duan, Xiao Liang, Shuai Lu, Yaoxiang Wang, Yelong Shen, Kai-Wei Chang, Ying Nian Wu, Mao Yang, Weizhu Chen, Yeyun Gong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Boyan Duan, Yaoxiang Wang, Yelong Shen, Kai-Wei Chang 0001, Ying Nian Wu, Mao Yang 0004, Weizhu Chen, Yeyun Gong
ACL (1)9
2026 Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability
abstract
Xiao Liang, Zhong-Zhi Li, Zhenghao Lin, Eric Hanchen Jiang, Hengyuan Zhang, Yelong Shen, Kai-Wei Chang, Ying Nian Wu, Yeyun Gong, Weizhu Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhongzhi Li, Zhenghao Lin, Eric Hanchen Jiang, Yelong Shen, Kai-Wei Chang 0001, Ying Nian Wu, Yeyun Gong, Weizhu Chen
ACL (1)10
2025 MTL-LoRA: Low-Rank Adaptation for Multi-Task Learning
abstract
Parameter-efficient fine-tuning (PEFT) has been widely employed for domain adaptation, with LoRA being one of the most prominent methods due to its simplicity and effectiveness. However, in multi-task learning (MTL) scenarios, LoRA tends to obscure the distinction between tasks by projecting sparse high-dimensional features from different tasks into the same dense low-dimensional intrinsic space. This leads to task interference and suboptimal performance for LoRA and its variants. To tackle this challenge, we propose MTL-LoRA, which retains the advantages of low-rank adaptation while significantly enhancing MTL capabilities. MTL-LoRA augments LoRA by incorporating additional task-adaptive parameters that differentiate task-specific information and capture shared knowledge across various tasks within low-dimensional spaces. This approach enables pretrained models to jointly adapt to different target domains with a limited number of trainable parameters. Comprehensive experimental results, including evaluations on public academic benchmarks for natural language understanding, commonsense reasoning, and image-text understanding, as well as real-world industrial text Ads relevance datasets, demonstrate that MTL-LoRA outperforms LoRA and its various variants with comparable or even fewer learnable parameters in MTL setting.
Yaming Yang 0001, Dilxat Muhtar, Yelong Shen, Yuefeng Zhan, Yujing Wang 0002, Hao Sun 0015, Feng Sun 0008, Qi Zhang 0066, Weizhu Chen, Yunhai Tong
AAAI11
2025 Key-Point-Driven Data Synthesis with Its Enhancement on Mathematical Reasoning
abstract
Large language models have shown great potential in complex reasoning tasks, yet their performance is often hampered by the scarcity of high-quality and reasoning-focused training datasets. Addressing this challenge, we propose Key-PointDriven Data Synthesis (KPDDS), a novel data synthesis framework that synthesizes question-answer pairs by leveraging key points and exemplar practices from authentic data sources. KPDDS ensures the generation of novel questions with rigorous quality control and substantial scalability. As a result, we present KPMath, an extensive synthetic dataset tailored for mathematical reasoning, comprising over 800K questionanswer pairs. Utilizing KPMath and augmenting it with additional reasoning-intensive corpora, we create the comprehensive KPMath-Plus dataset. Our experiments demonstrate that this dataset can enhance the mathematical reasoning performance of models across various architectures and sizes. The Qwen1.5-72B model, fine-tuned on KPMath-Plus, achieves 87.0% accuracy on GSM8K and 58.3% on MATH, surpassing competitors in the 7B to 72B range and best commercial models like GPT-4 across multiple math reasoning datasets.
Xiao Liu 0029, Yeyun Gong, Zhibin Gou, Yelong Shen, Nan Duan 0001, Weizhu Chen
AAAI7
2025 Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
abstract
Efficiently modeling sequences with infinite context length has long been a challenging problem. Previous approaches have either suffered from quadratic computational complexity or limited extrapolation ability in length generalization. In this work, we present Samba, a simple hybrid architecture that layer-wise combines Mamba, a selective State Space Model (SSM), with Sliding Window Attention (SWA). Samba selectively compresses a given sequence into recurrent hidden states while still maintaining the ability to precisely recall recent memories with the attention mechanism. We scale Samba up to 3.8B parameters with 3.2T training tokens and demonstrate that it significantly outperforms state-of-the-art models across a variety of benchmarks. Pretrained on sequences of 4K length, Samba shows improved perplexity in context lengths of up to 1M in zero-shot. When finetuned on 4K-length sequences, Samba efficiently extrapolates to a 256K context length with perfect memory recall on the Passkey Retrieval task, and exhibits superior retrieval extrapolation on the challenging Phonebook task compared to full-attention models. As a linear-time sequence model, Samba achieves a 3.73× higher throughput compared to Transformers with grouped-query attention for user prompts of 128K length, and a 3.64× speedup when generating 64K tokens with unlimited streaming.
Liliang Ren, Yang Liu 0003, Yadong Lu, Yelong Shen, Chen Liang 0006, Weizhu Chen
ICLR6
2025 LongRoPE2: Near-Lossless LLM Context Window Scaling
abstract
LongRoPE2 is a novel approach that extends the effective context window of pre-trained large language models (LLMs) to the target length, while preserving the performance on the original shorter context window. This is achieved by three contributions: (1) a hypothesis that insufficient training in higher RoPE dimensions contributes to the persistent out-of-distribution (OOD) issues observed in existing methods; (2) an effective RoPE rescaling algorithm that adopts evolutionary search guided by "needle-driven" perplexity to address the insufficient training problem; (3) a mixed context window training approach that fine-tunes model weights to adopt rescaled RoPE for long-context sequences while preserving the short-context performance with the original RoPE. Extensive experiments on LLaMA3-8B and Phi3-mini-3.8B across various benchmarks validate the hypothesis and demonstrate the effectiveness of LongRoPE2. Remarkably, LongRoPE2 extends LLaMA3-8B to achieve a 128K effective context length while retaining over 98.5% of short-context performance, using only 10B tokens – 80x fewer than Meta’s approach, which fails to reach the target effective context length.
Li Lyna Zhang, Siyuan Wang 0003, Gaokai Zhang, Gilsinia Lopez, Fan Yang 0024, Weizhu Chen, Mao Yang 0004
ICML7
2025 SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for training large language models (LLMs) on complex reasoning tasks, such as mathematical problem solving. A prerequisite for the scalability of RLVR is a high-quality problem set with precise and verifiable answers. However, the scarcity of well-crafted human-labeled math problems and limited-verification answers in existing distillation-oriented synthetic datasets limit their effectiveness in RL. Additionally, most problem synthesis strategies indiscriminately expand the problem set without considering the model’s capabilities, leading to low efficiency in generating useful questions. To mitigate this issue, we introduce a Self-aware Weakness-driven problem Synthesis framework (SwS) that systematically identifies model deficiencies and leverages them for problem augmentation. Specifically, we define weaknesses as questions that the model consistently fails to learn through its iterative sampling during RL training. We then extract the core concepts from these failure cases and synthesize new problems to strengthen the model's weak areas in subsequent augmented training, enabling it to focus on and gradually overcome its weaknesses. Without relying on external knowledge distillation, our framework enables robust generalization by empowering the model to self-identify and address its weaknesses in RL, yielding average performance gains of 10% and 7.7% on 7B and 32B models across eight mainstream reasoning benchmarks. Our code and data are available at https://anonymous.4open.science/r/SwS-E6F5/
Zhongzhi Li, Yeyun Gong, Yelong Shen, Ying Nian Wu, Weizhu Chen
NeurIPS8
2025 Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation
abstract
Recent advances in language modeling have demonstrated the effectiveness of State Space Models (SSMs) for efficient sequence modeling. While hybrid architectures such as Samba and the decoder-decoder architecture, YOCO, have shown promising performance gains over Transformers, prior works have not investigated the efficiency potential of representation sharing between SSM layers. In this paper, we introduce the Gated Memory Unit (GMU), a simple yet effective mechanism for efficient memory sharing across layers. We apply it to create SambaY, a decoder-hybrid-decoder architecture that incorporates GMUs in the cross-decoder to share memory readout states from a Samba-based self-decoder. SambaY significantly enhances decoding efficiency, preserves linear pre-filling time complexity, and boosts long-context performance, all while eliminating the need for explicit positional encoding. Through extensive scaling experiments, we demonstrate that our model exhibits a significantly lower irreducible loss compared to a strong YOCO baseline, indicating superior performance scalability under large-scale compute regimes. Our largest model enhanced with Differential Attention, Phi4-mini-Flash-Reasoning, achieves significantly better performance than Phi4-mini-Reasoning on reasoning tasks such as Math500, AIME24/25, and GPQA Diamond without any reinforcement learning, while delivering up to 10× higher decoding throughput on 2K-length prompts with 32K generation length under the vLLM inference framework. We release our training codebase on open-source data at https://github.com/microsoft/ArchScale.
Liliang Ren, Young Jin Kim 0006, Adam Atkinson, Zheng Zhan 0001, Jiankai Sun, Baolin Peng, Shuohang Wang, Hao Cheng 0002, Jianfeng Gao 0001, Weizhu Chen, Yelong Shen
NeurIPS13
2025 Reinforcement Learning for Reasoning in Large Language Models with One Training Example
abstract
We show that reinforcement learning with verifiable reward using one training example (1-shot RLVR) is effective in incentivizing the math reasoning capabilities of large language models (LLMs). Applying RLVR to the base model Qwen2.5-Math-1.5B, we identify a single example that elevates model performance on MATH500 from 36.0\% to 73.6\% (8.6\% improvement beyond format correction), and improves the average performance across six common mathematical reasoning benchmarks from 17.6\% to 35.7\% (7.0\% non-format gain). This result matches the performance obtained using the 1.2k DeepScaleR subset (MATH500: 73.6\%, average: 35.9\%), which contains the aforementioned example. Furthermore, RLVR with only two examples even slightly exceeds these results (MATH500: 74.8\%, average: 36.6\%). Similar substantial improvements are observed across various models (Qwen2.5-Math-7B, Llama3.2-3B-Instruct, DeepSeek-R1-Distill-Qwen-1.5B), RL algorithms (GRPO and PPO), and different math examples. In addition, we identify some interesting phenomena during 1-shot RLVR, including cross-category generalization, increased frequency of self-reflection, and sustained test performance improvement even after the training accuracy has saturated, a phenomenon we term \textit{post-saturation generalization}. Moreover, we verify that the effectiveness of 1-shot RLVR primarily arises from the policy gradient loss, distinguishing it from the "grokking" phenomenon. We also show the critical role of promoting exploration (e.g., by incorporating entropy loss with an appropriate coefficient) in 1-shot RLVR training. We also further discuss related observations about format correction, label robustness and prompt modification. These findings can inspire future work on RLVR efficiency and encourage a re-examination of recent progress and the underlying mechanisms in RLVR. Our code, models, and data are open source at https://github.com/ypwang61/One-Shot-RLVR.
Liliang Ren, Baolin Peng, Hao Cheng 0002, Xuehai He, Jianfeng Gao 0001, Weizhu Chen, Shuohang Wang, Simon S. Du, Yelong Shen
NeurIPS11
2024 Automatic Instruction Evolving for Large Language Models
abstract
Fine-tuning large pre-trained language models with Evol-Instruct has achieved encouraging results across a wide range of tasks.However, designing effective evolving methods for instruction evolution requires substantial human expertise.This paper proposes Auto Evol-Instruct, an end-to-end framework that evolves instruction datasets using large language models without any human effort.The framework automatically analyzes and summarizes suitable evolutionary strategies for the given instruction data and iteratively improves the evolving method based on issues exposed during the instruction evolution process.Our extensive experiments demonstrate that the best method optimized by Auto Evol-Instruct outperforms human-designed methods on various benchmarks, including MT-Bench, AlpacaEval, GSM8K, and HumanEval.
Can Xu 0002, Yingxiu Zhao, Jian-Guang Lou, Weizhu Chen
EMNLP5
2024 Seeking Neural Nuggets: Knowledge Transfer in Large Language Models from a Parametric Perspective
abstract
Large Language Models (LLMs) inherently encode a wealth of knowledge within their parameters through pre-training on extensive corpora. While prior research has delved into operations on these parameters to manipulate the underlying implicit knowledge — encompassing detection, editing, and merging — there remains an ambiguous understanding regarding their transferability across models with varying scales. In this paper, we seek to empirically investigate knowledge transfer from larger to smaller models through a parametric perspective. To achieve this, we employ sensitivity-based techniques to extract and align knowledge-specific parameters between different LLMs. Moreover, the LoRA module is used as the intermediary mechanism for injecting the extracted knowledge into smaller models. Evaluations across four benchmarks validate the efficacy of our proposed method. Our findings highlight the critical factors contributing to the process of parametric knowledge transfer, underscoring the transferability of model parameters across LLMs of different scales. Project website: https://maszhongming.github.io/ParaKnowTransfer.
Ming Zhong 0005, Chenxin An, Weizhu Chen, Jiawei Han 0001
ICLR3
2024 CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
abstract
Recent developments in large language models (LLMs) have been impressive. However, these models sometimes show inconsistencies and problematic behavior, such as hallucinating facts, generating flawed code, or creating offensive and toxic content. Unlike these models, humans typically utilize external tools to cross-check and refine their initial content, like using a search engine for fact-checking, or a code interpreter for debugging. Inspired by this observation, we introduce a framework called CRITIC that allows LLMs, which are essentially “black boxes” to validate and progressively amend their own outputs in a manner similar to human interaction with tools. More specifically, starting with an initial output, CRITIC interacts with appropriate tools to evaluate certain aspects of the text, and then revises the output based on the feedback obtained during this validation process. Comprehensive evaluations involving free-form question answering, mathematical program synthesis, and toxicity reduction demonstrate that CRITIC consistently enhances the performance of LLMs. Meanwhile, our research highlights the crucial importance of external feedback in promoting the ongoing self-improvement of LLMs.
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang 0001, Nan Duan 0001, Weizhu Chen
ICLR7
2024 ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
abstract
Large language models have made significant progress in various language tasks, yet they still struggle with complex mathematics. In this paper, we propose ToRA a series of Tool-integrated Reasoning Agents designed to solve challenging mathematical problems by seamlessly integrating natural language reasoning with the utilization of external tools (e.g., computation libraries and symbolic solvers), thereby amalgamating the analytical prowess of language and the computational efficiency of tools. To train ToRA, we curate interactive tool-use trajectories on mathematical datasets, apply imitation learning on the annotations, and propose output space shaping to further refine models' reasoning behavior. As a result, ToRA models significantly outperform open-source models on 10 mathematical reasoning datasets across all scales with 13%-19% absolute improvements on average. Notably, ToRA-7B reaches 44.6% on the competition-level dataset MATH, surpassing the best open-source model WizardMath-70B by 22% absolute. ToRA-34B is also the first open-source model that achieves an accuracy exceeding 50% on MATH, which significantly outperforms GPT-4's CoT result, and is competitive with GPT-4 solving problems with programs. Additionally, we conduct a comprehensive analysis of the benefits and remaining challenges of tool interaction for mathematical reasoning, providing valuable insights for future research.
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang 0001, Minlie Huang, Nan Duan 0001, Weizhu Chen
ICLR8
2024 LoftQ: LoRA-Fine-Tuning-aware Quantization for Large Language Models
abstract
Quantization is an indispensable technique for serving Large Language Models (LLMs) and has recently found its way into LoRA fine-tuning (Dettmers et al., 2023). In this work we focus on the scenario where quantization and LoRA fine- tuning are applied together on a pre-trained model. In such cases it is common to observe a consistent gap in the performance on downstream tasks between full fine-tuning and quantization plus LoRA fine-tuning approach. In response, we propose LoftQ (LoRA-Fine-Tuning-aware Quantization), a novel quantization framework that simultaneously quantizes an LLM and finds a proper low-rank initialization for LoRA fine-tuning. Such an initialization alleviates the discrep- ancy between the quantized and full-precision model and significantly improves the generalization in downstream tasks. We evaluate our method on natural lan- guage understanding, question answering, summarization, and natural language generation tasks. Experiments show that our method is highly effective and out- performs existing quantization methods, especially in the challenging 2-bit and 2/4-bit mixed precision regimes. We will release our code.
Yifan Yu 0008, Chen Liang 0006, Nikos Karampatziakis, Weizhu Chen, Tuo Zhao
ICLR6
2024 Supervised Knowledge Makes Large Language Models Better In-context Learners
abstract
Large Language Models (LLMs) exhibit emerging in-context learning abilities through prompt engineering. The recent progress in large-scale generative models has further expanded their use in real-world language applications. However, the critical challenge of improving the generalizability and factuality of LLMs in natural language understanding and question answering remains under-explored. While previous in-context learning research has focused on enhancing models to adhere to users' specific instructions and quality expectations, and to avoid undesired outputs, little to no work has explored the use of task-specific fine-tuned Language Models (SLMs) to improve LLMs' in-context learning during the inference stage. Our primary contribution is the establishment of a simple yet effective framework that enhances the reliability of LLMs as it: 1) generalizes out-of-distribution data, 2) elucidates how LLMs benefit from discriminative models, and 3) minimizes hallucinations in generative tasks. Using our proposed plug-in method, enhanced versions of Llama 2 and ChatGPT surpass their original versions regarding generalizability and factuality. We offer a comprehensive suite of resources, including 16 curated datasets, prompts, model checkpoints, and LLM outputs across 9 distinct tasks. Our empirical analysis sheds light on the advantages of incorporating discriminative models into LLMs and highlights the potential of our methodology in fostering more reliable LLMs.
Linyi Yang, Shuibai Zhang, Zhuohao Yu 0001, Guangsheng Bao, Yidong Wang 0003, Jindong Wang 0001, Ruochen Xu, Wei Ye 0004, Xing Xie 0001, Weizhu Chen, Yue Zhang 0004
ICLR10
2024 Make Your LLM Fully Utilize the Context
abstract
While many contemporary large language models (LLMs) can process lengthy input, they still struggle to fully utilize information within the long context, known as the *lost-in-the-middle* challenge. We hypothesize that it stems from insufficient explicit supervision during the long-context training, which fails to emphasize that any position in a long context can hold crucial information. Based on this intuition, our study presents **information-intensive (IN2) training**, a purely data-driven solution to overcome lost-in-the-middle. Specifically, IN2 training leverages a synthesized long-context question-answer dataset, where the answer requires (1) **fine-grained information awareness** on a short segment (~128 tokens) within a synthesized long context (4K-32K tokens), and (2) the **integration and reasoning** of information from two or more short segments. Through applying this information-intensive training on Mistral-7B, we present **FILM-7B** (FIll-in-the-Middle). To thoroughly assess the ability of FILM-7B for utilizing long contexts, we design three probing tasks that encompass various context styles (document, code, and structured-data context) and information retrieval patterns (forward, backward, and bi-directional retrieval). The probing results demonstrate that FILM-7B can robustly retrieve information from different positions in its 32K context window. Beyond these probing tasks, FILM-7B significantly improves the performance on real-world long-context tasks (e.g., 23.5->26.9 F1 score on NarrativeQA), while maintaining a comparable performance on short-context tasks (e.g., 59.3->59.2 accuracy on MMLU).
Shengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng 0001, Jian-Guang Lou, Weizhu Chen
NeurIPS6
2024 Not All Tokens Are What You Need for Pretraining
abstract
Previous language model pre-training methods have uniformly applied a next-token prediction loss to all training tokens. Challenging this norm, we posit that ''Not all tokens in a corpus are equally important for language model training''. Our initial analysis examines token-level training dynamics of language model, revealing distinct loss patterns for different tokens. Leveraging these insights, we introduce a new language model called Rho-1. Unlike traditional LMs that learn to predict every next token in a corpus, Rho-1 employs Selective Language Modeling (SLM), which selectively trains on useful tokens that aligned with the desired distribution. This approach involves scoring training tokens using a reference model, and then training the language model with a focused loss on tokens with higher scores. When continual continual pretraining on 15B OpenWebMath corpus, Rho-1 yields an absolute improvement in few-shot accuracy of up to 30% in 9 math tasks. After fine-tuning, Rho-1-1B and 7B achieved state-of-the-art results of 40.6% and 51.8% on MATH dataset, respectively - matching DeepSeekMath with only 3% of the pretraining tokens. Furthermore, when continual pretraining on 80B general tokens, Rho-1 achieves 6.8% average enhancement across 15 diverse tasks, increasing both data efficiency and performance of the language model pre-training.
Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu 0029, Yelong Shen, Ruochen Xu, Chen Lin 0001, Yujiu Yang 0001, Jian Jiao 0007, Nan Duan 0001, Weizhu Chen
NeurIPS11
2024 WizardArena: Post-training Large Language Models via Simulated Offline Chatbot Arena
abstract
Recent work demonstrates that, post-training large language models with open-domain instruction following data have achieved colossal success. Simultaneously, human Chatbot Arena has emerged as one of the most reasonable benchmarks for model evaluation and developmental guidance. However, the processes of manually curating high-quality training data and utilizing online human evaluation platforms are both expensive and limited. To mitigate the manual and temporal costs associated with post-training, this paper introduces a Simulated Chatbot Arena named WizardArena, which is fully based on and powered by open-source LLMs. For evaluation scenario, WizardArena can efficiently predict accurate performance rankings among different models based on offline test set. For training scenario, we simulate arena battles among various state-of-the-art models on a large scale of instruction data, subsequently leveraging the battle results to constantly enhance target model in both the supervised fine-tuning and reinforcement learning . Experimental results demonstrate that our WizardArena aligns closely with the online human arena rankings, and our models trained on offline extensive battle data exhibit significant performance improvements during SFT, DPO, and PPO stages.
Qingfeng Sun, Can Xu 0002, Pu Zhao 0004, Qingwei Lin, Jian-Guang Lou, Shifeng Chen, Yansong Tang, Weizhu Chen
NeurIPS9
2023 DSEE: Dually Sparsity-embedded Efficient Tuning of Pre-trained Language Models
abstract
Xuxi Chen, Tianlong Chen, Weizhu Chen, Ahmed Hassan Awadallah, Zhangyang Wang, Yu Cheng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Xuxi Chen, Tianlong Chen 0001, Weizhu Chen, Ahmed Awadallah 0001, Zhangyang Wang, Yu Cheng 0001
ACL (1)3
2023 Making Language Models Better Reasoners with Step-Aware Verifier
abstract
Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, Weizhu Chen. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yifei Li 0005, Zeqi Lin, Shizhuo Zhang, Qiang Fu 0015, Bei Chen 0008, Jian-Guang Lou, Weizhu Chen
ACL (1)7
2023 Skill-Based Few-Shot Selection for In-Context Learning
abstract
In-context learning is the paradigm that adapts large language models to downstream tasks by providing a few examples.Few-shot selectionselecting appropriate examples for each test instance separately-is important for in-context learning.In this paper, we propose SKILL-KNN, a skill-based few-shot selection method for in-context learning.The key advantages of SKILL-KNN include: (1) it addresses the problem that existing methods based on pre-trained embeddings can be easily biased by surface natural language features that are not important for the target task; (2) it does not require training or fine-tuning of any models, making it suitable for frequently expanding or changing example banks.The key insight is to optimize the inputs fed into the embedding model, rather than tuning the model itself.Technically, SKILL-KNN generates the skill-based descriptions for each test case and candidate example by utilizing a pre-processing few-shot prompting, thus eliminating unimportant surface features.Experimental results across five cross-domain semantic parsing datasets and six backbone models show that SKILL-KNN significantly outperforms existing methods.* Work done during the internship at Microsoft.DB Schema: employee (name, age, city, …) … Question: Which cities do more than one employee under age 30 come from?Prompting-Based Rewriting Off-the-Shelf Embedding Model Skill-Based Descriptions Input Query Candidate 1:This task requires the greater-than and less-than constraints.Candidate 2:This task requires to apply two constraints on one selected column.Candidate … … Example Bank Skill-Based Selection DB Schema: cinema (name, capacity, location, …) … Question: Find the locations that have more than one movie theater with capacity above 300.SQL Query: SELECT … WHERE capacity > 300 GROUP BY location HAVING count(*) > 1 DB Schema: endowment (school, donator, amount, …) … Question : Find the number of schools that have more than one donator whose donation amount is less than 8
Shengnan An, Zeqi Lin, Qiang Fu 0015, Bei Chen 0008, Nanning Zheng 0001, Weizhu Chen, Jian-Guang Lou
EMNLP7
2023 RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation
abstract
Fengji Zhang, Bei Chen, Yue Zhang, Jacky Keung, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, Weizhu Chen. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Fengji Zhang, Bei Chen 0008, Jacky W. Keung, Jin Liu 0016, Daoguang Zan, Jian-Guang Lou, Weizhu Chen
EMNLP9
2023 CodeT: Code Generation with Generated Tests
Bei Chen 0008, Fengji Zhang, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, Weizhu Chen
ICLR7
2023 DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
Jianfeng Gao 0001, Weizhu Chen
ICLR3
2023 Diffusion-GAN: Training GANs with Diffusion
Zhendong Wang 0005, Huangjie Zheng, Weizhu Chen, Mingyuan Zhou
ICLR4
2023 Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Yu Cheng 0001, Weizhu Chen, Tuo Zhao
ICLR6
2023 Truncated Diffusion Probabilistic Models and Diffusion-based Adversarial Auto-Encoders
Huangjie Zheng, Weizhu Chen, Mingyuan Zhou
ICLR3
2023 LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse Approximation
abstract
Transformer models have achieved remarkable results in various natural language tasks, but they are often prohibitively large, requiring massive memories and computational resources. To re- duce the size and complexity of these models, we propose LoSparse (Low-Rank and Sparse ap- proximation), a novel model compression tech- nique that approximates a weight matrix by the sum of a low-rank matrix and a sparse matrix. Our method combines the advantages of both low- rank approximations and pruning, while avoid- ing their limitations. Low-rank approximation compresses the coherent and expressive parts in neurons, while pruning removes the incoherent and non-expressive parts in neurons. Pruning enhances the diversity of low-rank approxima- tions, and low-rank approximation prevents prun- ing from losing too many expressive neurons. We evaluate our method on natural language under- standing, question answering, and natural lan- guage generation tasks. We show that it signif- icantly outperforms existing compression meth- ods. Our code is publicly available at https: //github.com/yxli2123/LoSparse
Yifan Yu 0008, Qingru Zhang, Chen Liang 0006, Weizhu Chen, Tuo Zhao
ICML6
2023 Less is More: Task-aware Layer-wise Distillation for Language Model Compression
abstract
Layer-wise distillation is a powerful tool to compress large models (i.e. teacher models) into small ones (i.e., student models). The student distills knowledge from the teacher by mimicking the hidden representations of the teacher at every intermediate layer. However, layer-wise distillation is difficult. Since the student has a smaller model capacity than the teacher, it is often under-fitted. Furthermore, the hidden representations of the teacher contain redundant information that the student does not necessarily need for the target task's learning. To address these challenges, we propose a novel Task-aware layEr-wise Distillation (TED). TED designs task-aware filters to align the hidden representations of the student and the teacher at each layer. The filters select the knowledge that is useful for the target task from the hidden representations. As such, TED reduces the knowledge gap between the two models and helps the student to fit better on the target task. We evaluate TED in two scenarios: continual pre-training and fine-tuning. TED demonstrates significant and consistent improvements over existing distillation methods in both scenarios. Code is available at https://github.com/cliang1453/task-aware-distillation.
Chen Liang 0006, Simiao Zuo, Qingru Zhang, Weizhu Chen, Tuo Zhao
ICML5
2023 Text Generation with Diffusion Language Models: A Pre-training Approach with Continuous Paragraph Denoise
abstract
In this paper, we introduce a novel dIffusion language modEl pre-training framework for text generation, which we call GENIE. GENIE is a large-scale pre-trained diffusion language model that consists of an encoder and a diffusion-based decoder, which can generate text by gradually transforming a random noise sequence into a coherent text sequence. To pre-train GENIE on a large-scale language corpus, we design a new continuous paragraph denoise objective, which encourages the diffusion-decoder to reconstruct a clean text paragraph from a corrupted version, while preserving the semantic and syntactic coherence. We evaluate GENIE on four downstream text generation benchmarks, namely XSum, CNN/DailyMail, Gigaword, and CommonGen. Our experimental results show that GENIE achieves comparable performance with the state-of-the-art autoregressive models on these benchmarks, and generates more diverse text samples. The code and models of GENIE are available at https://github.com/microsoft/ProphetNet/tree/master/GENIE.
Zhenghao Lin, Yeyun Gong, Yelong Shen, Zhihao Fan, Chen Lin 0001, Nan Duan 0001, Weizhu Chen
ICML8
2023 HyperTuning: Toward Adapting Large Language Models without Back-propagation
abstract
Fine-tuning large language models for different tasks can be costly and inefficient, and even methods that reduce the number of tuned parameters still require full gradient-based optimization. We propose HyperTuning, a novel approach to model adaptation that uses a hypermodel to generate task-specific parameters for a fixed downstream model. We demonstrate a simple setup for hypertuning with HyperT5, a T5-based hypermodel that produces soft prefixes or LoRA parameters for a frozen T5 model from few-shot examples. We train HyperT5 in two stages: first, hyperpretraining with a modified conditional language modeling objective that trains a hypermodel to generate parameters; second, multi-task fine-tuning (MTF) on a large number of diverse language tasks. We evaluate HyperT5 on P3, MetaICL and Super-NaturalInstructions datasets, and show that it can effectively generate parameters for unseen tasks. Moreover, we show that using hypermodel-generated parameters as initializations for further parameter-efficient fine-tuning improves performance. HyperTuning can thus be a flexible and efficient way to leverage large language models for diverse downstream applications.
Jason Phang, Weizhu Chen
ICML4
2023 Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language Models
abstract
Large language models can perform various reasoning tasks by using chain-of-thought prompting, which guides them to find answers through step-by-step demonstrations. However, the quality of the prompts depends on the demonstrations given to the models, and creating many of them by hand is costly. We introduce Synthetic prompting, a method that leverages a few handcrafted examples to prompt the model to generate more examples by itself, and selects effective demonstrations to elicit better reasoning. Our method alternates between a backward and forward process to generate new examples. The backward process generates a question that match a sampled reasoning chain, so that the question is solvable and clear. The forward process produces a more detailed reasoning chain for the question, improving the quality of the example. We evaluate our method on numerical, symbolic, and algorithmic reasoning tasks, and show that it outperforms existing prompting techniques.
Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan 0001, Weizhu Chen
ICML6
2023 Meet in the Middle: A New Pre-training Paradigm
abstract
Most language models (LMs) are trained and applied in an autoregressive left-to-right fashion, predicting the next token from the preceding ones. However, this ignores that the full sequence is available during training. In this paper, we introduce ``Meet in the Middle'' (MIM) a new pre-training paradigm that improves data efficiency by training in two directions, left-to-right and right-to-left, and encouraging the respective models to agree on their token distribution for each position. While the primary outcome is an improved left-to-right LM, we also obtain secondary benefits in the infilling task. There, we leverage the two pre-trained directions to propose an infilling procedure that builds the completion simultaneously from both sides. We conduct extensive experiments on both programming and natural languages and show that MIM significantly surpasses existing pre-training paradigms, in both left-to-right generation as well as infilling. Code and models available at https://github.com/microsoft/Meet-in-the-Middle
Nikos Karampatziakis, Weizhu Chen
NeurIPS3
2023 In-Context Learning Unlocked for Diffusion Models
abstract
We present Prompt Diffusion, a framework for enabling in-context learning in diffusion-based generative models. Given a pair of task-specific example images, such as depth from/to image and scribble from/to image, and a text guidance, our model automatically understands the underlying task and performs the same task on a new query image following the text guidance. To achieve this, we propose a vision-language prompt that can model a wide range of vision-language tasks and a diffusion model that takes it as input. The diffusion model is trained jointly on six different tasks using these prompts. The resulting Prompt Diffusion model becomes the first diffusion-based vision-language foundation model capable of in-context learning. It demonstrates high-quality in-context generation for the trained tasks and effectively generalizes to new, unseen vision tasks using their respective prompts. Our model also shows compelling text-guided image editing results. Our framework aims to facilitate research into in-context learning for computer vision. We share our code and pre-trained models at https://github.com/Zhendong-Wang/Prompt-Diffusion.
Zhendong Wang 0005, Yifan Jiang 0001, Yadong Lu, Yelong Shen, Weizhu Chen, Zhangyang Wang, Mingyuan Zhou
NeurIPS6
2023 Patch Diffusion: Faster and More Data-Efficient Training of Diffusion Models
abstract
Diffusion models are powerful, but they require a lot of time and data to train. We propose Patch Diffusion, a generic patch-wise training framework, to significantly reduce the training time costs while improving data efficiency, which thus helps democratize diffusion model training to broader users. At the core of our innovations is a new conditional score function at the patch level, where the patch location in the original image is included as additional coordinate channels, while the patch size is randomized and diversified throughout training to encode the cross-region dependency at multiple scales. Sampling with our method is as easy as in the original diffusion model. Through Patch Diffusion, we could achieve $\mathbf{\ge 2\times}$ faster training, while maintaining comparable or better generation quality. Patch Diffusion meanwhile improves the performance of diffusion models trained on relatively small datasets, $e.g.$, as few as 5,000 images to train from scratch. We achieve outstanding FID scores in line with state-of-the-art benchmarks: 1.77 on CelebA-64$\times$64, 1.93 on AFHQv2-Wild-64$\times$64, and 2.72 on ImageNet-256$\times$256. We share our code and pre-trained models at https://github.com/Zhendong-Wang/Patch-Diffusion.
Zhendong Wang 0005, Yifan Jiang 0001, Huangjie Zheng, Peihao Wang, Zhangyang Wang, Weizhu Chen, Mingyuan Zhou
NeurIPS7
2023 AR-Diffusion: Auto-Regressive Diffusion Model for Text Generation
abstract
Diffusion models have gained significant attention in the realm of image generation due to their exceptional performance. Their success has been recently expanded to text generation via generating all tokens within a sequence concurrently. However, natural language exhibits a far more pronounced sequential dependency in comparison to images, and the majority of existing language models are trained with a left-to-right auto-regressive approach. To account for the inherent sequential characteristic of natural language, we introduce Auto-Regressive Diffusion (AR-Diffusion). AR-Diffusion ensures that the generation of tokens on the right depends on the generated ones on the left, a mechanism achieved through employing a dynamic number of denoising steps that vary based on token position. This results in tokens on the left undergoing fewer denoising steps than those on the right, thereby enabling them to generate earlier and subsequently influence the generation of tokens on the right. In a series of experiments on various text generation tasks, including text summarization, machine translation, and common sense generation, AR-Diffusion clearly demonstrated its superiority over existing diffusion language models and that it can be $100\times\sim600\times$ faster when achieving comparable results. Our code is available at https://github.com/microsoft/ProphetNet/tree/master/AR-diffusion.
Zhihao Fan, Xiao Liu 0029, Hai-Tao Zheng 0002, Yeyun Gong, Yelong Shen, Jian Jiao 0007, Zhongyu Wei, Jian Guo 0016, Nan Duan 0001, Weizhu Chen
NeurIPS12
2022 XLM-K: Improving Cross-Lingual Language Model Pre-training with Multilingual Knowledge
abstract
Cross-lingual pre-training has achieved great successes using monolingual and bilingual plain text corpora. However, most pre-trained models neglect multilingual knowledge, which is language agnostic but comprises abundant cross-lingual structure alignment. In this paper, we propose XLM-K, a cross-lingual language model incorporating multilingual knowledge in pre-training. XLM-K augments existing multilingual pre-training with two knowledge tasks, namely Masked Entity Prediction Task and Object Entailment Task. We evaluate XLM-K on MLQA, NER and XNLI. Experimental results clearly demonstrate significant improvements over existing multilingual language models. The results on MLQA and NER exhibit the superiority of XLM-K in knowledge related tasks. The success in XNLI shows a better cross-lingual transferability obtained in XLM-K. What is more, we provide a detailed probing analysis to confirm the desired knowledge captured in our pre-training regimen. The code is available at https://github.com/microsoft/Unicoder/tree/master/pretraining/xlmk.
Xiaoze Jiang, Yaobo Liang, Weizhu Chen
AAAI3
2022 DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation
abstract
Wei Chen, Yeyun Gong, Song Wang, Bolun Yao, Weizhen Qi, Zhongyu Wei, Xiaowu Hu, Bartuer Zhou, Yi Mao, Weizhu Chen, Biao Cheng, Nan Duan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Wei Chen 0088, Yeyun Gong, Song Wang 0012, Bolun Yao, Weizhen Qi, Zhongyu Wei, Xiaowu Hu, Bartuer Zhou, Weizhu Chen, Biao Cheng, Nan Duan 0001
ACL (1)10
2022 A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models
abstract
Large pre-trained vision-language (VL) models can learn a new task with a handful of examples and generalize to a new task without fine-tuning.However, these VL models are hard to deploy for real-world applications due to their impractically huge sizes and slow inference speed.To solve this limitation, we study prompt-based low-resource learning of VL tasks with our proposed method, FEWVLM, relatively smaller than recent fewshot learners.For FEWVLM, we pre-train a sequence-to-sequence transformer model with prefix language modeling (PrefixLM) and masked language modeling (MaskedLM).Furthermore, we analyze the effect of diverse prompts for few-shot tasks.Experimental results on VQA show that FEWVLM with prompt-based learning outperforms Frozen (Tsimpoukelli et al., 2021) which is 31× larger than FEWVLM by 18.2% point and achieves comparable results to a 246× larger model, PICa (Yang et al., 2021).In our analysis, we observe that (1) prompts significantly affect zero-shot performance but marginally affect few-shot performance, (2) models with noisy prompts learn as quickly as hand-crafted prompts given larger training data, and (3) MaskedLM helps VQA tasks while PrefixLM boosts captioning performance.Our code is publicly available at https://github. com/woojeongjin/FewVLM * Work was mainly done while
Woojeong Jin 0001, Yu Cheng 0001, Yelong Shen, Weizhu Chen, Xiang Ren 0001
ACL (1)4
2022 CAMERO: Consistency Regularized Ensemble of Perturbed Language Models with Weight Sharing
abstract
Model ensemble is a popular approach to produce a low-variance and well-generalized model.However, it induces large memory and inference costs, which are often not affordable for real-world deployment.Existing work has resorted to sharing weights among models.However, when increasing the proportion of the shared weights, the resulting models tend to be similar, and the benefits of using model ensemble diminish.To retain ensemble benefits while maintaining a low memory cost, we propose a consistency-regularized ensemble learning approach based on perturbed models, named CAMERO.Specifically, we share the weights of bottom layers across all models and apply different perturbations to the hidden representations for different models, which can effectively promote the model diversity.Meanwhile, we apply a prediction consistency regularizer across the perturbed models to control the variance due to the model diversity.Our experiments using large language models demonstrate that CAMERO significantly improves the generalization performance of the ensemble model.Specifically, CAMERO outperforms the standard ensemble of 8 BERT-base models on the GLUE benchmark by 0.7 with a significantly smaller model size (114.2Mvs. 880.6M).* Work was done during an internship at Microsoft Azure AI.
Chen Liang 0006, Yelong Shen, Weizhu Chen, Tuo Zhao
ACL (1)4
2022 A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text Generation
abstract
Tianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, Bill Dolan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Tianyu Liu 0001, Yizhe Zhang 0002, Chris Brockett, Zhifang Sui, Weizhu Chen, William B. Dolan
ACL (1)6
2022 Scalable Learning to Optimize: A Learned Optimizer Can Train Big Models
Xuxi Chen, Tianlong Chen 0001, Yu Cheng 0001, Weizhu Chen, Ahmed Awadallah 0001, Zhangyang Wang
ECCV (23)4
2022 CodeRetriever: A Large Scale Contrastive Pre-Training Method for Code Search
abstract
Xiaonan Li, Yeyun Gong, Yelong Shen, Xipeng Qiu, Hang Zhang, Bolun Yao, Weizhen Qi, Daxin Jiang, Weizhu Chen, Nan Duan. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Yeyun Gong, Yelong Shen, Xipeng Qiu, Hang Zhang 0029, Bolun Yao, Weizhen Qi, Daxin Jiang, Weizhu Chen, Nan Duan 0001
EMNLP9
2022 Reasoning Like Program Executors
abstract
Reasoning over natural language is a longstanding goal for the research community.However, studies have shown that existing language models are inadequate in reasoning.To address the issue, we present POET, a novel reasoning pre-training paradigm.Through pretraining language models with programs and their execution results, POET empowers language models to harvest the reasoning knowledge possessed by program executors via a data-driven approach.POET is conceptually simple and can be instantiated by different kinds of program executors.In this paper, we showcase two simple instances POET-Math and POET-Logic, in addition to a complex instance, POET-SQL.Experimental results on six benchmarks demonstrate that POET can significantly boost model performance in natural language reasoning, such as numerical reasoning, logical reasoning, and multi-hop reasoning.POET opens a new gate on reasoningenhancement pre-training, and we hope our analysis would shed light on the future research of reasoning like program executors.
Xinyu Pi, Qian Liu 0033, Bei Chen 0008, Morteza Ziyadi, Zeqi Lin, Qiang Fu 0015, Yan Gao 0002, Jian-Guang Lou, Weizhu Chen
EMNLP9
2022 LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen Zhu, Yuanzhi Li, Shean Wang, Weizhu Chen
ICLR8
2022 No Parameters Left Behind: Sensitivity Guided Adaptive Learning Rate for Training Large Transformer Models
Chen Liang 0006, Haoming Jiang, Simiao Zuo, Xiaodong Liu 0003, Jianfeng Gao 0001, Weizhu Chen, Tuo Zhao
ICLR7
2022 TAPEX: Table Pre-training via Learning a Neural SQL Executor
Qian Liu 0033, Bei Chen 0008, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, Jian-Guang Lou
ICLR6
2022 Adversarial Retriever-Ranker for Dense Text Retrieval
Hang Zhang 0029, Yeyun Gong, Yelong Shen, Jiancheng Lv 0001, Nan Duan 0001, Weizhu Chen
ICLR6
2022 PLATON: Pruning Large Transformer Models with Upper Confidence Bound of Weight Importance
abstract
Large Transformer-based models have exhibited superior performance in various natural language processing and computer vision tasks. However, these models contain enormous amounts of parameters, which restrict their deployment to real-world applications. To reduce the model size, researchers prune these models based on the weights’ importance scores. However, such scores are usually estimated on mini-batches during training, which incurs large variability/uncertainty due to mini-batch sampling and complicated training dynamics. As a result, some crucial weights could be pruned by commonly used pruning methods because of such uncertainty, which makes training unstable and hurts generalization. To resolve this issue, we propose PLATON, which captures the uncertainty of importance scores by upper confidence bound of importance estimation. In particular, for the weights with low importance scores but high uncertainty, PLATON tends to retain them and explores their capacity. We conduct extensive experiments with several Transformer-based models on natural language understanding, question answering and image classification to validate the effectiveness of PLATON. Results demonstrate that PLATON manifests notable improvement under different sparsity levels. Our code is publicly available at https://github.com/QingruZhang/PLATON.
Qingru Zhang, Simiao Zuo, Chen Liang 0006, Alexander Bukharin, Weizhu Chen, Tuo Zhao
ICML6
2022 CERT: Continual Pre-training on Sketches for Library-oriented Code Generation
abstract
Code generation is a longstanding challenge, aiming to generate a code snippet based on a natural language description. Usually, expensive text-code paired data is essential for training a code generation model. Recently, thanks to the success of pre-training techniques, large language models are trained on large unlabelled code corpora and perform well in generating code. In this paper, we investigate how to leverage an unlabelled code corpus to train a model for library-oriented code generation. Since it is a common practice for programmers to reuse third-party libraries, in which case the text-code paired data are harder to obtain due to the huge number of libraries. We observe that library-oriented code snippets are more likely to share similar code sketches. Hence, we present CERT with two steps: a sketcher generates the sketch, then a generator fills the details in the sketch. Both the sketcher and generator are continually pre-trained upon a base model using unlabelled data. Also, we carefully craft two benchmarks to evaluate library-oriented code generation named PandasEval and NumpyEval. Experimental results have shown the impressive performance of CERT. For example, it surpasses the base model by an absolute 15.67% improvement in terms of pass@1 on PandasEval. Our work is available at https://github.com/microsoft/PyCodeGPT.
Daoguang Zan, Bei Chen 0008, Dejian Yang, Zeqi Lin, Bei Guan, Yongji Wang 0002, Weizhu Chen, Jian-Guang Lou
IJCAI8
2022 OmniTab: Pretraining with Natural and Synthetic Data for Few-shot Table-based Question Answering
abstract
Zhengbao Jiang, Yi Mao, Pengcheng He, Graham Neubig, Weizhu Chen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Zhengbao Jiang, Graham Neubig, Weizhu Chen
NAACL-HLT5
2022 MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation
abstract
Simiao Zuo, Qingru Zhang, Chen Liang, Pengcheng He, Tuo Zhao, Weizhu Chen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Simiao Zuo, Qingru Zhang, Chen Liang 0006, Tuo Zhao, Weizhu Chen
NAACL-HLT6
2021 UnitedQA: A Hybrid Approach for Open Domain Question Answering
abstract
Hao Cheng, Yelong Shen, Xiaodong Liu, Pengcheng He, Weizhu Chen, Jianfeng Gao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Hao Cheng 0002, Yelong Shen, Xiaodong Liu 0003, Weizhu Chen, Jianfeng Gao 0001
ACL/IJCNLP (1)5
2021 HiddenCut: Simple Data Augmentation for Natural Language Understanding with Better Generalizability
abstract
Jiaao Chen, Dinghan Shen, Weizhu Chen, Diyi Yang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jiaao Chen, Dinghan Shen, Weizhu Chen, Diyi Yang
ACL/IJCNLP (1)3
2021 Super Tickets in Pre-Trained Language Models: From Model Compression to Improving Generalization
abstract
Chen Liang, Simiao Zuo, Minshuo Chen, Haoming Jiang, Xiaodong Liu, Pengcheng He, Tuo Zhao, Weizhu Chen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Chen Liang 0006, Simiao Zuo, Minshuo Chen, Haoming Jiang, Xiaodong Liu 0003, Tuo Zhao, Weizhu Chen
ACL/IJCNLP (1)8
2021 Generation-Augmented Retrieval for Open-Domain Question Answering
abstract
Yuning Mao, Pengcheng He, Xiaodong Liu, Yelong Shen, Jianfeng Gao, Jiawei Han, Weizhu Chen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yuning Mao, Xiaodong Liu 0003, Yelong Shen, Jianfeng Gao 0001, Jiawei Han 0001, Weizhu Chen
ACL/IJCNLP (1)7
2021 Few-Shot Named Entity Recognition: An Empirical Baseline Study
abstract
Jiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose, Shobana Balakrishnan, Weizhu Chen, Baolin Peng, Jianfeng Gao, Jiawei Han. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Jiaxin Huang 0001, Chunyuan Li, Krishan Subudhi, Damien Jose, Shobana Balakrishnan, Weizhu Chen, Baolin Peng, Jianfeng Gao 0001, Jiawei Han 0001
EMNLP (1)6
2021 Finetuning Pretrained Transformers into RNNs
abstract
Jungo Kasai, Hao Peng, Yizhe Zhang, Dani Yogatama, Gabriel Ilharco, Nikolaos Pappas, Yi Mao, Weizhu Chen, Noah A. Smith. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Jungo Kasai, Hao Peng 0009, Yizhe Zhang 0002, Dani Yogatama, Gabriel Ilharco, Nikolaos Pappas 0002, Weizhu Chen, Noah A. Smith
EMNLP (1)8
2021 Adversarial Regularization as Stackelberg Game: An Unrolled Optimization Approach
abstract
Adversarial regularization has been shown to improve the generalization performance of deep learning models in various natural language processing tasks.Existing works usually formulate the method as a zero-sum game, which is solved by alternating gradient descent/ascent algorithms.Such a formulation treats the adversarial and the defending players equally, which is undesirable because only the defending player contributes to the generalization performance.To address this issue, we propose Stackelberg Adversarial Regularization (SALT), which formulates adversarial regularization as a Stackelberg game.This formulation induces a competition between a leader and a follower, where the follower generates perturbations, and the leader trains the model subject to the perturbations.Different from conventional approaches, in SALT, the leader is in an advantageous position.When the leader moves, it recognizes the strategy of the follower and takes the anticipated follower's outcomes into consideration.Such a leader's advantage enables us to improve the model fitting to the unperturbed data.The leader's strategic information is captured by the Stackelberg gradient, which is obtained using an unrolling algorithm.Our experimental results on a set of machine translation and natural language understanding tasks show that SALT outperforms existing adversarial regularization baselines across all tasks.Our code is publicly available.
Simiao Zuo, Chen Liang 0006, Haoming Jiang, Xiaodong Liu 0003, Jianfeng Gao 0001, Weizhu Chen, Tuo Zhao
EMNLP (1)7
2021 Deberta: decoding-Enhanced Bert with Disentangled Attention
Xiaodong Liu 0003, Jianfeng Gao 0001, Weizhu Chen
ICLR4
2021 MixKD: Towards Efficient Distillation of Large-scale Language Models
Kevin J. Liang, Weituo Hao, Dinghan Shen, Yufan Zhou 0001, Weizhu Chen, Changyou Chen, Lawrence Carin
ICLR5
2021 CoDA: Contrast-enhanced and Diversity-promoting Data Augmentation for Natural Language Understanding
Yanru Qu, Dinghan Shen, Yelong Shen, Sandra Sajeev, Weizhu Chen, Jiawei Han 0001
ICLR5
2021 BANG: Bridging Autoregressive and Non-autoregressive Generation with Large Scale Pretraining
abstract
In this paper, we propose BANG, a new pretraining model to Bridge the gap between Autoregressive (AR) and Non-autoregressive (NAR) Generation. AR and NAR generation can be uniformly regarded as to what extent previous tokens can be attended, and BANG bridges AR and NAR generation through designing a novel model structure for large-scale pre-training. A pretrained BANG model can simultaneously support AR, NAR, and semi-NAR generation to meet different requirements. Experiments on question generation (SQuAD 1.1), summarization (XSum), and dialogue generation (PersonaChat) show that BANG improves NAR and semi-NAR performance significantly as well as attaining comparable performance with strong AR pretrained models. Compared with the semi-NAR strong baselines, BANG achieves absolute improvements of 14.01 and 5.24 in the overall scores of SQuAD 1.1 and XSum, respectively. In addition, BANG achieves absolute improvements of 10.73, 6.39, and 5.90 in the overall scores of SQuAD, XSUM, and PersonaChat compared with the NAR strong baselines, respectively. Our code will be made publicly available.
Weizhen Qi, Yeyun Gong, Jian Jiao 0007, Weizhu Chen, Dayiheng Liu, Kewen Tang, Houqiang Li, Jiusheng Chen, Ruofei Zhang, Ming Zhou 0001, Nan Duan 0001
ICML5
2021 Poolingformer: Long Document Modeling with Pooling Attention
abstract
In this paper, we introduce a two-level attention schema, Poolingformer, for long document modeling. Its first level uses a smaller sliding window pattern to aggregate information from neighbors. Its second level employs a larger window to increase receptive fields with pooling attention to reduce both computational cost and memory consumption. We first evaluate Poolingformer on two long sequence QA tasks: the monolingual NQ and the multilingual TyDi QA. Experimental results show that Poolingformer sits atop three official leaderboards measured by F1, outperforming previous state-of-the-art models by 1.9 points (79.8 vs. 77.9) on NQ long answer, 1.9 points (79.5 vs. 77.6) on TyDi QA passage answer, and 1.6 points (67.6 vs. 66.0) on TyDi QA minimal answer. We further evaluate Poolingformer on a long sequence summarization task. Experimental results on the arXiv benchmark continue to demonstrate its superior performance.
Hang Zhang 0029, Yeyun Gong, Yelong Shen, Weisheng Li 0001, Jiancheng Lv 0001, Nan Duan 0001, Weizhu Chen
ICML7
2021 Contextual Bandit Applications in a Customer Support Bot
abstract
Virtual support agents have grown in popularity as a way for businesses to provide better and more accessible customer service. Some challenges in this domain include ambiguous user queries as well as changing support topics and user behavior (non-stationarity). We do, however, have access to partial feedback provided by the user (clicks, surveys, and other events) which can be leveraged to improve the user experience. Adaptable learning techniques, like contextual bandits, are a natural fit for this problem setting. In this paper, we discuss real-world implementations of contextual bandits (CB) for the Microsoft virtual agent. It includes intent disambiguation based on neural-linear bandits (NLB) and contextual recommendations based on a collection of multi-armed bandits (MAB). Our solutions have been deployed to production and have improved key business metrics of the Microsoft virtual agent, as confirmed by A/B experiments. Results include a relative increase of over 12% in problem resolution rate and relative decrease of over 4% in escalations to a human operator. While our current use cases focus on intent disambiguation and contextual recommendation for support bots, we believe our methods can be extended to other domains.
Sandra Sajeev, Jade Huang, Nikos Karampatziakis, Sebastian Kochman, Weizhu Chen
KDD6
2021 Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
abstract
Hyperparameter (HP) tuning in deep learning is an expensive process, prohibitively so for neural networks (NNs) with billions of parameters.We show that, in the recently discovered Maximal Update Parametrization ($\mu$P), many optimal HPs remain stable even as model size changes. This leads to a new HP tuning paradigm we call *$\mu$Transfer*: parametrize the target model in $\mu$P, tune the HP indirectly on a smaller model, and *zero-shot transfer* them to the full-sized model, i.e., without directly tuning the latter at all.We verify $\mu$Transfer on Transformer and ResNet. For example, 1) by transferring pretraining HPs from a model of 13M parameters, we outperform published numbers of BERT-large (350M parameters), with a total tuning cost equivalent to pretraining BERT-large once; 2) by transferring from 40M parameters, we outperform published numbers of the 6.7B GPT-3 model, with tuning cost only 7% of total pretraining cost. A Pytorch implementation of our technique can be found at github.com/microsoft/mup. See arxiv.org for the full, up-to-date version of this work.
Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu 0003, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, Jianfeng Gao 0001
NeurIPS9
2020 SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized Optimization
abstract
Transfer learning has fundamentally changed the landscape of natural language processing (NLP).Many state-of-the-art models are first pre-trained on a large text corpus and then fine-tuned on downstream tasks.However, due to limited data resources from downstream tasks and the extremely high complexity of pre-trained models, aggressive fine-tuning often causes the fine-tuned model to overfit the training data of downstream tasks and fail to generalize to unseen data.To address such an issue in a principled manner, we propose a new learning framework for robust and efficient fine-tuning for pre-trained models to attain better generalization performance.The proposed framework contains two important ingredients: 1. Smoothness-inducing regularization, which effectively manages the complexity of the model; 2. Bregman proximal point optimization, which is an instance of trustregion methods and can prevent aggressive updating.Our experiments show that the proposed framework achieves new state-of-the-art performance on a number of NLP tasks including GLUE, SNLI, SciTail and ANLI.Moreover, it also outperforms the state-of-the-art T5 model, which is the largest pre-trained model containing 11 billion parameters, on GLUE. 1
Haoming Jiang, Weizhu Chen, Xiaodong Liu 0003, Jianfeng Gao 0001, Tuo Zhao
ACL3
2020 Understanding the Difficulty of Training Transformers
abstract
Transformers have proved effective in many NLP tasks.However, their training requires non-trivial efforts regarding carefully designing cutting-edge optimizers and learning rate schedulers (e.g., conventional SGD fails to train Transformers effectively).Our objective here is to understand what complicates Transformer training from both empirical and theoretical perspectives.Our analysis reveals that unbalanced gradients are not the root cause of the instability of training.Instead, we identify an amplification effect that influences training substantially-for each layer in a multi-layer Transformer model, heavy dependency on its residual branch makes training unstable, since it amplifies small parameter perturbations (e.g., parameter updates) and results in significant disturbances in the model output.Yet we observe that a light dependency limits the model potential and leads to inferior trained models.Inspired by our analysis, we propose Admin (Adaptive model initialization) to stabilize the early stage's training and unleash its full potential in the late stage.Extensive experiments show that Admin is more stable, converges faster, and leads to better performance 1 .W (V 2 ) h , where f s is the row-wise softmax function and W (•) h are parameters.
Xiaodong Liu 0003, Jianfeng Gao 0001, Weizhu Chen, Jiawei Han 0001
EMNLP (1)4
2020 Exploiting Structured Knowledge in Text via Graph-Guided Representation Learning
abstract
In this work, we aim at equipping pre-trained language models with structured knowledge.We present two self-supervised tasks learning over raw text with the guidance from knowledge graphs.Building upon entity-level masked language models, our first contribution is an entity masking scheme that exploits relational knowledge underlying the text.This is fulfilled by using a linked knowledge graph to select informative entities and then masking their mentions.In addition, we use knowledge graphs to obtain distractors for the masked entities, and propose a novel distractor-suppressed ranking objective that is optimized jointly with masked language model.In contrast to existing paradigms, our approach uses knowledge graphs implicitly, only during pre-training, to inject language models with structured knowledge via learning from raw text.It is more efficient than retrieval-based methods that perform entity linking and integration during finetuning and inference, and generalizes more effectively than the methods that directly learn from concatenated graph triples.Experiments show that our proposed model achieves improved performance on five benchmarks, including question answering and knowledge base completion.
Tao Shen 0001, Guodong Long, Adam Trischler, Weizhu Chen
EMNLP (1)6
2020 On the Variance of the Adaptive Learning Rate and Beyond
Haoming Jiang, Weizhu Chen, Xiaodong Liu 0003, Jianfeng Gao 0001, Jiawei Han 0001
ICLR4
2019 Multi-Task Deep Neural Networks for Natural Language Understanding
abstract
In this paper, we present a Multi-Task Deep Neural Network (MT-DNN) for learning representations across multiple natural language understanding (NLU) tasks.MT-DNN not only leverages large amounts of cross-task data, but also benefits from a regularization effect that leads to more general representations to help adapt to new tasks and domains.MT-DNN extends the model proposed in Liu et al. (2015) by incorporating a pre-trained bidirectional transformer language model, known as BERT (Devlin et al., 2018).MT-DNN obtains new state-of-the-art results on ten NLU tasks, including SNLI, SciTail, and eight out of nine GLUE tasks, pushing the GLUE benchmark to 82.7% (2.2% absolute improvement) 1 .We also demonstrate using the SNLI and Sc-iTail datasets that the representations learned by MT-DNN allow domain adaptation with substantially fewer in-domain labels than the pre-trained BERT representations.The code and pre-trained models are publicly available at https://github.com/namisan/mt-dnn. * Equal Contribution. 1 As of February 25, 2019 on the latest GLUE test set.
Xiaodong Liu 0003, Weizhu Chen, Jianfeng Gao 0001
ACL (1)3
2019 Parameter-free Sentence Embedding via Orthogonal Basis
abstract
Ziyi Yang, Chenguang Zhu, Weizhu Chen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Weizhu Chen
EMNLP/IJCNLP (1)3
2019 DSCOVR: Randomized Primal-Dual Block Coordinate Algorithms for Asynchronous Distributed Optimization
abstract
Machine learning with big data often involves large optimization models. For distributed optimization over a cluster of machines, frequent communication and synchronization of all model parameters (optimization variables) can be very costly. A promising solution is to use parameter servers to store different subsets of the model parameters, and update them asynchronously at different machines using local datasets. In this paper, we focus on distributed optimization of large linear models with convex loss functions, and propose a family of randomized primal-dual block coordinate algorithms that are especially suitable for asynchronous distributed implementation with parameter servers. In particular, we work with the saddle-point formulation of such problems which allows simultaneous data and model partitioning, and exploit its structure by doubly stochastic coordinate optimization with variance reduction (DSCOVR). Compared with other first-order distributed algorithms, we show that DSCOVR may require less amount of overall computation and communication, and less or no synchronization. We discuss the implementation details of the DSCOVR algorithms, and present numerical experiments on an industrial distributed computing system.
Adams Wei Yu, Qihang Lin, Weizhu Chen
J. Mach. Learn. Res.4
2018 FusionNet: Fusing via Fully-aware Attention with Application to Machine Comprehension
Hsin-Yuan Huang, Yelong Shen, Weizhu Chen
ICLR (Poster)4
2017 ReasoNet: Learning to Stop Reading in Machine Comprehension
abstract
Teaching a computer to read and answer general questions pertaining to a document is a challenging yet unsolved problem. In this paper, we describe a novel neural network architecture called the Reasoning Network (ReasoNet) for machine comprehension tasks. ReasoNets make use of multiple turns to effectively exploit and then reason over the relation among queries, documents, and answers. Different from previous approaches using a fixed number of turns during inference, ReasoNets introduce a termination state to relax this constraint on the reasoning depth. With the use of reinforcement learning, ReasoNets can dynamically determine whether to continue the comprehension process after digesting intermediate results, or to terminate reading when it concludes that existing information is adequate to produce an answer. ReasoNets achieve superior performance in machine comprehension datasets, including unstructured CNN and Daily Mail datasets, the Stanford SQuAD dataset, and a structured Graph Reachability dataset.
Yelong Shen, Po-Sen Huang, Jianfeng Gao 0001, Weizhu Chen
KDD4
2017 Limited-memory Common-directions Method for Distributed Optimization and its Application on Empirical Risk Minimization
abstract
Distributed optimization has become an important research topic for dealing with extremely large volume of data available in the Internet companies nowadays. Additional machines make computation less expensive, but inter-machine communication becomes prominent in the optimization process, and efficient optimization methods should reduce the amount of the communication in order to achieve shorter overall running time. In this work, we utilize the advantages of the recently proposed, theoretically fast-convergent common-directions method, but tackle its main drawback of excessive spatial and computational costs to propose a limited-memory algorithm. The result is an efficient, linear-convergent optimization method for paraliel/distributed optimization. We further discuss how our method can exploit the problem structure to efficiently train regularized empirical risk minimization (ERM) models. Experimental results show that our method outperforms state-of-the-art distributed optimization methods for ERM problems.
Ching-Pei Lee, Po-Wei Wang, Weizhu Chen, Chih-Jen Lin
SDM3
2014 Transfer Understanding from Head Queries to Tail Queries
abstract
One of the biggest challenges of commercial search engines is how to handle tail queries, or queries that occur very infrequently. Frequent queries, also known as head queries, are easier to handle largely because their intents are evidenced by abundant click-through data (query logs). Tail queries have little historical data to rely on, which makes them difficult to be learned by ranking algorithms. In this paper, we leverage knowledge from two resources to fill the gap. The first is a general knowledgebase containing different granularities of concepts automatically harnessed from the Web. The second is the click-through data for head queries. From the click-through data, we obtain an understanding of queries that trigger clicks. Then, we show that by extracting single or multi-word expressions from both head and tail queries and mapping them to a common concept space defined by the knowledgebase, we are able to transfer the click information of the head queries to the tail queries. To validate our approach, we conduct large scale experiments on two real data sets. One is a mixture of head and tail queries, and the other contains pure tail queries. We show that our approach effectively improves tail query search relevance.
Yangqiu Song, Haixun Wang, Weizhu Chen, Shusen Wang
CIKM3
2014 Large-scale L-BFGS using MapReduce
Weizhu Chen, Jingren Zhou 0001
NIPS1
2012 Beyond ten blue links: enabling user click modeling in federated web search
abstract
Click models have been positioned as an effective approach to interpret user click behavior in search engines. Existing click models mostly focus on traditional Web search that considers only ten homogeneous Web HTML documents that appear on the first search-result page. However, in modern commercial search engines, more and more Web search results are federated from multiple sources and contain non-HTML results returned by other heterogeneous vertical engines, such as video or image search engines. In this paper, we study user click behavior in federated search. We observed that user click behavior in federated search is highly different from that in traditional Web search, making it difficult to interpret using existing click models. In response, we propose a novel federated click model (FCM) to interpret user click behavior in federated search. In particular, we take into considerations two new biases in FCM. The first comes from the observation that users tend to be attracted by vertical results and their visual attention on them may increase the examination probability of other nearby web results. The other illustrates that user click behavior on vertical results may lead to more clues of search relevance due to their presentation style in federated search. With these biases and an effective model to correct them, FCM is more accurate in characterizing user click behavior in federated search. Our extensive experimental results show that FCM can outperform other click models in interpreting user click behavior in federated search and achieve significant improvements in terms of both perplexity and log-likelihood.
Danqi Chen 0001, Weizhu Chen, Haixun Wang, Zheng Chen 0001, Qiang Yang 0001
WSDM2
2012 A noise-aware click model for web search
abstract
Recent advances in click model have established it as an attractive approach to infer document relevance. Most of these advances consider the user click/skip behavior as binary events but neglect the context in which a click happens. We show that real click behavior in industrial search engines is often noisy and not always a good indication of relevance. For a considerable percentage of clicks, users select what turn out to be irrelevant documents and these clicks should not be directly used as evidence for relevance inference. Thus in this paper, we put forward an observation that the relevance indication degree of a click is not a constant, but can be differentiated by user preferences and the context in which the user makes her click decision. In particular, to interpret the click behavior discriminatingly, we propose a Noise-aware Click Model (NCM) by characterizing the noise degree of a click, which indicates the quality of the click for inferring relevance. Specifically, the lower the click noise is, the more important the click is in its role for relevance inference. To verify the necessity of explicitly accounting for the uninformative noise in a user click, we conducted experiments on a billion-scale dataset. Extensive experimental results demonstrate that as compared with two state-of-the-art click models in Web Search, NCM can better interpret user click behavior and achieve significant improvements in terms of both perplexity and NDCG.
Weizhu Chen, Dong Wang 0022, Zheng Chen 0001, Adish Singla, Qiang Yang 0001
WSDM1
2012 Personalized click model through collaborative filtering
abstract
Click modeling aims to interpret the users' search click data in order to predict their clicking behavior. Existing models can well characterize the position bias of documents and snippets in relation to users' mainstream click behavior. Yet, current advances depict users' search actions only in a general setting by implicitly assuming that all users act in the same way, regardless of the fact that anyone, motivated with some individual interest, is more likely to click on a link than others. It is in light of this that we put forward a novel personalized click model to describe the user-oriented click preferences, which applies and extends matrix / tensor factorization from the view of collaborative filtering to connect users, queries and documents together. Our model serves as a generalized personalization framework that can be incorporated to the previously proposed click models and, in many cases, to their future extensions. Despite the sparsity of search click data, our personalized model demonstrates its advantage over the best click models previously discussed in the Web-search literature, supported by our large-scale experiments on a real dataset. A delightful bonus is the model's ability to gain insights into queries and documents through latent feature vectors, and hence to handle rare and even new query-document pairs much better than previous click models.
Si Shen, Botao Amber Hu, Weizhu Chen, Qiang Yang 0001
WSDM3
2011 A Whole Page Click Model to Better Interpret Search Engine Click Data
abstract
Recent advances in click modeling have established it as an attractive approach to interpret search click data. These advances characterize users’ search behavior either in advertisement blocks, or within an organic search block through probabilistic models. Yet, when searching for information on a search result page, one is often interacting with the search engine via an entire page instead of a single block. Consequently, previous works that exclusively modeled user behavior in a single block may sacrifice much useful user behavior information embedded in other blocks. To solve this problem, in this paper, we put forward a novel Whole Page Click (WPC) Model to characterize user behavior in multiple blocks. Specifically, WPC uses a Markov chain to learn the user transition probabilities among different blocks in the whole page. To compare our model with the best alternatives in the Web-Search literature, we run a large-scale experiment on a real dataset and demonstrate the advantage of the WPC model in terms of both the whole page and each block in the page. Especially, we find that WPC can achieve significant gain in interpreting the advertisement data, despite of the sparsity of the advertisement click data.
Weizhu Chen, Zhanglong Ji, Si Shen, Qiang Yang 0001
AAAI1
2011 Characterizing Inverse Time Dependency in Multi-class Learning
abstract
The training time of most learning algorithms increases as the size of training data increases. Yet, recent advances in linear binary SVM and LR challenge this commonsense by proposing an inverse dependency property, where the training time decreases as the size of training data increases. In this paper, we study the inverse dependency property of multi-class classification problem. We describe a general framework for multi-class classification problem with a single objective to achieve inverse dependency and extend it to three popular multi-class algorithms. We present theoretical results demonstrating its convergence and inverse dependency guarantee. We conduct experiments to empirically verify the inverse dependency of all the three algorithms on large-scale datasets as well as to ensure the accuracy.
Danqi Chen 0001, Weizhu Chen, Qiang Yang 0001
ICDM2
2011 Short Text Conceptualization Using a Probabilistic Knowledgebase
abstract
Most text mining tasks, including clustering and topic detection, are based on statistical methods that treat text as bags of words. Semantics in the text is largely ignored in the mining process, and mining results often have low interpretability. One particular challenge faced by such approaches lies in short text understanding, as short texts lack enough content from which statistical conclusions can be drawn easily. In this paper, we improve text understanding by using a probabilistic knowl-edgebase that is as rich as our mental world in terms of the concepts (of worldly facts) it contains. We then develop a Bayesian inference mechanism to conceptualize words and short text. We con-ducted comprehensive experiments on conceptual-izing textual terms, and clustering short pieces of text such as Twitter messages. Compared to purely statistical methods such as latent semantic topic modeling or methods that use existing knowledge-bases (e.g., WordNet, Freebase and Wikipedia), our approach brings significant improvements in short text understanding as reflected by the clustering ac-curacy. 1
Yangqiu Song, Haixun Wang, Zhongyuan Wang 0006, Hongsong Li, Weizhu Chen
IJCAI5
2011 User-click modeling for understanding and predicting search-behavior
abstract
Recent advances in search users' click modeling consider both users' search queries and click/skip behavior on documents to infer the user's perceived relevance. Most of these models, including dynamic Bayesian networks (DBN) and user browsing models (UBM), use probabilistic models to understand user click behavior based on individual queries. The user behavior is more complex when her actions to satisfy her information needs form a search session, which may include multiple queries and subsequent click behaviors on various items on search result pages. Previous research is limited to treating each query within a search session in isolation, without paying attention to their dynamic interactions with other queries in a search session.
Weizhu Chen, Dong Wang 0022, Qiang Yang 0001
KDD2
2011 Action prediction and identification from mining temporal user behaviors
abstract
Predicting user's action provides many monetization opportunities to web service providers. If a user's future action can be predicted and identified correctly in time or in advance, we cannot only satisfy user's current need, but also facilitate and simplify user's future online activities. Traditional works on user behavior modeling such as implicit feedback or personalization mainly investigate on users' immediate, short-term or aggregate behaviors. Hence, it is difficult to understand the diversity in temporal user behavior and predict user's future action. In this paper, we consider a forecasting problem of temporal user behavior modeling. Our first objective is able to capture relevant users that will perform an action. The second objective is able to identify whether a user has finished the action, even when the action happened offline. We propose an ensemble algorithm to achieve both objectives. The experiment compares several implementation methods and demonstrates the temporal user behavior modeling using the ensemble algorithm significantly outperforms other methods.
Dakan Wang, Gang Wang 0010, Xiaofeng Ke, Weizhu Chen
WSDM4
2011 Characterizing search intent diversity into click models
abstract
Modeling a user's click-through behavior in click logs is a challenging task due to the well-known position bias problem. Recent advances in click models have adopted the examination hypothesis which distinguishes document relevance from position bias. In this paper, we revisit the examination hypothesis and observe that user clicks cannot be completely explained by relevance and position bias. Specifically, users with different search intents may submit the same query to the search engine but expect different search results. Thus, there might be a bias between user search intent and the query formulated by the user, which can lead to the diversity in user clicks. This bias has not been considered in previous works such as UBM, DBN and CCM. In this paper, we propose a new intent hypothesis as a complement to the examination hypothesis. This hypothesis is used to characterize the bias between the user search intent and the query in each search session. This hypothesis is very general and can be applied to most of the existing click models to improve their capacities in learning unbiased relevance. Experimental results demonstrate that after adopting the intent hypothesis, click models can better interpret user clicks and achieve a significant NDCG improvement.
Botao Amber Hu, Weizhu Chen, Gang Wang 0010, Qiang Yang 0001
WWW3
2010 Explore click models for search ranking
abstract
Recent advances in click model have positioned it as an effective approach to estimate document relevance based on user behavior in web search. Yet, few works have been conducted to explore the use of click model to help web search ranking. In this paper, we focus on learning a ranking function by taking the results from a click model into account. Thus, besides the editorial relevance data arising from the explicit manually labeled search result by experts, we also have the estimated relevance data that is automatically inferred from click models based on user search behavior. We carry out extensive experiments on large-scale commercial datasets and demonstrate the effectiveness of the proposed methods.
Dong Wang 0022, Weizhu Chen, Gang Wang 0010, Botao Amber Hu
CIKM2
2010 Learning click models via probit bayesian inference
abstract
Recent advances in click models have positioned them as an effective approach to the improvement of interpreting click data, and some typical works include UBM, DBN, CCM, etc. After formulating the knowledge of user search behavior into a set of model assumptions, each click model developed an inference method to estimate its parameters. The inference method plays a critical role in terms of accuracy in interpreting clicks, and we observe that different inference methods for a click model can lead to significant accuracy differences. In this paper, we propose a novel Bayesian inference approach for click models. This approach regards click model under a unified framework, which has the following characteristics and advantages:
Dong Wang 0022, Gang Wang 0010, Weizhu Chen, Botao Amber Hu
CIKM4
2010 Incorporating post-click behaviors into a click model
abstract
Much work has attempted to model a user’s click-through behavior by mining the click logs. The task is not trivial due to the well-known position bias problem. Some breakthroughs have been made: two newly proposed click models, DBN and CCM, addressed this problem and improved document relevance estimation. However, to further improve the estimation, we need a model that can capture more sophisticated user behaviors. In particular, after clicking a search result, a user’s behavior (such as the dwell time on the clicked document, and whether there are further clicks on the clicked document) can be highly indicative of the relevance of the document. Unfortunately, such measures have not been incorporated in previous click models. In this paper, we introduce a novel click model, called the post-click click model (PCC), which provides an unbiased estimation of document relevance through leveraging both click behaviors on the search page and post-click behaviors beyond the search page. The PCC model is based on the Bayesian approach, and because of its incremental nature, it is highly scalable to large scale and constantly growing log data. Extensive experimental results illustrate that the proposed method significantly outperforms the state of the art methods merely relying on click logs. 1.
Feimin Zhong, Dong Wang 0022, Gang Wang 0010, Weizhu Chen, Zheng Chen 0001, Haixun Wang
SIGIR4
2010 A novel click model and its applications to online advertising
abstract
Recent advances in click model have positioned it as an attractive method for representing user preferences in web search and online advertising. Yet, most of the existing works focus on training the click model for individual queries, and cannot accurately model the tail queries due to the lack of training data. Simultaneously, most of the existing works consider the query, url and position, neglecting some other important attributes in click log data, such as the local time. Obviously, the click through rate is different between daytime and midnight. In this paper, we propose a novel click model based on Bayesian network, which is capable of modeling the tail queries because it builds the click model on attribute values, with those values being shared across queries. We called our work General Click Model (GCM) as we found that most of the existing works can be special cases of GCM by assigning different parameters. Experimental results on a large-scale commercial advertisement dataset show that GCM can significantly and consistently lead to better results as compared to the state-of-the-art works.
Zeyuan Allen Zhu, Weizhu Chen, Tom Minka, Zheng Chen 0001
WSDM2
2010 Co-optimization of multiple relevance metrics in web search
abstract
Several relevance metrics, such as NDCG, precision and pSkip, are proposed to measure search relevance, where different metrics try to characterize search relevance from different perspectives. Yet we empirically find that the direct optimization of one metric cannot always achieve the optimal ranking of another metric. In this paper, we propose two novel relevance optimization approaches, which take different metrics into a global consideration where the objective is to achieve an ideal tradeoff between different metrics. To achieve this objective, we propose to co-optimize multiple relevance metrics and show their effectiveness.
Dong Wang 0022, Weizhu Chen, Gang Wang 0010, Zheng Chen 0001
WWW3
2009 To divide and conquer search ranking by learning query difficulty
abstract
Learning to rank plays an important role in information retrieval. In most of the existing solutions for learning to rank, all the queries with their returned search results are learnt and ranked with a single model. In this paper, we demonstrate that it is highly beneficial to divide queries into multiple groups and conquer search ranking based on query difficulty. To this end, we propose a method which first characterizes a query using a variety of features extracted from user search behavior, such as the click entropy, the query reformulation probability. Next, a classification model is built on these extracted features to assign a score to represent how difficult a query is. Based on this score, our method automatically divides queries into groups, and trains a specific ranking model for each group to conquer search ranking. Experimental results on RankSVM and RankNet with a large-scale evaluation dataset show that the proposed method can achieve significant improvement in the task of web search ranking.
Zeyuan Allen Zhu, Weizhu Chen, Gang Wang 0010, Zheng Chen 0001
CIKM2
2009 A general magnitude-preserving boosting algorithm for search ranking
abstract
Traditional boosting algorithms for the ranking problems usually employ the pairwise approach and convert the document rating preference into a binary-value label, like RankBoost. However, such a pairwise approach ignores the information about the magnitude of preference in the learning process. In this paper, we present the directed distance function (DDF) as a substitute for binary labels in pairwise approach to preserve the magnitude of preference and propose a new boosting algorithm called MPBoost, which applies GentleBoost optimization and directly incorporates DDF into the exponential loss function. We give the boundedness property of MPBoost through theoretic analysis. Experimental results demonstrate that MPBoost not only leads to better NDCG accuracy as compared to state-of-the-art ranking solutions in both public and commercial datasets, but also has good properties of avoiding the overfitting problem in the task of learning ranking functions.
Weizhu Chen, Zeyuan Allen Zhu, Gang Wang 0010, Dong Wang 0022, Zheng Chen 0001
CIKM2
2009 P-packSVM: Parallel Primal grAdient desCent Kernel SVM
abstract
It is an extreme challenge to produce a nonlinear SVM classifier on very large scale data. In this paper we describe a novel P-packSVM algorithm that can solve the Support Vector Machine (SVM) optimization problem with an arbitrary kernel. This algorithm embraces the best known stochastic gradient descent method to optimize the primal objective, and has 1/ε dependency in complexity to obtain a solution of optimization error ε. The algorithm can be highly parallelized with a special packing strategy, and experiences sub-linear speed-up with hundreds of processors. We demonstrate that P-packSVM achieves accuracy sufficiently close to that of SVM-light, and overwhelms the state-of-the-art parallel SVM trainer PSVM in both accuracy and efficiency. As an illustration, our algorithm trains CCAT dataset with 800k samples in 13 minutes and 95% accuracy, while PSVM needs 5 hours but only has 92% accuracy. We at last demonstrate the capability of P-packSVM on 8 million training samples.
Zeyuan Allen Zhu, Weizhu Chen, Gang Wang 0010, Zheng Chen 0001
ICDM2
2009 Inverse Time Dependency in Convex Regularized Learning
abstract
In the conventional regularized learning, training time increases as the training set expands. Recent work on L2linear SVM challenges this common sense by proposing the inverse time dependency on the training set size. In this paper, we first put forward a Primal Gradient Solver (PGS) to effectively solve the convex regularized learning problem. This solver is based on the stochastic gradient descent method and the Fenchel conjugate adjustment, employing the well-known online strongly convex optimization algorithm with logarithmic regret. We then theoretically prove the inverse dependency property of our PGS, embracing the previous work of the L2linear SVM as a special case and enable the ¿p-norm optimization to run within a bounded sphere, which qualifies more convex loss functions in PGS. We further illustrate this solver in three examples: SVM, logistic regression and regularized least square. Experimental results substantiate the property of the inverse dependency on training data size.
Zeyuan Allen Zhu, Weizhu Chen, Gang Wang 0010, Haixun Wang, Zheng Chen 0001
ICDM2
2008 Mining Translations of Web Queries from Web Click-through Data
Weizhu Chen, Jian Hu 0001, Yansheng Lu, Zheng Chen 0001, Qiang Yang 0001
AAAI2
2008 Web query translation via web log mining
abstract
This paper describes a method to automatically acquire query translation pairs by mining web click-through data. The extraction requires no crawling or Chinese words segmentation, and can capture popular translations. Experimental results on a real click-through data show that only 17.4% of the extracted queries are in the dictionary, and our method can achieve 62.2% (in top-1) to 80.0% (in top-5) precision in translating web queries. Moreover, the extracted translations are semantically relevant to the source query, which is particularly useful for Cross-Lingual Information Retrieval (CLIR).
Weizhu Chen, Yansheng Lu, Zheng Chen 0001, Qiang Yang 0001
SIGIR2
2007 Mining Web Query Hierarchies from Clickthrough Data
Dou Shen, Weizhu Chen, Qiang Yang 0001, Zheng Chen 0001
AAAI3
2007 Document Transformation for Multi-label Feature Selection in Text Categorization
abstract
Feature selection on multi-label documents for automatic text categorization is an under-explored research area. This paper presents a systematic document transformation framework, whereby the multi-label documents are transformed into single-label documents before applying standard feature selection algorithms, to solve the multi-label feature selection problem. Under this framework, we undertake a comparative study on four intuitive document transformation approaches and propose a novel approach called entropy-based label assignment (ELA), which assigns the labels weights to a multi-label document based on label entropy. Three standard feature selection algorithms are utilized for evaluating the document transformation approaches in order to verify its impact on multi-class text categorization problems. Using a SVM classifier and two multi-label evaluation benchmark text collections, we show that the choice of document transformation approaches can significantly influence the performance of multi-class categorization and that our proposed document transformation approach ELA can achieve better performance than all other approaches.
Weizhu Chen, Jun Yan 0001, Benyu Zhang, Zheng Chen 0001, Qiang Yang 0001
ICDM1