Haitao Mi

dblp:19/70 · DBLP profile ↗
← Back
59ranked-venue papers
6as first author
37since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 59 · 6 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EconProver: Towards More Economical Test-Time Scaling for Automated Theorem Proving
abstract
Mukai Li, Linfeng Song, Zhenwen Liang, Jiahao Xu, Shansan Gong, Qi Liu, Haitao Mi, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mukai Li, Linfeng Song, Zhenwen Liang, Shansan Gong, Qi Liu 0049, Haitao Mi, Dong Yu 0001
ACL (1)7
2026 Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience
abstract
Zhenwen Liang, Ruosen Li, Yujun Zhou, Linfeng Song, Dian Yu, Xinya Du, Haitao Mi, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhenwen Liang, Ruosen Li, Yujun Zhou 0002, Linfeng Song, Dian Yu 0001, Xinya Du, Haitao Mi, Dong Yu 0001
ACL (1)7
2026 Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse Domains
abstract
Yi Su, Dian Yu, Linfeng Song, Juntao Li, Haitao Mi, Zhaopeng Tu, Min Zhang, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yi Su 0006, Dian Yu 0001, Linfeng Song, Juntao Li 0005, Haitao Mi, Zhaopeng Tu, Min Zhang 0005, Dong Yu 0001
ACL (1)5
2026 Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding
abstract
Wenkai Wang, Xiyun Li, Hongcan Guo, Wenhao Yu, Tianqing Fang, Haitao Mi, Dong Yu, Shengyu Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiyun Li, Hongcan Guo, Tianqing Fang, Haitao Mi, Shengyu Zhang 0001
ACL (1)6
2026 WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models
abstract
Rui Wang, Ce Zhang, Jun-Yu Ma, Jianshu Zhang, Hongru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Zhisong Zhang, Hongming Zhang, Haitao Mi, Dong Yu, Kam-Fai Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Rui Wang 0015, Ce Zhang 0009, Jun-Yu Ma, Hongru Wang 0003, Yi Chen 0007, Boyang Xue, Tianqing Fang, Zhisong Zhang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001, Kam-Fai Wong
ACL (1)11
2025 LiteSearch: Efficient Tree Search with Dynamic Exploration Budget for Math Reasoning
abstract
Recent research suggests that tree search algorithms (e.g. Monte Carlo Tree Search) can dramatically boost LLM performance on complex mathematical reasoning tasks. However, they often require more than 10 times the computational resources of greedy decoding due to wasteful search strategies, making them difficult to be deployed in practical applications. This study introduces a novel guided tree search algorithm with a goal-directed heuristic function and node-level exploration budget (maximum number of children) calculation to tackle this issue. By considering the search progress towards the final answer (history) and the guidance from a value network (future) trained without any step-wise annotations, our algorithm iteratively selects the most promising tree node before expanding it within the boundaries of the allocated computational budget. Experiments conducted on the GSM8K, TabMWP, and MATH datasets demonstrate that our method not only offers competitive performance but also enjoys significantly lower computational costs compared to baseline methods.
Ante Wang, Linfeng Song, Baolin Peng, Dian Yu 0001, Haitao Mi, Jinsong Su, Dong Yu 0001
AAAI6
2025 Low-Bit Quantization Favors Undertrained LLMs
abstract
Low-bit quantization improves machine learning model efficiency but surprisingly favors undertrained large language models (LLMs).Larger models or those trained on fewer tokens exhibit less quantization-induced degradation (QiD), while smaller, well-trained models face significant performance losses.To gain deeper insights into this trend, we study over 1500+ quantized LLM checkpoints of various sizes and at different training levels (undertrained or fully trained) in a controlled setting, deriving scaling laws for understanding the relationship between QiD and factors: the number of training tokens, model size and bit width.With our derived scaling laws, we propose a novel perspective that we can use QiD to measure an LLM's training levels and determine the number of training tokens required for fully training LLMs of various sizes.Moreover, we use the scaling laws to predict the quantization performance of different-sized LLMs trained with 100 trillion tokens.Our projection shows that the low-bit quantization performance of future models, which are expected to be trained with over 100 trillion tokens, may NOT be desirable.This poses a potential challenge for low-bit quantization in the future and highlights the need for awareness of a model's training level when evaluating lowbit quantization research.To facilitate future research on this problem, we release all the 1500+ quantized checkpoints used in this work at https://huggingface.co/Xu-Ouyang.
Xu Ouyang, Tao Ge 0001, Thomas Hartvigsen, Zhisong Zhang, Haitao Mi, Dong Yu 0001
ACL (1)5
2025 Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls
abstract
Recent advancements in tree search algorithms guided by verifiers have significantly enhanced the reasoning capabilities of large language models (LLMs), but at the cost of increased computational resources. In this work, we identify two key challenges contributing to this inefficiency: \textit{over-exploration} due to redundant states with semantically equivalent content, and \textit{under-exploration} caused by high variance in verifier scoring leading to frequent trajectory switching. To address these issues, we propose FETCH – an e{\bf f}fici{\bf e}nt {\bf t}ree sear{\bf ch} framework, which is a flexible, plug-and-play system compatible with various tree search algorithms.Our framework mitigates over-exploration by merging semantically similar states using agglomerative clustering of text embeddings obtained from a fine-tuned SimCSE model. To tackle under-exploration, we enhance verifiers by incorporating temporal difference learning with adjusted \lambda-returns during training to reduce variance, and employing a verifier ensemble to aggregate scores during inference. Experiments on GSM8K, GSM-Plus, and MATH datasets demonstrate that our methods significantly improve reasoning accuracy and computational efficiency across four different tree search algorithms, paving the way for more practical applications of LLM-based reasoning. The code is available at https://github.com/DeepLearnXMU/Fetch.
Ante Wang, Linfeng Song, Dian Yu 0001, Haitao Mi, Xiangyu Duan, Zhaopeng Tu, Jinsong Su, Dong Yu 0001
ACL (1)5
2025 Entropy Guided Extrapolative Decoding to Improve Factuality in Large Language Models
abstract
Large language models (LLMs) exhibit impressive natural language capabilities but suffer from hallucination – generating content ungrounded in the realities of training data. Recent work has focused on decoding techniques to improve factuality in decoding by leveraging LLMs’ hierarchical representation of factual knowledge, manipulating the predicted distributions at inference time. Current state-of-the-art approaches refine decoding by contrasting logits from a lower layer with the final layer to exploit information related factuality within the model forward procedure. However, such methods often assume the final layer is most reliable one and the lower layer selection process depends on it. In this work, we first propose logit extrapolation of critical token probabilities beyond the last layer for more accurate contrasting. We additionally employ layer-wise entropy-guided lower layer selection, decoupling the selection process from the final layer. Experiments demonstrate strong performance - surpassing state-of-the-art on multiple different datasets by large margins. Analyses show different kinds of prompts respond to different selection strategies.
Lifeng Jin, Linfeng Song, Haitao Mi, Baolin Peng, Dong Yu 0001
COLING4
2025 WebEvolver: Enhancing Web Agent Self-Improvement with Co-evolving World Model
abstract
Agent self-improvement, where agents autonomously train their underlying Large Language Model (LLM) on self-sampled trajectories, shows promising results but often stagnates in web environments due to limited exploration and under-utilization of pretrained web knowledge.To improve the performance of self-improvement, we propose a novel framework that introduces a co-evolving World Model LLM.This world model predicts the next observation based on the current observation and action within the web environment.The World Model serves dual roles: (1) as a virtual web server generating self-instructed training data to continuously refine the agent's policy, and (2) as an imagination engine during inference, enabling look-ahead simulation to guide action selection for the agent LLM.Experiments in real-world web environments (Mind2Web-Live, WebVoyager, and GAIAweb) show a 10% performance gain over existing self-evolving agents, demonstrating the efficacy and generalizability of our approach, without using any distillation from more powerful close-sourced models 1 .
Tianqing Fang, Hongming Zhang 0009, Zhisong Zhang, Kaixin Ma, Wenhao Yu 0002, Haitao Mi, Dong Yu 0001
EMNLP6
2025 Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and Extrapolation
abstract
Mamba's theoretical infinite-context potential is limited in practice when sequences far exceed training lengths.This work explores unlocking Mamba's long-context memory ability by a simple-yet-effective method, Recall with Reasoning (RwR), by distilling chain-ofthought (CoT) summarization from a teacher model.Specifically, RwR prepends these summarization as CoT prompts during fine-tuning, teaching Mamba to actively recall and reason over long contexts.Experiments on LONG-MEMEVAL and HELMET show that RwR outperforms existing long-term memory methods on the Mamba model.Furthermore, under similar pre-training conditions, RwR improves the long-context performance of Mamba relative to comparable Transformer/hybrid baselines while preserving short-context capabilities, all without changing the architecture.
Jun-Yu Ma, Tianqing Fang, Zhisong Zhang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001
EMNLP5
2025 DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
abstract
Enhancing the capability of large language models (LLMs) in reasoning has gained significant attention in recent years. Previous studies have demonstrated the effectiveness of various prompting strategies in aiding LLMs in reasoning (called "reasoning actions"), such as step-by-step thinking, reflecting before answering, solving with programs, and their combinations. However, these approaches often applied static, predefined reasoning actions uniformly to all questions, without considering the specific characteristics of each question or the capability of the task-solving LLM. In this paper, we propose DOTS, an approach enabling LLMs to reason Dynamically via Optimal reasoning Trajectories Search, tailored to the specific characteristics of each question and the inherent capability of the task-solving LLM. Our approach involves three key steps: i) defining atomic reasoning action modules that can be composed into various reasoning action trajectories; ii) searching for the optimal action trajectory for each training question through iterative exploration and evaluation for the specific task-solving LLM; and iii) using the collected optimal trajectories to train an LLM to plan for the reasoning trajectories of unseen questions. In particular, we propose two learning paradigms, i.e., fine-tuning an external LLM as a planner to guide the task-solving LLM, or directly fine-tuning the task-solving LLM with an internalized capability for reasoning actions planning. Our experiments across eight reasoning tasks show that our method consistently outperforms static reasoning techniques and the vanilla instruction tuning approach. Further analysis reveals that our method enables LLMs to adjust their computation based on problem complexity, allocating deeper thinking and reasoning to harder problems.
Murong Yue, Wenlin Yao, Haitao Mi, Dian Yu 0001, Ziyu Yao 0002, Dong Yu 0001
ICLR3
2025 Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
abstract
Reinforcement Learning with Human Feedback (RLHF) has achieved great success in aligning large language models (LLMs) with human preferences. Prevalent RLHF approaches are reward-based, following the Bradley-Terry (BT) model assumption, which may not fully capture the complexity of human preferences. In this paper, we explore RLHF under a general preference framework and approach it from a game-theoretic perspective. Specifically, we formulate the problem as a two-player game and propose a novel online algorithm, iterative Nash policy optimization (INPO). The key idea is to let the policy play against itself via no- regret learning, thereby approximating the Nash policy. Unlike previous methods, INPO bypasses the need for estimating the expected win rate for individual responses, which typically incurs high computational or annotation costs. Instead, we introduce a new loss objective that is directly minimized over a preference dataset. We provide theoretical analysis for our approach and demonstrate its effectiveness through experiments on various representative benchmarks. With an LLaMA-3-8B-based SFT model, INPO achieves a 42.6% length-controlled win rate on AlpacaEval 2.0 and a 37.8% win rate on Arena-Hard, showing substantial improvement over the state-of-the-art online RLHF algorithms.
Dian Yu 0001, Baolin Peng, Linfeng Song, Mingyue Huo, Nan Jiang 0008, Haitao Mi, Dong Yu 0001
ICLR8
2025 Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models
abstract
The remarkable performance of long reasoning models can be attributed to their ability to emulate human-like long-time thinking during inference. These models employ extended chain-of-thought (CoT) processes, exploring multiple strategies to enhance problem-solving capabilities. However, a critical question remains: How to intelligently and efficiently scale computational resources during testing. This paper presents the first comprehensive study on the prevalent issue of overthinking in these models, where long reasoning models generate redundant solutions that contribute minimally to accuracy and diversity, thereby wasting computational resources on simple problems with minimal benefit. We introduce novel efficiency metrics from both outcome and process perspectives to evaluate the rational use of computational resources by long reasoning models. Using a self-training paradigm, we propose strategies to mitigate overthinking, simplifying reasoning processes without compromising accuracy. Experimental results show that our approach successfully reduces computational overhead while preserving model performance across a range of testsets with varying difficulty levels, such as GSM8K, MATH500, GPQA, and AIME. Our code is open-source and available at https://github.com/galaxyChen/overthinking.
Zhiwei He 0002, Jianhui Pang, Dian Yu 0001, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang 0001, Rui Wang 0015, Zhaopeng Tu, Haitao Mi, Dong Yu 0001
ICML13
2025 UniGist: Towards General and Hardware-aligned Sequence-level Long Context Compression
abstract
Large language models are increasingly capable of handling long-context inputs, but the memory overhead of KV cache remains a major bottleneck for general-purpose deployment. While many compression strategies have been explored, sequence-level compression is particularly challenging due to its tendency to lose important details. We present UniGist, a gist token-based long context compression framework that removes the need for chunk-wise training, enabling the model to learn how to compress and utilize long-range context during training. To fully exploit the sparsity, we introduce a gist shift trick that transforms the attention layout into a right-aligned block structure and develop a block-table-free sparse attention kernel based on it. UniGist further supports one-pass training and flexible chunk sizes during inference, allowing efficient and adaptive context processing. Experiments across multiple long-context tasks show that UniGist significantly improves compression quality, with especially strong performance in recalling details and long-range dependency modeling.
Chenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li, Tianqing Fang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001, Zhicheng Dou
NeurIPS7
2025 The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
abstract
Improving the reasoning capabilities of large language models (LLMs) typically requires supervised fine-tuning with labeled data or computationally expensive sampling. We introduce Unsupervised Prefix Fine-Tuning (UPFT), which leverages the observation of Prefix Self-Consistency -- the shared initial reasoning steps across diverse solution trajectories -- to enhance LLM reasoning efficiency. By training exclusively on the initial prefix substrings (as few as 8 tokens), UPFT removes the need for labeled data or exhaustive sampling. Experiments on reasoning benchmarks show that UPFT matches the performance of supervised methods such as Rejection Sampling Fine-Tuning, while reducing training time by 75\% and sampling cost by 99\%. Further analysis reveals that errors tend to appear in later stages of the reasoning process and that prefix-based training preserves the model’s structural knowledge. This work demonstrates how minimal unsupervised fine-tuning can unlock substantial reasoning gains in LLMs, offering a scalable and resource-efficient alternative to conventional approaches.
Ke Ji, Qiuzhi Liu, Zhiwei He 0002, Benyou Wang, Zhaopeng Tu, Haitao Mi, Dong Yu 0001
NeurIPS11
2025 MPS-Prover: Advancing Stepwise Theorem Proving by Multi-Perspective Search and Data Curation
abstract
Automated Theorem Proving (ATP) in formal languages remains a formidable challenge in AI, demanding rigorous logical deduction and navigating vast search spaces. While large language models (LLMs) have shown promising performance, existing stepwise provers often suffer from biased search guidance, leading to inefficiencies and suboptimal proof strategies. This paper introduces the Multi-Perspective Search Prover (MPS-Prover), a novel stepwise ATP system designed to overcome these limitations. MPS-Prover incorporates two key innovations: a highly effective post-training data curation strategy that prunes approximately 40\% of redundant training data without sacrificing performance, and a multi-perspective tree search mechanism. This search integrates a learned critic model with strategically designed heuristic rules to diversify tactic selection, prevent getting trapped in unproductive states, and enhance search robustness. Extensive evaluations demonstrate that MPS-Prover achieves state-of-the-art performance on multiple challenging benchmarks, including miniF2F and ProofNet, outperforming prior 7B parameter models. Furthermore, our analyses reveal that MPS-Prover generates significantly shorter and more diverse proofs compared to existing stepwise and whole-proof methods, highlighting its efficiency and efficacy. Our work advances the capabilities of LLM-based formal reasoning and offers a robust framework and a comprehensive analysis for developing more powerful theorem provers.
Zhenwen Liang, Linfeng Song, Tao Yang 0033, Haitao Mi, Dong Yu 0001
NeurIPS5
2025 Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
abstract
Large Language Models (LLMs) show great promise in complex reasoning, with Reinforcement Learning with Verifiable Rewards (RLVR) being a key enhancement strategy. However, a prevalent issue is ``superficial self-reflection'', where models fail to robustly verify their own outputs. We introduce RISE (Reinforcing Reasoning with Self-Verification), a novel online RL framework designed to tackle this. RISE explicitly and simultaneously trains an LLM to improve both its problem-solving and self-verification abilities within a single, integrated RL process. The core mechanism involves leveraging verifiable rewards from an outcome verifier to provide on-the-fly feedback for both solution generation and self-verification tasks. In each iteration, the model generates solutions, then critiques its own on-policy generated solutions, with both trajectories contributing to the policy update. Extensive experiments on diverse mathematical reasoning benchmarks show that RISE consistently improves model's problem-solving accuracy while concurrently fostering strong self-verification skills. Our analyses highlight the advantages of online verification and the benefits of increased verification compute. Additionally, RISE models exhibit more frequent and accurate self-verification behaviors during reasoning. These advantages reinforce RISE as a flexible and effective path towards developing more robust and self-aware reasoners.
Zhiwei He 0002, Wenxuan Wang 0001, Pinjia He, Zhaopeng Tu, Haitao Mi, Dong Yu 0001
NeurIPS8
2025 Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
abstract
Mixture-of-Experts (MoE) architectures within Large Reasoning Models (LRMs) have achieved impressive reasoning capabilities by selectively activating experts to facilitate structured cognitive processes. Despite notable advances, existing reasoning models often suffer from cognitive inefficiencies like overthinking and underthinking. To address these limitations, we introduce a novel inference-time steering methodology called Reinforcing Cognitive Experts (RICE), designed to improve reasoning depth and efficiency without additional training or complex heuristics. Leveraging normalized Pointwise Mutual Information (nPMI), we systematically identify specialized experts, termed cognitive experts that orchestrate meta-level reasoning operations characterized by tokens like <think>. Empirical evaluations with leading MoE-based LRMs (DeepSeek-R1 and Qwen3-235B) on rigorous quantitative and scientific reasoning benchmarks (AIME and GPQA Diamond) demonstrate noticeable and consistent improvements in reasoning accuracy, cognitive efficiency, and cross-domain generalization. Crucially, our lightweight approach substantially outperforms prevalent reasoning-steering techniques, such as prompt design and decoding constraints, while preserving the model's general instruction-following skills. These results highlight reinforcing cognitive experts as a promising, practical, and interpretable direction to enhance cognitive efficiency within advanced reasoning models.
Yue Wang 0039, Zhiwei He 0002, Qiuzhi Liu, Yunzhi Yao, Wenxuan Wang 0001, Ruotian Ma, Haitao Mi, Ningyu Zhang 0001, Zhaopeng Tu, Dong Yu 0001
NeurIPS11
2025 Thoughts Are All Over the Place: On the Underthinking of Long Reasoning Models
abstract
Long reasoning models (LRMs) such as OpenAI's o1 and DeepSeek's R1 have demonstrated remarkable abilities in complex reasoning tasks by scaling test-time compute and exhibiting human-like deep thinking. However, we identify a phenomenon we term underthinking, where LRMs frequently switch between different reasoning thoughts without sufficiently exploring promising paths to reach a correct solution. This behavior leads to inadequate depth of reasoning and decreased performance, particularly on challenging mathematical problems. To systematically analyze this issue, we conduct experiments on three challenging test sets and two representative open-source LRMs, revealing that frequent thought switching correlates with incorrect responses. We introduce a novel metric to quantify underthinking by measuring token efficiency in incorrect answers. To address underthinking, we propose a decoding strategy with thought switching penalty (Tip) that discourages premature transitions between thoughts, encouraging deeper exploration of each reasoning path. Experimental results demonstrate that our approach improves accuracy across challenging datasets without requiring model fine-tuning. Our findings contribute to understanding reasoning inefficiencies in LRMs and offer a practical solution to enhance their problem-solving capabilities. Our code is open-source and available at https://github.com/wangyuenlp/underthinking.
Yue Wang 0039, Qiuzhi Liu, Zhiwei He 0002, Linfeng Song, Dian Yu 0001, Juntao Li 0005, Zhuosheng Zhang 0001, Rui Wang 0015, Zhaopeng Tu, Haitao Mi, Dong Yu 0001
NeurIPS13
2025 Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
abstract
Reinforcement learning from human feedback (RLHF) has demonstrated remarkable effectiveness in aligning large language models (LLMs) with human preferences. Many existing alignment approaches rely on the Bradley-Terry (BT) model assumption, which assumes the existence of a ground-truth reward for each prompt-response pair. However, this assumption can be overly restrictive when modeling complex human preferences. In this paper, we drop the BT model assumption and study LLM alignment under general preferences, formulated as a two-player game. Drawing on theoretical insights from learning in games, we integrate optimistic online mirror descent into our alignment framework to approximate the Nash policy. Theoretically, we demonstrate that our approach achieves an $\mathcal{O}(T^{-1})$ bound on the duality gap, improving upon the previous $\mathcal{O}(T^{-1/2})$ result. Meanwhile, it enjoys a linear convergence rate in the last iterate, a property not achieved by previous methods. More importantly, we implement our method and show through experiments that it outperforms state-of-the-art RLHF algorithms across multiple representative benchmarks.
Dian Yu 0001, Tao Ge 0001, Linfeng Song, Zhichen Zeng 0001, Haitao Mi, Nan Jiang 0008, Dong Yu 0001
NeurIPS6
2024 Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
abstract
Xiaoying Zhang, Baolin Peng, Ye Tian, Jingyan Zhou, Lifeng Jin, Linfeng Song, Haitao Mi, Helen Meng. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Baolin Peng, Jingyan Zhou, Lifeng Jin, Linfeng Song, Haitao Mi, Helen M. Meng
ACL (1)7
2024 A Knowledge Plug-and-Play Test Bed for Open-domain Dialogue Generation
abstract
Knowledge-based, open-domain dialogue generation aims to build chit-chat systems that talk to humans using mined support knowledge. Many types and sources of knowledge have previously been shown to be useful as support knowledge. Even in the era of large language models, response generation grounded in knowledge retrieved from additional up-to-date sources remains a practically important approach. While prior work using single-source knowledge has shown a clear positive correlation between the performances of knowledge selection and response generation, there are no existing multi-source datasets for evaluating support knowledge retrieval. Further, prior work has assumed that the knowledge sources available at test time are the same as during training. This unrealistic assumption unnecessarily handicaps models, as new knowledge sources can become available after a model is trained. In this paper, we present a high-quality benchmark named multi-source Wizard of Wikipedia (Ms.WoW) for evaluating multi-source dialogue knowledge selection and response generation. Unlike existing datasets, it contains clean support knowledge, grounded at the utterance level and partitioned into multiple knowledge sources. We further propose a new challenge, dialogue knowledge plug-and-play, which aims to test an already trained dialogue model on using new support knowledge from previously unseen sources in a zero-shot fashion.
Xiangci Li, Linfeng Song, Lifeng Jin, Haitao Mi, Jessica Ouyang 0001, Dong Yu 0001
LREC/COLING4
2024 The Trickle-down Impact of Reward Inconsistency on RLHF
abstract
Standard practice within Reinforcement Learning from Human Feedback (RLHF) involves optimizing against a Reward Model (RM), which itself is trained to reflect human preferences for desirable generations. A notable subject that is understudied is the (in-)consistency of RMs --- whether they can recognize the semantic changes to different prompts and appropriately adapt their reward assignments --- and their impact on the downstream RLHF model. In this paper, we visit a series of research questions relevant to RM inconsistency: (1) How can we measure the consistency of reward models? (2) How consistent are the existing RMs and how can we improve them? (3) In what ways does reward inconsistency influence the chatbots resulting from the RLHF model training? We propose **Contrast Instruction** -- a benchmarking strategy for the consistency of RM. Each example in **Contrast Instruction** features a pair of lexically similar instructions with different ground truth responses. A consistent RM is expected to rank the corresponding instruction and response higher than other combinations. We observe that current RMs trained with the standard ranking objective fail miserably on \contrast{} compared to average humans. To show that RM consistency can be improved efficiently without using extra training budget, we propose two techniques **ConvexDA** and **RewardFusion**, which enhance reward consistency through extrapolation during the RM training and inference stage, respectively. We show that RLHF models trained with a more consistent RM yield more useful responses, suggesting that reward inconsistency exhibits a trickle-down effect on the downstream RLHF process.
Lingfeng Shen, Linfeng Song, Lifeng Jin, Baolin Peng, Haitao Mi, Daniel Khashabi, Dong Yu 0001
ICLR6
2024 Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
abstract
Despite the impressive capabilities of Large Language Models (LLMs) on various tasks, they still struggle with scenarios that involves complex reasoning and planning. Self-correction and self-learning emerge as viable solutions, employing strategies that allow LLMs to refine their outputs and learn from self-assessed rewards. Yet, the efficacy of LLMs in self-refining its response, particularly in complex reasoning and planning task, remains dubious. In this paper, we introduce AlphaLLM for the self-improvements of LLMs, which integrates Monte Carlo Tree Search (MCTS) with LLMs to establish a self-improving loop, thereby enhancing the capabilities of LLMs without additional annotations. Drawing inspiration from the success of AlphaGo, AlphaLLM addresses the unique challenges of combining MCTS with LLM for self-improvement, including data scarcity, the vastness search spaces of language tasks, and the subjective nature of feedback in language tasks. AlphaLLM is comprised of prompt synthesis component, an efficient MCTS approach tailored for language tasks, and a trio of critic models for precise feedback. Our experimental results in mathematical reasoning tasks demonstrate that AlphaLLM significantly enhances the performance of LLMs without additional annotations, showing the potential for self-improvement in LLMs.
Baolin Peng, Linfeng Song, Lifeng Jin, Dian Yu 0001, Haitao Mi, Dong Yu 0001
NeurIPS7
2023 SafeConv: Explaining and Correcting Conversational Unsafe Behavior
abstract
One of the main challenges open-domain endto-end dialogue systems, or chatbots, face is the prevalence of unsafe behavior, such as toxic languages and harmful suggestions.However, existing dialogue datasets do not provide enough annotation to explain and correct such unsafe behavior.In this work, we construct a new dataset called SAFECONV for the research of conversational safety: (1) Besides the utterancelevel safety labels, SAFECONV also provides unsafe spans in an utterance, information able to indicate which words contribute to the detected unsafe behavior; (2) SAFECONV provides safe alternative responses to continue the conversation when unsafe behavior detected, guiding the conversation to a gentle trajectory.By virtue of the comprehensive annotation of SAFECONV, we benchmark three powerful models for the mitigation of conversational unsafe behavior, including a checker to detect unsafe utterances, a tagger to extract unsafe spans, and a rewriter to convert an unsafe response to a safe version.Moreover, we explore the huge benefits brought by combining the models for explaining the emergence of unsafe behavior and detoxifying chatbots.Experiments show that the detected unsafe behavior could be well explained with unsafe spans and popular chatbots could be detoxified by a huge extent.The dataset is available at https://github.com/mianzhang/SafeConv.
Lifeng Jin, Linfeng Song, Haitao Mi, Wenliang Chen, Dong Yu 0001
ACL (1)4
2023 Friend-training: Learning from Models of Different but Related Tasks
abstract
Current self-training methods such as standard self-training, co-training, tri-training, and others often focus on improving model performance on a single task, utilizing differences in input features, model architectures, and training processes.However, many tasks in natural language processing are about different but related aspects of language, and models trained for one task can be great teachers for other related tasks.In this work, we propose friendtraining, a cross-task self-training framework, where models trained to do different tasks are used in an iterative training, pseudo-labeling, and retraining process to help each other for better selection of pseudo-labels.With two dialogue understanding tasks, conversational semantic role labeling and dialogue rewriting, chosen for a case study, we show that the models trained with the friend-training framework achieve the best performance compared to strong baselines.
Lifeng Jin, Linfeng Song, Haitao Mi, Xiabing Zhou, Dong Yu 0001
EACL4
2023 More Than Spoken Words: Nonverbal Message Extraction and Generation
abstract
Nonverbal messages (NM) such as speakers' facial expressions and speed of speech are essential for face-to-face communication, and they can be regarded as implicit knowledge as they are usually not included in existing dialogue understanding or generation tasks.This paper introduces the task of extracting NMs in written text and generating NMs for spoken text.Previous studies merely focus on extracting NMs from relatively small-scale well-structured corpora such as movie scripts wherein NMs are enclosed in parentheses by scriptwriters, which greatly decreases the difficulty of extraction.To enable extracting NMs from unstructured corpora, we annotate the first NM extraction dataset for Chinese based on novels and develop three baselines to extract single-span or multi-span NM of a target utterance from its surrounding context.Furthermore, we use the extractors to extract 749K (context, utterance, NM) triples from Chinese novels and investigate whether we can use them to improve NM generation via semi-supervised learning.Experimental results demonstrate that the automatically extracted triples can serve as high-quality augmentation data of clean triples extracted from scripts to generate more relevant, fluent, valid, and factually consistent 1 NMs than the purely supervised generator, and the resulting generator can in turn help Chinese dialogue understanding tasks such as dialogue machine reading comprehension and emotion classification by simply adding the predicted "unspoken" NM to each utterance or narrative in inputs.
Dian Yu 0001, Xiaoyang Wang 0001, Wanshun Chen, Longyue Wang, Haitao Mi, Dong Yu 0001
EMNLP6
2023 Search-engine-augmented dialogue response generation with cheaply supervised query production
Ante Wang, Linfeng Song, Qi Liu 0049, Haitao Mi, Longyue Wang, Zhaopeng Tu, Jinsong Su, Dong Yu 0001
Artif. Intell.4
2023 Discover, Explain, Improve: An Automatic Slice Detection Benchmark for Natural Language Processing
abstract
Abstract Pretrained natural language processing (NLP) models have achieved high overall performance, but they still make systematic errors. Instead of manual error analysis, research on slice detection models (SDMs), which automatically identify underperforming groups of datapoints, has caught escalated attention in Computer Vision for both understanding model behaviors and providing insights for future model training and designing. However, little research on SDMs and quantitative evaluation of their effectiveness have been conducted on NLP tasks. Our paper fills the gap by proposing a benchmark named “Discover, Explain, Improve (DEIm)” for classification NLP tasks along with a new SDM Edisa. Edisa discovers coherent and underperforming groups of datapoints; DEIm then unites them under human-understandable concepts and provides comprehensive evaluation tasks and corresponding quantitative metrics. The evaluation in DEIm shows that Edisa can accurately select error-prone datapoints with informative semantic features that summarize error patterns. Detecting difficult datapoints directly boosts model performance without tuning any original model parameters, showing that discovered slices are actionable for users.1
Wenyue Hua, Lifeng Jin, Linfeng Song, Haitao Mi, Dong Yu 0001
Trans. Assoc. Comput. Linguistics4
2023 OpenFact: Factuality Enhanced Open Knowledge Extraction
abstract
Abstract We focus on the factuality property during the extraction of an OpenIE corpus named OpenFact, which contains more than 12 million high-quality knowledge triplets. We break down the factuality property into two important aspects—expressiveness and groundedness—and we propose a comprehensive framework to handle both aspects. To enhance expressiveness, we formulate each knowledge piece in OpenFact based on a semantic frame. We also design templates, extra constraints, and adopt human efforts so that most OpenFact triplets contain enough details. For groundedness, we require the main arguments of each triplet to contain linked Wikidata1 entities. A human evaluation suggests that the OpenFact triplets are much more accurate and contain denser information compared to OPIEC-Linked (Gashteovski et al., 2019), one recent high-quality OpenIE corpus grounded to Wikidata. Further experiments on knowledge base completion and knowledge base question answering show the effectiveness of OpenFact over OPIEC-Linked as supplementary knowledge to Wikidata as the major KG.
Linfeng Song, Ante Wang, Xiaoman Pan, Hongming Zhang 0009, Dian Yu 0001, Lifeng Jin, Haitao Mi, Jinsong Su, Yue Zhang 0004, Dong Yu 0001
Trans. Assoc. Comput. Linguistics7
2023 D$^{2}$PSG: Multi-Party Dialogue Discourse Parsing as Sequence Generation
abstract
Conversational discourse analysis aims to extract the interactions between dialogue turns, which is crucial for modeling complex multi-party dialogues. As the benchmarks are still limited in size and human annotations are costly, the current standard approaches apply pretrained language models, but they still require randomly initialized classifiers to make predictions. These classifiers usually require massive data to work smoothly with the pretrained encoder, causing severe data hunger issue. We propose two convenient strategies to formulate this task as a sequence generation problem, where classifier decisions are carefully converted into sequence of tokens. We then adopt a pretrained T5 1 model to solve this task so that no parameters are randomly initialized. We also leverage the descriptions of the discourse relations to help model understand their meanings. Experiments on two popular benchmarks show that our approach outperforms previous state-of-the-art models by a large margin, and it is also more robust in zero-shot and few-shot settings.
Ante Wang, Linfeng Song, Lifeng Jin, Junfeng Yao, Haitao Mi, Chen Lin 0001, Jinsong Su, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2022 Fast-R2D2: A Pretrained Recursive Neural Network based on Pruned CKY for Grammar Induction and Text Representation
abstract
Chart-based models have shown great potential in unsupervised grammar induction, running recursively and hierarchically, but requiring O(n 3 ) time-complexity.The Recursive Transformer based on Differentiable Trees (R2D2) makes it possible to scale to large language model pretraining even with a complex tree encoder, by introducing a heuristic pruning method.However, its rule-based pruning process suffers from local optima and slow inference.In this paper, we propose a unified R2D2 method that overcomes these issues.We use a top-down unsupervised parser as a modelguided pruning method, which also enables parallel encoding during inference.Our parser casts parsing as a split point scoring task by first scoring all split points for a given sentence and then using the highest-scoring one to recursively split a span into two parts.The reverse order of the splits is considered as the order of pruning in the encoder.We optimize the unsupervised parser by minimizing the Kullback-Leibler distance between tree probabilities from the parser and the R2D2 model.Our experiments show that our Fast-R2D2 significantly improves the grammar induction quality and achieves competitive results in downstream tasks. 1
Haitao Mi, Gerard de Melo
EMNLP2
2022 Learning a Grammar Inducer from Massive Uncurated Instructional Videos
abstract
Video-aided grammar induction aims to leverage video information for finding more accurate syntactic grammars for accompanying text.While previous work focuses on building systems for inducing grammars on text that are well-aligned with video content, we investigate the scenario, in which text and video are only in loose correspondence.Such data can be found in abundance online, and the weak correspondence is similar to the indeterminacy problem studied in language acquisition.Furthermore, we build a new model that can better learn video-span correlation without manually designed features adopted by previous work.Experiments show that our model trained only on large-scale YouTube data with no textvideo alignment reports strong and robust performances across three unseen datasets, despite domain shift and noisy label issues.Furthermore our model yields higher F1 scores than the previous state-of-the-art systems trained on in-domain data.
Songyang Zhang 0004, Linfeng Song, Lifeng Jin, Haitao Mi, Kun Xu 0005, Dong Yu 0001, Jiebo Luo 0001
EMNLP4
2022 COSSUM: Towards Conversation-Oriented Structured Summarization for Automatic Medical Insurance Assessment
abstract
In medical insurance industry, a lot of human labor is required to collect information of claimants. Human assessors need to converse with claimants in order to record key information and organize it into a structured summary. With the purpose of helping save human labor, we propose the task of conversation-oriented structured summarization which aims to automatically produce the desired structured summary from a conversation automatically. One major challenge of the task is that the structured summary contains multiple fields of different types. To tackle this problem, we propose a unified approach COSSUM based on prompting to generate the values of all fields simultaneously. By learning all fields together, our approach can capture the inherent relationship between them. Moreover, we propose a specially designed curriculum learning strategy for model training. Both automatic and human evaluations are performed, and the results show the effectiveness of our proposed approach.
Xiaojun Wan 0001, Sen Hu 0005, Mengdi Zhou, Teng Xu 0007, Haitao Mi
KDD7
2021 R2D2: Recursive Transformer based on Differentiable Tree for Interpretable Hierarchical Language Modeling
abstract
Xiang Hu, Haitao Mi, Zujie Wen, Yafang Wang, Yi Su, Jing Zheng, Gerard de Melo. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Haitao Mi, Zujie Wen, Yafang Wang, Gerard de Melo
ACL/IJCNLP (1)2
2021 IIAS: An Intelligent Insurance Assessment System through Online Real-time Conversation Analysis
abstract
With the development of Chinese medical insurance industry, the amount of claim cases is growing rapidly. Ultimately, more claims necessarily indicate that the insurance company has to spend much time assessing claims and decides how much compensation the claimant should receive, which is a highly professional process that involves many complex operations. Therefore, the insurance assessor's role is essential. However, for the junior assessor often lacking in practical experience, it is not easy to quickly handle such an online procedure. In order to alleviate assessors' cognitive workload, we propose an Intelligent Insurance Assessment System (IIAS) that helps effectively collect claimant information through online real-time conversation analysis. With the assistance of IIAS, the average time cost of the insurance assessment procedure is reduced from 55 minutes to 35 minutes.
Mengdi Zhou, Shuang Peng 0009, Haitao Mi, Zujie Wen, Teng Xu 0007
IJCAI7
2016 Sentence Similarity Learning by Lexical Decomposition and Composition
abstract
Most conventional sentence similarity methods only focus on similar parts of two input sentences, and simply ignore the dissimilar parts, which usually give us some clues and semantic meanings about the sentences. In this work, we propose a model to take into account both the similarities and dissimilarities by decomposing and composing lexical semantics over sentences. The model represents each word as a vector, and calculates a semantic matching vector for each word based on all words in the other sentence. Then, each word vector is decomposed into a similar component and a dissimilar component based on the semantic matching vector. After this, a two-channel CNN model is employed to capture features by composing the similar and dissimilar components. Finally, a similarity score is estimated over the composed feature vectors. Experimental results show that our model gets the state-of-the-art performance on the answer sentence selection task, and achieves a comparable result on the paraphrase identification task.
Haitao Mi, Abraham Ittycheriah
COLING2
2016 Semi-supervised Clustering for Short Text via Deep Representation Learning
abstract
In this work, we propose a semisupervised method for short text clustering, where we represent texts as distributed vectors with neural networks, and use a small amount of labeled data to specify our intention for clustering.We design a novel objective to combine the representation learning process and the kmeans clustering process together, and optimize the objective with both labeled data and unlabeled data iteratively until convergence through three steps: (1) assign each short text to its nearest centroid based on its representation from the current neural networks; (2) re-estimate the cluster centroids based on cluster assignments from step (1); (3) update neural networks according to the objective by keeping centroids and cluster assignments fixed.Experimental results on four datasets show that our method works significantly better than several other text clustering methods.
Haitao Mi, Abraham Ittycheriah
CoNLL2
2016 Coverage Embedding Models for Neural Machine Translation
abstract
In this paper, we enhance the attention-based neural machine translation (NMT) by adding explicit coverage embedding models to alleviate issues of repeating and dropping translations in NMT.For each source word, our model starts with a full coverage embedding vector to track the coverage status, and then keeps updating it with neural networks as the translation goes.Experiments on the large-scale Chinese-to-English task show that our enhanced model improves the translation quality significantly on various test sets over the strong large vocabulary NMT system.
Haitao Mi, Baskaran Sankaran, Abraham Ittycheriah
EMNLP1
2016 Supervised Attentions for Neural Machine Translation
abstract
In this paper, we improve the attention or alignment accuracy of neural machine translation by utilizing the alignments of training sentence pairs.We simply compute the distance between the machine attentions and the "true" alignments, and minimize this cost in the training procedure.Our experiments on large-scale Chinese-to-English task show that our model improves both translation and alignment qualities significantly over the large-vocabulary neural machine translation system, and even beats a state-of-the-art traditional syntax-based system.
Haitao Mi, Abraham Ittycheriah
EMNLP1
2015 Feature Optimization for Constituent Parsing via Neural Networks
abstract
Zhiguo Wang, Haitao Mi, Nianwen Xue. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Haitao Mi, Nianwen Xue
ACL (1)2
2015 Shift-Reduce Constituency Parsing with Dynamic Programming and POS Tag Lattice
abstract
We present the first dynamic programming (DP) algorithm for shift-reduce constituency parsing, which extends the DP idea of Huang and Sagae (2010) to context-free grammars. To alleviate the propagation of errors from part-of-speech tagging, we also extend the parser to take a tag lattice instead of a fixed tag sequence. Experiments on both English and Chinese treebanks show that our DP parser significantly improves parsing quality over non-DP baselines, and achieves the best accuracies among empirical linear-time parsers.
Haitao Mi, Liang Huang 0001
HLT-NAACL1
2014 A Structured Language Model for Incremental Tree-to-String Translation
Heng Yu 0006, Haitao Mi, Liang Huang 0001, Qun Liu 0001
COLING2
2013 Flexible and Efficient Hypergraph Interactions for Joint Hierarchical and Forest-to-String Decoding
abstract
Machine translation benefits from system combination.We propose flexible interaction of hypergraphs as a novel technique combining different translation models within one decoder.We introduce features controlling the interactions between the two systems and explore three interaction schemes of hiero and forest-to-string models-specification, generalization, and interchange.The experiments are carried out on large training data with strong baselines utilizing rich sets of dense and sparse features.All three schemes significantly improve results of any single system on four testsets.We find that specification-a more constrained scheme that almost entirely uses forest-to-string rules, but optionally uses hiero rules for shorter spans-comes out as the strongest, yielding improvement up to 0.9 (Ter-Bleu)/2 points.We also provide a detailed experimental and qualitative analysis of the results.
Martin Cmejrek, Haitao Mi
EMNLP2
2013 Max-Violation Perceptron and Forced Decoding for Scalable MT Training
abstract
While large-scale discriminative training has triumphed in many NLP problems, its definite success on machine translation has been largely elusive.Most recent efforts along this line are not scalable (training on the small dev set with features from top ∼100 most frequent words) and overly complicated.We instead present a very simple yet theoretically motivated approach by extending the recent framework of "violation-fixing perceptron", using forced decoding to compute the target derivations.Extensive phrase-based translation experiments on both Chinese-to-English and Spanish-to-English tasks show substantial gains in BLEU by up to +2.3/+2.0 on dev/test over MERT, thanks to 20M+ sparse features.This is the first successful effort of large-scale online discriminative training for MT.
Heng Yu 0006, Liang Huang 0001, Haitao Mi, Kai Zhao 0003
EMNLP3
2011 Rule Markov Models for Fast Tree-to-String Translation
Ashish Vaswani, Haitao Mi, Liang Huang 0001, David Chiang 0001
ACL2
2011 A novel dependency-to-string model for statistical machine translation
Haitao Mi, Qun Liu 0001
EMNLP2
2011 Bagging-based System Combination for Domain Adaption
Linfeng Song, Haitao Mi, Yajuan Lü, Qun Liu 0001
MTSummit2
2010 Forest-Based Semantic Role Labeling
abstract
Parsing plays an important role in semantic role labeling (SRL) because most SRL systems infer semantic relations from 1-best parses. Therefore, parsing errors inevitably lead to labeling mistakes. To alleviate this problem, we propose to use packed forest, which compactly encodes all parses for a sentence. We design an algorithm to exploit exponentially many parses to learn semantic relations efciently. Experimental results on the CoNLL-2005 shared task show that using forests achieves an absolute improvement of 1.2% in terms of F1 score over using 1-best parses and 0.6% over using 50-best parses.
Haitao Mi, Yang Liu 0005, Qun Liu 0001
AAAI2
2010 Constituency to Dependency Translation with Forests
Haitao Mi, Qun Liu 0001
ACL1
2010 Efficient Incremental Decoding for Tree-to-String Translation
Liang Huang 0001, Haitao Mi
EMNLP2
2010 Statistical Translation Model Based On Source Syntax Structure
Qun Liu 0001, Yang Liu 0005, Haitao Mi
PACLIC3
2009 Joint Decoding with Multiple Translation Models
Yang Liu 0005, Haitao Mi, Yang Feng 0004, Qun Liu 0001
ACL/IJCNLP2
2009 Lattice-based System Combination for Statistical Machine Translation
Yang Feng 0004, Yang Liu 0005, Haitao Mi, Qun Liu 0001, Yajuan Lü
EMNLP3
2008 Forest-Based Translation
Haitao Mi, Liang Huang 0001, Qun Liu 0001
ACL1
2008 Word Lattice Reranking for Chinese Word Segmentation and Part-of-Speech Tagging
Wenbin Jiang 0002, Haitao Mi, Qun Liu 0001
COLING2
2008 Forest-based Translation Rule Extraction
Haitao Mi, Liang Huang 0001
EMNLP1
2008 Refinements in BTG-based Statistical Machine Translation
Deyi Xiong, Min Zhang 0005, AiTi Aw, Haitao Mi, Qun Liu 0001, Shouxun Lin
IJCNLP4