EDBT 2026 Demo / reviewers in the wild / expert
Qiuzhi Liu
dblp:196/4163
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 58% Vision and language · 7% Trustworthy machine learning · 7% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.9 | 1 | 2025 | Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models · ICML 2025 |
Machine learning › Representation and self-supervised learning
contrastive estimation |
0.9 | 1 | 2025 | Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability · ICML 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.9 | 1 | 2025 | Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models · ICML 2025 |
Natural language and speech › Language models and text generation › large language model inference
inference-time steering |
0.9 | 1 | 2025 | Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model
large reasoning model |
0.9 | 1 | 2025 | Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training · NeurIPS 2025 |
Natural language and speech › Language models and text generation
mathematical reasoning |
0.9 | 1 | 2025 | Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability · ICML 2025 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.9 | 1 | 2025 | Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › multimodal trustworthiness
multimodal knowledge conflict |
0.9 | 1 | 2025 | Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs · ACL (1) 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model reasoning
overthinking |
0.9 | 1 | 2025 | Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models · ICML 2025 |
Natural language and speech › Language models and text generation › large language model
reasoning model |
0.9 | 1 | 2025 | Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models · ICML 2025 |
Natural language and speech › Language models and text generation › large language model inference
test-time compute |
0.9 | 1 | 2025 | Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models · ICML 2025 |
Natural language and speech › Machine translation › neural machine translation
nearest neighbor machine translation |
0.7 | 1 | 2023 | Simple and Scalable Nearest Neighbor Machine Translation · ICLR 2023 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.7 | 1 | 2023 | Simple and Scalable Nearest Neighbor Machine Translation · ICLR 2023 |
Methods — techniques the papers use, named apart from their topics
self-training · 0.9rollout sampling · 0.9pointwise mutual information · 0.9multimodal large language model · 0.9efficiency metrics · 0.9direct preference optimization · 0.9contrastive estimation · 0.9cognitive expert steering · 0.9k-nearest neighbor retrieval · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMsabstractXiaoyuan Liu, Wenxuan Wang, Youliang Yuan, Jen-tse Huang, Qiuzhi Liu, Pinjia He, Zhaopeng Tu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Wenxuan Wang 0001, Youliang Yuan, Jen-tse Huang 0001, Qiuzhi Liu, Pinjia He, Zhaopeng Tu |
ACL (1) | 5 |
| 2025 | Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning ModelsabstractThe remarkable performance of long reasoning models can be attributed to their ability to emulate human-like long-time thinking during inference. These models employ extended chain-of-thought (CoT) processes, exploring multiple strategies to enhance problem-solving capabilities. However, a critical question remains: How to intelligently and efficiently scale computational resources during testing. This paper presents the first comprehensive study on the prevalent issue of overthinking in these models, where long reasoning models generate redundant solutions that contribute minimally to accuracy and diversity, thereby wasting computational resources on simple problems with minimal benefit. We introduce novel efficiency metrics from both outcome and process perspectives to evaluate the rational use of computational resources by long reasoning models. Using a self-training paradigm, we propose strategies to mitigate overthinking, simplifying reasoning processes without compromising accuracy. Experimental results show that our approach successfully reduces computational overhead while preserving model performance across a range of testsets with varying difficulty levels, such as GSM8K, MATH500, GPQA, and AIME. Our code is open-source and available at https://github.com/galaxyChen/overthinking. Zhiwei He 0002, Jianhui Pang, Dian Yu 0001, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang 0001, Rui Wang 0015, Zhaopeng Tu, Haitao Mi, Dong Yu 0001 |
ICML | 8 |
| 2025 | Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning CapabilityabstractMathematical reasoning tasks pose significant challenges for large language models (LLMs) because they require precise logical deduction and sequence analysis. In this work, we introduce the concept of critical tokens – elements within reasoning trajectories that significantly influence incorrect outcomes. We present a novel framework for identifying these tokens through rollout sampling and demonstrate their substantial divergence from traditional error tokens. Through extensive experiments on datasets such as GSM8K and MATH500, we show that identifying and replacing critical tokens significantly improves model accuracy. We propose an efficient methodology for pinpointing these tokens in large-scale datasets using contrastive estimation and extend this framework to enhance model training processes with direct preference optimization (DPO). Experimental results on GSM8K and MATH500 benchmarks with the widely used models Llama-3 (8B and 70B) and Deepseek-math (7B) demonstrate the effectiveness of the proposed approach, cDPO. Our results underscore the potential of leveraging critical tokens to reduce errors in reasoning tasks, advancing the development of AI systems capable of robust logical deduction. Zicheng Lin, Qiuzhi Liu, Xing Wang 0007, Ruilin Luo, Chufan Shi, Siheng Li, Yujiu Yang 0001, Zhaopeng Tu |
ICML | 4 |
| 2025 | The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning ModelsabstractImproving the reasoning capabilities of large language models (LLMs) typically requires supervised fine-tuning with labeled data or computationally expensive sampling. We introduce Unsupervised Prefix Fine-Tuning (UPFT), which leverages the observation of Prefix Self-Consistency -- the shared initial reasoning steps across diverse solution trajectories -- to enhance LLM reasoning efficiency. By training exclusively on the initial prefix substrings (as few as 8 tokens), UPFT removes the need for labeled data or exhaustive sampling. Experiments on reasoning benchmarks show that UPFT matches the performance of supervised methods such as Rejection Sampling Fine-Tuning, while reducing training time by 75\% and sampling cost by 99\%. Further analysis reveals that errors tend to appear in later stages of the reasoning process and that prefix-based training preserves the model’s structural knowledge. This work demonstrates how minimal unsupervised fine-tuning can unlock substantial reasoning gains in LLMs, offering a scalable and resource-efficient alternative to conventional approaches. Ke Ji, Qiuzhi Liu, Zhiwei He 0002, Benyou Wang, Zhaopeng Tu, Haitao Mi, Dong Yu 0001 |
NeurIPS | 4 |
| 2025 | Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional TrainingabstractMixture-of-Experts (MoE) architectures within Large Reasoning Models (LRMs) have achieved impressive reasoning capabilities by selectively activating experts to facilitate structured cognitive processes. Despite notable advances, existing reasoning models often suffer from cognitive inefficiencies like overthinking and underthinking. To address these limitations, we introduce a novel inference-time steering methodology called Reinforcing Cognitive Experts (RICE), designed to improve reasoning depth and efficiency without additional training or complex heuristics. Leveraging normalized Pointwise Mutual Information (nPMI), we systematically identify specialized experts, termed cognitive experts that orchestrate meta-level reasoning operations characterized by tokens like <think>. Empirical evaluations with leading MoE-based LRMs (DeepSeek-R1 and Qwen3-235B) on rigorous quantitative and scientific reasoning benchmarks (AIME and GPQA Diamond) demonstrate noticeable and consistent improvements in reasoning accuracy, cognitive efficiency, and cross-domain generalization. Crucially, our lightweight approach substantially outperforms prevalent reasoning-steering techniques, such as prompt design and decoding constraints, while preserving the model's general instruction-following skills. These results highlight reinforcing cognitive experts as a promising, practical, and interpretable direction to enhance cognitive efficiency within advanced reasoning models. Yue Wang 0039, Zhiwei He 0002, Qiuzhi Liu, Yunzhi Yao, Wenxuan Wang 0001, Ruotian Ma, Haitao Mi, Ningyu Zhang 0001, Zhaopeng Tu, Dong Yu 0001 |
NeurIPS | 7 |
| 2025 | Thoughts Are All Over the Place: On the Underthinking of Long Reasoning ModelsabstractLong reasoning models (LRMs) such as OpenAI's o1 and DeepSeek's R1 have demonstrated remarkable abilities in complex reasoning tasks by scaling test-time compute and exhibiting human-like deep thinking. However, we identify a phenomenon we term underthinking, where LRMs frequently switch between different reasoning thoughts without sufficiently exploring promising paths to reach a correct solution. This behavior leads to inadequate depth of reasoning and decreased performance, particularly on challenging mathematical problems. To systematically analyze this issue, we conduct experiments on three challenging test sets and two representative open-source LRMs, revealing that frequent thought switching correlates with incorrect responses. We introduce a novel metric to quantify underthinking by measuring token efficiency in incorrect answers. To address underthinking, we propose a decoding strategy with thought switching penalty (Tip) that discourages premature transitions between thoughts, encouraging deeper exploration of each reasoning path. Experimental results demonstrate that our approach improves accuracy across challenging datasets without requiring model fine-tuning. Our findings contribute to understanding reasoning inefficiencies in LRMs and offer a practical solution to enhance their problem-solving capabilities. Our code is open-source and available at https://github.com/wangyuenlp/underthinking. Yue Wang 0039, Qiuzhi Liu, Zhiwei He 0002, Linfeng Song, Dian Yu 0001, Juntao Li 0005, Zhuosheng Zhang 0001, Rui Wang 0015, Zhaopeng Tu, Haitao Mi, Dong Yu 0001 |
NeurIPS | 2 |
| 2023 | Simple and Scalable Nearest Neighbor Machine Translation
Yuhan Dai, Zhirui Zhang, Qiuzhi Liu, Qu Cui, Yichao Du, Tong Xu 0001 |
ICLR | 3 |