VLDB 2026 Research / reviewers in the wild / expert
Yunhua Zhou
dblp:67/8389
· DBLP profile ↗
22ranked-venue papers
4as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 4 first-author · 21 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative AlignmentabstractYuming Yang, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuming Yang 0001, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao 0019, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 11 |
| 2025 | Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?abstractThe advent of test-time scaling in large language models (LLMs), exemplified by Ope-nAI's o1 series, has advanced reasoning capabilities by scaling computational resource allocation during inference.While successors like QwQ, Deepseek-R1 (R1) and LIMO replicate these advancements, whether these models truly possess test-time scaling capabilities remains underexplored.This study found that longer CoTs of these o1-like models do not consistently enhance accuracy; in fact, correct solutions are often shorter than incorrect ones for the same questions.Further investigation shows this phenomenon is closely related to models' self-revision capabilities -longer CoTs contain more self-revisions, which often lead to performance degradation.We then compare sequential and parallel scaling strategies on QwQ, R1 and LIMO, finding that parallel scaling achieves better coverage and scalability.Based on these insights, we propose Shortest Majority Vote, a method that combines parallel scaling strategies with CoT length characteristics, significantly improving models' test-time scalability compared to conventional majority voting approaches. Zhiyuan Zeng 0004, Qinyuan Cheng, Zhangyue Yin, Yunhua Zhou, Xipeng Qiu |
ACL (1) | 4 |
| 2025 | Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling LawabstractScaling law builds the relationship between training computation and validation loss, enabling researchers to effectively predict the loss trending of models across different levels of computation. However, a gap still remains between validation loss and the model’s downstream capabilities, making it untrivial to apply scaling law to direct performance prediction for downstream tasks. The loss typically represents a cumulative penalty for predicted tokens, which are implicitly considered to have equal importance. Nevertheless, our studies have shown evidence that when considering different training data distributions, we cannot directly model the relationship between downstream capability and computation or token loss. To bridge the gap between validation loss and downstream task capabilities, in this work, we introduce Capability Salience Vector, which decomposes the overall loss and assigns different importance weights to tokens to assess a specific meta-capability, aligning the validation loss with downstream task performance in terms of the model’s capabilities. Experiments on various popular benchmarks demonstrate that our proposed Capability Salience Vector could significantly improve the predictability of language model performance on downstream tasks. Qiming Ge, Shuhao Xing, Songyang Gao, Yunhua Zhou, Yicheng Zou, Songyang Zhang 0001, Zhi Chen 0006, Hang Yan 0001, Qi Zhang 0001, Qipeng Guo, Kai Chen 0026 |
ACL (1) | 4 |
| 2025 | Firewall Routing: Blocking Leads to Better Hybrid Inference for LLMsabstractThe rapid advancement of Large Language Models (LLMs) has significantly enhanced performance across various natural language processing (NLP) tasks, yet the high computational costs and latency associated with deploying such models continue to pose critical bottlenecks, limiting their broader applicability.To mitigate these challenges, we propose a dynamic hybrid inference framework, Firewall Routing, which efficiently selects between a strong and a weak LLMs based on the complexity of the query.A lightweight routing model is trained to optimize resource allocation by learning from response quality and preventing longtail queries, which are often too hard to solve by LLMs, from being routed to the stronger model.Moreover, our method incorporates multiple sampling to enhance query evaluation reliability while leveraging Hard Blocking and Soft Blocking to handle long-tail queries along with refining labels for model selection.Extensive experiments show our method outperforms existing routing strategies by up to 5.29% in APGR, demonstrating state-of-the-art performance across multiple benchmarks. Runyu Peng, Yunhua Zhou, Kai Lv 0001, Yang Gao 0042, Qipeng Guo, Xipeng Qiu |
EMNLP | 2 |
| 2025 | BitStack: Any-Size Compression of Large Language Models in Variable Memory EnvironmentsabstractLarge language models (LLMs) have revolutionized numerous applications, yet their deployment remains challenged by memory constraints on local devices. While scaling laws have enhanced LLM capabilities, the primary bottleneck has shifted from $\textit{capability}$ to $\textit{availability}$, emphasizing the need for efficient memory management. Traditional compression methods, such as quantization, often require predefined compression ratios and separate compression processes for each setting, complicating deployment in variable memory environments. In this paper, we introduce $\textbf{BitStack}$, a novel, training-free weight compression approach that enables megabyte-level trade-offs between memory usage and model performance. By leveraging weight decomposition, BitStack can dynamically adjust the model size with minimal transmission between running memory and storage devices. Our approach iteratively decomposes weight matrices while considering the significance of each parameter, resulting in an approximately 1-bit per parameter residual block in each decomposition iteration. These blocks are sorted and stacked in storage as basic transmission units, with different quantities loaded based on current memory availability. Extensive experiments across a wide range of tasks demonstrate that, despite offering fine-grained size control, BitStack consistently matches or surpasses strong quantization baselines, particularly at extreme compression ratios. To the best of our knowledge, this is the first decomposition-based method that effectively bridges the gap to practical compression techniques like quantization. Code is available at https://github.com/xinghaow99/BitStack. Pengyu Wang 0006, Bo Wang 0084, Yunhua Zhou, Xipeng Qiu |
ICLR | 5 |
| 2025 | Towards Universality: Studying Mechanistic Similarity Across Language Model ArchitecturesabstractThe hypothesis of \textit{Universality} in interpretability suggests that different neural networks may converge to
implement similar algorithms on similar tasks. In this work, we investigate two mainstream architectures
for language modeling, namely Transformers and Mambas, to explore the extent of their mechanistic similarity.
We propose to use Sparse Autoencoders (SAEs) to isolate interpretable features from these models and show
that most features are similar in these two models. We also validate the correlation between feature similarity
and~\univ. We then delve into the circuit-level analysis of Mamba models
and find that the induction circuits in Mamba are structurally analogous to those in Transformers. We also identify a nuanced difference we call \emph{Off-by-One motif}: The information of one token is written into the
SSM state in its next position. Whilst interaction between tokens in Transformers does not exhibit such trend. Junxuan Wang, Xuyang Ge, Wentao Shu, Qiong Tang, Yunhua Zhou, Zhengfu He, Xipeng Qiu |
ICLR | 5 |
| 2025 | Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling PerformanceabstractPretraining data of large language models composes multiple domains (e.g., web texts, academic papers, codes), whose mixture proportions crucially impact the competence of outcome models. While existing endeavors rely on heuristics or qualitative strategies to tune the proportions, we discover the quantitative predictability of model performance regarding the mixture proportions in function forms, which we refer to as the data mixing laws. Fitting such functions on sample mixtures unveils model performance on unseen mixtures before actual runs, thus guiding the selection of an ideal data mixture. Furthermore, we propose nested use of the scaling laws of training steps, model sizes, and our data mixing laws to predict the performance of large models trained on massive data under various mixtures with only small-scale training. Experimental results verify that our method effectively optimizes the training mixture of a 1B model trained for 100B tokens in RedPajama, reaching a performance comparable to the one trained for 48% more steps on the default mixture. Extending the application of data mixing laws to continual training accurately predicts the critical mixture proportion that avoids catastrophic forgetting and outlooks the potential for dynamic data schedules. Jiasheng Ye, Peiju Liu, Tianxiang Sun, Yunhua Zhou, Xipeng Qiu |
ICLR | 5 |
| 2025 | Pre-Trained Policy Discriminators are General Reward ModelsabstractWe offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guiding the training policy towards a target policy with desired behaviors. Based on this conceptual insight, we propose a scalable pre-training method named POLicy DiscriminAtive LeaRning (POLAR), which trains a reward model (RM) to discern identical policies and discriminate different ones. Unlike traditional reward modeling methods relying on absolute preferences, POLAR captures the relative difference between one policy and an arbitrary target policy, which is a scalable, high-level optimization objective suitable for modeling generic ranking relationships. Leveraging the POLAR pre-training paradigm, we present a series of RMs with parameter scales from 1.8B to 7B. Empirical results show that POLAR substantially outperforms traditional non-pre-trained methods, significantly enhancing RM performance.
For instance, POLAR-7B could improve preference accuracy from 54.8% to 81.0% on STEM tasks and from 57.9% to 85.5% on creative writing tasks compared to SOTA baselines.
POLAR also shows robust generalization capabilities in RLHF using Reinforcement Fine-tuning (RFT), providing reliable reward signals and markedly enhancing policy performance—improving LLaMa3.1-8B from an average of 47.36% to 56.33% and Qwen2.5-32B from 64.49% to 70.47% on 20 benchmarks.
Moreover, scaling experiments reveal a clear power-law relationship between computation and performance, supported by linear correlation coefficients approaching 0.99.
The impressive performance, strong generalization, and scaling properties suggest that POLAR is a promising direction for developing general and strong reward models. Shihan Dou, Shichun Liu, Yuming Yang 0001, Yicheng Zou, Yunhua Zhou, Shuhao Xing, Chenhao Huang, Qiming Ge, Haijun Lv, Demin Song, Songyang Gao, Chengqi Lyu, Enyu Zhou, Honglin Guo, Zhiheng Xi, Qipeng Guo, Tao Gui, Qi Zhang 0001, Xipeng Qiu, Xuanjing Huang 0001, Kai Chen 0026 |
NeurIPS | 5 |
| 2025 | Implicit Reward as the Bridge: A Unified View of SFT and DPO ConnectionsabstractPost-training processes are essential phases in grounding pre-trained language models to real-world tasks, with learning from demonstrations or preference signals playing a crucial role in this adaptation. We present a unified theoretical framework bridging Supervised Fine-Tuning (SFT) and preference learning in Large Language Model (LLM) post-training. Through rigorous mathematical derivation, we demonstrate that both SFT and preference learning methods like Direct Preference Optimization (DPO) operate within the same optimal policy-reward subspace, with SFT representing a special case of implicit reward learning. Our analysis reveals a critical limitation in conventional SFT: the KL divergence term in distribution matching becomes constant with respect to the policy during optimization, failing to constrain model updates. To address this, we propose a simple yet effective learning rate reduction approach that yields significant performance improvements (up to \textbf{25\%} relative gain and \textbf{6\%} absolute win rate increase in instruction following tasks. Additionally, we derive alternative SFT objectives from various f-divergence functions that preserve the KL term during optimization, further enhancing post-DPO model performance. Finally, we extend the theoretical relationship between LLM logits and Q-functions from preference learning to the SFT context, providing mathematical derivations and experimental validation. Bo Wang 0084, Qinyuan Cheng, Runyu Peng, Rong Bao, Peiji Li, Qipeng Guo, Linyang Li, Zhiyuan Zeng 0004, Yunhua Zhou, Xipeng Qiu |
NeurIPS | 9 |
| 2024 | DenoSent: A Denoising Objective for Self-Supervised Sentence Representation LearningabstractContrastive-learning-based methods have dominated sentence representation learning. These methods regularize the representation space by pulling similar sentence representations closer and pushing away the dissimilar ones and have been proven effective in various NLP tasks, e.g., semantic textual similarity (STS) tasks. However, it is challenging for these methods to learn fine-grained semantics as they only learn from the inter-sentence perspective, i.e., their supervision signal comes from the relationship between data samples. In this work, we propose a novel denoising objective that inherits from another perspective, i.e., the intra-sentence perspective. By introducing both discrete and continuous noise, we generate noisy sentences and then train our model to restore them to their original form. Our empirical evaluations demonstrate that this approach delivers competitive results on both semantic textual similarity (STS) and a wide range of transfer tasks, standing up well in comparison to contrastive-learning-based methods. Notably, the proposed intra-sentence denoising objective complements existing inter-sentence contrastive methodologies and can be integrated with them to further enhance performance. Our code is available at https://github.com/xinghaow99/DenoSent. Junliang He, Pengyu Wang 0006, Yunhua Zhou, Tianxiang Sun, Xipeng Qiu |
AAAI | 4 |
| 2024 | AnyGPT: Unified Multimodal LLM with Discrete Sequence ModelingabstractJun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, Hang Yan, Jie Fu, Tao Gui, Tianxiang Sun, Yu-Gang Jiang, Xipeng Qiu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Junqi Dai, Jiasheng Ye, Yunhua Zhou, Zhigeng Liu, Ruibin Yuan, Ge Zhang 0009, Linyang Li, Hang Yan 0001, Jie Fu 0001, Tao Gui, Tianxiang Sun, Yu-Gang Jiang 0001, Xipeng Qiu |
ACL (1) | 4 |
| 2024 | The Open-World Lottery Ticket Hypothesis for OOD Intent ClassificationabstractMost existing methods of Out-of-Domain (OOD) intent classification rely on extensive auxiliary OOD corpora or specific training paradigms. However, they are underdeveloped in the underlying principle that the models should have differentiated confidence in In- and Out-of-domain intent. In this work, we shed light on the fundamental cause of model overconfidence on OOD and demonstrate that calibrated subnetworks can be uncovered by pruning the overparameterized model. Calibrated confidence provided by the subnetwork can better distinguish In- and Out-of-domain, which can be a benefit for almost all post hoc methods. In addition to bringing fundamental insights, we also extend the Lottery Ticket Hypothesis to open-world scenarios. We conduct extensive experiments on four real-world datasets to demonstrate our approach can establish consistent improvements compared with a suite of competitive baselines. Yunhua Zhou, Pengyu Wang 0006, Peiju Liu, Yuxin Wang 0005, Xipeng Qiu |
LREC/COLING | 1 |
| 2024 | Turn Waste into Worth: Rectifying Top-k Router of MoEabstractZhiyuan Zeng, Qipeng Guo, Zhaoye Fei, Zhangyue Yin, Yunhua Zhou, Linyang Li, Tianxiang Sun, Hang Yan, Dahua Lin, Xipeng Qiu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Zhiyuan Zeng 0004, Qipeng Guo, Zhaoye Fei, Zhangyue Yin, Yunhua Zhou, Linyang Li, Tianxiang Sun, Hang Yan 0001, Dahua Lin, Xipeng Qiu |
EMNLP | 5 |
| 2024 | Memorize Step by Step: Efficient Long-Context Prefilling with Incremental Memory and Decremental ChunkabstractZhiyuan Zeng, Qipeng Guo, Xiaoran Liu, Zhangyue Yin, Wentao Shu, Mianqiu Huang, Bo Wang, Yunhua Zhou, Linlin Li, Qun Liu, Xipeng Qiu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Zhiyuan Zeng 0004, Qipeng Guo, Zhangyue Yin, Wentao Shu, Mianqiu Huang, Bo Wang 0084, Yunhua Zhou, Linlin Li 0001, Qun Liu 0001, Xipeng Qiu |
EMNLP | 8 |
| 2023 | UTC-IE: A Unified Token-pair Classification Architecture for Information ExtractionabstractInformation Extraction (IE) spans several tasks with different output structures, such as named entity recognition, relation extraction and event extraction.Previously, those tasks were solved with different models because of diverse task output structures.Through re-examining IE tasks, we find that all of them can be interpreted as extracting spans and span relations.They can further be decomposed into tokenpair classification tasks by using the start and end token of a span to pinpoint the span, and using the start-to-start and end-to-end token pairs of two spans to determine the relation.Based on the reformulation, we propose a Unified Token-pair Classification architecture for Information Extraction (UTC-IE), where we introduce Plusformer on top of the tokenpair feature matrix.Specifically, it models axis-aware interaction with plus-shaped selfattention and local interaction with Convolutional Neural Network over token pairs.Experiments show that our approach outperforms task-specific and unified models on all tasks in 10 datasets, and achieves better or comparable results on 2 joint IE datasets.Moreover, UTC-IE speeds up over state-of-the-art models on IE tasks significantly in most datasets, which verifies the effectiveness of our architecture.1 * Equal contribution. Hang Yan 0001, Yu Sun 0031, Yunhua Zhou, Xuanjing Huang 0001, Xipeng Qiu |
ACL (1) | 4 |
| 2023 | A Probabilistic Framework for Discovering New IntentsabstractDiscovering new intents is of great significance for establishing the Task-Oriented Dialogue System.Most prevailing approaches either cannot transfer prior knowledge inherent in known intents or fall into the dilemma of forgetting prior knowledge in the follow-up.Furthermore, such approaches fail to thoroughly explore the inherent structure of unlabeled data, thereby failing to capture the fundamental characteristics that define an intent in general sense.In this paper, starting from the intuition that discovering intents should be beneficial for identifying known intents, we propose a probabilistic framework for discovering intents where intent assignments are treated as latent variables.We adopt the Expectation Maximization framework for optimization.Specifically, In the Estep, we conduct intent discovery and explore the intrinsic structure of unlabeled data by the posterior of intent assignments.In the M-step, we alleviate the forgetting of prior knowledge transferred from known intents by optimizing the discrimination of labeled data.Extensive experiments conducted on three challenging real-world datasets demonstrate the generality and effectiveness of the proposed framework and implementation.Codes is publicly available.1 Yunhua Zhou, Guofeng Quan, Xipeng Qiu |
ACL (1) | 1 |
| 2023 | Two Birds One Stone: Dynamic Ensemble for OOD Intent ClassificationabstractOut-of-domain (OOD) intent classification is an active field of natural language understanding, which is of great practical significance for intelligent devices such as the Task-Oriented Dialogue System.It mainly contains two challenges: it requires the model to know what it knows and what it does not know.This paper investigates "overthinking" in the openworld scenario and its impact on OOD intent classification.Inspired by this, we propose a two-birds-one-stone method, which allows the model to decide whether to make a decision on OOD classification early during inference and can ensure accuracy and accelerate inference.At the same time, to adapt to the behavior of dynamic inference, we also propose a training method based on ensemble methods.In addition to bringing certain theoretical insights, we also conduct detailed experiments on three real-world intent datasets.Compared with the previous baselines, our method can not only improve inference speed, but also achieve significant performance improvements.Code is publicly available. Yunhua Zhou, Jianqiang Yang, Pengyu Wang 0006, Xipeng Qiu |
ACL (1) | 1 |
| 2023 | Graph Structure Learning via Lottery Hypothesis at Scale
Yuxin Wang 0005, Xiannian Hu, Jiaqing Xie, Zhangyue Yin, Yunhua Zhou, Xipeng Qiu, Xuanjing Huang 0001 |
ACML | 5 |
| 2022 | KNN-Contrastive Learning for Out-of-Domain Intent ClassificationabstractThe Out-of-Domain (OOD) intent classification is a basic and challenging task for dialogue systems.Previous methods commonly restrict the region (in feature space) of In-domain (IND) intent features to be compact or simplyconnected implicitly, which assumes no OOD intents reside, to learn discriminative semantic features.Then the distribution of the IND intent features is often assumed to obey a hypothetical distribution (Gaussian mostly) and samples outside this distribution are regarded as OOD samples.In this paper, we start from the nature of OOD intent classification and explore its optimization objective.We further propose a simple yet effective method, named KNN-contrastive learning.Our approach utilizes K-Nearest Neighbors (KNN) of IND intents to learn discriminative semantic features that are more conducive to OOD detection.Notably, the density-based novelty detection algorithm is so well-grounded in the essence of our method that it is reasonable to use it as the OOD detection algorithm without making any requirements for the feature distribution.Extensive experiments on four public datasets show that our approach can not only enhance the OOD detection performance substantially but also improve the IND intent classification while requiring no restrictions on feature distribution.Code is available.1 Yunhua Zhou, Peiju Liu, Xipeng Qiu |
ACL (1) | 1 |
| 2022 | BBTv2: Towards a Gradient-Free Future with Large Language ModelsabstractMost downstream adaptation methods tune all or part of the parameters of pre-trained models (PTMs) through gradient descent, where the tuning cost increases linearly with the growth of the model size.By contrast, gradient-free methods only require the forward computation of the PTM to tune the prompt, retaining the benefits of efficient tuning and deployment.Though, past work on gradient-free tuning often introduces gradient descent to seek a good initialization of prompt and lacks versatility across tasks and PTMs.In this paper, we present BBTv2, an improved version of Black-Box Tuning (Sun et al., 2022b), to drive PTMs for few-shot learning.We prepend continuous prompts to every layer of the PTM and propose a divide-and-conquer gradient-free algorithm to optimize the prompts at different layers alternately.Extensive experiments across various tasks and PTMs show that BBTv2 can achieve comparable performance to full model tuning and state-of-the-art parameter-efficient methods (e.g., Adapter, LoRA, BitFit, etc.) under few-shot settings while maintaining much fewer tunable parameters. Tianxiang Sun, Zhengfu He, Hong Qian, Yunhua Zhou, Xuanjing Huang 0001, Xipeng Qiu |
EMNLP | 4 |
| 2022 | What Dense Graph Do You Need for Self-Attention?abstractTransformers have made progress in miscellaneous tasks, but suffer from quadratic computational and memory complexities. Recent works propose sparse transformers with attention on sparse graphs to reduce complexity and remain strong performance. While effective, the crucial parts of how dense a graph needs to be to perform well are not fully explored. In this paper, we propose Normalized Information Payload (NIP), a graph scoring function measuring information transfer on graph, which provides an analysis tool for trade-offs between performance and complexity. Guided by this theoretical analysis, we present Hypercube Transformer, a sparse transformer that models token interactions in a hypercube and shows comparable or even better results with vanilla transformer while yielding $O(N\log N)$ complexity with sequence length $N$. Experiments on tasks requiring various sequence lengths lay validation for our graph function well. Yuxin Wang 0005, Chu-Tak Lee, Qipeng Guo, Zhangyue Yin, Yunhua Zhou, Xuanjing Huang 0001, Xipeng Qiu |
ICML | 5 |
| 2019 | Mention Recommendation in Twitter with Cooperative Multi-Agent Reinforcement LearningabstractIn Twitter-like social networking services, the "@'' symbol can be used with the tweet to mention users whom the user wants to alert regarding the message. An automatic suggestion to the user of a small list of candidate names can improve communication efficiency. Previous work usually used several most recent tweets or randomly select historical tweets to make an inference about this preferred list of names. However, because there are too many historical tweets by users and a wide variety of content types, the use of several tweets cannot guarantee the desired results. In this work, we propose the use of a novel cooperative multi-agent approach to mention recommendation, which incorporates dozens of more historical tweets than earlier approaches. The proposed method can effectively select a small set of historical tweets and cooperatively extract relevant indicator tweets from both the user and mentioned users. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods. Tao Gui, Qi Zhang 0001, Minlong Peng, Yunhua Zhou, Xuanjing Huang 0001 |
SIGIR | 6 |