VLDB 2026 Research / reviewers in the wild / expert
Zhenglin Wang
dblp:40/5946
· DBLP profile ↗
14ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-5812-2989ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CATS: Category-Aware Token-level Steering for Training-Free Redundancy Reduction in Large Reasoning ModelsabstractWhile Large Reasoning Models (LRMs) exhibit remarkable capabilities in complex tasks, they often suffer from excessive redundancy in their chain-of-thought reasoning. This significantly reduces inference efficiency and increases computational costs. We identify that LRM redundancy is not uniformly homogeneous but can be taxonomized according to whether it is destructive to the final answer: destructive redundancy (e.g., logical drift, hallucination amplification) versus non-destructive redundancy (e.g., repetition, over-elaboration). Moreover, LRM's redundant and concise responses exhibit a significant distinction in their hidden layer representation spaces. Based on these insights, we propose CATS (Category-Aware Token-level Steering), a training-free and lightweight method to reduce the redundancy phenomenon. CATS decomposes redundancy into six semantically interpretable characteristic dimensions. By flexibly weighting and combining the differential vectors corresponding to these dimensions, CATS synthesizes a composite intervention vector, enabling zero-parameter intervention in the hidden layers. Experiments across three LRM models and five mathematical reasoning datasets demonstrate that CATS reduces reasoning length by an average of 25% while maintaining or even slightly improving task accuracy. CATS offers a pluggable, training-free, and lightweight solution, making it particularly beneficial for users in low-resource environments. Mengfei Zhang, Zhenglin Wang |
AAAI | 2 |
| 2026 | When KV Cache Reuse Fails in Multi-Agent Systems: Cross-Candidate Interaction is Crucial for LLM JudgesabstractMulti-agent LLM systems routinely generate multiple candidate responses that are aggregated by an LLM judge.To reduce the dominant prefill cost in such pipelines, recent work advocates KV cache reuse across partially shared contexts and reports substantial speedups for generation agents.In this work, we show that these efficiency gains do not transfer uniformly to judge-centric inference.Across GSM8K, MMLU, and HumanEval, we find that reuse strategies that are effective for execution agents can severely perturb judge behavior: end-task accuracy may appear stable, yet the judge's selection becomes highly inconsistent with dense prefill.We quantify this risk using Judge Consistency Rate (JCR) and provide diagnostics showing that reuse systematically weakens cross-candidate attention, especially for later candidate blocks.Our ablation further demonstrates that explicit crosscandidate interaction is crucial for preserving dense-prefill decisions.Overall, our results identify a previously overlooked failure mode of KV cache reuse and highlight judge-centric inference as a distinct regime that demands dedicated, risk-aware system design.1 Sichu Liang, Zhenglin Wang, Jiajia Chu, Hui Zang |
ACL (1) | 2 |
| 2025 | SCOPE: Optimizing Key-Value Cache Compression in Long-context GenerationabstractKey-Value (KV) cache has become a bottleneck of LLMs for long-context generation.Despite the numerous efforts in this area, the optimization for the decoding phase is generally ignored.However, we believe such optimization is crucial, especially for long-output generation tasks based on the following two observations: (i) Excessive compression during the prefill phase which requires specific full context, impairs the comprehension of the reasoning task; (ii) Deviation of heavy hitters 1 occurs in the reasoning tasks with long outputs.Therefore, SCOPE, a simple yet efficient framework that separately performs KV cache optimization during the prefill and decoding phases, is introduced.Specifically, the KV cache during the prefill phase is preserved to maintain the essential information, while a novel strategy based on sliding is proposed to select essential heavy hitters for the decoding phase.Memory usage and memory transfer are further optimized using adaptive and discontinuous strategies.Extensive experiments on LONGGENBENCH show the effectiveness and generalization of SCOPE and its compatibility as a plug-in to other prefill-only KV compression methods. 2 Jialong Wu 0007, Zhenglin Wang, Linhai Zhang, Yilong Lai, Yulan He 0001 |
ACL (1) | 2 |
| 2025 | WebWalker: Benchmarking LLMs in Web TraversalabstractRetrieval-augmented generation (RAG) demonstrates remarkable performance across tasks in open-domain question-answering. However, traditional search engines may retrieve shallow content, limiting the ability of LLMs to handle complex, multi-layered information. To address this, we introduce WebWalkerQA, a benchmark designed to assess the ability of LLMs to perform web traversal. It evaluates the capacity of LLMs to traverse a website’s subpages to extract high-quality data systematically. We propose WebWalker, which is a multi-agent framework that mimics human-like web navigation through an explore-critic paradigm. Extensive experimental results show that WebWalkerQA is challenging and demonstrates the effectiveness of RAG combined with WebWalker, through this horizontal and vertical integration in real-world scenarios. Jialong Wu 0007, Wenbiao Yin, Yong Jiang 0005, Zhenglin Wang, Zekun Xi, Runnan Fang, Linhai Zhang, Yulan He 0001, Pengjun Xie, Fei Huang 0002 |
ACL (1) | 4 |
| 2025 | VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward MechanismabstractCongzhi Zhang, Jiawei Peng, Zhenglin Wang, Yilong Lai, Haowen Sun, Heng Chang, Fei Ma, Weijiang Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Congzhi Zhang, Zhenglin Wang, Yilong Lai, Heng Chang, Weijiang Yu |
ACL (1) | 3 |
| 2025 | SEED: Accelerating Reasoning Tree Construction via Scheduled Speculative DecodingabstractLarge Language Models (LLMs) demonstrate remarkable emergent abilities across various tasks, yet fall short of complex reasoning and planning tasks. The tree-search-based reasoning methods address this by encouraging the exploration of intermediate steps, surpassing the capabilities of chain-of-thought prompting. However, significant inference latency is introduced due to the systematic exploration and evaluation of multiple thought paths. This paper introduces SEED, a novel and efficient inference framework to improve both runtime speed and GPU memory management concurrently. Based on a scheduled speculative execution, SEED efficiently handles multiple iterations for thought generation and state evaluation, leveraging a rounds-scheduled strategy to manage draft model dispatching. Extensive experimental evaluations on three reasoning datasets demonstrate the superior speedup performance of SEED. Zhenglin Wang, Jialong Wu 0007, Yilong Lai, Congzhi Zhang |
COLING | 1 |
| 2025 | AdaRewriter: Unleashing the Power of Prompting-based Conversational Query Reformulation via Test-Time AdaptationabstractPrompting-based conversational query reformulation has emerged as a powerful approach for conversational search, refining ambiguous user queries into standalone search queries.Bestof-N reformulation over the generated candidates via prompting shows impressive potential scaling capability.However, both the previous tuning methods (training time) and adaptation approaches (test time) can not fully unleash their benefits.In this paper, we propose AdaRewriter, a novel framework for query reformulation using an outcome-supervised reward model via test-time adaptation.By training a lightweight reward model with contrastive ranking loss, AdaRewriter selects the most promising reformulation during inference.Notably, it can operate effectively in black-box systems, including commercial LLM APIs.Experiments on five conversational search datasets show that AdaRewriter significantly outperforms the existing methods across most settings, demonstrating the potential of test-time adaptation for conversational query reformulation. 1 Yilong Lai, Jialong Wu 0007, Zhenglin Wang |
EMNLP | 3 |
| 2025 | Large Language Models Have Intrinsic Meta-Cognition, but Need a Good LensabstractPrevious research has primarily focused on the cognitive error detection capabilities of Large Language Models (LLMs), often prompting them to analyze mistakes in reasoning chains.However, few studies have examined the metacognitive abilities of LLMs (e.g., their selfawareness of step errors), which are crucial for their reliability.While studies on LLM self-evaluation present some measures, such as perplexity, which can reflect the answer correctness and be viewed as the lens of metacognition, they lack step-level analysis and adaptation.This paper studies the evaluation of LLM meta-cognition using the current lenses and how to improve these lenses.Specifically, we propose AutoMeco, an Automated Metacognition Evaluation framework for benchmarking the existing lenses.Furthermore, a training-free Markovian Intrinsic Reward Adjustment strategy, MIRA, is proposed to boost current meta-cognition lenses.Experimental results on three mathematical reasoning datasets and three LLMs show the reasonableness of AutoMeco by comparing it with Best-of-N verification.Moreover, the meta-cognition ability of LLMs can be better evaluated using MIRA. 1 Qingyue Yuan, Zhenglin Wang |
EMNLP | 3 |
| 2025 | WebDancer: Towards Autonomous Information Seeking AgencyabstractAddressing intricate real-world problems necessitates in-depth information seeking and multi-step reasoning.
Recent progress in agentic systems, exemplified by Deep Research, underscores the potential for autonomous multi-step research.
In this work, we present a cohesive paradigm for building end-to-end agentic information seeking agents from a data-centric and training-stage perspective.
Our approach consists of four key stages: (1) browsing data construction, (2) trajectories sampling, (3) supervised fine-tuning for effective cold start, and (4) reinforcement learning for enhanced generalisation.
We instantiate this framework in a web agent based on the ReAct format, WebDancer.
Empirical evaluations on the challenging GAIA and WebWalkerQA benchmarks demonstrate the strong performance of WebDancer, achieving considerable results and highlighting the efficacy of our training paradigm.
Further analysis of agent training provides valuable insights and actionable, systematic pathways for developing more capable agentic models. Jialong Wu 0007, Baixuan Li, Runnan Fang, Wenbiao Yin, Zhenglin Wang, Zhengwei Tao, Dingchu Zhang, Zekun Xi, Robert Tang, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Jingren Zhou 0001 |
NeurIPS | 6 |
| 2025 | Design of a Blockchain Plugin for Secure Electronic Medical Records
Zhenglin Wang, Umair Ullah Tariq, Md. Mamunur Rashid 0001, Yufeng Lin, Fariza Sabrina, Salahuddin A. Azad |
PDCAT | 1 |
| 2025 | DimSum: Disentangling representation with automatically generated multi-category summary templates for fine-grained opinion summarization
Yanyue Zhang, Yilong Lai, Zhenglin Wang |
Expert Syst. Appl. | 3 |
| 2024 | Opinions Are Not Always Positive: Debiasing Opinion Summarization with Model-Specific and Model-Agnostic MethodsabstractAs in the existing opinion summary data set, more than 70% are positive texts, the current opinion summarization approaches are reluctant to generate the negative opinion summary given the input of negative opinions. To address such sentiment bias, two approaches are proposed through two perspectives: model-specific and model-agnostic. For the model-specific approach, a variational autoencoder is proposed to disentangle the input representation into sentiment-relevant and sentiment-irrelevant components through adversarial loss. Therefore, the sentiment information in the input is kept and employed for the following decoding which avoids interference of content information with emotional signals. To further avoid relying on some specific opinion summarization frameworks, a model-agnostic approach based on counterfactual data augmentation is proposed. A dataset with a more balanced emotional polarity distribution is constructed using a large pre-trained language model based on some pairwise and mini-edited principles. Experimental results show that the sentiment consistency of the generated summaries is significantly improved using the proposed approaches, while their semantics quality is unaffected. Yanyue Zhang, Yilong Lai, Zhenglin Wang, Yulan He 0001 |
LREC/COLING | 3 |
| 2024 | M2A: A model-agnostic and metadata-free adversarial framework for unsupervised opinion summarization
Yanyue Zhang, Zhenglin Wang, Yilong Lai |
Comput. Speech Lang. | 3 |
| 2022 | ConnPrompt: Connective-cloze Prompt Learning for Implicit Discourse Relation RecognitionabstractImplicit Discourse Relation Recognition (IDRR) is to detect and classify relation sense between two text segments without an explicit connective. Vanilla pre-train and fine-tuning paradigm builds upon a Pre-trained Language Model (PLM) with a task-specific neural network. However, the task objective functions are often not in accordance with that of the PLM. Furthermore, this paradigm cannot well exploit some linguistic evidence embedded in the pre-training process. The recent pre-train, prompt, and predict paradigm selects appropriate prompts to reformulate downstream tasks, so as to utilizing the PLM itself for prediction. However, for its success applications, prompts, verbalizer as well as model training should still be carefully designed for different tasks. As the first trial of using this new paradigm for IDRR, this paper develops a Connective-cloze Prompt (ConnPrompt) to transform the relation prediction task as a connective-cloze task. Specifically, we design two styles of ConnPrompt template: Insert-cloze Prompt (ICP) and Prefix-cloze Prompt (PCP) and construct an answer space mapping to the relation senses based on the hierarchy sense tags and implicit connectives. Furthermore, we use a multi-prompt ensemble to fuse predictions from different prompting results. Experiments on the PDTB corpus show that our method significantly outperforms the state-of-the-art algorithms, even with fewer training data. Wei Xiang 0005, Zhenglin Wang, Bang Wang 0001 |
COLING | 2 |