VLDB 2026 Research / reviewers in the wild / expert
Wenyu Du
dblp:38/10657 · also Wenyu Derek Du
· DBLP profile ↗
14ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel ReasoningabstractParallel reasoning enhances Large Reasoning Models (LRMs) but incurs prohibitive costs due to futile paths caused by early errors.To mitigate this, path pruning at the prefix level is essential, yet existing research remains fragmented without a standardized framework.In this work, we propose the first systematic taxonomy of path pruning, categorizing methods by their signal source (internal vs. external) and learnability (learnable vs. non-learnable).This classification reveals the unexplored potential of learnable internal methods, motivating our proposal of STOP (Super TOken for Pruning).Extensive evaluations across LRMs ranging from 1.5B to 20B parameters demonstrate that STOP achieves superior effectiveness and efficiency compared to existing baselines.Furthermore, we rigorously validate the scalability of STOP under varying compute budgets-for instance, boosting GPT-OSS-20B accuracy on AIME25 from 84% to nearly 90% under fixed compute budgets.Finally, we distill our findings into formalized empirical guidelines to facilitate optimal real-world deployment.Code, data and models are available at https://bijiaxihh.github.io/STOP. Jiaxi Bi, Tongxu Luo, Wenyu Du, Zhengyang Tang, Benyou Wang |
ACL (1) | 3 |
| 2025 | Finite State Automata Inside Transformers with Chain-of-Thought: A Mechanistic Study on State TrackingabstractChain-of-thought (CoT) significantly enhances the performance of large language models (LLMs) across a wide range of tasks, and prior research shows that CoT can theoretically increase expressiveness. However, there is limited mechanistic understanding of the algorithms that Transformer+CoT can learn. Our key contributions are: (1) We evaluate the state tracking capabilities of Transformer+CoT and its variants, confirming the effectiveness of CoT. (2) Next, we identify the circuit (a subset of model components, responsible for tracking the world state), indicating that late-layer MLP neurons play a key role. We propose two metrics, compression and distinction, and show that the neuron sets for each state achieve nearly 100% accuracy, providing evidence of an implicit finite state automaton (FSA) embedded within the model. (3) Additionally, we explore three challenging settings: skipping intermediate steps, introducing data noises, and testing length generalization. Our results demonstrate that Transformer+CoT learns robust algorithms (FSAs), highlighting its resilience in challenging scenarios. Our code is available at https://github.com/IvanChangPKU/FSA. Yifan Zhang 0004, Wenyu Du, Dongming Jin, Jie Fu 0001, Zhi Jin 0001 |
ACL (1) | 2 |
| 2025 | Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit AnalysisabstractFine-tuning significantly improves the performance of Large Language Models (LLMs), yet its underlying mechanisms remain poorly understood. This paper aims to provide an in-depth interpretation of the fine-tuning process through circuit analysis, a popular tool in *Mechanistic Interpretability (MI)*. Unlike previous studies (Prakash et al. 2024, Chhabra et al. 2024) that focus on tasks where pre-trained models already perform well, we develop a set of mathematical tasks where fine-tuning yields substantial performance gains, bringing the setup closer to real-world scenarios. In our experiments, we identify circuits at various checkpoints during fine-tuning and examine the interplay between circuit analysis, fine-tuning methods, and task complexities. First, we find that while circuits maintain high node similarity before and after fine-tuning, their edges undergo significant changes, contrasting with previous work (Prakash et al. 2024, Chhabra et al. 2024) that reported only small circuit additions after fine-tuning. Based on these observations, we develop a **circuit-aware Low-Rank Adaptation (LoRA)** method that assigns ranks to layers according to edge changes in the circuits. Experimental results demonstrate that our circuit-based LoRA achieves an average improvement of 2.46% over standard LoRA with comparable parameter sizes. Furthermore, we explore how combining circuits from subtasks can enhance fine-tuning in compositional tasks, offering new insights into task design and deepening our understanding of circuit dynamics and fine-tuning mechanisms. Xu Wang 0033, Wenyu Du, Reynold Cheng, Benyou Wang, Difan Zou |
ICML | 3 |
| 2025 | Thinker: Learning to Think Fast and SlowabstractRecent studies show that the reasoning capabilities of Large Language Models (LLMs) can be improved by applying Reinforcement Learning (RL) to question-answering (QA) tasks in areas such as math and coding. With a long context length, LLMs may learn to perform search, as indicated by the self-correction behavior observed in DeepSeek R1. However, this search behavior is often imprecise and lacks confidence, resulting in long, redundant responses and highlighting deficiencies in intuition and verification. Inspired by the Dual Process Theory in psychology, we introduce a simple modification to the QA task that includes four stages: Fast Thinking, where the LLM must answer within a strict token budget; Verification, where the model evaluates its initial response; Slow Thinking, where it refines the initial response with more deliberation; and Summarization, where it distills the refinement from the previous stage into precise steps. Our proposed task improves average accuracy from 25.6% to 27.3% for Qwen2.5-1.5B, and from 45.9% to 51.0% for DeepSeek-R1-Qwen-1.5B. Notably, for Qwen2.5-1.5B, the Fast Thinking mode alone achieves 25.2% accuracy using fewer than 1000 tokens, demonstrating substantial inference efficiency gains. These findings suggest that intuition and deliberative reasoning are distinct, complementary systems benefiting from targeted training. Additionally, we have open-sourced both the trained models and the source code. Stephen Chung, Wenyu Du, Jie Fu 0001 |
NeurIPS | 2 |
| 2024 | Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-TrainingabstractLLMs are computationally expensive to pre-train due to their large scale.
Model growth emerges as a promising approach by leveraging smaller models to accelerate the training of larger ones.
However, the viability of these model growth methods in efficient LLM pre-training remains underexplored.
This work identifies three critical $\underline{\textit{O}}$bstacles: ($\textit{O}$1) lack of comprehensive evaluation, ($\textit{O}$2) untested viability for scaling, and ($\textit{O}$3) lack of empirical guidelines.
To tackle $\textit{O}$1, we summarize existing approaches into four atomic growth operators and systematically evaluate them in a standardized LLM pre-training setting.
Our findings reveal that a depthwise stacking operator, called $G_{\text{stack}}$, exhibits remarkable acceleration in training, leading to decreased loss and improved overall performance on eight standard NLP benchmarks compared to strong baselines.
Motivated by these promising results, we conduct extensive experiments to delve deeper into $G_{\text{stack}}$ to address $\textit{O}$2 and $\textit{O}$3.
For $\textit{O}$2 (untested scalability), our study shows that $G_{\text{stack}}$ is scalable and consistently performs well, with experiments up to 7B LLMs after growth and pre-training LLMs with 750B tokens.
For example, compared to a conventionally trained 7B model using 300B tokens, our $G_{\text{stack}}$ model converges to the same loss with 194B tokens, resulting in a 54.6\% speedup.
We further address $\textit{O}$3 (lack of empirical guidelines) by formalizing guidelines to determine growth timing and growth factor for $G_{\text{stack}}$, making it practical in general LLM pre-training.
We also provide in-depth discussions and comprehensive ablation studies of $G_{\text{stack}}$.
Our code and pre-trained model are available at https://llm-stacking.github.io/. Wenyu Du, Tongxu Luo, Zihan Qiu, Yikang Shen, Reynold Cheng, Yike Guo, Jie Fu 0001 |
NeurIPS | 1 |
| 2023 | Graphix-T5: Mixing Pre-trained Transformers with Graph-Aware Layers for Text-to-SQL ParsingabstractThe task of text-to-SQL parsing, which aims at converting natural language questions into executable SQL queries, has garnered increasing attention in recent years. One of the major challenges in text-to-SQL parsing is domain generalization, i.e., how to generalize well to unseen databases. Recently, the pre-trained text-to-text transformer model, namely T5, though not specialized for text-to-SQL parsing, has achieved state-of-the-art performance on standard benchmarks targeting domain generalization. In this work, we explore ways to further augment the pre-trained T5 model with specialized components for text-to-SQL parsing. Such components are expected to introduce structural inductive bias into text-to-SQL parsers thus improving the model’s capacity on (potentially multi-hop) reasoning, which is critical for generating structure-rich SQLs. To this end, we propose a new architecture GRAPHIX-T5, a mixed model with the standard pre-trained transformer model augmented by specially-designed graph-aware layers. Extensive experiments and analysis demonstrate the effectiveness of GRAPHIX-T5 across four text-to-SQL benchmarks: SPIDER, SYN, REALISTIC and DK. GRAPHIX-T5 surpasses all other T5-based parsers with a significant margin, achieving new state-of-the-art performance. Notably, GRAPHIX-T5-large reaches performance superior to the original T5-large by 5.7% on exact match (EM) accuracy and 6.6% on execution accuracy (EX). This even outperforms the T5-3B by 1.2% on EM and 1.5% on EX Jinyang Li 0003, Binyuan Hui, Reynold Cheng, Bowen Qin, Chenhao Ma 0001, Nan Huo, Fei Huang 0002, Wenyu Du, Luo Si, Yongbin Li 0001 |
AAAI | 8 |
| 2023 | f-Divergence Minimization for Sequence-Level Knowledge DistillationabstractKnowledge distillation (KD) is the process of transferring knowledge from a large model to a small one.It has gained increasing attention in the natural language processing community, driven by the demands of compressing evergrowing language models.In this work, we propose an f -DISTILL framework, which formulates sequence-level knowledge distillation as minimizing a generalized f -divergence function.We propose four distilling variants under our framework and show that existing SeqKD and ENGINE approaches are approximations of our f -DISTILL methods.We further derive step-wise decomposition for our f -DISTILL, reducing intractable sequence-level divergence to word-level losses that can be computed in a tractable manner.Experiments across four datasets show that our methods outperform existing KD approaches, and that our symmetric distilling losses can better force the student to learn from the teacher distribution.1 Yuqiao Wen, Zichao Li 0001, Wenyu Du, Lili Mou |
ACL (1) | 3 |
| 2021 | End-to-End AMR Corefencence ResolutionabstractQiankun Fu, Linfeng Song, Wenyu Du, Yue Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Qiankun Fu, Linfeng Song, Wenyu Du, Yue Zhang 0004 |
ACL/IJCNLP (1) | 3 |
| 2021 | Linguistic Dependencies and Statistical DependenceabstractAre pairs of words that tend to occur together also likely to stand in a linguistic dependency?This empirical question is motivated by a long history of literature in cognitive science, psycholinguistics, and NLP.In this work we contribute an extensive analysis of the relationship between linguistic dependencies and statistical dependence between words.Improving on previous work, we introduce the use of large pretrained language models to compute contextualized estimates of the pointwise mutual information between words (CPMI).For multiple models and languages, we extract dependency trees which maximize CPMI, and compare to gold standard linguistic dependencies.Overall, we find that CPMI dependencies achieve an unlabelled undirected attachment score of at most ≈ 0.5.While far above chance, and consistently above a non-contextualized PMI baseline, this score is generally comparable to a simple baseline formed by connecting adjacent words.We analyze which kinds of linguistic dependencies are best captured in CPMI dependencies, and also find marked differences between the estimates of the large pretrained language models, illustrating how their different training schemes affect the type of dependencies they capture. Jacob Hoover Vigly, Wenyu Du, Alessandro Sordoni, Timothy J. O'Donnell |
EMNLP (1) | 2 |
| 2020 | Exploiting Syntactic Structure for Better Language Modeling: A Syntactic Distance ApproachabstractIt is commonly believed that knowledge of syntactic structure should improve language modeling.However, effectively and computationally efficiently incorporating syntactic structure into neural language models has been a challenging topic.In this paper, we make use of a multi-task objective, i.e., the models simultaneously predict words as well as ground truth parse trees in a form called "syntactic distances", where information between these two separate objectives shares the same intermediate representation.Experimental results on the Penn Treebank and Chinese Treebank datasets show that when ground truth parse trees are provided as additional training signals, the model is able to achieve lower perplexity and induce trees with better quality. Wenyu Du, Zhouhan Lin, Yikang Shen, Timothy J. O'Donnell, Yoshua Bengio, Yue Zhang 0004 |
ACL | 1 |
| 2019 | A Multi-Task Learning Approach for Answer Selection: A Study and a Chinese Law DatasetabstractIn this paper, we propose a Multi-Task learning approach for Answer Selection (MTAS), motivated by the fact that humans have no difficulty performing such task because they possess capabilities of multiple domains (tasks). Specifically, MTAS consists of two key components: (i) A category classification model that learns rich category-aware document representation; (ii) An answer selection model that provides the matching scores of question-answer pairs. These two tasks work on a shared document encoding layer, and they cooperate to learn a high-quality answer selection system. In addition, a multi-head attention mechanism is proposed to learn important information from different representation subspaces at different positions. We manually annotate the first Chinese question answering dataset in law domain (denoted as LawQA) to evaluate the effectiveness of our model. The experimental results show that our model MTAS consistently outperforms the compared methods.1 Wenyu Du, Baocheng Li, Min Yang 0007, Qiang Qu 0001, Ying Shen 0001 |
AAAI | 1 |
| 2019 | Effective organizational improvisation in information systems development: Insights from the Tencent messaging system development
Wenyu Du, Junjie Wu 0002, Shanshi Liu, Ray Hackney |
Inf. Manag. | 1 |
| 2019 | Affordances, experimentation and actualization of FinTech: A blockchain implementation study
Wenyu Du, Shan Ling Pan, Dorothy E. Leidner, Wenchi Ying |
J. Strateg. Inf. Syst. | 1 |
| 2018 | Developing and maintaining clients' trust through institutional mechanisms in online service markets for digital entrepreneurs: A process model
Wenyu Du, Ji-Ye Mao |
J. Strateg. Inf. Syst. | 1 |