VLDB 2026 Research / reviewers in the wild / expert
Zhoutong Wu
dblp:291/3917
· DBLP profile ↗
3ranked-venue papers
3as first author
3since 2021 · last 2025
0009-0005-6137-5492ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Deep learning architectures and training · 36% Optimization for machine learning · 30% Efficient and distributed learning · 23% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
attention mechanism |
0.9 | 1 | 2025 | Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression |
0.9 | 1 | 2025 | Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.9 | 1 | 2025 | Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › equilibrium models
deep equilibrium model |
0.8 | 1 | 2024 | Separation and Bias of Deep Equilibrium Models on Expressivity and Learning Dynamics · NeurIPS 2024 |
Machine learning › Optimization for machine learning
gradient-based optimization |
0.8 | 1 | 2024 | Designing Universally-Approximating Deep Neural Networks: A First-Order Optimization Approach · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Optimization for machine learning
gradient flow |
0.8 | 1 | 2024 | Separation and Bias of Deep Equilibrium Models on Expressivity and Learning Dynamics · NeurIPS 2024 |
Machine learning › Optimization for machine learning
implicit regularization |
0.8 | 1 | 2024 | Separation and Bias of Deep Equilibrium Models on Expressivity and Learning Dynamics · NeurIPS 2024 |
Machine learning › Learning theory › approximation theory › neural network approximation
universal approximation |
0.8 | 1 | 2024 | Designing Universally-Approximating Deep Neural Networks: A First-Order Optimization Approach · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Machine learning › Deep learning architectures and training
feedforward neural network |
0.2 | 1 | 2024 | Separation and Bias of Deep Equilibrium Models on Expressivity and Learning Dynamics · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
skip connections · 0.9multi-latent attention · 0.9group-query attention · 0.9normalization · 0.8implicit regularization · 0.8gradient flow · 0.8gradient descent · 0.8first-order optimization · 0.8downsampling · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Model Representation and Reducing KV Cache via Skip Connections with First Value HeadsabstractTransformer models have driven breakthroughs across various language tasks by their strong capability to learn rich contextual representations. Scaling them to improve representation, however, often demands substantial memory and compute costs, such as the Key-Value (KV) cache used during auto-regressive decoding. Skip connections offer a promising way to improve representation without bloating resource usage, yet most prior works either improve expressivity while leaving KV costs unchanged, or reduce memory at the cost of weaker representation. In this work, we propose SkipV1Former, a Transformer variant that uses skip connections from the first layer's Value heads to strengthen model representation and reduce KV cache. Specifically, from the second block onward, each layer reuses half of its Value heads from the very first layer, while computing the other half as usual-cutting Value projections and V cache by nearly 50 \%. Theoretically, we show that routing uncompressed first-layer Values into deeper layers restores information lost to compression and accelerates the model’s implicit mesa-optimization-a key pattern of Transformer in auto-regressive tasks. Empirically, across different model scales, SkipV1Former delivers consistent reductions of approximately 25 \% in KV cache while improving perplexity relative to standard Multi-Head Attention (MHA) Transformers and some advanced variants. Moreover, we propose a recipe for uptraining existing MHA Transformer checkpoints to SkipV1Former with only 10-15\% additional compute. Finally, SkipV1Former can seamlessly combine advanced methods like Group-Query Attention and Multi-Latent Attention to achieve further KV cache savings and performance improvement. When combined with YOCO, it cuts KV cache size by nearly 50 \% while still improving performance. The code is
available at: https://github.com/Zhoutong-Wu/SkipV1Former. Zhoutong Wu, Yiming Dong, Chenheng Zhang, Cong Fang 0001, Zhouchen Lin |
NeurIPS | 1 |
| 2024 | Separation and Bias of Deep Equilibrium Models on Expressivity and Learning DynamicsabstractThe deep equilibrium model (DEQ) generalizes the conventional feedforward neural network by fixing the same weights for each layer block and extending the number of layers to infinity. This novel model directly finds the fixed points of such a forward process as features for prediction. Despite empirical evidence showcasing its efficacy
compared to feedforward neural networks, a theoretical understanding for its separation and bias is still limited. In this paper, we take a step
by proposing some separations and studying the bias of DEQ in its expressive power and learning dynamics. The results include: (1) A general separation is proposed, showing the existence of a width-$m$ DEQ that any fully connected neural networks (FNNs) with depth $O(m^{\alpha})$ for $\alpha \in (0,1)$ cannot
approximate unless its width is sub-exponential in $m$; (2) DEQ with polynomially bounded size and magnitude can efficiently approximate certain steep functions (which has very large derivatives) in $L^{\infty}$ norm, whereas FNN with bounded depth and exponentially bounded width cannot unless its weights magnitudes are exponentially large; (3) The implicit regularization caused by gradient flow from a diagonal linear DEQ is characterized, with specific examples showing the benefits brought by such regularization.
From the overall study, a high-level conjecture from our analysis and empirical validations is that DEQ has potential advantages in learning certain high-frequency components. Zhoutong Wu, Yimu Zhang, Cong Fang 0001, Zhouchen Lin |
NeurIPS | 1 |
| 2024 | Designing Universally-Approximating Deep Neural Networks: A First-Order Optimization ApproachabstractUniversal approximation capability, also referred to as universality, is an important property of deep neural networks, endowing them with the potency to accurately represent the underlying target function in learning tasks. In practice, the architecture of deep neural networks largely influences the performance of the models. However, most existing methodologies for designing neural architectures, such as the heuristic manual design or neural architecture search, ignore the universal approximation property, thus losing a potential safeguard about the performance. In this paper, we propose a unified framework to design the architectures of deep neural networks with a universality guarantee based on first-order optimization algorithms, where the forward pass is interpreted as the updates of an optimization algorithm. The (explicit or implicit) network is designed by replacing each gradient term in the algorithm with a learnable module similar to a two-layer network or its derivatives. Specifically, we explore the realm of width-bounded neural networks, a common practical scenario, showcasing their universality. Moreover, adding operations of normalization, downsampling, and upsampling does not hurt the universality. To the best of our knowledge, this is the first work that width-bounded networks with universal approximation guarantee can be designed in a principled way. Our framework can inspire a variety of neural architectures including some renowned structures such as ResNet and DenseNet, as well as novel innovations. The experimental results on image classification problems demonstrate that the newly inspired networks are competitive and surpass the baselines of ResNet, DenseNet, as well as the advanced ConvNeXt and ViT, testifying to the effectiveness of our framework. Zhoutong Wu, Mingqing Xiao 0002, Cong Fang 0001, Zhouchen Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |