Zhoutong Wu

dblp:291/3917 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2025
0009-0005-6137-5492ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Deep learning architectures and training · 36% Optimization for machine learning · 30% Efficient and distributed learning · 23%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
attention mechanism
0.912025
Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads · NeurIPS 2025
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression
0.912025
Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads · NeurIPS 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads · NeurIPS 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads · NeurIPS 2025
Machine learning › Deep learning architectures and training › equilibrium models
deep equilibrium model
0.812024
Separation and Bias of Deep Equilibrium Models on Expressivity and Learning Dynamics · NeurIPS 2024
Machine learning › Optimization for machine learning
gradient-based optimization
0.812024
Designing Universally-Approximating Deep Neural Networks: A First-Order Optimization Approach · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Optimization for machine learning
gradient flow
0.812024
Separation and Bias of Deep Equilibrium Models on Expressivity and Learning Dynamics · NeurIPS 2024
Machine learning › Optimization for machine learning
implicit regularization
0.812024
Separation and Bias of Deep Equilibrium Models on Expressivity and Learning Dynamics · NeurIPS 2024
Machine learning › Learning theory › approximation theory › neural network approximation
universal approximation
0.812024
Designing Universally-Approximating Deep Neural Networks: A First-Order Optimization Approach · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Deep learning architectures and training
feedforward neural network
0.212024
Separation and Bias of Deep Equilibrium Models on Expressivity and Learning Dynamics · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

skip connections · 0.9multi-latent attention · 0.9group-query attention · 0.9normalization · 0.8implicit regularization · 0.8gradient flow · 0.8gradient descent · 0.8first-order optimization · 0.8downsampling · 0.8
YearPublicationVenuePosition
2025 Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
abstract
Transformer models have driven breakthroughs across various language tasks by their strong capability to learn rich contextual representations. Scaling them to improve representation, however, often demands substantial memory and compute costs, such as the Key-Value (KV) cache used during auto-regressive decoding. Skip connections offer a promising way to improve representation without bloating resource usage, yet most prior works either improve expressivity while leaving KV costs unchanged, or reduce memory at the cost of weaker representation. In this work, we propose SkipV1Former, a Transformer variant that uses skip connections from the first layer's Value heads to strengthen model representation and reduce KV cache. Specifically, from the second block onward, each layer reuses half of its Value heads from the very first layer, while computing the other half as usual-cutting Value projections and V cache by nearly 50 \%. Theoretically, we show that routing uncompressed first-layer Values into deeper layers restores information lost to compression and accelerates the model’s implicit mesa-optimization-a key pattern of Transformer in auto-regressive tasks. Empirically, across different model scales, SkipV1Former delivers consistent reductions of approximately 25 \% in KV cache while improving perplexity relative to standard Multi-Head Attention (MHA) Transformers and some advanced variants. Moreover, we propose a recipe for uptraining existing MHA Transformer checkpoints to SkipV1Former with only 10-15\% additional compute. Finally, SkipV1Former can seamlessly combine advanced methods like Group-Query Attention and Multi-Latent Attention to achieve further KV cache savings and performance improvement. When combined with YOCO, it cuts KV cache size by nearly 50 \% while still improving performance. The code is available at: https://github.com/Zhoutong-Wu/SkipV1Former.
Zhoutong Wu, Yiming Dong, Chenheng Zhang, Cong Fang 0001, Zhouchen Lin
NeurIPS1
2024 Separation and Bias of Deep Equilibrium Models on Expressivity and Learning Dynamics
abstract
The deep equilibrium model (DEQ) generalizes the conventional feedforward neural network by fixing the same weights for each layer block and extending the number of layers to infinity. This novel model directly finds the fixed points of such a forward process as features for prediction. Despite empirical evidence showcasing its efficacy compared to feedforward neural networks, a theoretical understanding for its separation and bias is still limited. In this paper, we take a step by proposing some separations and studying the bias of DEQ in its expressive power and learning dynamics. The results include: (1) A general separation is proposed, showing the existence of a width-$m$ DEQ that any fully connected neural networks (FNNs) with depth $O(m^{\alpha})$ for $\alpha \in (0,1)$ cannot approximate unless its width is sub-exponential in $m$; (2) DEQ with polynomially bounded size and magnitude can efficiently approximate certain steep functions (which has very large derivatives) in $L^{\infty}$ norm, whereas FNN with bounded depth and exponentially bounded width cannot unless its weights magnitudes are exponentially large; (3) The implicit regularization caused by gradient flow from a diagonal linear DEQ is characterized, with specific examples showing the benefits brought by such regularization. From the overall study, a high-level conjecture from our analysis and empirical validations is that DEQ has potential advantages in learning certain high-frequency components.
Zhoutong Wu, Yimu Zhang, Cong Fang 0001, Zhouchen Lin
NeurIPS1
2024 Designing Universally-Approximating Deep Neural Networks: A First-Order Optimization Approach
abstract
Universal approximation capability, also referred to as universality, is an important property of deep neural networks, endowing them with the potency to accurately represent the underlying target function in learning tasks. In practice, the architecture of deep neural networks largely influences the performance of the models. However, most existing methodologies for designing neural architectures, such as the heuristic manual design or neural architecture search, ignore the universal approximation property, thus losing a potential safeguard about the performance. In this paper, we propose a unified framework to design the architectures of deep neural networks with a universality guarantee based on first-order optimization algorithms, where the forward pass is interpreted as the updates of an optimization algorithm. The (explicit or implicit) network is designed by replacing each gradient term in the algorithm with a learnable module similar to a two-layer network or its derivatives. Specifically, we explore the realm of width-bounded neural networks, a common practical scenario, showcasing their universality. Moreover, adding operations of normalization, downsampling, and upsampling does not hurt the universality. To the best of our knowledge, this is the first work that width-bounded networks with universal approximation guarantee can be designed in a principled way. Our framework can inspire a variety of neural architectures including some renowned structures such as ResNet and DenseNet, as well as novel innovations. The experimental results on image classification problems demonstrate that the newly inspired networks are competitive and surpass the baselines of ResNet, DenseNet, as well as the advanced ConvNeXt and ViT, testifying to the effectiveness of our framework.
Zhoutong Wu, Mingqing Xiao 0002, Cong Fang 0001, Zhouchen Lin
IEEE Trans. Pattern Anal. Mach. Intell.1