Zipeng Xiao

dblp:359/3378 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 42% Deep learning architectures and training · 37% Language models and text generation · 21%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational science and engineering · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational science and engineering
scientific machine learning
1.522024
Amortized Fourier Neural Operators · NeurIPS 2024
Improved Operator Learning by Orthogonal Attention · ICML 2024
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression
0.912025
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection · ICLR 2025
Natural language and speech › Language models and text generation
large language model inference
0.912025
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection · ICLR 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection · ICLR 2025
Machine learning › Deep learning architectures and training › neural operator
fourier neural operator
0.812024
Amortized Fourier Neural Operators · NeurIPS 2024
Machine learning › Deep learning architectures and training
neural operator
0.812024
Amortized Fourier Neural Operators · NeurIPS 2024
Computational science and engineering › scientific machine learning
neural operator
0.812024
Improved Operator Learning by Orthogonal Attention · ICML 2024
Computational science and engineering › scientific machine learning › physics-informed machine learning › physics-informed neural networks
partial differential equation solving
0.812024
Amortized Fourier Neural Operators · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

orthogonal embedding · 1.5kolmogorov-arnold network · 1.5principal component analysis · 0.9orthogonal projection · 0.9matryoshka learning · 0.9low-rank projection · 0.9orthogonalization · 0.8multilayer perceptron · 0.8multi-layer perceptron · 0.8attention mechanism · 0.8
YearPublicationVenuePosition
2025 MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
abstract
KV cache has become a *de facto* technique for the inference of large language models (LLMs), where tensors of shape (layer number, head number, sequence length, feature dimension) are introduced to cache historical information for self-attention. As the size of the model and data grows, the KV cache can, yet, quickly become a bottleneck within the system in both storage and memory transfer. To address this, prior studies usually focus on the first three axes of the cache tensors for compression. This paper supplements them, focusing on the feature dimension axis, by utilizing low-rank projection matrices to transform the cache features into spaces with reduced dimensions. We begin by investigating the canonical orthogonal projection method for data compression through principal component analysis (PCA). We identify the drawback of PCA projection that model performance degrades rapidly under relatively low compression rates (less than 60%). This phenomenon is elucidated by insights derived from the principles of attention mechanisms. To bridge the gap, we propose to directly tune the orthogonal projection matrix on the continual pre-training or supervised fine-tuning datasets with an elaborate Matryoshka learning strategy. Thanks to such a strategy, we can adaptively search for the optimal compression rates for various layers and heads given varying compression budgets. Compared to Multi-head Latent Attention (MLA), our method can easily embrace pre-trained LLMs and hold a smooth tradeoff between performance and compression rate. We witness the high data efficiency of our training procedure and find that our method can sustain over 90\% performance with an average KV cache compression rate of 60% (and up to 75% in certain extreme scenarios) for popular LLMs like LLaMA2 and Mistral.
Bokai Lin, Zihao Zeng, Zipeng Xiao, Siqi Kou, Xiaofeng Gao 0001, Zhijie Deng
ICLR3
2024 Improved Operator Learning by Orthogonal Attention
abstract
This work presents orthogonal attention for constructing neural operators to serve as surrogates to model the solutions of a family of Partial Differential Equations (PDEs). The motivation is that the kernel integral operator, which is usually at the core of neural operators, can be reformulated with orthonormal eigenfunctions. Inspired by the success of the neural approximation of eigenfunctions (Deng et al., 2022), we opt to directly parameterize the involved eigenfunctions with flexible neural networks (NNs), based on which the input function is then transformed by the rule of kernel integral. Surprisingly, the resulting NN module bears a striking resemblance to regular attention mechanisms, albeit without softmax. Instead, it incorporates an orthogonalization operation that provides regularization during model training and helps mitigate overfitting, particularly in scenarios with limited data availability. In practice, the orthogonalization operation can be implemented with minimal additional overheads. Experiments on six standard neural operator benchmark datasets comprising both regular and irregular geometries show that our method can outperform competing baselines with decent margins.
Zipeng Xiao, Zhongkai Hao, Bokai Lin, Zhijie Deng, Hang Su 0006
ICML1
2024 Amortized Fourier Neural Operators
abstract
Fourier Neural Operators (FNOs) have shown promise for solving partial differential equations (PDEs). Typically, FNOs employ separate parameters for different frequency modes to specify tunable kernel integrals in Fourier space, which, yet, results in an undesirably large number of parameters when solving high-dimensional PDEs. A workaround is to abandon the frequency modes exceeding a predefined threshold, but this limits the FNOs' ability to represent high-frequency details and poses non-trivial challenges for hyper-parameter specification. To address these, we propose AMortized Fourier Neural Operator (AM-FNO), where an amortized neural parameterization of the kernel function is deployed to accommodate arbitrarily many frequency modes using a fixed number of parameters. We introduce two implementations of AM-FNO, based on the recently developed, appealing Kolmogorov–Arnold Network (KAN) and Multi-Layer Perceptrons (MLPs) equipped with orthogonal embedding functions respectively. We extensively evaluate our method on diverse datasets from various domains and observe up to 31\% average improvement compared to competing neural operator baselines.
Zipeng Xiao, Siqi Kou, Zhongkai Hao, Bokai Lin, Zhijie Deng
NeurIPS1