VLDB 2026 Research / reviewers in the wild / expert
Fanxu Meng 0003
dblp:256/5536-3
· DBLP profile ↗
6ranked-venue papers
4as first author
4since 2021 · last 2026
0009-0009-2382-556XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Efficient and distributed learning · 74% Language models and text generation · 15% Transfer learning and domain adaptation · 7% |
Topics — the 9 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression |
1.9 | 2 | 2026 | TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill & Decode Inference · ASPLOS (2) 2026 TransMLA: Migrating GQA Models to MLA with Full DeepSeek Compatibility and Speedup · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation |
1.6 | 2 | 2025 | LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades · NeurIPS 2025 PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
1.6 | 2 | 2025 | LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades · NeurIPS 2025 PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › inference efficiency
LLM inference optimization |
1.0 | 1 | 2026 | TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill & Decode Inference · ASPLOS (2) 2026 |
Machine learning › Efficient and distributed learning › distributed training › model parallelism
tensor parallelism |
1.0 | 1 | 2026 | TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill & Decode Inference · ASPLOS (2) 2026 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades · NeurIPS 2025 |
Machine learning › Transfer learning and domain adaptation
model adaptation |
0.9 | 1 | 2025 | LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 2 | 2020 | Pruning Filter in Filter · NeurIPS 2020 Filter Grafting for Deep Neural Networks · CVPR 2020 |
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
channel pruning |
0.4 | 1 | 2020 | Pruning Filter in Filter · NeurIPS 2020 |
Methods — techniques the papers use, named apart from their topics
multi-head latent attention · 1.9quantization · 1.6hadamard transform · 1.0allreduce · 1.0PCA · 1.0transfer matrix · 0.9low-rank compression · 0.9grouped-query attention · 0.9cosine similarity · 0.9centered kernel alignment · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill & Decode InferenceabstractMulti-Head Latent Attention (MLA), introduced in DeepSeek-V2, compresses key–value states into a low-rank latent vector cKV, caching only this vector to reduce memory. In tensor parallelism (TP), however, attention heads are computed across multiple devices, and each device must load the full cKV, eroding the advantage of MLA over Grouped Query Attention (GQA). We present TPLA, a scheme that partitions both the latent representation and each head's input dimension across devices, performs attention independently on each shard, and aggregates the results with an all-reduce. Unlike GLA, every attention head in TPLA still attends to the full latent space, preserving MLA's representational capacity while reducing the per-device KV cache. To make TPLA drop-in compatible with MLA checkpoints, we further derive orthogonal reparameterizations of RMSNorm and softmax---instantiated with Hadamard and PCA transforms---that mitigate cross-shard discrepancies when slicing latent vectors across devices. Finally, we introduce a prefill-decode separation scheme that keeps the MLA form during compute-bound prefilling and switches to TPLA during memory-bound decoding, minimizing conversion-induced error. By reducing the per-device KV cache for DeepSeek-V3 and Kimi-K2, we achieve 1.79x) and 1.93) speedups respectively, at a 32K-token context length while maintaining accuracy on commonsense and LongBench benchmarks. TPLA can be further implemented on top of FlashAttention-3, enabling practical end-to-end acceleration. Xiaojuan Tang, Fanxu Meng 0003, Pingzhi Tang, Yuxuan Wang 0012, Xing Sun 0001, Muhan Zhang |
ASPLOS (2) | 2 |
| 2025 | LoRASuite: Efficient LoRA Adaptation Across Large Language Model UpgradesabstractAs Large Language Models (LLMs) are frequently updated, LoRA weights trained on earlier versions quickly become obsolete. The conventional practice of retraining LoRA weights from scratch on the latest model is costly, time-consuming, and environmentally detrimental, particularly as the diversity of LLMs and downstream tasks expands. This motivates a critical question: "How can we efficiently leverage existing LoRA weights to adapt to newer model versions?" To address this, we propose LoRASuite, a modular approach tailored specifically to various types of LLM updates. First, we compute a transfer matrix utilizing known parameters from both old and new LLMs. Next, we allocate corresponding layers and attention heads based on centered kernel alignment and cosine similarity metrics, respectively. A subsequent small-scale, skillful fine-tuning step ensures numerical stability. Experimental evaluations demonstrate that LoRASuite consistently surpasses small-scale vanilla LoRA methods. Notably, on backbone LLMs such as MiniCPM and Qwen, LoRASuite even exceeds the performance of full-scale LoRA retraining, with average improvements of +1.4 and +6.6 points on math tasks, respectively. Additionally, LoRASuite significantly reduces memory consumption by 5.5 GB and computational time by 78.23%. Fanxu Meng 0003, Muhan Zhang, Shiai Zhu, Shangguang Wang, Mengwei Xu 0001 |
NeurIPS | 2 |
| 2025 | TransMLA: Migrating GQA Models to MLA with Full DeepSeek Compatibility and SpeedupabstractModern large-language models often face communication bottlenecks on current hardware rather than computational limitations.
*Multi-head latent attention (MLA)* addresses this by compressing the key-value cache using low-rank matrices, while the Absorb operation prevents the KV cache from reverting to its original size, significantly boosting both training and inference speed.
Despite the success of DeepSeek V2/V3/R1, most model providers have heavily invested in optimizing GQA-based models and, therefore, lack strong incentives to retrain MLA-based models from scratch.
This paper demonstrates that MLA provides superior expressive power compared to GQA with the same KV cache overhead, thereby offering a rationale for transitioning from GQA to MLA.
In addition, we introduce TransMLA, a framework that seamlessly converts any GQA-based pre-trained model (e.g., LLaMA, Qwen, Gemma, Mistral/Mixtral) into an MLA-based model.
For the first time, our method enables *direct conversion of these models into a format compatible with DeepSeek's codebase*, allowing them to fully leverage the existing, highly-optimized support for the DeepSeek architecture within inference engines like vLLM and SGlang.
By compressing 93\% of the KV cache in LLaMA-2-7B, we achieve a **10x speedup** with an 8K context length while maintaining meaningful output.
Moreover, the model requires only **6B tokens** for fine-tuning to recover comparable performance across multiple benchmarks.
TransMLA provides a practical path for migrating GQA-based models to the MLA structure, and when combined with DeepSeek’s advanced optimizations—such as FP8 quantization and Multi-Token Prediction—further inference acceleration can be achieved. Fanxu Meng 0003, Pingzhi Tang, Zengwei Yao, Xing Sun 0001, Muhan Zhang |
NeurIPS | 1 |
| 2024 | PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language ModelsabstractTo parameter-efficiently fine-tune (PEFT) large language models (LLMs), the low-rank adaptation (LoRA) method approximates the model changes $\Delta W \in \mathbb{R}^{m \times n}$ through the product of two matrices $A \in \mathbb{R}^{m \times r}$ and $B \in \mathbb{R}^{r \times n}$, where $r \ll \min(m, n)$, $A$ is initialized with Gaussian noise, and $B$ with zeros. LoRA **freezes the original model $W$** and **updates the "Noise \& Zero" adapter**, which may lead to slow convergence. To overcome this limitation, we introduce **P**r**i**ncipal **S**ingular values and **S**ingular vectors **A**daptation (PiSSA). PiSSA shares the same architecture as LoRA, but initializes the adaptor matrices $A$ and $B$ with the principal components of the original matrix $W$, and put the remaining components into a residual matrix $W^{res} \in \mathbb{R}^{m \times n}$ which is frozen during fine-tuning.
Compared to LoRA, PiSSA **updates the principal components** while **freezing the "residual" parts**, allowing faster convergence and enhanced performance. Comparative experiments of PiSSA and LoRA across 11 different models, ranging from 184M to 70B, encompassing 5 NLG and 8 NLU tasks, reveal that PiSSA consistently outperforms LoRA under identical experimental setups. On the GSM8K benchmark, Gemma-7B fine-tuned with PiSSA achieves an accuracy of 77.7\%, surpassing LoRA's 74.53\% by 3.25\%. Due to the same architecture, PiSSA is also compatible with quantization to further reduce the memory requirement of fine-tuning. Compared to QLoRA, QPiSSA (PiSSA with 4-bit quantization) exhibits smaller quantization errors in the initial stages. Fine-tuning LLaMA-3-70B on GSM8K, QPiSSA attains an accuracy of 86.05\%, exceeding the performances of QLoRA at 81.73\%. Leveraging a fast SVD technique, PiSSA can be initialized in only a few seconds, presenting a negligible cost for transitioning from LoRA to PiSSA. Fanxu Meng 0003, Muhan Zhang |
NeurIPS | 1 |
| 2020 | Filter Grafting for Deep Neural NetworksabstractThis paper proposes a new learning paradigm called filter grafting, which aims to improve the representation capability of Deep Neural Networks (DNNs). The motivation is that DNNs have unimportant (invalid) filters (e.g., l1norm close to 0). These filters limit the potential of DNNs since they are identified as having little effect on the network. While filter pruning removes these invalid filters for efficiency consideration, filter grafting re-activates them from an accuracy boosting perspective. The activation is processed by grafting external information (weights) into invalid filters. To better perform the grafting process, we develop an entropy-based criterion to measure the information of filters and an adaptive weighting strategy for balancing the grafted information among networks. After the grafting operation, the network has very few invalid filters compared with its untouched state, empowering the model with more representation capacity. We also perform extensive experiments on the classification and recognition tasks to show the superiority of our method. For example, the grafted MobileNetV2 outperforms the non-grafted MobileNetV2 by about 7 percent on CIFAR-100 dataset. Fanxu Meng 0003, Hao Cheng 0012, Ke Li 0015, Zhixin Xu, Rongrong Ji, Xing Sun 0001, Guangming Lu 0002 |
CVPR | 1 |
| 2020 | Pruning Filter in FilterabstractPruning has become a very powerful and effective technique to compress and accelerate modern neural networks. Existing pruning methods can be grouped into two categories: filter pruning (FP) and weight pruning (WP). FP wins at hardware compatibility but loses at the compression ratio compared with WP. To converge the strength of both methods, we propose to prune the filter in the filter. Specifically, we treat a filter F, whose size is CKK, as KK stripes, i.e., 11 filters, then by pruning the stripes instead of the whole filter, we can achieves finer granularity than traditional FP while being hardware friendly. We term our method as SWP (Stripe-Wise Pruning). SWP is implemented by introducing a novel learnable matrix called Filter Skeleton, whose values reflect the optimal shape of each filter. As some recent work has shown that the pruned architecture is more crucial than the inherited important weights, we argue that the architecture of a single filter, i.e., the Filter Skeleton, also matters. Through extensive experiments, we demonstrate that SWP is more effective compared to the previous FP-based methods and achieves the state-of-art pruning ratio on CIFAR-10 and ImageNet datasets without obvious accuracy drop. Fanxu Meng 0003, Hao Cheng 0012, Ke Li 0015, Huixiang Luo, Guangming Lu 0002, Xing Sun 0001 |
NeurIPS | 1 |