Pingzhi Tang

dblp:408/6122 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0001-7958-7144ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 65% Language models and text generation · 18% Learning paradigms · 8%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression
1.922026
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill & Decode Inference · ASPLOS (2) 2026
TransMLA: Migrating GQA Models to MLA with Full DeepSeek Compatibility and Speedup · NeurIPS 2025
Machine learning › Learning paradigms › continual learning › pre-trained model continual learning
continual knowledge adaptation
1.012026
Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation · ACL (1) 2026
Natural language and speech › Language models and text generation
knowledge editing
1.012026
Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation · ACL (1) 2026
Machine learning › Efficient and distributed learning › inference efficiency
LLM inference optimization
1.012026
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill & Decode Inference · ASPLOS (2) 2026
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
skill transfer
1.012026
Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation · ACL (1) 2026
Machine learning › Efficient and distributed learning › distributed training › model parallelism
tensor parallelism
1.012026
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill & Decode Inference · ASPLOS (2) 2026
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
attention head pruning
0.912025
CLOVER: Cross-Layer Orthogonal Vectors Pruning · ICML 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
CLOVER: Cross-Layer Orthogonal Vectors Pruning · ICML 2025
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.912025
HD-PiSSA: High-Rank Distributed Orthogonal Adaptation · EMNLP 2025
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning
0.912025
CLOVER: Cross-Layer Orthogonal Vectors Pruning · ICML 2025
Natural language and speech › Language models and text generation
large language model fine-tuning
0.312025
HD-PiSSA: High-Rank Distributed Orthogonal Adaptation · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

multi-head latent attention · 1.9supervised fine-tuning · 1.0skill vector injection · 1.0reinforcement learning · 1.0hadamard transform · 1.0allreduce · 1.0PCA · 1.0orthogonal decomposition · 0.9low-rank adaptation · 0.9data parallelism · 0.9
YearPublicationVenuePosition
2026 Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation
abstract
Large Language Models (LLMs) face the "knowledge cutoff" challenge, where their frozen parametric memory prevents direct internalization of new information.While Supervised Fine-Tuning (SFT) is commonly used to update model knowledge, it often updates factual content without reliably improving the model's ability to use the newly incorporated information for question answering or decisionmaking.Reinforcement Learning (RL) is essential for acquiring reasoning skills; however, its high computational cost makes it impractical for efficient online adaptation.We empirically observe that the parameter updates induced by SFT and RL are nearly orthogonal.Based on this observation, we propose Parametric Skill Transfer (PaST), a framework that supports modular skill transfer for efficient and effective knowledge adaptation.By extracting a domain-agnostic Skill Vector from a source domain, we can linearly inject knowledge manipulation skills into a target model after it has undergone lightweight SFT on new data.Experiments on knowledge-incorporation QA (SQuAD, LooGLE) and agentic tool-use benchmarks (ToolBench) demonstrate the effectiveness of our method.On SQuAD, PaST outperforms the state-of-the-art self-editing SFT baseline by up to 9.9 points.PaST further scales to long-context QA on LooGLE with an 8.0-point absolute accuracy gain, and improves zero-shot ToolBench success rates by +10.3 points on average with consistent gains across tool categories, indicating strong scalability and crossdomain transferability of the Skill Vector. Motivation: Functional Disconnect between Knowledge and ReasoningI'm trying to download a post and a reel from Instagram.Can you provide me with the download links for the post and reel?The post link is [post link] and the reel link is [reel link]. Standard SFT (Knowledge Only)Thought: I should call the 'posts_for_instagram_ reels_and_post_downloader' function with the argument 'link' set to [post link]… Action: posts_for_instagram_reels_and_post_ downloader Action Input: {"
Pingzhi Tang, Muhan Zhang
ACL (1)1
2026 TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill & Decode Inference
abstract
Multi-Head Latent Attention (MLA), introduced in DeepSeek-V2, compresses key–value states into a low-rank latent vector cKV, caching only this vector to reduce memory. In tensor parallelism (TP), however, attention heads are computed across multiple devices, and each device must load the full cKV, eroding the advantage of MLA over Grouped Query Attention (GQA). We present TPLA, a scheme that partitions both the latent representation and each head's input dimension across devices, performs attention independently on each shard, and aggregates the results with an all-reduce. Unlike GLA, every attention head in TPLA still attends to the full latent space, preserving MLA's representational capacity while reducing the per-device KV cache. To make TPLA drop-in compatible with MLA checkpoints, we further derive orthogonal reparameterizations of RMSNorm and softmax---instantiated with Hadamard and PCA transforms---that mitigate cross-shard discrepancies when slicing latent vectors across devices. Finally, we introduce a prefill-decode separation scheme that keeps the MLA form during compute-bound prefilling and switches to TPLA during memory-bound decoding, minimizing conversion-induced error. By reducing the per-device KV cache for DeepSeek-V3 and Kimi-K2, we achieve 1.79x) and 1.93) speedups respectively, at a 32K-token context length while maintaining accuracy on commonsense and LongBench benchmarks. TPLA can be further implemented on top of FlashAttention-3, enabling practical end-to-end acceleration.
Xiaojuan Tang, Fanxu Meng 0003, Pingzhi Tang, Yuxuan Wang 0012, Xing Sun 0001, Muhan Zhang
ASPLOS (2)3
2025 HD-PiSSA: High-Rank Distributed Orthogonal Adaptation
abstract
Existing parameter-efficient fine-tuning (PEFT) methods for large language models (LLMs), such as LoRA and PiSSA, constrain model updates to low-rank subspaces, limiting their expressiveness and leading to suboptimal performance on complex tasks.To address this, we introduce High-rank Distributed PiSSA (HD-PiSSA 1 ), a distributed PEFT approach that initializes orthogonal adapters across different devices and aggregates their delta updates collectively on W for fine-tuning.Unlike Data Parallel LoRA or PiSSA, which maintain identical adapters across all devices, HD-PiSSA assigns different principal components of the pre-trained weights to each GPU, significantly expanding the range of update directions.This results in over 16× higher effective updated ranks than data-parallel LoRA or PiSSA when fine-tuning on 8 GPUs with the same per-device adapter rank.Empirically, we evaluate HD-PiSSA across various challenging downstream tasks, including mathematics, code generation, and multi-task learning.In the multi-task setting, HD-PiSSA achieves average gains of 10.0 absolute points (14.63%) over LoRA and 4.98 points (6.60%) over PiSSA across 12 benchmarks, demonstrating its benefits from the extra optimization flexibility.
Pingzhi Tang, Muhan Zhang
EMNLP5
2025 CLOVER: Cross-Layer Orthogonal Vectors Pruning
abstract
Decoder-only models generate tokens autoregressively by caching key/value vectors, but as the cache grows, inference becomes memory-bounded. To address this challenge, we introduce CLOVER (Cross-Layer Orthogonal Vectors) pruning, a novel approach that treats pairs of components of the attention mechanism as low-rank decompositions. CLOVER applies Singular Value Decomposition (SVD) to the Q-K and V-O pairs within each attention head. The resulting singular values, in turn, guide pruning and further serve as trainable parameters for efficient fine-tuning, ultimately enabling the model to recover its performance to the level before pruning.After pruning and fine-tuning, these values are reintegrated into the model without increasing its parameter count. Visualizations across various models show that CLOVER effectively removes linear redundancies within attention heads, greatly improving pruning efficiency. For example, pruning 70% of the Q-K head dimension in GPT-2 XL results in a perplexity comparable to that of pruning just 8% using vanilla pruning. The combination of CLOVER and TransMLA achieves a speedup of up to 11.1$\times$ over LLaMA-2-7B.
Pingzhi Tang, Muhan Zhang
ICML2
2025 TransMLA: Migrating GQA Models to MLA with Full DeepSeek Compatibility and Speedup
abstract
Modern large-language models often face communication bottlenecks on current hardware rather than computational limitations. *Multi-head latent attention (MLA)* addresses this by compressing the key-value cache using low-rank matrices, while the Absorb operation prevents the KV cache from reverting to its original size, significantly boosting both training and inference speed. Despite the success of DeepSeek V2/V3/R1, most model providers have heavily invested in optimizing GQA-based models and, therefore, lack strong incentives to retrain MLA-based models from scratch. This paper demonstrates that MLA provides superior expressive power compared to GQA with the same KV cache overhead, thereby offering a rationale for transitioning from GQA to MLA. In addition, we introduce TransMLA, a framework that seamlessly converts any GQA-based pre-trained model (e.g., LLaMA, Qwen, Gemma, Mistral/Mixtral) into an MLA-based model. For the first time, our method enables *direct conversion of these models into a format compatible with DeepSeek's codebase*, allowing them to fully leverage the existing, highly-optimized support for the DeepSeek architecture within inference engines like vLLM and SGlang. By compressing 93\% of the KV cache in LLaMA-2-7B, we achieve a **10x speedup** with an 8K context length while maintaining meaningful output. Moreover, the model requires only **6B tokens** for fine-tuning to recover comparable performance across multiple benchmarks. TransMLA provides a practical path for migrating GQA-based models to the MLA structure, and when combined with DeepSeek’s advanced optimizations—such as FP8 quantization and Multi-Token Prediction—further inference acceleration can be achieved.
Fanxu Meng 0003, Pingzhi Tang, Zengwei Yao, Xing Sun 0001, Muhan Zhang
NeurIPS2