VLDB 2026 Research / reviewers in the wild / expert
Peng Xu 0046
dblp:84/586-46
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0006-4482-9760ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Efficient and distributed learning · 75% Language models and text generation · 25% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
GPUs and heterogeneous computing · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › model compression › quantization
KV cache quantization |
0.9 | 1 | 2025 | MILLION: MasterIng Long-Context LLM Inference Via Outlier-Immunized KV Product QuaNtization · DAC 2025 |
Natural language and speech › Language models and text generation › large language model inference
long-context inference |
0.9 | 1 | 2025 | MILLION: MasterIng Long-Context LLM Inference Via Outlier-Immunized KV Product QuaNtization · DAC 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | MILLION: MasterIng Long-Context LLM Inference Via Outlier-Immunized KV Product QuaNtization · DAC 2025 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.9 | 1 | 2025 | MILLION: MasterIng Long-Context LLM Inference Via Outlier-Immunized KV Product QuaNtization · DAC 2025 |
GPUs and heterogeneous computing › deep learning on GPUs
GPU inference |
0.3 | 1 | 2025 | MILLION: MasterIng Long-Context LLM Inference Via Outlier-Immunized KV Product QuaNtization · DAC 2025 |
Methods — techniques the papers use, named apart from their topics
sparse computation · 1.7product quantization · 1.7non-uniform quantization · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploiting Differential-Based Data Encoding for Enhanced Query EfficiencyabstractStoring large-scale high-dimensional data, which is rapidly generated by both industry and academia, poses substantial challenges, primarily in terms of storage and maintenance costs. While data compression techniques offer a potential solution to these challenges, they must overcome two critical hurdles: 1) preserving data integrity within lossless bounds and 2) maintaining query performance on compressed data. Fangxin Liu, Zongwu Wang, Peng Xu 0046, Shiyuan Huang 0004, Li Jiang 0002 |
ASP-DAC | 3 |
| 2025 | MILLION: MasterIng Long-Context LLM Inference Via Outlier-Immunized KV Product QuaNtizationabstractLarge language models (LLMs) are increasingly utilized for complex tasks requiring longer context lengths, with some models supporting up to 128 K or 1 M tokens. This trend, however, presents significant challenges in inference speed and memory management. The primary bottleneck in long-context LLM inference is the quadratic computational complexity of attention mechanisms, causing substantial slowdowns as sequence length increases. KV cache mechanism alleviates this issue by storing pre-computed data, but introduces memory requirements that scale linearly with context length, hindering efficient LLM deployment. Quantization emerges as a promising approach to address the widening gap between LLM size and memory capacity. However, traditional quantization schemes often yield suboptimal compression results for KV caches due to two key factors: i) On-the-fly quantization and de-quantization, causing significant performance overhead; ii) Prevalence of outliers in KV values, challenging low-bitwidth uniform quantization. To this end, we propose MILLION, a novel quantization framework achieving low-bitwidth KV cache through product quantization. First, we conduct a thorough analysis of KV cache distribution, revealing the limitations of existing quantization schemes. Second, we introduce a non-uniform quantization algorithm based on product quantization, which efficiently compresses data while preserving accuracy. Third, we develop a high-performance GPU inference framework with efficient attention kernel and pipeline design for MILLION that leverages sparse computation and asynchronous quantization, significantly enhancing inference speed. Comprehensive evaluation results demonstrate that MILLION can achieve 4 bits quantization with trivial perplexity and accuracy loss, and achieve 2.09 x end-to-end performance gains at 32 K context length. Code is released at https://github.com/ZongwuWang/MILLION. Zongwu Wang, Peng Xu 0046, Fangxin Liu, Qingxiao Sun, Gezi Li, Li Jiang 0002, Haibing Guan |
DAC | 2 |
| 2025 | EVASION: Efficient KV CAche CompreSsion vIa PrOduct QuaNtizationabstractLarge language models (LLMs) are increasingly utilized for complex tasks requiring longer context lengths, with some models supporting up to 128K or 1M tokens. This trend, however, presents significant challenges in inference speed and memory management. The primary bottleneck in long-context LLM inference is the quadratic computational complexity of attention mechanisms, causing substantial slowdowns as sequence length increases. KV cache mechanism alleviates this issue by storing pre-computed data, but introduces memory requirements that scale linearly with context length, hindering efficient LLM deployment. Quantization emerges as a promising approach to address the widening gap between LLM size and memory capacity. However, traditional quantization schemes often yield suboptimal compression results for KV caches due to two key factors: i) On-the-fly quantization and de-quantization, causing significant performance overhead; ii) Prevalence of outliers in KV values, challenging low-bitwidth uniform quantization. To this end, we propose EVASION, a novel quantization framework achieving low-bitwidth KV cache through product quantization. First, we conduct a thorough analysis of KV cache distribution, revealing the limitations of existing quantization schemes. Second, we introduce a non-uniform quantization algorithm based on product quantization, which efficiently compresses data while preserving accuracy. Third, we develop a high-performance GPU inference framework for EVASION that leverages sparse computation and asynchronous quantization, significantly enhancing inference speed. Comprehensive evaluation results demonstrate that EVASION can achieve 4 bits quantization trivial perplexity and accuracy loss. Zongwu Wang, Fangxin Liu, Peng Xu 0046, Qingxiao Sun, Junping Zhao, Li Jiang 0002 |
DATE | 3 |