VLDB 2026 Research / reviewers in the wild / expert
Haimeng Ren
dblp:348/5022
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0001-2287-7503ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 54% Memory systems · 22% GPUs and heterogeneous computing · 14% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
1.7 | 2 | 2025 | Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM · HPCA 2025 COMET: Towards Practical W4A4KV4 LLMs Serving · ASPLOS (2) 2025 |
GPUs and heterogeneous computing
GPU kernel optimization |
0.9 | 1 | 2025 | COMET: Towards Practical W4A4KV4 LLMs Serving · ASPLOS (2) 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
KV cache compression |
0.9 | 1 | 2025 | COMET: Towards Practical W4A4KV4 LLMs Serving · ASPLOS (2) 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM inference accelerator |
0.9 | 1 | 2025 | Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM · HPCA 2025 |
Cloud and datacenter computing › inference serving
LLM serving |
0.9 | 1 | 2025 | COMET: Towards Practical W4A4KV4 LLMs Serving · ASPLOS (2) 2025 |
Hardware accelerators and domain-specific architectures › quantization
mixed-precision quantization |
0.9 | 1 | 2025 | COMET: Towards Practical W4A4KV4 LLMs Serving · ASPLOS (2) 2025 |
Memory systems › processing-in-memory
near-data processing |
0.9 | 1 | 2025 | Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM · HPCA 2025 |
Memory systems
processing-in-memory |
0.9 | 1 | 2025 | Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM · HPCA 2025 |
GPUs and heterogeneous computing
GPU scheduling |
0.3 | 1 | 2025 | COMET: Towards Practical W4A4KV4 LLMs Serving · ASPLOS (2) 2025 |
Methods — techniques the papers use, named apart from their topics
software pipelining · 0.9online scheduling · 0.9neuron partition · 0.9fine-grained mixed-precision quantization · 0.9dequantization · 0.9activation sparsity · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | COMET: Towards Practical W4A4KV4 LLMs ServingabstractQuantization is a widely-used compression technology to reduce the overhead of serving large language models (LLMs) on terminal devices and in cloud data centers. However, prevalent quantization methods, such as 8-bit weight-activation or 4-bit weight-only quantization, achieve limited performance improvements due to poor support for low-precision (e.g., 4-bit) activation. This work, for the first time, realizes practical W4A4KV4 serving for LLMs, fully utilizing the INT4 tensor cores on modern GPUs and reducing the memory bottleneck caused by the KV cache. Specifically, we propose a novel fine-grained mixed-precision quantization algorithm (FMPQ) that compresses most activations into 4-bit with negligible accuracy loss. To support mixed-precision matrix multiplication for W4A4 and W4A8, we develop a highly optimized W4Ax kernel. Our approach introduces a novel mixed-precision data layout to facilitate access and fast dequantization for activation and weight tensors, utilizing the GPU's software pipeline to hide the overhead of data loading and conversion. Additionally, we propose fine-grained streaming multiprocessor (SM) scheduling to achieve load balance across different SMs. We integrate the optimized W4Ax kernel into our inference framework, COMET, and provide efficient management to support popular LLMs such as LLaMA-3-70B. Extensive evaluations demonstrate that, when running LLaMA family models on a single A100-80G-SMX4, COMET achieves a kernel-level speedup of 2.88x over cuBLAS and a 2.02x throughput improvement compared to TensorRT-LLM from an end-to-end framework perspective. Long Cheng 0003, Haimeng Ren, Zhaohui Xu, Yudong Pan, Mengdi Wang 0004, Xiaowei Li 0001, Yinhe Han 0001, Ying Wang 0001 |
ASPLOS (2) | 3 |
| 2025 | Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMMabstractThe billion-scale Large Language Models (LLMs) necessitate deployment on expensive server-grade GPUs with large-storage HBMs and abundant computation capability. As LLM-assisted services become popular, achieving cost-effective LLM inference on budget-friendly hardware becomes the current trend. This has sparked extensive research into relocating LLM parameters from expensive GPUs to external host memory. However, the restricted bandwidth between the host and GPU memory limits the inference performance of existing solutions. This work introduces Hermes, a budget-friendly system that leverages the near-data processing units (NDP) within commodity DRAM DIMMs to enhance the performance of a single consumer-grade GPU, achieving efficient LLM inference. We recognize that the inherent activation sparsity in LLMs naturally divides weight parameters into two categories, termed “hot” and “cold” neurons, respectively. Hot neurons, which consist of only approximately 20% of all weight parameters, account for 80% of the total computational load. In contrast, cold neurons make up the other 80% of parameters but are responsible for just 20% of the computational workload. Leveraging this observation, we propose a heterogeneous computing strategy: mapping hot neurons to a single computation-efficient GPU without large-capacity HBMs, while offloading cold neurons to NDP-DIMMs, which offer large memory size but limited computation capabilities. In addition, the dynamic nature of activation sparsity necessitates a real-time partition of hot and cold neurons and adaptive remapping of cold neurons across multiple NDP-DIMM modules. To tackle these issues, we introduce a lightweight predictor that ensures optimal real-time neuron partition and adjustment between GPU and NDP-DIMMs. Furthermore, we utilize a window-based online scheduling mechanism to maintain load balance among multiple NDP-DIMM modules. In summary, Hermes facilitates the deployment of LLaMA2-70B on consumer-grade hardware at a rate of 13.75 tokens/s and realizes an average 75.24 × speedup over the state-of-the-art offloading-based inference system on popular LLMs. Bing Li 0017, Haimeng Ren, Zhaohui Xu, Mengdi Wang 0004, Xiaowei Li 0001, Yinhe Han 0001, Ying Wang 0001 |
HPCA | 4 |