Haimeng Ren

dblp:348/5022 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0001-2287-7503ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 54% Memory systems · 22% GPUs and heterogeneous computing · 14%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.722025
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM · HPCA 2025
COMET: Towards Practical W4A4KV4 LLMs Serving · ASPLOS (2) 2025
GPUs and heterogeneous computing
GPU kernel optimization
0.912025
COMET: Towards Practical W4A4KV4 LLMs Serving · ASPLOS (2) 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator
KV cache compression
0.912025
COMET: Towards Practical W4A4KV4 LLMs Serving · ASPLOS (2) 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM inference accelerator
0.912025
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM · HPCA 2025
Cloud and datacenter computing › inference serving
LLM serving
0.912025
COMET: Towards Practical W4A4KV4 LLMs Serving · ASPLOS (2) 2025
Hardware accelerators and domain-specific architectures › quantization
mixed-precision quantization
0.912025
COMET: Towards Practical W4A4KV4 LLMs Serving · ASPLOS (2) 2025
Memory systems › processing-in-memory
near-data processing
0.912025
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM · HPCA 2025
Memory systems
processing-in-memory
0.912025
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM · HPCA 2025
GPUs and heterogeneous computing
GPU scheduling
0.312025
COMET: Towards Practical W4A4KV4 LLMs Serving · ASPLOS (2) 2025

Methods — techniques the papers use, named apart from their topics

software pipelining · 0.9online scheduling · 0.9neuron partition · 0.9fine-grained mixed-precision quantization · 0.9dequantization · 0.9activation sparsity · 0.9
YearPublicationVenuePosition
2025 COMET: Towards Practical W4A4KV4 LLMs Serving
abstract
Quantization is a widely-used compression technology to reduce the overhead of serving large language models (LLMs) on terminal devices and in cloud data centers. However, prevalent quantization methods, such as 8-bit weight-activation or 4-bit weight-only quantization, achieve limited performance improvements due to poor support for low-precision (e.g., 4-bit) activation. This work, for the first time, realizes practical W4A4KV4 serving for LLMs, fully utilizing the INT4 tensor cores on modern GPUs and reducing the memory bottleneck caused by the KV cache. Specifically, we propose a novel fine-grained mixed-precision quantization algorithm (FMPQ) that compresses most activations into 4-bit with negligible accuracy loss. To support mixed-precision matrix multiplication for W4A4 and W4A8, we develop a highly optimized W4Ax kernel. Our approach introduces a novel mixed-precision data layout to facilitate access and fast dequantization for activation and weight tensors, utilizing the GPU's software pipeline to hide the overhead of data loading and conversion. Additionally, we propose fine-grained streaming multiprocessor (SM) scheduling to achieve load balance across different SMs. We integrate the optimized W4Ax kernel into our inference framework, COMET, and provide efficient management to support popular LLMs such as LLaMA-3-70B. Extensive evaluations demonstrate that, when running LLaMA family models on a single A100-80G-SMX4, COMET achieves a kernel-level speedup of 2.88x over cuBLAS and a 2.02x throughput improvement compared to TensorRT-LLM from an end-to-end framework perspective.
Long Cheng 0003, Haimeng Ren, Zhaohui Xu, Yudong Pan, Mengdi Wang 0004, Xiaowei Li 0001, Yinhe Han 0001, Ying Wang 0001
ASPLOS (2)3
2025 Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
abstract
The billion-scale Large Language Models (LLMs) necessitate deployment on expensive server-grade GPUs with large-storage HBMs and abundant computation capability. As LLM-assisted services become popular, achieving cost-effective LLM inference on budget-friendly hardware becomes the current trend. This has sparked extensive research into relocating LLM parameters from expensive GPUs to external host memory. However, the restricted bandwidth between the host and GPU memory limits the inference performance of existing solutions. This work introduces Hermes, a budget-friendly system that leverages the near-data processing units (NDP) within commodity DRAM DIMMs to enhance the performance of a single consumer-grade GPU, achieving efficient LLM inference. We recognize that the inherent activation sparsity in LLMs naturally divides weight parameters into two categories, termed “hot” and “cold” neurons, respectively. Hot neurons, which consist of only approximately 20% of all weight parameters, account for 80% of the total computational load. In contrast, cold neurons make up the other 80% of parameters but are responsible for just 20% of the computational workload. Leveraging this observation, we propose a heterogeneous computing strategy: mapping hot neurons to a single computation-efficient GPU without large-capacity HBMs, while offloading cold neurons to NDP-DIMMs, which offer large memory size but limited computation capabilities. In addition, the dynamic nature of activation sparsity necessitates a real-time partition of hot and cold neurons and adaptive remapping of cold neurons across multiple NDP-DIMM modules. To tackle these issues, we introduce a lightweight predictor that ensures optimal real-time neuron partition and adjustment between GPU and NDP-DIMMs. Furthermore, we utilize a window-based online scheduling mechanism to maintain load balance among multiple NDP-DIMM modules. In summary, Hermes facilitates the deployment of LLaMA2-70B on consumer-grade hardware at a rate of 13.75 tokens/s and realizes an average 75.24 × speedup over the state-of-the-art offloading-based inference system on popular LLMs.
Bing Li 0017, Haimeng Ren, Zhaohui Xu, Mengdi Wang 0004, Xiaowei Li 0001, Yinhe Han 0001, Ying Wang 0001
HPCA4