VLDB 2026 Research / reviewers in the wild / expert
Zhuoming Chen
dblp:226/5729
· DBLP profile ↗
17ranked-venue papers
2as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative DecodingabstractModern large language model (LLM) applications exhibit diverse service-level objectives (SLOs), from low-latency requirements in interactive coding assistants to more relaxed constraints in data wrangling tasks. Existing LLM serving systems, which rely on uniform batching and scheduling strategies, often fail to meet these heterogeneous SLOs concurrently. We present AdaServe, the first LLM serving system designed to support efficient multi-SLO serving through SLO-customized speculative decoding. AdaServe formulates multi-SLO serving as a constrained optimization problem and introduces a hardware-aware algorithm that constructs a speculation tree tailored to each request's latency target. It features a speculate-select-verify pipeline that enables fine-grained control over decoding speed while maximizing system throughput. AdaServe further adapts to workload variation by dynamically adjusting speculation parameters. Evaluations across diverse workloads show that AdaServe reduces SLO violations by up to 4.3X and improves goodput by up to 1.9X compared to the best-performing baselines, highlighting its effectiveness in multi-SLO serving. Zikun Li, Zhuofu Chen, Remi Delacourt, Gabriele Oliaro, Qinghan Chen, Shuhuai Lin, April Yang, Zhihao Zhang 0001, Zhuoming Chen, Yi-Hsiang Lai, Xinhao Cheng, Xupeng Miao |
EuroSys | 10 |
| 2025 | Mithril: A Scalable System for Deep GNN TrainingabstractCommunication is a key bottleneck for distributed graph neural network (GNN) training. Existing GNN training systems fail to scale to deep GNNs because of the tremendous amount of inter-GPU communication. This paper proposes Mithril, a new approach that significantly scales the distributed full-graph deep GNN training. Being the first to use layer-level model parallelism for GNN training, Mithril partitions GNN layers among GPUs, each device performs the computation for a disjoint subset of consecutive GNN layers on the whole graph. Compared to graph parallelism with each GPU handling a graph partition, Mithril reduces the communication volume by a factor of the number of GNN layers to scale to deep models. Mithril overcomes the unique challenges for pipelined layer-level model parallelism on the whole graph by partitioning it into dependent chunks, breaking the dependencies with embedding speculation, and applying specific training techniques to ensure convergence. We also propose a hybrid approach by combining Mithril with graph parallelism to handle large graphs, achieve better computer resource utilization and ensure model convergence. We build a general GNN training system supporting all three parallelism settings. Extensive experiments show that Mithril reduces the perepoch communication volume by up to $22.89 \times$ (on average $6.78 \times$). It achieves a maximum training time speedup of $2.34 \times$ (on average $1.49 \times$) on a GPU cluster with a high-performance InfiniBand network. On another cluster with a commodity Ethernet, Mithril outperforms the baseline by up to $10.21 \times$ (on average $7.16 \times$). Mithril also achieves a comparable level of model accuracy and convergence speed compared to graph parallelism. Jingji Chen, Zhuoming Chen, Xuehai Qian |
HPCA | 2 |
| 2025 | MagicPIG: LSH Sampling for Efficient LLM GenerationabstractLarge language models (LLMs) with long context windows have gained significant attention. However, the KV cache, stored to avoid re-computation, becomes a bottleneck. Various dynamic sparse or TopK-based attention approximation methods have been proposed to leverage the common insight that attention is sparse. In this paper, we first show that TopK attention itself suffers from quality degradation in certain downstream tasks because attention is not always as sparse as expected. Rather than selecting the keys and values with the highest attention scores, sampling with theoretical guarantees can provide a better estimation for attention output. To make the sampling-based approximation practical in LLM generation, we propose MagicPIG, a heterogeneous system based on Locality Sensitive Hashing (LSH). MagicPIG significantly reduces the workload of attention computation while preserving high accuracy for diverse tasks. MagicPIG stores the LSH hash tables and runs the attention computation on the CPU, which allows it to serve longer contexts and larger batch sizes with high approximation accuracy. MagicPIG can improve decoding throughput by up to $5\times$ across various GPU hardware and achieve 54ms decoding latency on a single RTX 4090 for Llama-3.1-8B-Instruct model with a context of 96k tokens. Zhuoming Chen, Ranajoy Sadhukhan, Zihao Ye 0001, Niklas Nolte, Yuandong Tian, Matthijs Douze, Léon Bottou, Beidi Chen |
ICLR | 1 |
| 2025 | MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative DecodingabstractLarge Language Models (LLMs) have become more prevalent in long-context applications such as interactive chatbots, document analysis, and agent workflows, but it is challenging to serve long-context requests with low latency and high throughput. Speculative decoding (SD) is a widely used technique to reduce latency losslessly, but the conventional wisdom suggests that its efficacy is limited to small batch sizes. In MagicDec, we show that surprisingly SD can achieve speedup even for a high throughput inference regime for moderate to long sequences. More interestingly, an intelligent drafting strategy can achieve better speedup with increasing batch size based on our rigorous analysis. MagicDec first identifies the bottleneck shifts with increasing batch size and sequence length, and uses these insights to deploy SD more effectively for high throughput inference. We leverage draft model with sparse KV cache to address the KV bottleneck, which scales with both sequence length and batch size. Additionally, we propose a theoretical model to select the optimal drafting strategy for maximum speedup. Our work highlights the broad applicability of speculative decoding in long-context serving, as it can enhance throughput and reduce latency without compromising accuracy. For moderate to long sequences, we demonstrate up to 2.51x speedup for LLaMA-3.1-8B when serving batch sizes ranging from 32 to 256 on various types of hardware and tasks. Ranajoy Sadhukhan, Zhuoming Chen, Vashisth Tiwari, Ruihang Lai, Jinyuan Shi, Ian En-Hsu Yen, Avner May, Tianqi Chen 0001, Beidi Chen |
ICLR | 3 |
| 2025 | GSM-∞: How Do your LLMs Behave over Infinitely Increasing Reasoning Complexity and Context Length?abstractRecently, long-context large language models (LLMs) have shown strong performance in information retrieval and long-document QA. However, to tackle the most challenging intellectual problems, LLMs must reason effectively in long and complex contexts (e.g., frontier mathematical research). Studying how LLMs handle increasing reasoning complexity and context length is essential, yet existing benchmarks lack a solid basis for quantitative evaluation. Inspired by the abstraction of GSM-8K problems as computational graphs—and the ability to introduce noise by adding unnecessary nodes and edges—we develop a grade-school math problem generator capable of producing arithmetic problems with infinite difficulty and context length under fine-grained control. Using our newly synthesized GSM-$\infty$ benchmark, we comprehensively evaluate existing LLMs. We find a consistent sigmoid decline in reasoning performance as complexity increases, along with a systematic inference scaling trend: exponentially increasing inference computation yields only linear performance gains. These findings underscore the fundamental limitations of current long-context LLMs and the key challenges in scaling reasoning capabilities. Our GSM-$\infty$ benchmark provides a scalable and controllable testbed for systematically studying and advancing LLM reasoning in long and complex contexts. Zhuoming Chen, Yuandong Tian, Beidi Chen |
ICML | 3 |
| 2025 | Kinetics: Rethinking Test-Time Scaling LawabstractWe rethink test-time scaling laws from a practical efficiency perspective, revealing that the effectiveness of smaller models is significantly overestimated. Prior work, grounded in compute-optimality, overlooks critical memory access bottlenecks introduced by inference-time strategies (e.g., Best-of-N, long CoTs). Our holistic analysis, spanning models from 0.6B to 32B parameters, reveals a new Kinetics Scaling Law that better guides resource allocation by incorporating both computation and memory access costs. The Kinetics Scaling Law suggests that test-time compute is more effective when used on models above a threshold (14B) than on smaller ones. A key reason is that in test-time scaling, attention—rather than parameter count—emerges as the dominant cost factor. Motivated by this, we propose a new scaling paradigm centered on sparse attention, which lowers per-token cost and enables longer generations and more parallel samples within the same resource budget. Empirically, we show that sparse attention models consistently outperform dense counterparts, achieving over 60-point gains in low-cost regimes and over 5-point gains in high-cost regimes for problem-solving accuracy on AIME and LiveCodeBench. These results suggest that sparse attention is essential for realizing the full potential of test-time scaling because, unlike training where parameter scaling saturates, test-time accuracy continues to improve through increased generation. Ranajoy Sadhukhan, Zhuoming Chen, Haizhong Zheng, Beidi Chen |
NeurIPS | 2 |
| 2024 | SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and VerificationabstractThis paper introduces SpecInfer, a system that accelerates generative large language model (LLM) serving with tree-based speculative inference and verification. The key idea behind SpecInfer is leveraging small speculative models to predict the LLM's outputs; the predictions are organized as a token tree, whose nodes each represent a candidate token sequence. The correctness of all candidate token sequences represented by a token tree is verified against the LLM in parallel using a novel tree-based parallel decoding mechanism. SpecInfer uses an LLM as a token tree verifier instead of an incremental decoder, which significantly reduces the end-to-end latency and computational requirement for serving generative LLMs while provably preserving model quality. Our evaluation shows that SpecInfer outperforms existing LLM serving systems by 1.5-2.8× for distributed LLM inference and by 2.6-3.5× for offloading-based LLM inference, while preserving the same generative performance. SpecInfer is publicly available at https://github.com/flexflow/FlexFlow/ Xupeng Miao, Gabriele Oliaro, Zhihao Zhang 0001, Xinhao Cheng, Rae Ying Yee Wong, Alan Zhu 0001, Lijie Yang 0003, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar |
ASPLOS (3) | 12 |
| 2024 | Acoustic changes in speech prosody produced by children with autism after robot-assisted speech trainingabstractInterspeech 2024, 1-5 September 2024, Kos, Greece Bruce Xiao Wang, Yitian Hong, Angel Chan, Po-yi Tang, Bin Li 0003, Chunyi Wen, James Cheung, Zhuoming Chen |
INTERSPEECH | 11 |
| 2024 | Sequoia: Scalable and Robust Speculative DecodingabstractAs the usage of large language models (LLMs) grows, it becomes increasingly important to serve them quickly and efficiently. While speculative decoding has recently emerged as a promising direction for accelerating LLM serving, existing methods are limited in their ability to scale to larger speculation budgets and adapt to different hyperparameters. This paper introduces Sequoia, a scalable and robust algorithm for speculative decoding. To improve scalability, Sequoia introduces a dynamic programming algorithm to find an optimal tree structure for the speculated tokens. To achieve robust speculative decoding, Sequoia uses a novel sampling and verification method that outperforms prior work across different decoding temperatures. Sequoia improves the decoding speed of Llama2-7B, Llama2-13B, and Vicuna-33B on an A100 GPU by up to $4.04\times$, $3.73\times$, and $2.27 \times$. To serve Llama3-70B-Instruct on a single L40 GPU through offloading, Sequoia reduces the per-token decoding latency to 0.60 s/token, $9.5\times$ faster than DeepSpeed-Zero-Inference. Zhuoming Chen, Avner May, Ruslan Svirschevski, Yuhsun Huang, Max Ryabinin, Beidi Chen |
NeurIPS | 1 |
| 2024 | Mini-Sequence Transformers: Optimizing Intermediate Memory for Long Sequences TrainingabstractWe introduce Mini-Sequence Transformer (MsT), a simple and effective methodology for highly efficient and accurate LLM training with extremely long sequences. MsT partitions input sequences and iteratively processes mini-sequences to reduce intermediate memory usage. Integrated with activation recomputation, it enables significant memory savings in both forward and backward passes. In experiments with the Llama3-8B model, with MsT, we measure no degradation in throughput or convergence even with 12x longer sequences than standard implementations. MsT is fully general, implementation-agnostic, and requires minimal code changes to integrate with existing LLM training frameworks. Integrated with the huggingface library, MsT successfully extends the maximum context length of Qwen, Mistral, and Gemma-2 by 12-24x. Zhuoming Chen, Beidi Chen, Anima Anandkumar |
NeurIPS | 3 |
| 2024 | SpecExec: Massively Parallel Speculative Decoding For Interactive LLM Inference on Consumer DevicesabstractAs large language models gain widespread adoption, running them efficiently becomes a crucial task. Recent works on LLM inference use speculative decoding to achieve extreme speedups. However, most of these works implicitly design their algorithms for high-end datacenter hardware. In this work, we ask the opposite question: how fast can we run LLMs on consumer machines? Consumer GPUs can no longer fit the largest available models and must offload them to RAM or SSD. With parameter offloading, hundreds or thousands of tokens can be processed in batches within the same time as just one token, making it a natural fit for speculative decoding. We propose SpecExec (Speculative Execution), a simple parallel decoding method that can generate up to 20 tokens per target model iteration for popular LLM families. SpecExec takes the most probable continuations from the draft model to build a "cache" tree for the target model, which then gets validated in a single pass. Using SpecExec, we demonstrate inference of 50B+ parameter LLMs on consumer GPUs with RAM offloading at 4--6 tokens per second with 4-bit quantization or 2--3 tokens per second with 16-bit weights. Our code is available at https://github.com/yandex-research/specexec . Ruslan Svirschevski, Avner May, Zhuoming Chen, Beidi Chen, Max Ryabinin |
NeurIPS | 3 |
| 2024 | SIRIUS : Contexual Sparisty with Correction for Efficient LLMsabstractWith the blossom of large language models (LLM), inference efficiency becomes increasingly important. Various approximate methods are proposed to reduce the cost at inference time. Contextual Sparsity (CS) is appealing for its training-free nature and its ability to reach a higher compression ratio seemingly without significant performance degradation. However, after a comprehensive evaluation of contextual sparsity methods on various complex generation tasks, we find that although CS succeeds in prompt-understanding tasks, it significantly degrades the model performance for reasoning, deduction, and knowledge-based tasks. Despite the gap in end-to-end accuracy, we observed that sparse models and original models often share the general problem-solving logic and require only a few token corrections to recover the original model performance. This paper introduces SIRIUS, an efficient correction mechanism, which significantly boosts CS models on reasoning tasks while maintaining its efficiency gain. SIRIUS is evaluated on 6 models with 8 difficult generation tasks in reasoning, deduction, and coding and shows consistent effectiveness and efficiency. Also, we carefully develop a system implementation for SIRIUS and show that SIRIUS delivers theoretical latency reduction with roughly a 20% reduction in latency for 8B model on-chip and a 35% reduction in latency for 70B model offloading. We open-source our implementation of Sirius at https://github.com/Infini-AI-Lab/Sirius.git. Zhuoming Chen, Zhaozhuo Xu, Xi Victoria Lin, Beidi Chen |
NeurIPS | 2 |
| 2022 | A Dataset for Falling Risk Assessment of the Elderly using Wearable Plantar PressureabstractFalling is characterized by high incidence and great harm among the elderly. Timely assessing falling risk in daily life is helpful for reducing the occurrence of severe health outcomes. Establishing dataset for falling risk assessment based on wearable devices in the elderly is important work. However, current existing datasets might not reflect the natural gait of the subject due to the discomfort in wearing. Relevant data processing methods based on these datasets have limited practicability and might not be applied to real scenes in daily life. To make daily falling risk assessment possible, we proposed a novel approach to set up a continuous and wearable plantar pressure dataset of 48 older adults along with falling risk labels. The dataset was collected by plantar pressure monitoring shoes which were suitable for daily living spaces. Moreover, the Conv-LSTM algorithm was applied on the dataset, and the average classification result was up to 95.57%, reflecting the effectiveness of this dataset. The dataset is helpful for the studies of falling risk assessment and health monitoring among the elderly. Guohua Hu, Jianxiu Jin, Shibin Wu, Junan Xie, Jianlin Ou, Zhuoming Chen, Xiangmin Xu 0001 |
BIBM | 8 |
| 2022 | Quantized Training of Gradient Boosting Decision TreesabstractRecent years have witnessed significant success in Gradient Boosting Decision Trees (GBDT) for a wide range of machine learning applications. Generally, a consensus about GBDT's training algorithms is gradients and statistics are computed based on high-precision floating points. In this paper, we investigate an essentially important question which has been largely ignored by the previous literature - how many bits are needed for representing gradients in training GBDT? To solve this mystery, we propose to quantize all the high-precision gradients in a very simple yet effective way in the GBDT's training algorithm. Surprisingly, both our theoretical analysis and empirical studies show that the necessary precisions of gradients without hurting any performance can be quite low, e.g., 2 or 3 bits. With low-precision gradients, most arithmetic operations in GBDT training can be replaced by integer operations of 8, 16, or 32 bits. Promisingly, these findings may pave the way for much more efficient training of GBDT from several aspects: (1) speeding up the computation of gradient statistics in histograms; (2) compressing the communication cost of high-precision statistical information during distributed training; (3) the inspiration of utilization and development of hardware architectures which well support low-precision computation for GBDT training. Benchmarked on CPUs, GPUs, and distributed clusters, we observe up to 2$\times$ speedup of our simple quantization strategy compared with SOTA GBDT systems on extensive datasets, demonstrating the effectiveness and potential of the low-precision training of GBDT. The code will be released to the official repository of LightGBM. Guolin Ke, Zhuoming Chen, Shuxin Zheng, Tie-Yan Liu |
NeurIPS | 3 |
| 2021 | LFMB-3DFB: A Large-scale Finger Multi-Biometric Database and Benchmark for 3D Finger BiometricsabstractFinger contains several discriminative biometric traits, including fingerprint, finger vein, finger knuckle, and finger shape, which are complementary in identity information. However, in most current researches and practical applications, only a single or several traits are utilized, which are prone to unsatisfactory recognition performance and easy forgery. Our work is the first attempt to collect and study all biometric traits on the finger. Firstly, a novel multi-view, multi-spectral 3D finger imaging system is designed. To the best of our knowledge, it is the first biometric imaging system that can capture almost all finger-based traits. With this 3D finger imaging system, we scanned numerous fingers, acquiring their external skin images and internal vein images from 6 different views. Then 3D finger models with skin and vein textures are reconstructed by space carving, mesh regularization, and texture mapping algorithms. Secondly, we establish a benchmark dataset, namely the Large- scale Finger Multi-Biometric database and benchmark for 3D Finger Biometrics (LFMB-3DFB). LFMB-3DFB contains 695 fingers, and each finger is captured 10 times. Then, 6 finger skin images and 6 finger vein images are obtained for each acquisition, and final 83,400 images and 6,950 3D finger models are obtained. Besides, we designed a more rigorous and comprehensive evaluation protocol for both identification and verification tasks. Finally, we designed corresponding baselines for 2D finger traits recognition, multi-view finger traits recognition, 3D finger traits recognition, and score-level fusion. Rigorous experiments have been conducted to verify the significance and usefulness of the proposed LFMB-3DFB. Weili Yang, Zhuoming Chen, Junduan Huang, Wenxiong Kang |
IJCB | 2 |
| 2021 | Mutual-Collision-Avoidance Scheme Synthesized by Neural Networks for Dual Redundant Robot Manipulators Executing Cooperative TasksabstractCollision between dual robot manipulators during working process will lead to task failure and even robot damage. To avoid mutual collision of dual robot manipulators while doing collaboration tasks, a novel recurrent neural network (RNN)-based mutual-collision-avoidance (MCA) scheme for solving the motion planning problem of dual manipulators is proposed and exploited. Because of the high accuracy and low computation complexity, the linear variational inequality-based primal-dual neural network is used to solve the proposed scheme. The proposed scheme is applied to the collaboration trajectory tracking and cup-stacking tasks, and shows its effectiveness for avoiding collision between the dual robot manipulators. Through network iteration and online learning, the dual robot manipulators will learn the ability of MCA. Moreover, a line-segment-based distance measure algorithm is proposed to calculate the minimum distance between the dual manipulators. If the computed minimum distance is less than the first safe-related distance threshold, a speed brake operation is executed and guarantees that the robot cannot exceed the second safe-related distance threshold. Furthermore, the proposed MCA strategy is formulated as a standard quadratic programming problem, which is further solved by an RNN. Computer simulations and a real dual robot experiment further verify the effectiveness, accuracy, and physical realizability of the RNN-based MCA scheme when manipulators cooperatively execute the end-effector tasks. Zhijun Zhang 0003, Lunan Zheng, Zhuoming Chen, Lingdong Kong, Hamid Reza Karimi |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | A Captcha Design Based on Visual ReasoningabstractCAPTCHA is a reverse Turing test to distinguish humans from machines. It is widely used in the internet industry for cyber security. A good CAPTCHA is supposed to be easy for humans but difficult for machines. Many existing CAPTCHA implementations leverage the inability of automatic visual recognition, e.g., recognizing the text or other objects in an image. These CAPTCHAs are becoming more and more vulnerable recently, due to the rapid development of visual recognition techniques. This paper presents our study of using visual reasoning in CAPTCHA design. This CAPTCHA asks the users to find specific object(s) in an image according to a given text query. It is generally easy for humans to understand the text query and make sophisticated reasoning about the image, but still remains difficult and computationally expensive for machines. We describe the CAPTCHA design, provide usability analysis and present security experiments. Moreover, we show that the security can be further improved by the use of neural style transfer. Zhuoming Chen, Renjia Wei |
ICASSP | 3 |