Zheng Wang 0075

dblp:181/2834-75 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-8575-9432ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Yggdrasil: Bridging Dynamic Speculation and Static Runtime for Latency-Optimal Tree-Based LLM Decoding
abstract
Speculative decoding improves LLM inference by generating and verifying multiple tokens in parallel, but existing systems suffer from suboptimal performance due to a mismatch between dynamic speculation and static runtime assumptions. We present Yggdrasil, a co-designed system that enables latency-optimal speculative decoding through context-aware tree drafting and compiler-friendly execution. Yggdrasil introduces an equal-growth tree structure for static graph compatibility, a latency-aware optimization objective for draft selection, and stage-based scheduling to reduce overhead. Yggdrasil supports unmodified LLMs and achieves up to $3.98\times$ speedup over state-of-the-art baselines across multiple hardware setups.
Yue Guan 0003, Changming Yu, Shihan Fang, Weiming Hu 0005, Zaifeng Pan, Zheng Wang 0075, Zihan Liu 0002, Yangjie Zhou 0001, Yufei Ding 0001, Minyi Guo, Jingwen Leng
NeurIPS6
2025 WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
Zheng Wang 0075, Anna Cai, Xinfeng Xie, Zaifeng Pan, Yue Guan 0003, Weiwei Chu, Jie Wang 0022, Shikai Li, Chris Cai, Yuchen Hao, Yufei Ding 0001
OSDI1
2025 GMI-DRL: Empowering Multi-GPU DRL with Adaptive-Grained Parallelism
Boyuan Feng, Zheng Wang 0075, Guyue Huang, Tong Geng, Ang Li 0006, Yufei Ding 0001
USENIX ATC3
2024 ZENO: A Type-based Optimization Framework for Zero Knowledge Neural Network Inference
abstract
Zero knowledge Neural Networks draw increasing attention for guaranteeing computation integrity and privacy of neural networks (NNs) based on zero-knowledge Succinct Non-interactive ARgument of Knowledge (zkSNARK) security scheme. However, the performance of zkSNARK NNs is far from optimal due to the million-scale circuit computation with heavy scalar-level dependency. In this paper, we propose a type-based optimizing framework for efficient zero-knowledge NN inference, namely ZENO (ZEro knowledge Neural network Optimizer). We first introduce ZENO language construct to maintain high-level semantics and the type information (e.g., privacy and tensor) for allowing more aggressive optimizations. We then propose privacy-type driven and tensor-type driven optimizations to further optimize the generated zkSNARK circuit. Finally, we design a set of NN-centric system optimizations to further accelerate zkSNARK NNs. Experimental results show that ZENO achieves up to 8.5× end-to-end speedup than state-of-the-art zkSNARK NNs. We reduce proof time for VGG16 from 6 minutes to 48 seconds, which makes zkSNARK NNs practical.
Boyuan Feng, Zheng Wang 0075, Yufei Ding 0001
ASPLOS (1)2
2024 RAP: Resource-aware Automated GPU Sharing for Multi-GPU Recommendation Model Training and Input Preprocessing
abstract
Ensuring high-quality recommendations for newly onboarded users requires the continuous retraining of Deep Learning Recommendation Models (DLRMs) with freshly generated data. To serve the online DLRM retraining, existing solutions use hundreds of CPU computing nodes designated for input preprocessing, causing significant power consumption that surpasses even the power usage of GPU trainers.
Zheng Wang 0075, Jiaqi Deng 0002, Da Zheng 0004, Ang Li 0006, Yufei Ding 0001
ASPLOS (2)1
2024 OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation Model
Zheng Wang 0075, Boyuan Feng, Guyue Huang, Dheevatsa Mudigere, Bharath Muthiah, Ang Li 0006, Yufei Ding 0001
USENIX ATC1
2023 ECSSD: Hardware/Data Layout Co-Designed In-Storage-Computing Architecture for Extreme Classification
abstract
With the rapid growth of classification scale in deep learning systems, the final classification layer becomes extreme classification with a memory footprint exceeding the main memory capacity of the CPU or GPU. The emerging in-storage-computing technique offers an opportunity on account of the fact that SSD has enough storage capacity for the parameters of extreme classification. However, the limited performance of naive in-storage-computing schemes is insufficient to support the heavy workload of extreme classification.
Siqi Li 0013, Fengbin Tu, Liu Liu 0017, Jilan Lin, Zheng Wang 0075, Yangwook Kang, Yufei Ding 0001, Yuan Xie 0001
ISCA5
2023 MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU Platforms
Boyuan Feng, Zheng Wang 0075, Tong Geng, Kevin J. Barker, Ang Li 0006, Yufei Ding 0001
OSDI3
2023 TC-GNN: Bridging Sparse GNN Computation and Dense Tensor Cores on GPUs
Boyuan Feng, Zheng Wang 0075, Guyue Huang, Yufei Ding 0001
USENIX ATC3
2022 EL-Rec: Efficient Large-Scale Recommendation Model Training via Tensor-Train Embedding Table
abstract
Deep learning Recommendation Models (DLRMs) plays an important role in various application domains. However, existing DLRM training systems require a large number of GPUs due to the memory-intensive embedding tables. To this end, we propose EL-Rec, an efficient computing framework harnessing the Tensor-train (TT) technique to democratize the training of large-scale DLRMs with limited GPU resources. Specifically, EL-Rec optimizes TT decomposition based on key computation primitives of embedding tables and implements a high-performance compressed embedding table which is a drop-in replacement of Pytorch API. EL-Rec introduces an index reordering technique to harvest the performance gains from both local and global information of training inputs. EL-Rec also highlights a pipeline training paradigm to eliminate the communication overhead between the host memory and the training worker. Comprehensive experiments demonstrate that EL-Rec can handle the largest publicly available DLRM dataset with a single GPU and achieves 3× speedup over the state-of-the-art DLRM frameworks.
Zheng Wang 0075, Boyuan Feng, Dheevatsa Mudigere, Bharath Muthiah, Yufei Ding 0001
SC1
2022 Faith: An Efficient Framework for Transformer Verification on GPUs
Boyuan Feng, Tianqi Tang 0001, Zhaodong Chen 0001, Zheng Wang 0075, Yuan Xie 0001, Yufei Ding 0001
USENIX ATC5