EDBT 2026 Demo / reviewers in the wild / expert
Zheng Wang 0075
dblp:181/2834-75
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-8575-9432ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Yggdrasil: Bridging Dynamic Speculation and Static Runtime for Latency-Optimal Tree-Based LLM DecodingabstractSpeculative decoding improves LLM inference by generating and verifying multiple tokens in parallel, but existing systems suffer from suboptimal performance due to a mismatch between dynamic speculation and static runtime assumptions. We present Yggdrasil, a co-designed system that enables latency-optimal speculative decoding through context-aware tree drafting and compiler-friendly execution. Yggdrasil introduces an equal-growth tree structure for static graph compatibility, a latency-aware optimization objective for draft selection, and stage-based scheduling to reduce overhead. Yggdrasil supports unmodified LLMs and achieves up to $3.98\times$ speedup over state-of-the-art baselines across multiple hardware setups. Yue Guan 0003, Changming Yu, Shihan Fang, Weiming Hu 0005, Zaifeng Pan, Zheng Wang 0075, Zihan Liu 0002, Yangjie Zhou 0001, Yufei Ding 0001, Minyi Guo, Jingwen Leng |
NeurIPS | 6 |
| 2025 | WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
Zheng Wang 0075, Anna Cai, Xinfeng Xie, Zaifeng Pan, Yue Guan 0003, Weiwei Chu, Jie Wang 0022, Shikai Li, Chris Cai, Yuchen Hao, Yufei Ding 0001 |
OSDI | 1 |
| 2025 | GMI-DRL: Empowering Multi-GPU DRL with Adaptive-Grained Parallelism
Boyuan Feng, Zheng Wang 0075, Guyue Huang, Tong Geng, Ang Li 0006, Yufei Ding 0001 |
USENIX ATC | 3 |
| 2024 | ZENO: A Type-based Optimization Framework for Zero Knowledge Neural Network InferenceabstractZero knowledge Neural Networks draw increasing attention for guaranteeing computation integrity and privacy of neural networks (NNs) based on zero-knowledge Succinct Non-interactive ARgument of Knowledge (zkSNARK) security scheme. However, the performance of zkSNARK NNs is far from optimal due to the million-scale circuit computation with heavy scalar-level dependency. In this paper, we propose a type-based optimizing framework for efficient zero-knowledge NN inference, namely ZENO (ZEro knowledge Neural network Optimizer). We first introduce ZENO language construct to maintain high-level semantics and the type information (e.g., privacy and tensor) for allowing more aggressive optimizations. We then propose privacy-type driven and tensor-type driven optimizations to further optimize the generated zkSNARK circuit. Finally, we design a set of NN-centric system optimizations to further accelerate zkSNARK NNs. Experimental results show that ZENO achieves up to 8.5× end-to-end speedup than state-of-the-art zkSNARK NNs. We reduce proof time for VGG16 from 6 minutes to 48 seconds, which makes zkSNARK NNs practical. Boyuan Feng, Zheng Wang 0075, Yufei Ding 0001 |
ASPLOS (1) | 2 |
| 2024 | RAP: Resource-aware Automated GPU Sharing for Multi-GPU Recommendation Model Training and Input PreprocessingabstractEnsuring high-quality recommendations for newly onboarded users requires the continuous retraining of Deep Learning Recommendation Models (DLRMs) with freshly generated data. To serve the online DLRM retraining, existing solutions use hundreds of CPU computing nodes designated for input preprocessing, causing significant power consumption that surpasses even the power usage of GPU trainers. Zheng Wang 0075, Jiaqi Deng 0002, Da Zheng 0004, Ang Li 0006, Yufei Ding 0001 |
ASPLOS (2) | 1 |
| 2024 | OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation Model
Zheng Wang 0075, Boyuan Feng, Guyue Huang, Dheevatsa Mudigere, Bharath Muthiah, Ang Li 0006, Yufei Ding 0001 |
USENIX ATC | 1 |
| 2023 | ECSSD: Hardware/Data Layout Co-Designed In-Storage-Computing Architecture for Extreme ClassificationabstractWith the rapid growth of classification scale in deep learning systems, the final classification layer becomes extreme classification with a memory footprint exceeding the main memory capacity of the CPU or GPU. The emerging in-storage-computing technique offers an opportunity on account of the fact that SSD has enough storage capacity for the parameters of extreme classification. However, the limited performance of naive in-storage-computing schemes is insufficient to support the heavy workload of extreme classification. Siqi Li 0013, Fengbin Tu, Liu Liu 0017, Jilan Lin, Zheng Wang 0075, Yangwook Kang, Yufei Ding 0001, Yuan Xie 0001 |
ISCA | 5 |
| 2023 | MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU Platforms
Boyuan Feng, Zheng Wang 0075, Tong Geng, Kevin J. Barker, Ang Li 0006, Yufei Ding 0001 |
OSDI | 3 |
| 2023 | TC-GNN: Bridging Sparse GNN Computation and Dense Tensor Cores on GPUs
Boyuan Feng, Zheng Wang 0075, Guyue Huang, Yufei Ding 0001 |
USENIX ATC | 3 |
| 2022 | EL-Rec: Efficient Large-Scale Recommendation Model Training via Tensor-Train Embedding TableabstractDeep learning Recommendation Models (DLRMs) plays an important role in various application domains. However, existing DLRM training systems require a large number of GPUs due to the memory-intensive embedding tables. To this end, we propose EL-Rec, an efficient computing framework harnessing the Tensor-train (TT) technique to democratize the training of large-scale DLRMs with limited GPU resources. Specifically, EL-Rec optimizes TT decomposition based on key computation primitives of embedding tables and implements a high-performance compressed embedding table which is a drop-in replacement of Pytorch API. EL-Rec introduces an index reordering technique to harvest the performance gains from both local and global information of training inputs. EL-Rec also highlights a pipeline training paradigm to eliminate the communication overhead between the host memory and the training worker. Comprehensive experiments demonstrate that EL-Rec can handle the largest publicly available DLRM dataset with a single GPU and achieves 3× speedup over the state-of-the-art DLRM frameworks. Zheng Wang 0075, Boyuan Feng, Dheevatsa Mudigere, Bharath Muthiah, Yufei Ding 0001 |
SC | 1 |
| 2022 | Faith: An Efficient Framework for Transformer Verification on GPUs
Boyuan Feng, Tianqi Tang 0001, Zhaodong Chen 0001, Zheng Wang 0075, Yuan Xie 0001, Yufei Ding 0001 |
USENIX ATC | 5 |