EDBT 2026 Demo / reviewers in the wild / expert
Leping Wang
dblp:80/7791
· DBLP profile ↗
12ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimizing Deep Learning Inference Efficiency through Block Dependency AnalysisabstractInter-operator optimization in deep neural networks (DNNs) relies on accurate data dependency analysis. Traditional machine learning compilers (MLCs) perform static data dependency analysis at the element and operator levels, leading to two key limitations: complex dependencies that hinder efficient inter-operator optimizations, and overlooked parallelizable computations that underutilize GPU resources. We introduce BlockDepend, a novel MLC framework that addresses these issues through block-level dependency analysis. By examining the lower-level phases of compilation, BlockDepend extracts crucial block-level dependency information, simplifying complex relationships between operators and uncovering hidden parallelization opportunities. This allows for targeted optimization strategies that enhance memory access efficiency and improve GPU utilization. Our experiments demonstrate BlockDepend's effectiveness, achieving speedups of 1.71× and 2.88× compared to NVIDIA TensorRT and AMD MIGraphX, respectively, across various workloads. Zhanyuan Di, Leping Wang, En Shao, Zhaojia Ma, Ziyi Ren, Feng Hua, Lixian Ma, Jie Zhao 0002, Guangming Tan, Ninghui Sun |
ASPLOS (2) | 2 |
| 2025 | Magneto: Accelerating Parallel Structures in DNNs via Co-Optimization of OperatorsabstractDeep neural networks (DNNs) increasingly rely on parallel structures to enhance performance and efficiency. However, existing machine learning compilers (MLCs) face challenges in optimizing these structures due to limited parallel fusion scopes and insufficient consideration of intra-operator information. This paper introduces Magneto, a novel framework designed to accelerate parallel structures in DNNs through the co-optimization of parallel operators. By expanding the scope of parallel operator fusion and introducing a dedicated co-tuning algorithm, Magneto unlocks new opportunities for co-optimization. Experimental results demonstrate that Magneto outperforms NVIDIA TensorRT and AMD MIGraphX, achieving speedups of 3.02× and 4.19×, respectively. Zhanyuan Di, Leping Wang, Ziyi Ren, En Shao, Jie Zhao 0002, Siyuan Feng 0007, Dingwen Tao, Guangming Tan, Ninghui Sun |
PPoPP | 2 |
| 2025 | VastPipe: A High-Throughput Inference System via Adaptive Space-Division Multiplexing for Diverse Accelerators
Lixian Ma, Leping Wang, En Shao, Rongyu Cao, Guangming Tan |
J. Comput. Sci. Technol. | 2 |
| 2025 | Accelerating Parallel Structures in DNNs via Parallel Fusion and Operator Co-OptimizationabstractParallel structures have become a key pattern in deep neural networks (DNNs), offering improved efficiency and scalability. However, existing machine learning compilers (MLCs) face challenges in optimizing these structures due to limited parallel fusion scope and insufficient analysis of intra-operator characteristics. This article introduces Magneto, a framework designed to accelerate DNN inference by co-optimizing parallel operators. Magneto broadens the fusion scope and incorporates a specialized co-tuning algorithm to optimize operators jointly. Our approach addresses the unique challenges inherent in optimizing parallel structures, enabling significant performance improvements across various hardware platforms. Experimental results show that Magneto outperforms state-of-the-art NVIDIA TensorRT and AMD MIGraphX, achieving geometric mean speedups of 2.27× and 2.88×, respectively. Zhanyuan Di, Leping Wang, Zhaojia Ma, En Shao, Jie Zhao 0002, Ziyi Ren, Siyuan Feng 0007, Dingwen Tao, Guangming Tan, Ninghui Sun |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | ElasticRoom: Multi-Tenant DNN Inference Engine via Co-design with Resource-constrained Compilation and Strong Priority SchedulingabstractGPU partition mechanisms in run-time software have been widely used in job scheduler and multi-tenant computing system to improve resource utilization and throughput. The latency requirements of different DNN requests, such as real-time and best-effort requests, often exhibit variations in computational systems that handle batch tasks for DNN inference. However, the existing GPU partition mechanisms and state-of-the-art scheduling approaches face challenges in effectively promising both high throughput and low latency for real-time requests. The current limitation lies in the inability of existing GPU partition mechanisms to enhance GPU resource utilization and ensure job priority simultaneously. Lixian Ma, Haoruo Chen, En Shao, Leping Wang, Quan Chen 0002, Guangming Tan |
HPDC | 4 |
| 2024 | AsymFB: Accelerating LLM Training Through Asymmetric Model Parallelism
En Shao, Leping Wang, Guangming Tan, Ninghui Sun |
NPC (1) | 3 |
| 2024 | POSTER: FineCo: Fine-grained Heterogeneous Resource Management for Concurrent DNN InferencesabstractCo-locating multiple DNN servings to share GPU resource is widely used to improve resource utilization while guaranteeing user QoS. Existing GPU sharing mechanism is restricted to model level, and fluctuations in kernel-level resource demands highlight a suboptimal utilization of the current sharing mechanism. We design a multi-DNN serving system, FineCo, that leverages a novel fine-grained resource sharing mechanism to optimize concurrent inference without modifications to the hardware or operating system. Our prototype implementation demonstrates that FineCo achieves up to 40% throughput improvement over the state-of-the-art work. Lixian Ma, Haoruo Chen, En Shao, Leping Wang, Quan Chen 0002, Guangming Tan |
PPoPP | 4 |
| 2024 | FILL: a heterogeneous resource scheduling system addressing the low throughput problem in GROMACS
Yueyuan Zhou, Ziyi Ren, En Shao, Lixian Ma, Leping Wang, Guangming Tan |
CCF Trans. High Perform. Comput. | 6 |
| 2023 | Falcon: Fast OLTP Engine for Persistent Cache and Non-Volatile MemoryabstractNon-volatile memory(NVM) has the properties of both byte addressable and persistence, which provides new opportunities for building on-line transaction processing (OLTP) engines. Recently, a new feature called eADR puts CPU cache also in the persistence domain. Existing OLTP engines are based on volatile cache and now have the opportunity to improve performance further and reduce programming complexity with persistent cache. Kang Chen 0001, Leping Wang, Yongwei Wu 0001 |
SOSP | 3 |
| 2010 | Power-efficient workload distribution for virtualized server clustersabstractWith growing cost of electricity, the power management of server clusters has become an important problem. Most existing research, however, either do not apply to virtualized environments or do not focus on power-efficient workload distribution. To fill in this research gap, we propose a workload distribution algorithm for virtualized server clusters to reduce their power consumptions and provide quality of service (QoS). Built upon optimization, queuing theory and control theory techniques, our approach achieves the design goal, where QoS is provided to a larger number of requests with a smaller amount of power consumption. Leping Wang, Ying Lu 0002 |
HiPC | 1 |
| 2010 | An Efficient Threshold-Based Power Management Mechanism for Heterogeneous Soft Real-Time ClustersabstractWith growing cost of electricity, the power management (PM) of server clusters has become an important problem. However, most previous researchers only address the challenge in homogeneous environments. Considering the increasing popularity of heterogeneous systems, this paper proposes an efficient algorithm for PM of heterogeneous soft real-time clusters. It is built on simple but effective mathematical models. When deployed to a new platform, the software incurs low configuration cost because no extensive performance measurements and profiling are required. To strive for efficiency, a threshold-based approach is adopted. In this paper, we systematically study this approach and its design decisions. Leping Wang, Ying Lu 0002 |
IEEE Trans. Ind. Informatics | 1 |
| 2008 | Efficient Power Management of Heterogeneous Soft Real-Time ClustersabstractWith growing cost of electricity, the power management of server clusters has become an important problem. However, most previous researchers only address the challenge in homogeneous environments. Considering the increasing popularity of heterogeneous systems, this paper proposes an efficient algorithm for power management of heterogeneous soft real-time clusters. It is built on simple but effective mathematical models. When deployed to a new platform, the software incurs low configuration cost because no extensive performance measurements and profiling are required. To strive for efficiency, a threshold-based approach is adopted. In this paper, we systematically study this approach and its design decisions. Leping Wang, Ying Lu 0002 |
RTSS | 1 |