EDBT 2026 Demo / reviewers in the wild / expert
En Shao
dblp:193/6720
· DBLP profile ↗
22ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0002-9678-7228ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 3 first-author · 17 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM-SYCL: Automated SYCL Generation from CUDA via Search-Driven LLM Translation
Zhou Liang, Yuanbo Wen 0001, En Shao, Guangming Tan |
APPT | 4 |
| 2026 | SYCL-MLU: unifying SIMT and SIMD in heterogeneous programming
Runyu Zhou, Yijin Li, En Shao, Ziyan Xie, Huimin Cui |
CCF Trans. High Perform. Comput. | 5 |
| 2025 | Optimizing Deep Learning Inference Efficiency through Block Dependency AnalysisabstractInter-operator optimization in deep neural networks (DNNs) relies on accurate data dependency analysis. Traditional machine learning compilers (MLCs) perform static data dependency analysis at the element and operator levels, leading to two key limitations: complex dependencies that hinder efficient inter-operator optimizations, and overlooked parallelizable computations that underutilize GPU resources. We introduce BlockDepend, a novel MLC framework that addresses these issues through block-level dependency analysis. By examining the lower-level phases of compilation, BlockDepend extracts crucial block-level dependency information, simplifying complex relationships between operators and uncovering hidden parallelization opportunities. This allows for targeted optimization strategies that enhance memory access efficiency and improve GPU utilization. Our experiments demonstrate BlockDepend's effectiveness, achieving speedups of 1.71× and 2.88× compared to NVIDIA TensorRT and AMD MIGraphX, respectively, across various workloads. Zhanyuan Di, Leping Wang, En Shao, Zhaojia Ma, Ziyi Ren, Feng Hua, Lixian Ma, Jie Zhao 0002, Guangming Tan, Ninghui Sun |
ASPLOS (2) | 3 |
| 2025 | ORION: Optimizing OLAP Query Execution with Proactive Caching and Separate OperatorsabstractCurrent work leverages data caching and operator execution accelerations to reduce the Online Analytical Processing (OLAP) query execution time on the disaggregated architecture with computation, cache, GPU, and storage clusters.However, their optimizations rely heavily on the OLAP engine, thus have defects of passive data fetching and integrated operator executions, leading to poor OLAP query execution performance.To resolve the above problems, we propose the ORION manager to take over the data and operator management capabilities from the OLAP engine for reducing OLAP query execution time.ORION consists of * Zhixin Tong and Jiuchen Shi contributed equally to this work. Zhixin Tong, Jiuchen Shi, Quan Chen 0002, Pu Pang, Shixuan Sun, En Shao, Minyi Guo |
ICS | 8 |
| 2025 | Magneto: Accelerating Parallel Structures in DNNs via Co-Optimization of OperatorsabstractDeep neural networks (DNNs) increasingly rely on parallel structures to enhance performance and efficiency. However, existing machine learning compilers (MLCs) face challenges in optimizing these structures due to limited parallel fusion scopes and insufficient consideration of intra-operator information. This paper introduces Magneto, a novel framework designed to accelerate parallel structures in DNNs through the co-optimization of parallel operators. By expanding the scope of parallel operator fusion and introducing a dedicated co-tuning algorithm, Magneto unlocks new opportunities for co-optimization. Experimental results demonstrate that Magneto outperforms NVIDIA TensorRT and AMD MIGraphX, achieving speedups of 3.02× and 4.19×, respectively. Zhanyuan Di, Leping Wang, Ziyi Ren, En Shao, Jie Zhao 0002, Siyuan Feng 0007, Dingwen Tao, Guangming Tan, Ninghui Sun |
PPoPP | 4 |
| 2025 | VastPipe: A High-Throughput Inference System via Adaptive Space-Division Multiplexing for Diverse Accelerators
Lixian Ma, Leping Wang, En Shao, Rongyu Cao, Guangming Tan |
J. Comput. Sci. Technol. | 3 |
| 2025 | Accelerating Parallel Structures in DNNs via Parallel Fusion and Operator Co-OptimizationabstractParallel structures have become a key pattern in deep neural networks (DNNs), offering improved efficiency and scalability. However, existing machine learning compilers (MLCs) face challenges in optimizing these structures due to limited parallel fusion scope and insufficient analysis of intra-operator characteristics. This article introduces Magneto, a framework designed to accelerate DNN inference by co-optimizing parallel operators. Magneto broadens the fusion scope and incorporates a specialized co-tuning algorithm to optimize operators jointly. Our approach addresses the unique challenges inherent in optimizing parallel structures, enabling significant performance improvements across various hardware platforms. Experimental results show that Magneto outperforms state-of-the-art NVIDIA TensorRT and AMD MIGraphX, achieving geometric mean speedups of 2.27× and 2.88×, respectively. Zhanyuan Di, Leping Wang, Zhaojia Ma, En Shao, Jie Zhao 0002, Ziyi Ren, Siyuan Feng 0007, Dingwen Tao, Guangming Tan, Ninghui Sun |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | ElasticRoom: Multi-Tenant DNN Inference Engine via Co-design with Resource-constrained Compilation and Strong Priority SchedulingabstractGPU partition mechanisms in run-time software have been widely used in job scheduler and multi-tenant computing system to improve resource utilization and throughput. The latency requirements of different DNN requests, such as real-time and best-effort requests, often exhibit variations in computational systems that handle batch tasks for DNN inference. However, the existing GPU partition mechanisms and state-of-the-art scheduling approaches face challenges in effectively promising both high throughput and low latency for real-time requests. The current limitation lies in the inability of existing GPU partition mechanisms to enhance GPU resource utilization and ensure job priority simultaneously. Lixian Ma, Haoruo Chen, En Shao, Leping Wang, Quan Chen 0002, Guangming Tan |
HPDC | 3 |
| 2024 | AsymFB: Accelerating LLM Training Through Asymmetric Model Parallelism
En Shao, Leping Wang, Guangming Tan, Ninghui Sun |
NPC (1) | 2 |
| 2024 | POSTER: FineCo: Fine-grained Heterogeneous Resource Management for Concurrent DNN InferencesabstractCo-locating multiple DNN servings to share GPU resource is widely used to improve resource utilization while guaranteeing user QoS. Existing GPU sharing mechanism is restricted to model level, and fluctuations in kernel-level resource demands highlight a suboptimal utilization of the current sharing mechanism. We design a multi-DNN serving system, FineCo, that leverages a novel fine-grained resource sharing mechanism to optimize concurrent inference without modifications to the hardware or operating system. Our prototype implementation demonstrates that FineCo achieves up to 40% throughput improvement over the state-of-the-art work. Lixian Ma, Haoruo Chen, En Shao, Leping Wang, Quan Chen 0002, Guangming Tan |
PPoPP | 3 |
| 2024 | Mille-feuille: A Tile-Grained Mixed Precision Single-Kernel Conjugate Gradient Solver on GPUsabstractConjugate gradient (CG) and biconjugate gradient stabilized (BiCGSTAB) are effective methods used for solving sparse linear systems. We in this paper propose Mille-feuille, a new solver for accelerating CG and BiCGSTAB on GPUs. We first analyze the two methods and list three findings related to the use of mixed precision, the reduction of kernel synchronization costs, and the awareness of partial convergence during the iteration steps. Then, (1) to enable tile-grained mixed precision, we develop a tiled sparse format; (2) to reduce synchronization costs, we leverage atomic operations that make the whole solving procedure work within a single GPU kernel; (3) to support a partial convergence-aware mixed precision strategy, we enable tile-wise on-chip dynamic precision conversion within the single kernel at runtime. The experimental results on an NVIDIA A100 and an AMD MI210 show that the Mille-feuille solver outperforms baseline implementations using the vendor-support cuSPARSE/hipSPARSE as well as two state-of-the-art libraries PETSc and Ginkgo by a factor of on average 3.03x/2.68x, 5.37 x, 4.36x (up to $8.77 \mathrm{x} / 7.14 x$, 16.54x, 15.69x) in CG, on average 2.65x/2.32x, 3.57x, 3.78x (up to 7.51x/6.63x, 16.64x, 11.73x) in BiCGSTAB, on average 3.82x/3.47x (up to 40.38x/47.75x) in preconditioned CG (PCG), on average 1.79x/1.63x (up to 45.63x/44.34x) in preconditioned BiCGSTAB (PBiCGSTAB), respectively. Dechuang Yang, Yiduo Niu, Weile Jia, En Shao, Weifeng Liu 0002, Guangming Tan, Zhou Jin 0001 |
SC | 5 |
| 2024 | FILL: a heterogeneous resource scheduling system addressing the low throughput problem in GROMACS
Yueyuan Zhou, Ziyi Ren, En Shao, Lixian Ma, Leping Wang, Guangming Tan |
CCF Trans. High Perform. Comput. | 3 |
| 2023 | DeletePop: A DLT Execution Time Predictor Based on Comprehensive Modeling
Yongzhe He, Yueyuan Zhou, En Shao, Guangming Tan, Ninghui Sun |
ICA3PP (7) | 3 |
| 2023 | JetEsti: A New DLT Job Scheduling Simulator Based on Fine-Grained Process ModelingabstractLarge-scale Deep Learning Training(DLT) jobs consume a large amount of time and are usually carried out in a distributed cluster environment. However, existing DLT framework like TensorFlow does not contain adhoc optimizations at parallelism and scheduling, which results in seriously low efficiency. Due to this problem, researchers need to choose appropriate scheduling algorithms for cluster jobs. Consider the expensiveness of hardware resources, using job scheduling simulator(JSS) to verify the performance of different scheduling algorithms in advance is necessary. Yongzhe He, Yueyuan Zhou, En Shao, Guangming Tan, Ninghui Sun |
ICDCS | 3 |
| 2021 | WidePipe: High-Throughput Deep Learning Inference System on a Cluster of Neural Processing UnitsabstractThe wide application of machine learning technology promotes the generation of ML-as-a-Service(MLaaS), which is a serverless computing paradigm for rapidly deploying a trained model as a serving. However, it is a challenge to design an inference system that is capable of coping with large traffic for low latency and heterogeneous neural networks. It is difficult to adaptively configure multilevel parallelism in existing cloud inference systems for machine learning servings, particularly if the cluster has accelerators, such as GPUs, NPUs, FPGAs, etc. These issues lead to poor resource utilization and limit the system throughput. In this paper, we propose and implement a high-throughput inference system called WidePipe, which WidePipe leverages reinforcement learning to co-adapt resource allocation and batch size of request according to device status. We evaluated the performance of WidePipe for a large cluster with 1000 neural processing units in 250 nodes. Our experimental results show that WidePipe has a 2.11× higher throughput than current inference systems when deploying heterogeneous machine learning servings, meeting the service-level objectives for the response time. Lixian Ma, En Shao, Yueyuan Zhou, Guangming Tan |
ICCD | 2 |
| 2021 | Building Agile Workflow Microservice System for HPC Applications Based on Fast-start OSvabstractThe advances of containers have significantly promoted the development of microservice architecture. This architecture splits a monolithic application into multiple independent components and the container orchestrator manages these components by the container in the cloud environment. The feasibility of deploying high performance computing(HPC) applications as microservices has been proven, but the existing container orchestrator incurs a large performance overhead as there is interference between different containers on the same physical host. In this paper, we design an agile workflow microservice system for HPC applications with fast-start OSv. We consider improving HPC workflow performance from two aspects: single OSv startup time optimization and workflow orchestration optimization. For single OSv startup time optimization, we design a fast-start OSv by analyzing the process of OSv startup and finding an optimization by modifying OSv source code. In this way, we get nearly 50% improvement of startup time. For workflow orchestration optimization, we propose four optimization techniques to speed up the execution of workflow by jointly considering OSv and workflow features, namely: node fusion, node merge, image preload, boot delay. Furthermore, we utilize our fast startup OSv to design an orchestration system for efficiently building an agile HPC workflow microservice by Kubevirt. Our experimental results optimization microservice system reduces the execution time by 30% compared with the original deployment with docker. Lixian Ma, En Shao, Guangming Tan |
ICPADS | 3 |
| 2021 | Deep Reinforcement Agent for Failure-aware Job scheduling in High-Performance ComputingabstractJob scheduling is crucial in high-performance computing (HPC), which is dedicated to deciding when and which jobs are allocated to the system and placing the jobs on which resources, by considering multiple scheduling goals. Along with the incremental of various resources and dazzling deep learning training (DLT) workloads, job failure becomes a quite common issue in HPC, which will affect user satisfaction and cluster utilization. To alleviate the influence of hardware and software errors as much as possible, in this paper, we aim to tackle the problem of failure-aware job scheduling in HPC clusters. Inspired by the success of previous studies of deep reinforcement learning-driven job scheduling, we propose a novel HPC scheduling agent named FARS (Failure-aware RL-based scheduler) by considering the effects of job failures. On the one hand, a neural network is applied to map the information of raw cluster and job states to job placement decisions. On the other hand, to consider the influence of job failure for user satisfaction and cluster utilization, FARS leverages make-span of the entire workload as the training objective. Additionally, effective exploration and experience replay techniques are applied to obtain effectively converged agent. To evaluate the capability of FARS, we design extensive trace-based simulation experiments with the popular DLT workloads. The experimental results show that, compared with the best baseline model, FARS obtains 5.69% improvement of average make-span under different device error rates. Together, our FARS is an ideal candidate for failure-aware job scheduler in HPC clusters. Rongyu Cao, Yueyuan Zhou, En Shao, Guangming Tan |
ICPADS | 5 |
| 2021 | A New Optoelectronic Hybrid Network Based on Scheduling Optimization of Optical LinksabstractThe emergence of exascale computers will represent a milestone in high-performance computing (HPC). Optoelectronic interconnections and configurable switches will change the traditional supercomputer architecture. However, new hardware is not easily adapted to dynamic running conditions. Based on scheduling optimization of optical links, we propose a new optoelectronic hybrid network, the software-defined network accelerator (sDNA), for an exascale computer. Our scheduling optimization contains an optical interconnection method and an adaptive routing method. The main contribution of our work is an extended edge forwarding index (E-EFI) optical interconnection method based on slow-switching optical devices. The optical link connections are established by evaluating the traffic offloading revenue for each optical link candidate. To support optical interconnection, sDNA selects a suitable routing strategy according to the job-schedule information and prior HPC application knowledge. We tested sDNA in a network simulator and a prototype exascale computer system using both the US Department of Energy (DOE) application and real-world communication benchmarks. The verification results for traffic offloading reveal that our optical interconnection method not only offloads traffic from electrical links to optical links but also avoids the congestion inherent to electrical links. sDNA maintains a throughput of more than 80 percent bandwidth and reduces the communication delay by 10 percent in our real prototype system and simulator. Thus, sDNA is an ideal candidate for accelerating the communication performance of exascale computers. En Shao, Guangming Tan, Zhan Wang 0003, Guojun Yuan, Zheng Cao 0003, Ninghui Sun |
IEEE Trans. Computers | 1 |
| 2020 | Reducing the Time of Live Container Migration in a Workflow
Zhanyuan Di, En Shao, Mujun He |
NPC | 2 |
| 2019 | A New Traffic Offloading Method with Slow Switching Optical Device in Exascale ComputerabstractThe expected exascale computer will comprise tens of thousands of computing nodes and nearly 5000 interconnected nodes in years to come. Such a large-scale system will represent a milestone in the progress of High-Performance Computing (HPC). The more efficient network hardware, like optoelectronic interconnection and configurable switches, is reforming the traditional architecture of supercomputers. However, the present architecture containing new hardware is not easy to adapt to the dynamically running condition, because the newly developed hardware is normally unable to effectively improve overall performance. Here, we propose a new accelerated system called Software Defined Network Accelerator (sDNA) for the exascale computer. Inspired by edge forwarding index (EFI), the main contribution of our work is that it presents an extended EFI-based optical interconnection method with slow switching optical device. The optical link is connected by the evaluation of each optical link candidate's traffic offloading revenue. As the supporting method for optical interconnection, sDNA selects the most suitable routing configuration according to the job-schedule information and the prior-knowledge of HPC applications. We tested sDNA in a network simulator and a prototype system for the exascale computer, using both DOE application benchmarks and a real-world communication benchmark. From the result of verification of traffic offloading, we found that our optical interconnection method based on our extended EFI evaluation is not only able to offload the traffic from an electrical link to an optical link but is also able to avoid congestion inherent to electrical link. Furthermore, our experimental results show that sDNA maintains the throughput of more than 80% bandwidth and reduced the communication delay by 10% in our real prototype system and simulator. Together, our sDNA is an ideal candidate for accelerating communication performance of the exascale computer. En Shao, Guangming Tan, Zhan Wang 0003, Guojun Yuan, Ninghui Sun |
ICCD | 1 |
| 2019 | Wormhole optical network: a new architecture to solve long diameter problem in exascale computer
En Shao, Zhan Wang 0003, Guojun Yuan, Guangming Tan, Ninghui Sun |
CCF Trans. High Perform. Comput. | 1 |
| 2016 | Modeling Traffic of Big Data Platform for Large Scale Datacenter NetworksabstractPrior to deployment, network designers often use simulators to pre-evaluate the performance of designed network with artificial network traffic. The traditional way of separating network design from real applications will not only result in over-designed network configurations, wasting money and energy, but also miss the real network demands of applications, degrading system performance. In this paper, we provide a method to model the network traffic of current popular big data platforms, which can observably improve the matching between network design and applications. The new method extracts communication behavior from the popular big data applications and replays the behavior instead of the packet traces. Experiments show that the traffic generated by the model is almost match the real traffic and the model can easily scale to thousands of nodes. Zheng Cao 0003, Zhan Wang 0003, Dawei Zang, En Shao, Ninghui Sun |
ICPADS | 5 |