EDBT 2026 Demo / reviewers in the wild / expert
Shulai Zhang
dblp:251/9029
· DBLP profile ↗
10ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0002-0802-7203ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 7 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-DesignabstractEfficiently training large-scale models (LMs) in GPU clusters involves two separate avenues: inter-job dynamic scheduling and intra-job adaptive parallelism (AP). However, existing dynamic schedulers struggle with large-model scheduling due to the mismatch between static parallelism (SP)-aware scheduling and AP-based execution, leading to cluster inefficiencies such as degraded throughput and prolonged job queuing. This paper presents Arena, a large-model training system that co-designs dynamic scheduling and adaptive parallelism to achieve high cluster efficiency. To reduce scheduling costs while improving decision quality, Arena designs low-cost, disaggregated profiling and AP-tailored, load-aware performance estimation, while unifying them by sharding the joint scheduling-parallelism optimization space via a grid abstraction. Building on this, Arena dynamically schedules profiled jobs in elasticity and heterogeneity dimensions, and executes them using efficient AP with pruned search space. Evaluated on heterogeneous testbeds and production workloads, Arena reduces job completion time by up to 49.3% and improves cluster throughput by up to 1.6X. Chunyu Xue, Weihao Cui, Quan Chen 0002, Chen Chen 0067, Han Zhao 0005, Shulai Zhang, Linmei Wang, Limin Xiao 0001, Weifeng Zhang 0003, Jing Yang 0017, Bingsheng He, Minyi Guo |
EuroSys | 6 |
| 2026 | MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing
Chunyu Xue, Yi Pan 0001, Weihao Cui, Quan Chen 0002, Shulai Zhang, Bingsheng He, Minyi Guo |
NSDI | 5 |
| 2025 | Improving GPU Sharing Performance through Adaptive Bubbleless Spatial-Temporal SharingabstractData centers now allow multiple applications that have lightweight workloads to share a GPU. Existing temporal or spatial sharing systems struggle to provide efficient and accurate quota assignments. We observe that the performance of the multi-user system is often underestimated because of the existence of unused GPU "bubbles" and can be enhanced by squeezing the bubbles. Based on this observation, we design Bless, a bubble-less spatial-temporal sharing GPU system that fine-tunes the GPU resource allocation to improve multi-user performance. Bless leverages precise computing resource management and fine-grained kernel scheduling to ensure stringent quota guarantees and reduce latency fairly for applications with varying GPU quotas. We implement and evaluate Bless with multiple applications and workloads. Our result shows that Bless achieves 21.1% - 37.3% average latency reduction over the state-of-the-art while guaranteeing the promised quota for all applications. Shulai Zhang, Quan Chen 0002, Weihao Cui, Han Zhao 0005, Chunyu Xue, Zhen Zheng, Wei Lin 0016, Minyi Guo |
EuroSys | 1 |
| 2025 | Efficient Performance-Aware GPU Sharing with Compatibility and Isolation through Kernel Space Interception
Shulai Zhang, Quan Chen 0002, Han Zhao 0005, Weihao Cui, Limin Xiao 0001, Minyi Guo |
USENIX ATC | 1 |
| 2025 | ARACHNE: Optimizing Distributed Parallel Applications with Reduced Inter-Process CommunicationabstractIn high-performance computing (HPC), parallelization is essential for improving computational efficiency as data and computation scales exceed single-node capacity. Existing methods, such as the polyhedral model used in Pluto -Distmem, focus on loop and array optimizations within shared memory but struggle with high communication overheads and inflexibility in distributed environments. These methods often fail to effectively partition computation and manage data across nodes, leading to suboptimal performance. This paper presents Arachne , an innovative system designed to address these shortcomings by generating distributed parallel code with minimized communication overhead. The system introduces a dynamic programming algorithm to optimally distribute computational tasks across multiple processes, ensuring minimal communication costs. It also incorporates user-friendly compiler directives, allowing programmers to influence code generation easily and accommodate a broader range of parallelization scenarios without needing in-depth knowledge of parallel architectures. Arachne significantly reduces the learning curve and need for extensive code modifications, making parallel programming more accessible and efficient. Evaluation of various HPC benchmarks demonstrates that Arachne outperforms existing methods by reducing communication overhead, lowering memory requirements, and supporting more complex parallel logic, thus enhancing the overall scalability and efficiency of HPC applications. Yifu He, Han Zhao 0005, Weihao Cui, Shulai Zhang, Quan Chen 0002, Minyi Guo |
ACM Trans. Archit. Code Optim. | 4 |
| 2023 | Maximizing the Utilization of GPUs Used by Cloud Gaming through Adaptive Co-location with ComboabstractCloud vendors are now providing cloud gaming services with GPUs. GPUs in cloud gaming experience periods of idle because not every frame in a game always keeps the GPU busy for rendering. Previous works temporally co-locate games with best-effort applications to harvest these idle cycles. However, these works ignore the spatial sharing of GPUs, leading to not maximized throughput improvement. The newly introduced RT (ray tracing) Cores inside GPU SMs for ray tracing exacerbate the situation. Binghao Chen, Han Zhao 0005, Weihao Cui, Yifu He, Shulai Zhang, Quan Chen 0002, Zijun Li 0001, Minyi Guo |
SoCC | 5 |
| 2022 | PAME: precision-aware multi-exit DNN serving for reducing latencies of batched inferencesabstractIn emerging DNN serving systems, queries are usually batched to fully leverage hardware resources, and all the queries in a batch run through the complete model and return at the same time. According to our findings, some queries only need to pass through a portion of the DNN model to attain sufficient precision in a DNN service. These queries can have shorter latencies if they can return early in the middle of a model. Therefore, we propose precision-aware multi-exit inference serving, PAME, to achieve the above purpose. PAME provides a holistic scheme to build a multi-exit DNN model and a corresponding system-level design of the inference engine. We use representative CV and NLP benchmarks to evaluate PAME. PAME is adaptive to various DNN tasks and service loads. Experimental results show that PAME reduces 39.9% average latency without increasing the tail latency, while maintaining 99.68% precision of the original single-exit DNN models on average. Shulai Zhang, Weihao Cui, Quan Chen 0002, Zhengnian Zhang, Yue Guan 0003, Jingwen Leng, Chao Li 0009, Minyi Guo |
ICS | 1 |
| 2021 | Dubhe: Towards Data Unbiasedness with Homomorphic Encryption in Federated Learning Client SelectionabstractFederated learning (FL) is a distributed machine learning paradigm that allows clients to collaboratively train a model over their own local data. FL promises the privacy of clients and its security can be strengthened by cryptographic methods such as additively homomorphic encryption (HE). However, the efficiency of FL could seriously suffer from the statistical heterogeneity in both the data distribution discrepancy among clients and the global distribution skewness. We mathematically demonstrate the cause of performance degradation in FL and examine the performance of FL over various datasets. To tackle the statistical heterogeneity problem, we propose a pluggable system-level client selection method named Dubhe, which allows clients to proactively participate in training, meanwhile preserving their privacy with the assistance of HE. Experimental results show that Dubhe is comparable with the optimal greedy method on the classification accuracy, with negligible encryption and communication overhead. Shulai Zhang, Quan Chen 0002, Wenli Zheng, Jingwen Leng, Minyi Guo |
ICPP | 1 |
| 2020 | A General Difficulty Control Algorithm for Proof-of-Work Based BlockchainsabstractDesigning an efficient difficulty control algorithm is an essential problem in Proof-of-Work (PoW) based blockchains because the network hash rate is randomly changing. This paper proposes a general difficulty control algorithm and provides insights for difficulty adjustment rules for PoW based blockchains. The proposed algorithm consists a two-layer neural network. It has low memory cost, meanwhile satisfying the fast-updating and low volatility requirements for difficulty adjustment. Real data from Ethereum are used in the simulations to prove that the proposed algorithm has better performance for the control of the block difficulty. Shulai Zhang, Xiaoli Ma |
ICASSP | 1 |
| 2019 | Exploiting Caching and Prediction to Promote User Experience for a Real-Time Wireless VR ServiceabstractIn this paper, we propose a novel wireless virtual reality scheme by caching resource on head-mounted displays (HMD) and exploiting head movement prediction to improve user experience. We render images for the future request based on the user's head movement prediction at the server-end to reduce the motion-to- photon latency and cache those rendered images on the HMD to form an image pool. We then develop a corresponding algorithm to select and warp the exact image from the image pool for display. Furthermore, a performance metric - warping distance is defined and used to evaluate the image quality of the proposed scheme. Finally, real dataset-driven results show that the proposed scheme is able to provide higher image quality as well as experience consistency compared with existed schemes. Shulai Zhang, Meixia Tao, Zhiyong Chen 0002 |
GLOBECOM | 1 |