EDBT 2026 Demo / reviewers in the wild / expert
Xiao Shi 0003
dblp:179/9306-3
· DBLP profile ↗
6ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-7105-8355ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective ElasticityabstractServerless computing, with its ease of management, auto-scaling, and cost-effectiveness, is widely adopted by deep learning (DL) applications. DL workloads, especially with large language models, require substantial GPU resources to ensure QoS. However, it is prone to produce GPU fragments (e.g., 15%-94%) in serverless DL systems due to the dynamicity of workloads and coarse-grained static GPU allocation mechanisms, gradually eroding the profits offered by serverless elasticity. Different from classical serverless systems that only scale horizontally, we present introspective elasticity (IE), a fine-grained and adaptive two-dimensional co-scaling mechanism to support GPU resourcing-on-demand for serverless DL tasks. Based on this insight, we build Dilu, a cross-layer and GPU-based serverless DL system with IE support. First, Dilu provides multi-factor profiling for DL tasks with efficient pruning search methods. Second, Dilu adheres to the resourcing-complementary principles in scheduling to improve GPU utilization with QoS guarantees. Third, Dilu adopts an adaptive 2D co-scaling method to enhance the elasticity of GPU provisioning in real time. Evaluations show that it can dynamically adjust the resourcing of various DL functions with low GPU fragmentation (10%-46% GPU defragmentation), high throughput (up to 1.8× inference and 1.1× training throughput increment) and QoS guarantees (11%-71% violation rate reduction), compared to the SOTA baselines. Cunchi Lv, Xiao Shi 0003, Zhengyu Lei, Jinyue Huang, Wenting Tan, Xiaohui Zheng |
ASPLOS (1) | 2 |
| 2024 | SpecInF: Exploiting Idle GPU Resources in Distributed DL Training via Speculative Inference Filling
Cunchi Lv, Xiao Shi 0003, Wenting Tan |
NPC (1) | 2 |
| 2023 | Chitu: Accelerating Serverless Workflows with Asynchronous State Replication PipelinesabstractServerless workflows are characterized as multi-stage computing, while downstream functions require accessing intermediate states or the output of upstream functions for running. The workflow's performance can be easily affected due to the inefficiency of data access. Studies accelerate data access with various policies, such as direct and indirect methods. However, these methods may fail due to various limitations such as resource availability. Zhengyu Lei, Xiao Shi 0003, Cunchi Lv |
SoCC | 2 |
| 2022 | SMPI: Scalable Serverless MPI ComputingabstractRunning HPC in the cloud has gained more and more practice. As a new cloud paradigm, serverless is highly attractive for HPC service providers due to its distinctive benefits such as scalability. However, it is difficult for serverless to meet the demands of MPI programming and running, resulting in that MPI programs cannot scale with serverless functions. We introduce a serverless parallel function model to solve the problems. It divides parallelism at function and worker levels to bridge gaps in programming and running between serverless and MPI. Then we present the SMPI framework atop the model. For programming, SMPI redefines the function generation pipeline for parallel functions to prepare metadata for MPI parallel functions. For running, SMPI employs the parallel function gateway and scheduler to realize parallel function invocation and instantiation for MPI parallel functions. It is implemented and evaluated with OpenFaaS. Experiments show that SMPI supports MPI programming and running in a complete serverless manner. Compared to server-centric methods, it reduces efforts on cluster maintenance, provides scalable serverless MPI computing with competitive performance (0.559–1.048s slower of start-up time, and 0.145s–0.945s slower of computing time than best-behaved baseline), and is potential to scale on multiple clusters for higher scalability. Yuxin Yuan, Xiao Shi 0003, Zhengyu Lei |
IPCCC | 2 |
| 2022 | TrainFlow: A Lightweight, Programmable ML Training Framework via Serverless Paradigm
Wenting Tan, Xiao Shi 0003, Zhengyu Lei, Cunchi Lv |
NPC | 2 |
| 2020 | Spark-based parallel calculation of 3D fourier shell correlation for macromolecule structure local resolution estimationabstractBACKGROUND: Resolution estimation is the main evaluation criteria for the reconstruction of macromolecular 3D structure in the field of cryoelectron microscopy (cryo-EM). At present, there are many methods to evaluate the 3D resolution for reconstructed macromolecular structures from Single Particle Analysis (SPA) in cryo-EM and subtomogram averaging (SA) in electron cryotomography (cryo-ET). As global methods, they measure the resolution of the structure as a whole, but they are inaccurate in detecting subtle local changes of reconstruction. In order to detect the subtle changes of reconstruction of SPA and SA, a few local resolution methods are proposed. The mainstream local resolution evaluation methods are based on local Fourier shell correlation (FSC), which is computationally intensive. However, the existing resolution evaluation methods are based on multi-threading implementation on a single computer with very poor scalability. RESULTS: This paper proposes a new fine-grained 3D array partition method by key-value format in Spark. Our method first converts 3D images to key-value data (K-V). Then the K-V data is used for 3D array partitioning and data exchange in parallel. So Spark-based distributed parallel computing framework can solve the above scalability problem. In this distributed computing framework, all 3D local FSC tasks are simultaneously calculated across multiple nodes in a computer cluster. Through the calculation of experimental data, 3D local resolution evaluation algorithm based on Spark fine-grained 3D array partition has a magnitude change in computing speed compared with the mainstream FSC algorithm under the condition that the accuracy remains unchanged, and has better fault tolerance and scalability. CONCLUSIONS: In this paper, we proposed a K-V format based fine-grained 3D array partition method in Spark to parallel calculating 3D FSC for getting a 3D local resolution density map. 3D local resolution density map evaluates the three-dimensional density maps reconstructed from single particle analysis and subtomogram averaging. Our proposed method can significantly increase the speed of the 3D local resolution evaluation, which is important for the efficient detection of subtle variations among reconstructed macromolecular structures. Yongchun Lü, Xinhui Tian, Xiao Shi 0003, Xiaohui Zheng, Xin Gao 0001, Min Xu 0009 |
BMC Bioinform. | 4 |