EDBT 2026 Demo / reviewers in the wild / expert
Jiuchen Shi
dblp:272/2785
· DBLP profile ↗
15ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0002-5470-210XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 6 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM ServingabstractMultiple Low-Rank Adapters (Multi-LoRA) are gaining popularity for task-specific Large Language Model (LLM) applications. For Multi-LoRA serving, caching hot LoRAs and KV caches in the GPU memory can improve inference performance. However, existing Multi-LoRA inference systems fail to optimize serving performance like Time-To-First-Token (TTFT), neglecting usage dependencies when caching LoRAs and KV caches. We therefore propose ELORA, a Multi-LoRA caching system to optimize the serving performance. ELORA comprises a dependency-aware cache manager and a performancedriven cache swapper. The cache manager maintains the usage dependencies between LoRAs and KV caches during inference with a unified caching pool. The cache swapper determines the swap-in or swap-out of LoRAs and KV caches based on a unified cost model, when the GPU memory is idle or busy, respectively. Experimental results show that ELORA reduces the TTFT by$\mathbf{4 5. 7 \%}$on average, compared to state-of-the-art works. Jiuchen Shi, Quan Chen 0002, Yizhou Shan, Kaihua Fu, Wei Wang 0030, Minyi Guo |
HPCA | 1 |
| 2026 | Delphinus: Improving Resource Efficiency of Applications with Shared Microservices and Diverse QueriesabstractMicroservices are widely shared in production user-facing applications. These shared microservices have various resource usage patterns when queries from different call graphs of different services access them. However, existing microservice management works fail to efficiently scale resources for them, mainly due to the lack of fine-grained scheduling of diverse queries. We therefore propose Delphinus , a runtime system that efficiently manages resources for shared microservices while ensuring the Quality-of-Service (QoS). Delphinus comprises a group-oriented query scheduler and a borrowing-based load adapter . The query scheduler identifies diverse queries, groups the containers of shared microservices, and schedules the queries into separate groups. The load adapter efficiently scales resources for shared microservices, and fully utilizes the idle containers among groups when the loads of diverse queries change. Results show that Delphinus reduces CPU and memory usage by 40.1% and 36.4% for shared microservices, respectively, compared to state-of-the-art works. Jiuchen Shi, Jinyuan Chen, Quan Chen 0002, Kaihua Fu, Fanrong Du, Zijun Li 0001, Deze Zeng, Jiannong Cao 0001, Shuo Quan, Jie Wu 0001, Minyi Guo |
ACM Trans. Archit. Code Optim. | 1 |
| 2026 | QoS Awareness and Improved Throughput of Point Cloud Services With Dynamic WorkloadsabstractDeep learning on 3D point clouds plays a vital role in a wide range of applications such as AR/VR visualization, 3D cloth virtual try-on, and game rendering. As some applications require low latency, the point cloud services are also deployed on datacenter with powerful GPUs. While the queries of point cloud services show various workload change patterns due to different degrees of sparsity, current batching-based serving schemes result in either long latency or low throughput. We propose a scheme called Volans to address the above challenges and effectively support point cloud services. Volans comprises a workload predictor, a topology deployer, and a progress-aware scheduler. The predictor grids the input query and estimates the workload changes. Afterward, the deployer splits the model into several stages and determines the batch size for each stage based on the workload changes. The scheduler reduces the QoS violation when queries run slower due to unpredicted workload spikes. Experiments show that Volans enhances the peak supported throughput by up to 31.1% while maintaining the required 99%-ile latencies compared to state-of-the-art techniques. Kaihua Fu, Jiuchen Shi, Yao Chen 0008, Quan Chen 0002, Weng-Fai Wong, Wei Wang 0030, Bingsheng He, Minyi Guo |
IEEE Trans. Computers | 2 |
| 2025 | Veyth: Adaptive Container Placement for Optimizing Cross-Server Network Traffic of Microservice Applications
Jinyuan Chen, Jiuchen Shi, Quan Chen 0002, Lin Gu 0002, Minyi Guo |
APPT | 2 |
| 2025 | Comber: QoS-Aware and Efficient Deployment for Co-located Microservices and Best-Effort Tasks in Disaggregated Datacenters
Ruogang Ma, Jiuchen Shi, Quan Chen 0002, Minyi Guo |
APPT | 2 |
| 2025 | Generating Microservice Graphs with Production Characteristics for Efficient Resource ScalingabstractA production microservice application can have multiple services with varying call graphs, and a microservice may be shared across different call graphs.Improving resource efficiency in such complex applications requires proper benchmarks, but production traces are often too large to be used in experiments.To this end, we propose a Service Dependency Graph Generator (DGG) that comprises a Data Handler and a Graph Generator, to generate service dependency graphs of benchmarks that incorporate production-level characteristics from traces.The data handler constructs fine-grained call graphs with dynamic interface and repeated calling features from the trace, and then clusters these call graphs based on the topological and invocation types.The graph generator uses a random graph model to simulate real microservice invocations, generating multiple call graphs and merging them into small-scale service dependency graphs with production-level characteristics.Case studies show that * Fanrong Du and Jiuchen Shi contributed equally to this work. Fanrong Du, Jiuchen Shi, Quan Chen 0002, Pu Pang, Li Li 0012, Minyi Guo |
ICS | 2 |
| 2025 | ORION: Optimizing OLAP Query Execution with Proactive Caching and Separate OperatorsabstractCurrent work leverages data caching and operator execution accelerations to reduce the Online Analytical Processing (OLAP) query execution time on the disaggregated architecture with computation, cache, GPU, and storage clusters.However, their optimizations rely heavily on the OLAP engine, thus have defects of passive data fetching and integrated operator executions, leading to poor OLAP query execution performance.To resolve the above problems, we propose the ORION manager to take over the data and operator management capabilities from the OLAP engine for reducing OLAP query execution time.ORION consists of * Zhixin Tong and Jiuchen Shi contributed equally to this work. Zhixin Tong, Jiuchen Shi, Quan Chen 0002, Pu Pang, Shixuan Sun, En Shao, Minyi Guo |
ICS | 2 |
| 2024 | Adaptive QoS-Aware Microservice Deployment With Excessive Loads via Intra- and Inter-Datacenter SchedulingabstractUser-facing applications often experience excessive loads and are shifting towards the microservice architecture. To fully utilize heterogeneous resources, current datacenters have adopted the disaggregated storage and compute architecture, where the storage and compute clusters are suitable to deploy the stateful and stateless microservices, respectively. Moreover, when the local datacenter has insufficient resources to host excessive loads, a reasonable solution is moving some microservices to remote datacenters. However, it is nontrivial to decide the appropriate microservice deployment inside the local datacenter and identify the appropriate migration decision to remote datacenters, as microservices show different characteristics, and the local datacenter shows different resource contention situations. We therefore propose ELIS, an intra- and inter-datacenter scheduling system that ensures the Quality-of-Service (QoS) of the microservice application, while minimizing the network bandwidth usage and computational resource usage. ELIS comprises aresource manager, across-cluster microservice deployer, and areward-based microservice migrator. The resource manager allocates near-optimal resources for microservices while ensuring QoS. The microservice deployer deploys the microservices between the storage and compute clusters in the local datacenter, to minimize the network bandwidth usage while satisfying the microservice resource demand. The microservice migrator migrates some microservices to remote datacenters when local resources cannot afford the excessive loads. Experimental results show that ELIS ensures the QoS of user-facing applications. Meanwhile, it reduces the public network bandwidth usage, the remote computational resource usage, and the local network bandwidth usage by 49.6%, 48.5%, and 60.7% on average, respectively. Jiuchen Shi, Kaihua Fu, Quan Chen 0002, Deze Zeng, Minyi Guo |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | BLAD: Adaptive Load Balanced Scheduling and Operator Overlap Pipeline For Accelerating The Dynamic GNN TrainingabstractDynamic graph networks are widely used for learning time-evolving graphs, but prior work on training these networks is inefficient due to communication overhead, long synchronization, and poor resource usage. Our investigation shows that communication and synchronization can be reduced by carefully scheduling the workload. And the execution order of operators in GNNs can be adjusted without hurting training convergence. We propose a system called BLAD to consider the above factors, comprising a two-level load scheduler and an overlap-aware topology manager. The scheduler allocates each snapshot group to a GPU, alleviating cross-GPU communication. The snapshots in a group are then carefully allocated to processes on a GPU, enabling overlap of compute-intensive NN operators and memory-intensive graph operators. The topology manager adjusts the operators' execution order to maximize the overlap. Experiments show that BLAD achieves 27.2% speed up on training time on average without affecting final accuracy, compared to state-of-the-art solutions. Kaihua Fu, Quan Chen 0002, Yuzhuo Yang, Jiuchen Shi, Chao Li 0009, Minyi Guo |
SC | 4 |
| 2023 | Nodens: Enabling Resource Efficient and Fast QoS Recovery of Dynamic Microservice Applications in Datacenters
Jiuchen Shi, Zhixin Tong, Quan Chen 0002, Kaihua Fu, Minyi Guo |
USENIX ATC | 1 |
| 2022 | Characterizing and orchestrating VM reservation in geo-distributed clouds to improve the resource efficiencyabstractCloud providers often build a geo-distributed cloud from multiple datacenters in different geographic regions, to serve tenants at different locations. The tenants that run large scale applications often reserve resources based on their peak loads in the region close to the end users to handle the ever changing application load, wasting a large amount of resources. We therefore characterize the VM request patterns of the top tenants in our production public geo-distributed cloud, and open-source the VM request traces in four months from the top 20 tenants of our cloud. The characterization shows that the resource usage of large tenants has various temporal and spatial patterns on the dimensions of time series, regions, and VM types, and has the potential of peak shaving between different tenants to further reduce the resource reservation cost. Based on the findings, we propose a resource reservation and VM request scheduling scheme named ROS to minimize the resource reservation cost while satisfying the VM allocation requests. Our experiments show that ROS reduces the overall deployment cost by 75.4% and the reservation resources by 60.1%, compared to the tenant-specified reservation strategy. Jiuchen Shi, Kaihua Fu, Quan Chen 0002, Changpeng Yang, Mosong Zhou, Jieru Zhao, Chen Chen 0067, Minyi Guo |
SoCC | 1 |
| 2022 | QoS-awareness of Microservices with Excessive Loads via Inter-Datacenter SchedulingabstractUser-facing applications often experience excessive loads and are shifting towards microservice software architecture. While the local datacenter may not have enough resources to host the excessive loads, a reasonable solution is moving some microservices of the applications to remote datacenters. However, it is nontrivial to identify the appropriate migration decision, as the microservices show different characteristics, and the local datacenter also shows different resource contention situations. We therefore propose ELIS, an inter-datacenter scheduling system that ensures the required Quality-of-Service (QoS) of the microservice application with excessive loads, while minimizing the resource usage of the remote datacenter. ELIS comprises a resource manager and a reward-based microservice migrator. The resource manager finds the near-optimal resource configurations for different microservices to minimize resource usage while ensuring QoS. The microservice migrator migrates some microservices to remote datacenters when local resources cannot afford the excessive loads. Our experimental results show that ELIS ensures the required QoS of user-facing applications at excessive loads. Meanwhile, it reduces overall/remote resource usage by 13.1% and 58.1% on average, respectively. Jiuchen Shi, Kaihua Fu, Quan Chen 0002, Deze Zeng, Minyi Guo |
IPDPS | 1 |
| 2022 | QoS-Aware Irregular Collaborative Inference for Improving Throughput of DNN ServicesabstractWith collaborative DNN inference, part of queries run on their source edge device to reduce latencies. Because edges show diverse performance and network conditions, different layers should run on different devices, and queries on the datacenter show irregular structures. However, emerging schemes are not able to process such irregular queries. We propose ICE, a collaborative inference service scheme that effectively supports irregular queries. ICE comprises a query slicer, a query manager, and a lag enhancer. The query slicer maps the execution of queries based on the edges' performance and network conditions. The query manager batches irregular queries adaptively and schedules the irregular queries based on their progress. The lag enhancer reduces the QoS violation when queries run slower due to interference on the edge. Experiments show that ICE improves the supported peak load of the datacenter by 43.2% on average while guaranteeing the required 99%-ile latencies compared with state-of-the-art techniques. Kaihua Fu, Jiuchen Shi, Quan Chen 0002, Ningxin Zheng, Wei Zhang 0149, Deze Zeng, Minyi Guo |
SC | 2 |
| 2022 | Reliability and Incentive of Performance Assessment for Decentralized Clouds
Jiuchen Shi, Xiaoqing Cai, Wenli Zheng, Quan Chen 0002, Deze Zeng, Tatsuhiro Tsuchiya, Minyi Guo |
J. Comput. Sci. Technol. | 1 |
| 2020 | OVERSEE: Outsourcing Verification to Enable Resource Sharing in Edge EnvironmentabstractMulti-tenant (or colocation) data centers are good solutions to support edge computing, since each enterprise or organization usually has limited servers at an edge site. When any data center tenant faces a burst of workload, renting resources from the other tenants in the same data center can provide the required resources while keeping the merits of edge computing, but its challenges of reliability and performance are daunting. In this paper, we propose OVERSEE, an outsourcing verification mechanism that enables resource sharing in multi-tenant data centers, fully exploiting the benefits of edge computing. OVERSEE addresses the above two challenges by making skillful use of Intel SGX suite. OVERSEE consists of two sub schemes, the Report-Proof mechanism and Sampling-Challenging mechanism. The Report-Proof mechanism guarantees a task outsourced by a tenant can be executed correctly, i.e., completely and without modification, in the operating environment provided by another tenant. The Sampling-Challenging mechanism can be used to verify that sufficient computing capacity is provided to achieve the required QoS according to the resource lease agreement between the tenants. The theoretical analysis shows the effectiveness of OVERSEE and the experimental results show that it brings minimal overhead. Xiaoqing Cai, Jiuchen Shi, Wenli Zheng, Quan Chen 0002, Chao Li 0009, Jingwen Leng, Minyi Guo |
ICPP | 2 |