VLDB 2026 Research / reviewers in the wild / expert
Kaihua Fu
dblp:42/9897
· DBLP profile ↗
17ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0001-5117-7162ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 5 first-author · 15 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM ServingabstractMultiple Low-Rank Adapters (Multi-LoRA) are gaining popularity for task-specific Large Language Model (LLM) applications. For Multi-LoRA serving, caching hot LoRAs and KV caches in the GPU memory can improve inference performance. However, existing Multi-LoRA inference systems fail to optimize serving performance like Time-To-First-Token (TTFT), neglecting usage dependencies when caching LoRAs and KV caches. We therefore propose ELORA, a Multi-LoRA caching system to optimize the serving performance. ELORA comprises a dependency-aware cache manager and a performancedriven cache swapper. The cache manager maintains the usage dependencies between LoRAs and KV caches during inference with a unified caching pool. The cache swapper determines the swap-in or swap-out of LoRAs and KV caches based on a unified cost model, when the GPU memory is idle or busy, respectively. Experimental results show that ELORA reduces the TTFT by$\mathbf{4 5. 7 \%}$on average, compared to state-of-the-art works. Jiuchen Shi, Quan Chen 0002, Yizhou Shan, Kaihua Fu, Wei Wang 0030, Minyi Guo |
HPCA | 6 |
| 2026 | Delphinus: Improving Resource Efficiency of Applications with Shared Microservices and Diverse QueriesabstractMicroservices are widely shared in production user-facing applications. These shared microservices have various resource usage patterns when queries from different call graphs of different services access them. However, existing microservice management works fail to efficiently scale resources for them, mainly due to the lack of fine-grained scheduling of diverse queries. We therefore propose Delphinus , a runtime system that efficiently manages resources for shared microservices while ensuring the Quality-of-Service (QoS). Delphinus comprises a group-oriented query scheduler and a borrowing-based load adapter . The query scheduler identifies diverse queries, groups the containers of shared microservices, and schedules the queries into separate groups. The load adapter efficiently scales resources for shared microservices, and fully utilizes the idle containers among groups when the loads of diverse queries change. Results show that Delphinus reduces CPU and memory usage by 40.1% and 36.4% for shared microservices, respectively, compared to state-of-the-art works. Jiuchen Shi, Jinyuan Chen, Quan Chen 0002, Kaihua Fu, Fanrong Du, Zijun Li 0001, Deze Zeng, Jiannong Cao 0001, Shuo Quan, Jie Wu 0001, Minyi Guo |
ACM Trans. Archit. Code Optim. | 4 |
| 2026 | QoS Awareness and Improved Throughput of Point Cloud Services With Dynamic WorkloadsabstractDeep learning on 3D point clouds plays a vital role in a wide range of applications such as AR/VR visualization, 3D cloth virtual try-on, and game rendering. As some applications require low latency, the point cloud services are also deployed on datacenter with powerful GPUs. While the queries of point cloud services show various workload change patterns due to different degrees of sparsity, current batching-based serving schemes result in either long latency or low throughput. We propose a scheme called Volans to address the above challenges and effectively support point cloud services. Volans comprises a workload predictor, a topology deployer, and a progress-aware scheduler. The predictor grids the input query and estimates the workload changes. Afterward, the deployer splits the model into several stages and determines the batch size for each stage based on the workload changes. The scheduler reduces the QoS violation when queries run slower due to unpredicted workload spikes. Experiments show that Volans enhances the peak supported throughput by up to 31.1% while maintaining the required 99%-ile latencies compared to state-of-the-art techniques. Kaihua Fu, Jiuchen Shi, Yao Chen 0008, Quan Chen 0002, Weng-Fai Wong, Wei Wang 0030, Bingsheng He, Minyi Guo |
IEEE Trans. Computers | 1 |
| 2025 | FaaSGNN: Enabling Memory Efficient and Low Latency GNN Inference Services with Serverless ComputingabstractWhile GNN-based services often experience load fluctuation, applying serverless computing to serve GNN inference reduces the cost and allows elastic resource scaling. However, GNN serverless shows poor performance due to heavy data fetching latency and long cold startup overhead, and our observation indicates opportunities for reducing data redundancy and mitigating cold startup latency. In this paper, we present FaaSGNN, a serverless GNN inference framework that enables low latency and memory efficient GNN serving through three key designs: (i) serverless-native on-demand graph fetching strategy that enables lightweight in-container graph sampling with full dataset resides in remote; (ii) memory-aware adaptive feature caching policy, which facilitates data reuse between requests to reduce redundant fetching; and (iii) load-aware request scheduler, which reschedules requests to bypass cold start and achieve load balance between containers. Experimental results show that FaaSGNN achieves a 5.6x lower end-to-end latency and 57.1% less memory usage on average compared to state-of-the-art works. Yuzhuo Yang, Kaihua Fu, Quan Chen 0002, Deze Zeng, Shuo Quan, Jie Wu 0001, Minyi Guo |
SoCC | 2 |
| 2024 | Adaptive QoS-Aware Microservice Deployment With Excessive Loads via Intra- and Inter-Datacenter SchedulingabstractUser-facing applications often experience excessive loads and are shifting towards the microservice architecture. To fully utilize heterogeneous resources, current datacenters have adopted the disaggregated storage and compute architecture, where the storage and compute clusters are suitable to deploy the stateful and stateless microservices, respectively. Moreover, when the local datacenter has insufficient resources to host excessive loads, a reasonable solution is moving some microservices to remote datacenters. However, it is nontrivial to decide the appropriate microservice deployment inside the local datacenter and identify the appropriate migration decision to remote datacenters, as microservices show different characteristics, and the local datacenter shows different resource contention situations. We therefore propose ELIS, an intra- and inter-datacenter scheduling system that ensures the Quality-of-Service (QoS) of the microservice application, while minimizing the network bandwidth usage and computational resource usage. ELIS comprises aresource manager, across-cluster microservice deployer, and areward-based microservice migrator. The resource manager allocates near-optimal resources for microservices while ensuring QoS. The microservice deployer deploys the microservices between the storage and compute clusters in the local datacenter, to minimize the network bandwidth usage while satisfying the microservice resource demand. The microservice migrator migrates some microservices to remote datacenters when local resources cannot afford the excessive loads. Experimental results show that ELIS ensures the QoS of user-facing applications. Meanwhile, it reduces the public network bandwidth usage, the remote computational resource usage, and the local network bandwidth usage by 49.6%, 48.5%, and 60.7% on average, respectively. Jiuchen Shi, Kaihua Fu, Quan Chen 0002, Deze Zeng, Minyi Guo |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | BLAD: Adaptive Load Balanced Scheduling and Operator Overlap Pipeline For Accelerating The Dynamic GNN TrainingabstractDynamic graph networks are widely used for learning time-evolving graphs, but prior work on training these networks is inefficient due to communication overhead, long synchronization, and poor resource usage. Our investigation shows that communication and synchronization can be reduced by carefully scheduling the workload. And the execution order of operators in GNNs can be adjusted without hurting training convergence. We propose a system called BLAD to consider the above factors, comprising a two-level load scheduler and an overlap-aware topology manager. The scheduler allocates each snapshot group to a GPU, alleviating cross-GPU communication. The snapshots in a group are then carefully allocated to processes on a GPU, enabling overlap of compute-intensive NN operators and memory-intensive graph operators. The topology manager adjusts the operators' execution order to maximize the overlap. Experiments show that BLAD achieves 27.2% speed up on training time on average without affecting final accuracy, compared to state-of-the-art solutions. Kaihua Fu, Quan Chen 0002, Yuzhuo Yang, Jiuchen Shi, Chao Li 0009, Minyi Guo |
SC | 1 |
| 2023 | Nodens: Enabling Resource Efficient and Fast QoS Recovery of Dynamic Microservice Applications in Datacenters
Jiuchen Shi, Zhixin Tong, Quan Chen 0002, Kaihua Fu, Minyi Guo |
USENIX ATC | 5 |
| 2022 | Astraea: towards QoS-aware and resource-efficient multi-stage GPU servicesabstractMulti-stage user-facing applications on GPUs are widely-used nowa- days, and are often implemented to be microservices. Prior re- search works are not applicable to ensuring QoS of GPU-based microservices due to the different communication patterns and shared resource contentions. We propose Astraea to manage GPU microservices considering the above factors. In Astraea, a microser- vice deployment policy is used to maximize the supported peak service load while ensuring the required QoS. To adaptively switch the communication methods between microservices according to different deployments, we propose an auto-scaling GPU communi- cation framework. The framework automatically scales based on the currently used hardware topology and microservice location, and adopts global memory-based techniques to reduce intra-GPU communication. Astraea increases the supported peak load by up to 82.3% while achieving the desired 99%-ile latency target compared with state-of-the-art solutions. Wei Zhang 0149, Quan Chen 0002, Kaihua Fu, Ningxin Zheng, Zhiyi Huang 0001, Jingwen Leng, Minyi Guo |
ASPLOS | 3 |
| 2022 | Characterizing and orchestrating VM reservation in geo-distributed clouds to improve the resource efficiencyabstractCloud providers often build a geo-distributed cloud from multiple datacenters in different geographic regions, to serve tenants at different locations. The tenants that run large scale applications often reserve resources based on their peak loads in the region close to the end users to handle the ever changing application load, wasting a large amount of resources. We therefore characterize the VM request patterns of the top tenants in our production public geo-distributed cloud, and open-source the VM request traces in four months from the top 20 tenants of our cloud. The characterization shows that the resource usage of large tenants has various temporal and spatial patterns on the dimensions of time series, regions, and VM types, and has the potential of peak shaving between different tenants to further reduce the resource reservation cost. Based on the findings, we propose a resource reservation and VM request scheduling scheme named ROS to minimize the resource reservation cost while satisfying the VM allocation requests. Our experiments show that ROS reduces the overall deployment cost by 75.4% and the reservation resources by 60.1%, compared to the tenant-specified reservation strategy. Jiuchen Shi, Kaihua Fu, Quan Chen 0002, Changpeng Yang, Mosong Zhou, Jieru Zhao, Chen Chen 0067, Minyi Guo |
SoCC | 2 |
| 2022 | QoS-awareness of Microservices with Excessive Loads via Inter-Datacenter SchedulingabstractUser-facing applications often experience excessive loads and are shifting towards microservice software architecture. While the local datacenter may not have enough resources to host the excessive loads, a reasonable solution is moving some microservices of the applications to remote datacenters. However, it is nontrivial to identify the appropriate migration decision, as the microservices show different characteristics, and the local datacenter also shows different resource contention situations. We therefore propose ELIS, an inter-datacenter scheduling system that ensures the required Quality-of-Service (QoS) of the microservice application with excessive loads, while minimizing the resource usage of the remote datacenter. ELIS comprises a resource manager and a reward-based microservice migrator. The resource manager finds the near-optimal resource configurations for different microservices to minimize resource usage while ensuring QoS. The microservice migrator migrates some microservices to remote datacenters when local resources cannot afford the excessive loads. Our experimental results show that ELIS ensures the required QoS of user-facing applications at excessive loads. Meanwhile, it reduces overall/remote resource usage by 13.1% and 58.1% on average, respectively. Jiuchen Shi, Kaihua Fu, Quan Chen 0002, Deze Zeng, Minyi Guo |
IPDPS | 3 |
| 2022 | QoS-Aware Irregular Collaborative Inference for Improving Throughput of DNN ServicesabstractWith collaborative DNN inference, part of queries run on their source edge device to reduce latencies. Because edges show diverse performance and network conditions, different layers should run on different devices, and queries on the datacenter show irregular structures. However, emerging schemes are not able to process such irregular queries. We propose ICE, a collaborative inference service scheme that effectively supports irregular queries. ICE comprises a query slicer, a query manager, and a lag enhancer. The query slicer maps the execution of queries based on the edges' performance and network conditions. The query manager batches irregular queries adaptively and schedules the irregular queries based on their progress. The lag enhancer reduces the QoS violation when queries run slower due to interference on the edge. Experiments show that ICE improves the supported peak load of the datacenter by 43.2% on average while guaranteeing the required 99%-ile latencies compared with state-of-the-art techniques. Kaihua Fu, Jiuchen Shi, Quan Chen 0002, Ningxin Zheng, Wei Zhang 0149, Deze Zeng, Minyi Guo |
SC | 1 |
| 2022 | Toward QoS-Awareness and Improved Utilization of Spatial Multitasking GPUsabstractDatacenters use GPUs to provide the significant computing throughput required by emerging user-facing services. The diurnal user access pattern of user-facing services provides a strong incentive to co-located applications for better GPU utilization, and prior work has focused on enabling co-location on multicore processors and traditional non-preemptive GPUs. However, current GPUs are evolving towards spatial multitasking and introduce a new set of challenges to eliminate QoS violations. To address this open problem, we explore the underlying causes of QoS violation on spatial multitasking GPUs. In response to these causes, we propose C-Laius, a runtime system that carefully allocates the computation resource to co-located applications for maximizing the throughput of batch applications while guaranteeing the required QoS of user-facing services. C-Laius not only allows co-locating one user-facing application with multiple batch applications, but also supports the co-location of multiple user-facing applications with batch applications. In the case of a single co-located user-facing application, our evaluation on an Nvidia RTX 2080Ti GPU shows that C-Laius improves the utilization of spatial multitasking GPUs by 20.8 percent, while achieving the 99%-ile latency target for user-facing services. As to the case of multiple co-located user-facing applications, C-Laius ensures no violation of QoS while improving the accelerator utilization by 35.9 percent on average. Wei Zhang 0149, Quan Chen 0002, Ningxin Zheng, Weihao Cui, Kaihua Fu, Minyi Guo |
IEEE Trans. Computers | 5 |
| 2022 | Adaptive Resource Efficient Microservice Deployment in Cloud-Edge ContinuumabstractUser-facing services are now evolving towards the microservice architecture where a service is built by connecting multiple microservice stages. Since the entire service is heavy, the microservice architecture shows the opportunity to only offload some microservice stages to the edge devices that are close to the end users. However, emerging techniques often result in the violation of Quality-of-Service (QoS) of microservice-based services in cloud-edge continuum, as they do not consider the communication overhead or the resource contention between microservices and external co-located tasks. We propose Nautilus, a runtime system that effectively deploys microservice-based user-facing services in cloud-edge continuum. Nautilus ensures the QoS of microservice-based user-facing services while minimizing the required computational resources, which is comprised of a communication-aware microservice mapper, a contention-aware resource manager and an IO-sensitive and load-aware microservice migration scheduler. The mapper divides the microservice graph into multiple partitions based on the communication overhead and maps the partitions to appropriate nodes. On each node, the resource manager determines the optimal resource allocation for its microservices based on reinforcement learning that may capture the complex contention behaviors. Once the microservices are suffered from external IO pressure, the IO-sensitive microservice scheduler migrates the critical one to idle nodes. Furthermore, when the load of microservices changes dynamically, the load-aware microservice scheduler migrates microservices from busy nodes to idle ones to ensure the QoS goal of the entire service. Our experimental results show that Nautilus can guarantee the required QoS target under external shared resources contention while the state-of-the-art suffers from QoS violations. Meanwhile, Nautilus reduces the computational resource usage by 23.9% and the network bandwidth usage by 53.4%, while achieving the required 99%-ile latency. Kaihua Fu, Wei Zhang 0149, Quan Chen 0002, Deze Zeng, Minyi Guo |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | CHARM: Collaborative Host and Accelerator Resource Management for GPU DatacentersabstractEmerging latency-critical (LC) services often have both CPU and GPU stages (e.g. DNN-assisted services) and require short response latency. Co-locating best-effort (BE) applications on the both CPU side and GPU side with the LC service improves resource utilization. However, resource contention often results in the QoS violation of LC services. We therefore present CHARM, a collaborative host-accelerator resource management system. CHARM ensures the required QoS target of DNN-assisted LC services, while maximizing the resource utilization of both the host and accelerator. CHARM is comprised of a BE-aware QoS target allocator, a unified heterogeneous resource manager, and a collaborative accelerator-side QoS compensator. The QoS target allocator determines the time limit of an LC service running on the host side and the accelerator side. The resource manager allocates the shared resources on both host side and accelerator side. The QoS compensator allocates more resources to the LC service to speed up its execution, if it runs slower than expected. Experimental results on an Nvidia GPU RTX 2080Ti show that CHARM improves the resource utilization by 43.2%, while ensuring the required QoS target compared with state-of-the-art solutions. Wei Zhang 0149, Kaihua Fu, Ningxin Zheng, Quan Chen 0002, Chao Li 0009, Wenli Zheng, Minyi Guo |
ICCD | 2 |
| 2021 | QoS-Aware and Resource Efficient Microservice Deployment in Cloud-Edge ContinuumabstractUser-facing services are now evolving towards the microservice architecture where a service is built by connecting multiple microservice stages. While an entire service is heavy, the microservice architecture shows the opportunity to only offload some microservice stages to the edge devices that are close to the end users. However, emerging techniques often result in the violation of Quality-of-Service (QoS) of microservice-based services in cloud-edge continuum, as they do not consider the communication overhead or the resource contention between microservices.We propose Nautilus, a runtime system that effectively deploys microservice-based user-facing services in cloud-edge continuum. It ensures the QoS of microservice-based user-facing services while minimizing the required computational resources. Nautilus is comprised of a communication-aware microservice mapper, a contention-aware resource manager and a load-aware microservice scheduler. The mapper divides the microservice graph into multiple partitions based on the communication overhead and maps the partitions to the nodes. On each node, the resource manager determines the optimal resource allocation for its microservices based on reinforcement learning that may capture the complex contention behaviors. The microservice scheduler monitors the QoS of the entire service, and migrates microservices from busy nodes to idle ones at runtime. Our experimental results show that Nautilus reduces the computational resource usage by 23.9% and the network bandwidth usage by 53.4%, while achieving the required 99%-ile latency. Kaihua Fu, Wei Zhang 0149, Quan Chen 0002, Deze Zeng, Xin Peng 0001, Wenli Zheng, Minyi Guo |
IPDPS | 1 |
| 2019 | Laius: Towards latency awareness and improved utilization of spatial multitasking accelerators in datacentersabstractDatacenters use accelerators to provide the significant compute throughput required by emerging user-facing services. The diurnal user access pattern of user-facing services provides a strong incentive to co-located applications for better accelerator utilization, and prior work has focused on enabling co-location on multicore processors and traditional non-preemptive accelerators. However, current accelerators are evolving towards spatial multitasking and introduce a new set of challenges to eliminate QoS violation. To address this open problem, we explore the underlying causes of QoS violation on spatial multitasking accelerators. In response to these causes, we propose Laius, a runtime system that carefully allocates the computation resource to co-located applications for maximizing the throughput of batch applications while guaranteeing the required QoS of user-facing services. Our evaluation on a Nvidia RTX 2080Ti GPU shows that Laius improves the utilization of spatial multitasking accelerators by 20.8%, while achieving the 99%-ile latency target for user-facing services. Wei Zhang 0149, Weihao Cui, Kaihua Fu, Quan Chen 0002, Daniel Mawhirter, Bo Wu 0002, Chao Li 0009, Minyi Guo |
ICS | 3 |
| 2007 | An Extended 3-D Radiosity-Graphics Combined Model for Studying Thermal-Emission Directionality of Crop CanopyabstractRadiosity-graphics combined model (RGM) has been proposed to calculate the radiation regime and bidirectional reflectance distribution function of complex 3D scene, which is limited in visible and near-infrared wavelength (0.3-3 mum) region. In this paper, RGM is extended to thermal region (named as TRGM) based on thermal-radiosity theory and thermal-emission directionality of vegetation canopy. The TRGM has been implemented on Microsoft Windows platform, and a parameterization scheme for crop canopies is introduced in this paper. It is then evaluated by comparing with two row-crop directional thermal emission models and one thermal radiative-transfer model. Field experiment data has been used to validate the TRGM for row structural wheat and maize canopies. The root mean square error of directional brightness temperature (DBT) is smaller than 1.0degC for the wheat canopy and 0.5degC for the maize canopy while the canopy DBTs vary more than 4degC. Model sensitivity analyses have also been conducted to illustrate influences of component temperature distribution, component emissivity, incident atmospheric radiation, and canopy structure on the crop canopy DBT. Qinhuo Liu, Huaguo Huang, Wenhan Qin, Kaihua Fu, Xiaowen Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |