VLDB 2026 Research / reviewers in the wild / expert
Zijun Li 0001
dblp:44/10301-1
· DBLP profile ↗
15ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0003-4706-8451ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 6 first-author · 14 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Resource-Efficient Serverless LLM Inference with SLINFERabstractThe rise of LLMs has driven demand for private serverless deployments, characterized by moderate-sized models and infrequent requests. While existing serverless solutions follow exclusive GPU allocation, we take a step back to explore modern platforms and find that: Emerging CPU architectures with built-in accelerators are capable of serving LLMs but remain underutilized, and both CPUs and GPUs can accommodate multiple LLMs simultaneously. We propose SLINFER, a resource-efficient serverless inference scheme tailored for small- to mid-sized LLMs that enables elastic and on-demand sharing across heterogeneous hardware. SLINFER tackles three fundamental challenges: (1) precise, fine-grained compute resource allocation at token-level to handle fluctuating computational demands; (2) a coordinated and forward-looking memory scaling mechanism to detect out-ofmemory hazards and reduce operational overhead; and (3) a dual approach that consolidates fragmented instances through proactive preemption and reactive bin-packing. Experimental results on 4 32-core CPUs and 4 A100 GPUs show that SLINFER improves serving capacity by 47% - 62% through sharing, while further leveraging CPUs boosts this to 86% - 154%. Chuhao Xu, Zijun Li 0001, Quan Chen 0002, Han Zhao 0005, Xueyan Tang, Minyi Guo |
HPCA | 2 |
| 2026 | LEGO: Supporting LLM-Enhanced Games with One Gaming GPUabstractArtificial intelligence (AI) has been increasingly applied to gaming, with large language models (LLMs) playing a key role in character control. However, efficiently co-locating game rendering and LLM inference on one GPU presents challenges due to resource constraints, diverse latency requirements, and fine-grained task scheduling. We propose LEGO, an algorithm-system co-design that enables the efficient co-location of LLM inference and game rendering tasks. Algorithmwise, LEGO features a resource-oriented layer-skipping adaptor, which distills knowledge from skipped layers to reduce computational demand while maintaining inference accuracy. System-wise, LEGO proposes a headroom-maximizing LLM scheduler, which dynamically partitions inference tasks to utilize available rendering headroom. Evaluations on an Nvidia RTX 4090 show that LEGO meets latency targets in all scenarios, improves rendering headroom utilization by up to 28.6 %, and reduces LLM inference accuracy loss by up to 86.3 % compared to current layer-skipping approaches. Han Zhao 0005, Weihao Cui, Zeshen Zhang, Jiangtong Li, Quan Chen 0002, Pu Pang, Zijun Li 0001, Zhenhua Han, Yuqing Yang 0001, Minyi Guo |
HPCA | 8 |
| 2026 | Delphinus: Improving Resource Efficiency of Applications with Shared Microservices and Diverse QueriesabstractMicroservices are widely shared in production user-facing applications. These shared microservices have various resource usage patterns when queries from different call graphs of different services access them. However, existing microservice management works fail to efficiently scale resources for them, mainly due to the lack of fine-grained scheduling of diverse queries. We therefore propose Delphinus , a runtime system that efficiently manages resources for shared microservices while ensuring the Quality-of-Service (QoS). Delphinus comprises a group-oriented query scheduler and a borrowing-based load adapter . The query scheduler identifies diverse queries, groups the containers of shared microservices, and schedules the queries into separate groups. The load adapter efficiently scales resources for shared microservices, and fully utilizes the idle containers among groups when the loads of diverse queries change. Results show that Delphinus reduces CPU and memory usage by 40.1% and 36.4% for shared microservices, respectively, compared to state-of-the-art works. Jiuchen Shi, Jinyuan Chen, Quan Chen 0002, Kaihua Fu, Fanrong Du, Zijun Li 0001, Deze Zeng, Jiannong Cao 0001, Shuo Quan, Jie Wu 0001, Minyi Guo |
ACM Trans. Archit. Code Optim. | 6 |
| 2025 | ServScale: Concurrency-Aware Serverless Execution and Scaling Paradigm
Zichen Xu 0008, Zijun Li 0001, Quan Chen 0002, Minyi Guo |
NPC (1) | 2 |
| 2025 | EDAS: Enabling Fast Data Loading for GPU Serverless ComputingabstractIntegrating GPUs into serverless computing platforms is crucial for improving efficiency. Many GPU functions, such as DNN inferences and scientific services, benefit from GPU usage, which requires only tens to hundreds of milliseconds for pure computation. Under these circumstances, fast data loading is imperative for function performance. However, existing GPU serverless systems face significant data stall issues, leading to extremely low GPU efficiency. Faced with the above problems, we observe opportunities to optimize data loading, such as data preloading and deduplicated data loading. However, these optimizations are impossible in existing GPU serverless systems due to the lack of insights into data information, such as data sizes and read-write attributes of function inputs. To address this, we propose a novel GPU serverless system, EDAS. EDAS first enhances user request specifications, allowing users to annotate data retrieved by GPU functions from the database with additional attributes. Based on this, EDAS takes over data loading from GPU functions and proposes two innovative data loading management schemes: a parallelized data loading scheme and a multi-stage resource exit scheme. Our experimental results show that EDAS reduces function duration by 16.2× and improves system throughput by 1.91× compared with the state-of-the-art serverless platform. Han Zhao 0005, Weihao Cui, Quan Chen 0002, Zijun Li 0001, Zhenhua Han, Yu Feng 0007, Jieru Zhao, Chen Chen 0067, Jingwen Leng, Minyi Guo |
ACM Trans. Archit. Code Optim. | 4 |
| 2025 | Lightweight and Holistic-Scalable Serverless Secure Container Runtime for High-Density Deployment and High-Concurrency StartupabstractThe secure container that hosts a single container in a micro virtual machine (VM) is now used in serverless computing, as the containers are isolated through the microVMs. There are high demands on the high-density container deployment and high-concurrency container startup to improve both the resource utilization and user experience, as user functions are fine-grained in serverless platforms. Our investigation shows that the entire software stacks, containing the cgroups in the host operating system, the guest operating system, and the containerrootfsfor the function workload, together result in low deployment density and slow startup performance at high-concurrency.We propose a lightweight and holistic-scalable secure container runtime, named RunD-V, to resolve above problems in serverless computing. RunD-V proposes a guest-to-host runtime template for microVM scaling-out, and CR-bind feature in guest kernel for microVM scaling-up. Using guest-to-host runtime template, over 200 secure containers can be launched within 1son a node equipped with 104 vCPUs. It also enables more than 2,500 secure containers to be deployed on a node with 384GB of memory. The vertical scaling mechanism CR-bind further enhances both startup concurrency and deployment density. Zijun Li 0001, Chuhao Xu, Quan Chen 0002, Shuo Quan, Bin Zha, Weidong Han 0003, Jie Wu 0001, Minyi Guo |
IEEE Trans. Computers | 1 |
| 2024 | FaaSGraph: Enabling Scalable, Efficient, and Cost-Effective Graph Processing with Serverless ComputingabstractGraph processing is widely used in cloud services; however, current frameworks face challenges in efficiency and cost-effectiveness when deployed under the Infrastructure-as-a-Service model due to its limited elasticity. In this paper, we present FaaSGraph, a serverless-native graph computing scheme that enables efficient and economical graph processing through the co-design of graph processing frameworks and serverless computing systems. Specifically, we design a data-centric serverless execution model to efficiently power heavy computing tasks. Furthermore, we carefully design a graph processing paradigm to seamlessly cooperate with the data-centric model. Our experiments show that FaaS-Graph improves end-to-end performance by up to 8.3X and reduces memory usage by up to 52.4% compared to state-of-the-art IaaS-based methods. Moreover, FaaSGraph delivers steady 99%-ile performance in highly fluctuated workloads and reduces monetary cost by 85.7%. Yushi Liu 0003, Shixuan Sun, Zijun Li 0001, Quan Chen 0002, Bingsheng He, Chao Li 0009, Minyi Guo |
ASPLOS (2) | 3 |
| 2024 | FaaSMem: Improving Memory Efficiency of Serverless Computing with Memory Pool ArchitectureabstractIn serverless computing, an idle container is not recycled directly, in order to mitigate time-consuming cold container startup. These idle containers still occupy the memory, exasperating the memory shortage of today's data centers. By offloading their cold memory to remote memory pool could potentially resolve this problem. However, existing offloading policies either hurt the Quality of Service (QoS) or are too coarse-grained in serverless computing scenarios. Chuhao Xu, Yiyu Liu, Zijun Li 0001, Quan Chen 0002, Han Zhao 0005, Deze Zeng, Xueqi Wu, Senbo Fu, Minyi Guo |
ASPLOS (3) | 3 |
| 2023 | DataFlower: Exploiting the Data-flow Paradigm for Serverless Workflow OrchestrationabstractServerless computing that runs functions with auto-scaling is a popular task execution pattern in the cloud-native era. By connecting serverless functions into workflows, tenants can achieve complex functionality. Prior research adopts the control-flow paradigm to orchestrate a serverless workflow. However, the control-flow paradigm inherently results in long response latency, due to the heavy data persistence overhead, sequential resource usage, and late function triggering. Zijun Li 0001, Chuhao Xu, Quan Chen 0002, Jieru Zhao, Chen Chen 0067, Minyi Guo |
ASPLOS (4) | 1 |
| 2023 | Maximizing the Utilization of GPUs Used by Cloud Gaming through Adaptive Co-location with ComboabstractCloud vendors are now providing cloud gaming services with GPUs. GPUs in cloud gaming experience periods of idle because not every frame in a game always keeps the GPU busy for rendering. Previous works temporally co-locate games with best-effort applications to harvest these idle cycles. However, these works ignore the spatial sharing of GPUs, leading to not maximized throughput improvement. The newly introduced RT (ray tracing) Cores inside GPU SMs for ray tracing exacerbate the situation. Binghao Chen, Han Zhao 0005, Weihao Cui, Yifu He, Shulai Zhang, Quan Chen 0002, Zijun Li 0001, Minyi Guo |
SoCC | 7 |
| 2023 | Microless: Cost-Efficient Hybrid Deployment of Microservices on IaaS VMs and ServerlessabstractMicroservices have gained popularity as an architectural approach for developing scalable and modular applications. Traditionally, microservice deployment relies on virtual machines (VMs) from Infrastructure-as-a-Service (IaaS) computing. However, the emerging serverless computing offers new possibilities for more scalable microservice deployment. In this paper, we provide insights into the optimal scenarios for IaaS VMs and serverless, and investigate the challenges in the programming model and invocation pattern. We propose Microless, a framework that achieves the hybrid deployment of microservices on serverless and IaaS VMs and overcomes the challenges. In Microless, the steady workload is processed on IaaS VMs, ensuring optimal resource utilization and run-time performance. For the fluctuating workload, serverless can rapidly scale out resources to handle burst requests, minimizing response latency and enhancing cost-effectiveness. Experimental results validate the effectiveness of Microless in runtime performance and deployment cost. Jiagan Cheng, Zijun Li 0001, Quan Chen 0002, Weihao Cui, Minyi Guo |
ICPADS | 3 |
| 2022 | FaaSFlow: enable efficient workflow execution for function-as-a-serviceabstractServerless computing (Function-as-a-Service) provides fine-grain resource sharing by running functions (or Lambdas) in containers. Data-dependent functions are required to be invoked following a pre-defined logic, which is known as serverless workflows. However, our investigation shows that the traditional master-worker based workflow execution architecture performs poorly in serverless context. One significant overhead results from the master-side workflow schedule pattern, with which the functions are triggered in the master node and assigned to worker nodes for execution. Besides, the data movement between workers also reduces the throughput. Zijun Li 0001, Yushi Liu 0003, Linsong Guo, Quan Chen 0002, Jiagan Cheng, Wenli Zheng, Minyi Guo |
ASPLOS | 1 |
| 2022 | RunD: A Lightweight Secure Container Runtime for High-density Deployment and High-concurrency Startup in Serverless Computing
Zijun Li 0001, Jiagan Cheng, Quan Chen 0002, Eryu Guan, Zizheng Bian, Bin Zha, Weidong Han 0003, Minyi Guo |
USENIX ATC | 1 |
| 2022 | Help Rather Than Recycle: Alleviating Cold Startup in Serverless Computing Through Inter-Function Container Sharing
Zijun Li 0001, Linsong Guo, Quan Chen 0002, Jiagan Cheng, Chuhao Xu, Deze Zeng, Tao Ma 0006, Yong Yang 0013, Chao Li 0009, Minyi Guo |
USENIX ATC | 1 |
| 2020 | Amoeba: QoS-Awareness and Reduced Resource Usage of Microservices with Serverless ComputingabstractWhile microservices that have stringent Quality-of-Service constraints are deployed in the Clouds, the long-term rented infrastructures that host the microservices are under-utilized except peak hours due to the diurnal load pattern. It is resource efficient for Cloud vendors and cost efficient for service maintainers to deploy the microservices in the long-term infrastructure at high load and in the serverless computing platform at low load. However, prior work fails to take advantage of the opportunity, because the contention between microservices on the serverless platform seriously affects their response latencies.Our investigation shows that the load of a microservice, the shared resource contentions on the serverless platform, and its sensitivities to the contention together affect the response latency of the microservice on the platform. To this end, we propose Amoeba, a runtime system that dynamically switches the deployment of a microservice. Amoeba is comprised of a contention-aware deployment controller, a hybrid execution engine, and a multi-resource contention monitor. The deployment controller predicts the tail latency of a microservice based on its load and the contention on the serverless platform, and determines the appropriate deployment of the microservice. The hybrid execution engine enables the quick switch of the two deploy modes. The contention monitor periodically quantifies the contention on multiple types of shared resources. Experimental results show that Amoeba is able to significantly reduce up to 72.9% of CPU usage and up to 84.9% of memory usage compared with the traditional pure IaaS-based deployment, while ensuring the required latency target. Zijun Li 0001, Quan Chen 0002, Tao Ma 0006, Yong Yang 0013, Minyi Guo |
IPDPS | 1 |