VLDB 2026 Research / reviewers in the wild / expert
Zinuo Cai
dblp:299/5175
· DBLP profile ↗
18ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0001-9373-8474ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 9 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QFI-Opt: Communication-Efficient Quantum Federated Learning via Quantum Fisher InformationabstractABSTRACT Background Quantum federated learning presents a promising paradigm for privacy‐preserving collaborative training across distributed quantum devices. However, its scalability is hindered by the significant communication overhead associated with transmitting high‐dimensional, high‐precision quantum model parameters over classical networks. Methods To address this bottleneck, this paper proposes QFI‐Opt (Quantum Fisher Information‐guided Adaptive Optimization), a quantum adaptive communication optimization framework based on Quantum Fisher Information. QFI‐Opt establishes a “sensing‐compression‐regulation” pipeline that achieves communication efficiency while preserving quantum model fidelity. The framework uses QFI as a physically interpretable metric to dynamically assess quantum state sensitivity to parameter perturbations, enabling a progressive pruning strategy that removes low‐sensitivity parameters during training while retaining critical quantum features. Additionally, a dynamic bit‐width quantization mechanism adapts precision based on parameter importance, maximizing compression without compromising numerical stability. This is further complemented by a physics‐aware aggregation method that weighs client updates based on both local data volume and quantum information quality derived from QFI scores, improving global model robustness. Results Extensive evaluation on quantum convolutional neural networks demonstrates that QFI‐Opt significantly reduces per‐round communication overhead compared to baseline methods. Conclusions Simultaneously, the proposed framework maintains competitive model accuracy and convergence performance across diverse quantum architectures. Rui Zhang 0087, Zinuo Cai, Yicheng Di, Jiayu Bao, Jiansong Fan, Zhongle Qu |
Softw. Pract. Exp. | 4 |
| 2026 | HeShare: Energy-Aware and Efficient Multi-Task GPU Sharing in Heterogeneous GPU-Based Computing SystemsabstractWith the rapid growth of artificial intelligence and large-scale model computing, the demand for GPUs in datacenters continues to increase, especially for large-scale training and inference tasks. Heterogeneous multi-GPU systems, which integrate GPUs with varying types and computational capabilities, have become critical computing resources. This leads to two main challenges. First, due to the differences in GPU performance and power consumption, task scheduling involves a complex multi-objective optimization to balance energy efficiency and performance. More importantly, the lack of coordinated mechanisms for multi-task sharing and energy-efficient resource management across heterogeneous GPUs can result in GPU overload or underutilization, leading to wasted resources and potential system risks.To address these challenges, we propose HESHARE, an energy-aware and efficient heterogeneous GPU framework for datacenters. First, we design an energy-aware task scheduling strategy that optimizes task allocation across different GPUs to achieve a balance between energy consumption and performance. Second, we introduce a GPU sharing optimization mechanism that adaptively configures MPS and DVFS settings for each GPU, enhancing resource utilization, reducing overall energy consumption, and ensuring task performance. Compared to the state-of-the-art framework, we reduce average energy costs by 26% and improve job completion time by 31%, achieving a balance between energy efficiency and performance. Zhuolong Jiang, Zinuo Cai, Baoheng Zhang, Yiming Qiang, Ruhui Ma, Haibing Guan, Rajkumar Buyya |
IEEE Trans. Computers | 2 |
| 2026 | FaaShare: Enabling Generic Low-Latency and SLO-Aware GPU Sharing in Serverless Computing
Zhuolong Jiang, Zinuo Cai, Yumou Liu, Ruhui Ma, Rajkumar Buyya |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | AARC: Automated Affinity-aware Resource Configuration for Serverless WorkflowsabstractServerless computing is increasingly adopted for its ability to manage complex, event-driven workloads without the need for infrastructure provisioning. However, traditional resource allocation in serverless platforms couples CPU and memory, which may not be optimal for all functions. Existing decoupling approaches, while offering some flexibility, are not designed to handle the vast configuration space and complexity of serverless workflows. In this paper, we propose AARC, an innovative, automated framework that decouples CPU and memory resources to provide more flexible and efficient provisioning for serverless workloads. AARC is composed of two key components: Graph-Centric Scheduler, which identifies critical paths in workflows, and Priority Configurator, which applies priority scheduling techniques to optimize resource allocation. Our experimental evaluation demonstrates that AARC achieves substantial improvements over state-of-the-art methods, with total search time reductions of $85.8 \%$ and $89.6 \%$, and cost savings of $49.6 \%$ and $61.7 \%$, respectively, while maintaining SLO compliance. Lingxiao Jin, Zinuo Cai, Ruhui Ma |
DAC | 2 |
| 2025 | HWDSQP: A Historical Weighted and Dynamic Scheduling Quantum Protocol to Enhance Communication ReliabilityabstractQuantum computing holds the promise of solving problems difficult for classical computers. However, we are still in the era of Noisy Intermediate-Scale Quantum (NISQ) computers, necessary to establish effective distributed quantum communication protocols to distribute complex quantum computing tasks across different quantum computers for execution. Significant progress has been made in quantum communication technology, particularly in quantum path creation and resource scheduling. The establishment of quantum paths relies on quantum entanglement and quantum relay technologies, achieving long-distance, high-fidelity quantum state transmission through entanglement swapping between multiple relays. However, resources in quantum communication networks are limited and expensive, making efficient resource scheduling strategies crucial for improving overall network efficiency. To address these issues, we design a network protocol that includes the Historical Weighted Fidelity Routing (HWFR) algorithm and the Dynamic Multi-Priority Quantum Scheduling (DMPQS) algorithm to enhance communication reliability across quantum computers. Both algorithms aim to enhance the reliability of quantum links, optimize resource utilization, and adapt to dynamic changes in the links. The former algorithm dynamically selects the optimal path by considering factors such as link length, noise level, entanglement success rate, and quantum relay resource constraints, ensuring high-fidelity and reliable quantum communication. The latter dynamically adjusts request priorities based on the urgency of quantum service requests and fidelity requirements, optimizing resource utilization. Experimental results show that the proposed protocol performs excellently in terms of an average response time of requests and link utilization, effectively improving the utilization efficiency of network resources and the overall performance of the system. Rongbo Ma, Zejian Wang, Zinuo Cai, Haochen Xu, Baoheng Zhang, Ruhui Ma, Rajkumar Buyya |
IEEE J. Sel. Areas Commun. | 4 |
| 2025 | Ephemera: Accelerating I/O-Intensive Serverless Workloads with a Harvested In-memory File SystemabstractServerless computing has gained popularity for its ability to shift the burden of server management from developers to cloud providers, which allows providers to exercise greater control over resource management, optimizing configurations to enhance efficiency and performance. The diversity of serverless computing tasks, from short-lived, event-driven tasks to more complex workloads, highlights the growing importance of efficient file I/O performance for I/O-intensive workloads, yet effectively handling ephemeral storage for I/O-intensive tasks remains a challenge. Traditional file system approaches often introduce substantial latency and fail to fully leverage available memory resources within the execution environment, limiting performance and efficiency. Our work stems from the observation of the under-utilization of memory resources in serverless computing platforms and the potential efficiency improvement of I/O operations using an in-memory file system. Based on this observation, we propose Ephemera , a system designed to enhance ephemeral storage efficiency and memory utilization. Ephemera satisfies three design goals: transparent memory I/O integration , heterogeneous tasks resource synergy , and harmonized cluster workload orchestration . Ephemera integrates three components: the Runtime Daemon, responsible for managing a container’s in-memory file system; the Tenant Manager, facilitating memory configuration sharing across containers; and the Cluster Controller, optimizing workload balancing. Our experiments demonstrate that Ephemera significantly improves performance for I/O-intensive tasks compared to traditional file systems. Specifically, Ephemera decreases I/O processing time by 50% on average and reduces latency by up to 95.73% in certain scenarios with negligible overhead. Lingxiao Jin, Zinuo Cai, Haoxin Wang 0005, Zongpu Zhang, Ruhui Ma, Haibing Guan, Yuan Liu 0021, Rajkumar Buyya |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | MemoriaNova: Optimizing Memory-Aware Model Inference for Edge ComputingabstractIn recent years, deploying deep learning models on edge devices has become pervasive, driven by the increasing demand for intelligent edge computing solutions across various industries. From industrial automation to intelligent surveillance and healthcare, edge devices are being leveraged for real-time analytics and decision-making. Existing methods face two challenges when deploying machine learning models on edge devices. The first challenge is handling the execution order of operators with a simple strategy, which can lead to a potential waste of memory resources when dealing with directed acyclic graph structure models. The second challenge is that they usually process operators of a model one by one to optimize the inference latency, which may lead to the optimization problem getting trapped in local optima. We present MemoriaNova, comprising BTSearch and GenEFlow, to solve these two problems. BTSearch is a graph state backtracking algorithm with efficient pruning and hashing strategies designed to minimize memory overhead during inference and enlarge latency optimization search space. GenEFlow, based on genetic algorithms (GA), integrates latency modeling, and memory constraints to optimize distributed inference latency. This innovative approach considers a comprehensive search space for model partitioning, ensuring robust and adaptable solutions. We implement BTSearch and GenEFlow and test them on 11 deep-learning models with different structures and scales. The results show that BTSearch can reach 12% memory optimization compared with the widely used random execution strategy. At the same time, GenEFlow reduces inference latency by 33.9% in distributed systems with four-edge devices. Renjun Zhang, Tianming Zhang, Zinuo Cai, Dongmei Li 0008, Ruhui Ma, Rajkumar Buyya |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | FasDL: An Efficient Serverless-Based Training Architecture With Communication Optimization and Resource ConfigurationabstractDeploying distributed training workloads of deep learning models atop serverless architecture alleviates the burden of managing servers from deep learning practitioners. However, when supporting deep model training, the current serverless architecture faces the challenges of inefficient communication patterns and rigid resource configuration that incur subpar and unpredictable training performance. In this paper, we proposeFasDL, an efficient serverless-based deep learning training architecture to solve these two challenges.FasDLadopts a novel training frameworkK-REDUCEto release the communication overhead and accelerate the training. Additionally, FasDL builds a lightweight mathematical model forK-REDUCEtraining, offering predictable performance and supporting subsequent resource configuration. It achieves the optimal resource configuration by formulating an optimization problem related to system-level and application-level parameters and solving it with a pruning-based heuristic search algorithm. Extensive experiments on AWS Lambda verify a prediction accuracy over 94% and demonstrate performance and cost advantages over the state-of-art architecture LambdaML by up to 16.8% and 28.3% respectively. Xinglei Chen, Zinuo Cai, Hanwen Zhang 0019, Ruhui Ma, Rajkumar Buyya |
IEEE Trans. Computers | 2 |
| 2025 | SMore: Enhancing GPU Utilization in Deep Learning Clusters by Serverless-Based Co-Location Scheduling
Junhan Liu, Zinuo Cai, Yumou Liu, Hao Li 0142, Zongpu Zhang, Ruhui Ma, Rajkumar Buyya |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | Hermes: Memory-Efficient Pipeline Inference for Large Models on Edge DevicesabstractThe application of Transformer-based large models has achieved numerous success in recent years. However, the exponential growth in the parameters of large models introduces formidable memory challenge for edge deployment. Prior works to address this challenge mainly focus on optimizing the model structure and adopting memory swapping methods. However, the former reduces the inference accuracy, and the latter raises the inference latency. This paper introduces PIPELoAD, a novel memory-efficient pipeline execution mechanism. It reduces memory usage by incorporating dynamic memory management and minimizes inference latency by employing parallel model loading. Based on PIPELoAD mechanism, we present Hermes, a framework optimized for large model inference on edge devices. We evaluate Hermes on Transformer-based models of different sizes. Our experiments illustrate that Hermes achieves up to 4.24 x increase in inference speed and 86.7% lower memory consumption than the state-of-the-art pipeline mechanism for BERT and ViT models, 2.58 x increase in inference speed and 90.3% lower memory consumption for GPT-style models. Xueyuan Han, Zinuo Cai, Yichu Zhang, Chongxin Fan, Junhan Liu, Ruhui Ma, Rajkumar Buyya |
ICCD | 2 |
| 2024 | SMSS: Stateful Model Serving in Metaverse With Serverless Computing and GPU SharingabstractWith the rapid development of information technology, the concept of the Metaverse has swept the world and set off a new wave of the industrial revolution. The construction of living and manufacturing scenes based on the Metaverse requires the joint participation of scientists and engineers from various fields where “human” is at the core. In the Metaverse, predicting human behavior and response based on the deep learning model is meaningful because the prediction results can provide more satisfactory services for participants. Therefore, how to deploy a multi-stage machine learning reasoning model has become the bottleneck to improving the development level of Metaverse. Thanks to its scalability and pay-as-you-go billing model, the emerging serverless computing can effectively cope with the workload of machine learning inference. However, the statelessness of serverless computing and the lack of good GPU resource-sharing support make it difficult to deploy the machine learning model directly on the serverless computing platform to play its advantages. Therefore, we propose SMSS, a stateful model inference service, which is deployed on a serverless computing platform that supports GPU sharing. Since the serverless computing platform does not support stateful workflow execution, SMSS adopts log-based workflow runtime support. We also design a mechanism of two-layer GPU sharing to fully explore the potential of inter-model and intra-model GPU sharing. We evaluate the effectiveness of SMSS with real workloads. Our experimental results show that log-based stateful workflow operation support can ensure the stateful execution of tasks with low overhead but facilitate error location and recovery. Two-layer GPU Sharing can reduce the cold start time of inference tasks to two orders of magnitude at most. Zinuo Cai, Ruhui Ma, Haibing Guan |
IEEE J. Sel. Areas Commun. | 1 |
| 2024 | SLoB: Suboptimal Load Balancing Scheduling in Local Heterogeneous GPU Clusters for Large Language Model InferenceabstractLarge language models (LLMs) are becoming powerful engines for social productivity in the manufacturing lifecycle. Existing application-level LLMs inference services focus on large datacenter and small edge intelligence (EI) scenarios, adopting iteration-level batch schedulers to solve resource utilization and inference speed problems. However, these services are incompatible with the scene of medium-sized local heterogeneous graphics processing unit (GPU) clusters with specific patterns, whose scale is between the two aforementioned scenarios. This type of scene proposes tradeoff problems for inference resource and speed, as well as user satisfaction problems for the semisparse frequency of queries with streaming responses. We propose suboptimal load balancing (SLoB), a distributed LLMs inference service scheduler in medium-sized local heterogeneous GPU clusters. SLoB leverages a multilevel adapter to accommodate LLMs usage patterns of scenes and balance resource utilization with inference efficiency. For semisparse problems, it adopts a mixed-priority pipeline scheduler with the least-padding principle to improve users’ satisfaction, a metric considering the weights of different tokens in streaming responses. Based on the system prototype, our experiments under simulated workloads demonstrate that SLoB gains a maximum improvement of 29.4$\times$under the satisfaction metric compared with the traditional run-to-completion scheduling solution while improving by up to 3.0$\times$compared with the state-of-the-art (SOTA) solution Orca. Peiwen Jiang, Haoxin Wang 0005, Zinuo Cai, Lintao Gao, Weishan Zhang, Ruhui Ma, Xiaokang Zhou |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | SPSC: Stream Processing Framework Atop Serverless Computing for Industrial Big DataabstractWith the advance of smart manufacturing and information technologies, the volume of data to process is increasing accordingly. Current solutions for big data processing resort to distributed stream processing systems, such as Apache Flink and Spark. However, such frameworks face challenges of resource underutilization and high latency in big data application scenarios. In this article, we propose SPSC, a serverless-based stream computing framework where events are discretized into the atomic stream and stateless Lambda functions are taken as context-irrelevant operators, achieving task parallelism and inherent data parallelism in processing. Also, we implement a prototype of the framework on Amazon Web service (AWS) using AWS Lambda, AWS simple queue service, and AWS DynamoDB. The evaluation shows that compared with Alibaba's real-time computing Flink version, SPSC outperforms by 10.12% when the overhead is close. Zinuo Cai, Xinglei Chen, Ruhui Ma, Haibing Guan, Rajkumar Buyya |
IEEE Trans. Cybern. | 1 |
| 2024 | RIDIC: Real-Time Intelligent Transportation System With Dispersed ComputingabstractModern transportation big data features high Volume, Velocity, and Variety, making it more and more challenging to develop an intelligent transportation system for data analysis. Current transportation systems resort to cloud computing to deploy their applications but face two bottlenecks, lack of real-time processing and under-utilization of intelligent roadside devices. We observe that dispersed computing—an emerging paradigm of cloud computing—well fits the requirements of modern transportation systems. It provides real-time response by alleviating data transmission between data sources and cloud servers and fully utilizes smart devices by exploring their computing capacity. Therefore, we design RIDIC, an intelligent transportation system with dispersed computing to provide a real-time response when processing transportation big data. RIDIC abstracts all the heterogeneous smart roadside devices as actors, and its workflow consists of three stages, Actor Registration, Resource Application and Task Execution. We conduct experiments on two real-life traffic scenarios—road vehicle detection and traffic signal recognition—and the results show that RIDIC can utilize edge devices to process transportation big data faster while reducing the demand for device computing resources. Zinuo Cai, Quanmin Xie, Ruhui Ma, Haibing Guan |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Sustainable Serverless Computing With Cold-Start Optimization and Automatic Workflow Resource SchedulingabstractIn recent years, serverless computing has garnered significant attention owing to its high scalability, pay-as-you-go billing model, and efficient resource management provided by cloud service providers. Optimal resource scheduling of serverless computing has become imperative to reduce energy consumption and enable sustainable computing. However, existing serverless platforms encounter two significant challenges: the cold-start problem of containers and the absence of an effective resource allocation strategy for serverless workflows. Existing pre-warm strategies are associated with high computational overhead, while current resource scheduling techniques inadequately account for the intricate structure of serverless workflows. To address these challenges, we present SSC, a pre-warming and automatic resource allocation framework designed explicitly for serverless workflows. We introduce an innovative gradient-based algorithm for pre-warming containers, significantly reducing cold start hit rates. Moreover, leveraging a critical path and priority queue-based algorithm, SSC enables efficient allocation of resources for serverless workflows. In our experimental evaluation, SSC reduces the cold start hit rate by nearly$50\%$and achieves substantial cost savings of approximately$30\%$. Shanxing Pan, Zinuo Cai, Dongmei Li 0008, Ruhui Ma, Haibing Guan |
IEEE Trans. Sustain. Comput. | 3 |
| 2023 | Towards Variance Reduction for Reinforcement Learning of Industrial Decision-making Tasks: A Bi-Critic based Demand-Constraint Decoupling ApproachabstractLearning to plan and schedule receives increasing attention due to its efficiency in problem-solving and potential to outperform heuristics. In particular, actor-critic-based reinforcement learning (RL) has been widely adopted for uncertain environments. Yet one standing challenge for applying RL to real-world industrial decision-making problems is the high variance during training. Existing efforts design novel value functions to alleviate the issue but still suffer. In this paper, we address this issue from the perspective of adjusting the actor-critic paradigm. We start by making an observation ignored in many industrial problems---the environmental dynamics for an agent consist of two parts physically independent of each other: the exogenous task demand over time and the hard constraint for action. And we theoretically show that decoupling these two effects in the actor-critic technique would reduce variance. Accordingly, we propose to decouple and model them separately in the state transition of the Markov decision process (MDP). In the demand-encoding process, the temporal task demand, e.g., the passengers for elevator scheduling is encoded followed by a critic for scoring. While in the constraint-encoding process, an actor-critic module is adopted for action embedding, and the two critics are then used for a revised advantaged function calculation. Experimental results show that our method can adaptively handle different dynamic planning and scheduling tasks and outperform recent learning-based models and traditional heuristic algorithms. Jianyong Yuan, Jiayi Zhang 0003, Zinuo Cai, Junchi Yan |
KDD | 3 |
| 2023 | GUARDIAN: A Hardware-Assisted Distributed Framework to Enhance Deep Learning SecurityabstractThe ubiquity of artificial intelligence (AI) has led to its extensive research and application in various fields, such as computer vision, natural language processing, and medical image analysis. However, responsible AI faces severe security challenges, including the leakage of pretrained models and valuable training data. The existing solutions adopt new algorithm designs (such as federated learning) or cryptography (such as homomorphic encryption) to prevent possible security vulnerabilities. We observe that hardware-assisted trusted execution environments (TEEs) can further improve machine learning responsibility. Intel Software Guard Extension (SGX) is a popular, trusted execution hardware that enables users’ programs to run in an untrusted execution environment, such as a malicious operating system, but ensures the confidentiality and integrity of data. Therefore, we have designed GUARDIAN, a hardware-assisted secure machine learning training framework that protects data security during the training process. We have analyzed the typical characteristics of machine learning applications and characterized GUARDIAN through extensive experiments. Our findings demonstrate that introducing security guarantees causes performance degradation, which provides a feasible optimization direction in the near future. Zinuo Cai, Bojun Ren, Ruhui Ma, Haibing Guan, Mengke Tian, Yong Wang 0085 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2021 | Themis: A Fair Evaluation Platform for Computer Vision CompetitionsabstractIt has become increasingly thorny for computer vision competitions to preserve fairness when participants intentionally fine-tune their models against the test datasets to improve their performance. To mitigate such unfairness, competition organizers restrict the training and evaluation process of participants' models. However, such restrictions introduce massive computation overheads for organizers and potential intellectual property leakage for participants. Thus, we propose Themis, a framework that trains a noise generator jointly with organizers and participants to prevent intentional fine-tuning by protecting test datasets from surreptitious manual labeling. Specifically, with the carefully designed noise generator, Themis adds noise to perturb test sets without twisting the performance ranking of participants' models. We evaluate the validity of Themis with a wide spectrum of real-world models and datasets. Our experimental results show that Themis effectively enforces competition fairness by precluding manual labeling of test sets and preserving the performance ranking of participants' models. Zinuo Cai, Jianyong Yuan, Yang Hua 0001, Tao Song 0003, Hao Wang 0022, Zhengui Xue, Ningxin Hu, Jonathan Ding, Ruhui Ma, Mohammad R. Haghighat, Haibing Guan |
IJCAI | 1 |