VLDB 2026 Research / reviewers in the wild / expert
Binlei Cai
dblp:182/7719
· DBLP profile ↗
11ranked-venue papers
9as first author
7since 2021 · last 2024
0000-0002-0591-7032ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 first-author · 2 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A self-stabilizing and auto-provisioning orchestration for microservices in edge-cloud continuum
Binlei Cai, Meihong Yang |
Comput. Networks | 1 |
| 2023 | Balanar: Balancing deadline guarantee and Jain's fairness for inter-datacenter transfers
Xiaodong Dong, Binlei Cai |
Comput. Networks | 2 |
| 2023 | AutoMan: Resource-efficient provisioning with tail latency guarantees for microservices
Binlei Cai, Meihong Yang |
Future Gener. Comput. Syst. | 1 |
| 2023 | AutoInfer: Self-Driving Management for Resource-Efficient, SLO-Aware Machine=Learning Inference in GPU ClustersabstractAs Internet of Things (IoT) keeps growing, IoT-side intelligence services, such as intelligent personal assistant, healthcare surveillance, and smart home service, offload more and more complex machine-learning (ML) inference workloads to cloud clusters. GPUs have been widely adopted to accelerate the execution of these ML inference workloads. However, current cluster management systems guarantee low tail latency for ML inferences using resource over-provisioning and small batch sizes, resulting in a serious waste of GPU resources and increasing the service costs greatly. To mitigate poor GPU utilization, we present AutoInfer, a self-driving cluster management system for ML inference serving in GPU clusters, where users express only the latency and accuracy requirements for their workloads without needing to specify the model variant, GPU provisioning strategy, and batching mechanism. AutoInfer extends the matrix factorization model to automatically recommend model variants for each new incoming ML inference workload with respect to latency and accuracy requirements, by identifying similarities to previously scheduled workloads. During runtime, AutoInfer leverages online telemetry data and deep reinforcement learning to adaptively adjust the GPU allocation and batch size to account for load variations while minimizing the effects on tail latency service level objectives (SLOs). Testbed experiments show that AutoInfer is able to improve the average GPU utilization by up to 77% and keep the tail latency SLO violations under 5.5%. Binlei Cai, Xiaodong Dong |
IEEE Internet Things J. | 1 |
| 2022 | LraSched: Admitting More Long-Running Applications via Auto-Estimating Container Size and AffinityabstractAbstract Many long-running applications (LRAs) are increasingly using containerization in shared production clusters. To achieve high resource efficiency and LRA performance, one of the key decisions made by existing cluster schedulers is the placement of LRA containers within a cluster. However, they fail to account for estimating the size and affinity of LRA containers before executing placement. We present LraSched, a cluster scheduler that places LRA containers onto machines based on their sizes and affinities while providing consistently high performance. LraSched introduces an automated method that leverages historical data and collects new information to estimate container size and affinity for an LRA. Specifically, it uses an online machine learning method to map a new incoming LRA to previous workloads from which we can transfer experience and recommends the amount of resources (size) and the degree of collocation (affinity) for the containers of the new incoming LRA. By means of recommendations, LraSched adapts the heuristic for vector bin packing to LRA scheduling and places LRA containers in a manner that both maximizes the number of LRAs deployed and minimizes the resource fragmentation, but without affecting LRA performance. Testbed and simulation experiments show that LraSched can improve the resource utilization by up to 6.2% while meeting performance constraints for LRAs. Binlei Cai, Junfeng Yu |
Comput. J. | 1 |
| 2022 | Slardar: Scheduling information incomplete inter-datacenter deadline-aware coflows with a decentralized framework
Xiaodong Dong, Binlei Cai |
Comput. Networks | 2 |
| 2022 | Less Provisioning: A Hybrid Resource Scaling Engine for Long-Running Services With Tail Latency GuaranteesabstractModern resource management frameworks guarantee low tail latency for long-running services using the resource over-provisioning method, resulting in serious waste of resources and increasing the service costs greatly. To reduce the over-provisioning cost, we present HRSE, a hybrid resource scaling engine that enables much more efficient resource provisioning for both periodic and non-periodic workloads of long-running services while guaranteeing the tail latency Service Level Objective (SLO). HRSE employs a convolution-based time series analysis to identify periodic patterns in workloads. If periodic patterns are discovered, HRSE estimates the just-right amount of resources based on the periodic features through atop-$K$based collaborative filtering approach. Otherwise, it leverages wavelet-clustering to capture the short-term patterns in non-periodic workloads and predict the resource demands for the near future. To further enforce the tail latency SLO, HRSE uses an online reprovisioning mechanism that dynamically adjusts the resources to mitigate the performance uncertainty due to workload burstinesses. We fully implement HRSE on top of Docker and conduct extensive experiments using traces from production systems. Testbed experiments show that HRSE is able to increase the average resource utilization to 43 and 45 percent for periodic and non-periodic workloads respectively while guaranteeing the same tail latency objective. Binlei Cai, Keqiu Li, Laiping Zhao, Rongqi Zhang |
IEEE Trans. Cloud Comput. | 1 |
| 2019 | SLO-aware colocation: Harvesting transient resources from latency-critical services
Binlei Cai, Keqiu Li |
J. Syst. Archit. | 1 |
| 2019 | On evaluating the resource usage effectiveness of multi-tenant cloud storage
Binlei Cai, Laiping Zhao, Xiaobo Zhou 0003, Rongqi Zhang, Keqiu Li |
J. Syst. Archit. | 1 |
| 2018 | Less Provisioning: A Fine-grained Resource Scaling Engine for Long-running Services with Tail Latency GuaranteesabstractModern resource management frameworks guarantee low tail latency for long-running services using the resource over-provisioning method, resulting in serious waste of resource and increasing the service costs greatly. To reduce the over-provisioning cost, we present EFRA, an elastic and fine-grained resource allocator that enables much more efficient resource provisioning while guaranteeing the tail latency Service Level Objective (SLO). EFRA achieves this through the cooperation of three key components running on a containerized platform: The period detector identifies the period features of the workload through a convolution-based time series analysis. The resource reservation component estimates the just-right amount of resources based on the period analysis through a top-K based collaborative filtering approach. The online reprovisioning component dynamically adjusts the resources for further enforcing the tail latency SLO. Testbed experiments show that EFRA is able to increase the average resource utilization to 43%, and save up to 66% resources while guaranteeing the same tail latency objective. Binlei Cai, Rongqi Zhang, Laiping Zhao, Keqiu Li |
ICPP | 1 |
| 2017 | Experience Availability: Tail-Latency Oriented Availability in Software-Defined Cloud Computing
Binlei Cai, Rongqi Zhang, Xiaobo Zhou 0003, Laiping Zhao, Keqiu Li |
J. Comput. Sci. Technol. | 1 |