VLDB 2026 Research / reviewers in the wild / expert
Luyao Luo
dblp:176/7592
· DBLP profile ↗
14ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-6255-4370ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 4 first-author · 7 since 2021Systems, architecture and hardware · 6 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Focus: Fault-Tolerant Online Cloud Utility Scheduling with Imperfect-Data-Driven StrategiesabstractAchieving load balancing under dynamic workloads remains a critical challenge in cloud computing. However, due to inherent prediction inaccuracies and task load volatility, existing static or dynamic scheduling approaches struggle to reconcile real-time responsiveness with long-term resource efficiency, particularly when facing imprecise prediction, declared uncertainty, and data drift. To address these challenges, we propose Focus, which introduces three key innovations to achieve fault-tolerant scheduling: (1) a dynamic reservation algorithm that leverages cluster-level predictive thresholds to maintain flexible buffer capacity for sudden workload spikes, thus ensuring efficient scheduling under peak demand; (2) a multi-level load-queue architecture that enhances tolerance to task load fluctuations through probabilistic machine selection; and (3) a past information conservative mechanism combined with dynamic weighted drift detection to mitigate prediction degradation. Through the evaluation of 40 million task executions across real-world traces, Focus demonstrates superior load balancing compared to state-of-the-art algorithms, reducing load imbalance by 23.3 % under prediction errors up to 20 %. Luyao Luo, Yu-e Sun, He Huang 0001 |
ICPADS | 2 |
| 2025 | Mocas: Affinity-Aware Moldable Scheduling for Containers in Heterogeneous ClustersabstractContainer scheduling is a critical issue in cloud computing. However, existing container scheduling algorithms focus on resource allocation under multiple resource constraints to maximize resource utilization while overlooking the need to balance the load of the containers and fail to mitigate resource fragmentation during task scheduling, resulting in suboptimal use of compute resources and increased makespan. This oversight often translates into unnecessarily prolonged system completion times. To solve the issue, this paper introduces the Container Scheduling with Moldable Tasks (CSMT) problem, addressing these inefficiencies through a novel framework Mocas. Mocas includes Affallo algorithm with an affinity function to achieve load balance and a fully polynomial time approximation scheme (FPTAS) DoubleShelves with an approximate ratio of$(\mathbf{1}+\mathbf{3} \epsilon)$to eliminate resource fragmentation. Extensive experiments on the real-world datasets demonstrate that the Mocas outperforms state-of-the-art solutions, achieving up to 67 % improvement in makespan compared to conventional rigid scheduling methods. Xun Hu, Luyao Luo, Yu-e Sun, He Huang 0001 |
ICPADS | 2 |
| 2025 | Towards Real-Time Job Scheduling for Load Balancing in CloudsabstractJob scheduling is essential for optimizing resource utilization and enhancing user experience in cloud computing. However, dynamic cloud workloads, which are susceptible to sudden traffic surges, often lead to cluster load imbalance. Traditional scheduling methods, such as heuristic algorithms, can alleviate this imbalance but typically incur significant schedule delay due to the need for extensive job scheduling. This issue is particularly unacceptable for real-time scheduling, especially in containerized environments where jobs have short lifecycles. By incorporating a schedule delay constraint, the integer linear programming (ILP) scheduling problem solved via a solver can effectively maintain low schedule delay. However, these algorithms still suffer from high computational complexity, resulting in long decision delay and outdated scheduling actions. To address these challenges, we propose a scheduling algorithm that effectively achieves load balancing and maintains low schedule delay. Moreover, it mitigates the issue of scheduling decisions becoming outdated due to high decision delay. The proposed algorithm ensures that the schedule delay is bounded by an approximation factor of$\frac{4 \log n}{\alpha}+3$, while the server CPU capacity constraint is preserved within an approximation factor of$\frac{3 \log n}{\alpha}+3$. Experimental results based on real-world datasets show that our scheduling algorithm decreases the schedule delay by 52% to 75% and improves CPU load balance by 15% to 35% compared with the state-of-the-art solutions. Qiuyang Zhu, Luyao Luo, Yu-e Sun |
ICPADS | 2 |
| 2024 | Non-Idle Machine-Aware Worker Placement for Efficient Distributed Training in GPU ClustersabstractDistributed training (DT) has emerged as a solution to address the growing computational resource demands of training large-scale machine learning models. To meet this need, major cloud providers typically build GPU clusters to accommodate DT jobs. Specifically, for an incoming DT job request, cloud providers need to determine in which GPUs place workers (i.e., worker placement). Existing approaches usually place workers on as few idle machines as possible to minimize communication time. However, this scheme will lead to resource fragmentation problem, which degrades the efficiency of the GPU cluster and increases training costs for cloud providers. In this paper, we propose Titan, a novel worker placement scheme that mitigates the influence of resource fragmentation by enhancing the utilization of non-idle machines. Titan formulates a multi-objectives non-linear optimization problem that incorporates the collective communication constraint and proves its NP-hardness. To solve the problem, Titan presents an effective submodular-based greedy algorithm with a tight approximation ratio ($1-\frac{1}{e}$). We evaluate Titan with a large-scale simulation employing real-world job traces and a small-scale testbed consisting of 8 servers with 32 logical GPUs. Experimental results show that Titan can achieve near-optimal training throughput while improving the efficiency of the cluster by 74.9% compared to the state-of-the-art solutions. Gongming Zhao, Hongli Xu 0001, Luyao Luo, An Xie |
ICNP | 4 |
| 2024 | SMART: Dual-channel Southbound Message Delivery in Clouds with Rate EstimationabstractDriving southbound messages from a cloud control plane down to the distributed data plane on every compute node is one of the critical challenges in public clouds. Existing message delivery solutions solely based on remote procedure call (RPC) or message queue (MQ) tend to overlook strict resource constraints, e.g., network bandwidth and CPU capacity. This often results in extensive overhead in the control plane or message redundancy in the data plane, especially when a cloud receives highly concurrent user requests or experiences a rapid expansion. To this end, we design a dual-channel southbound message delivery framework, namely SMART, which combines an RPC channel with an MQ channel, to maximize the resource utilization in the cloud network. In the control plane, we implement a message parsing mechanism and propose a delivery channel selection algorithm based on the deep reinforcement learning (DRL) approach to support efficient dual-channel delivery under resource constraints. In the data plane, we design a message agent on each compute node to ensure the order preservation and state consistency of southbound messages. Both experimental and large-scale simulation results show that SMART demonstrates a reduction in control plane overhead by 64% compared to RPC and redundant messages by 45% compared to MQ, respectively. Luyao Luo, Gongming Zhao, Hongli Xu 0001, Chun-Jen Chung, Liguang Xie |
IWQoS | 1 |
| 2024 | Achieving Cost Optimization for Tenant Task Placement in Geo-Distributed CloudsabstractCloud infrastructure has gradually displayed a tendency of geographical distribution in order to provide anywhere, anytime connectivity to tenants all over the world. The tenant task placement in geo-distributed clouds comes with three critical and coupled factors:regional diversity in electricity prices,access delay for tenants, andtraffic demand among tasks. However, existing works disregard either the regional difference in electricity prices or the tenant requirements in geo-distributed clouds, resulting in increased operating costs or low user QoS. To bridge the gap, we design a cost optimization framework for tenant task placement in geo-distributed clouds, called TanGo. However, it is non-trivial to achieve an optimization framework while meeting all the tenant requirements. To this end, we first formulate the electricity cost minimization for task placement problem as a constrained mixed-integer non-linear programming problem. We then propose a near-optimal algorithm with a tight approximation ratio$(1-1/e)$using an effective submodular-based method. Results of in-depth simulations based on real-world datasets show the effectiveness of our algorithm as well as the overall 10%-30% reduction in electricity expenses compared to commonly-adopted alternatives. Luyao Luo, Gongming Zhao, Hongli Xu 0001, Zhuolong Yu, Liguang Xie |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | TanGo: A Cost Optimization Framework for Tenant Task Placement in Geo-distributed CloudsabstractCloud infrastructure has gradually displayed a tendency of geographical distribution in order to provide anywhere, anytime connectivity to tenants all over the world. The tenant task placement in geo-distributed clouds comes with three critical and coupled factors: regional diversity in electricity prices, access delay for tenants, and traffic demand among tasks. However, existing works disregard either the regional difference in electricity prices or the tenant requirements in geo-distributed clouds, resulting in increased operating costs or low user QoS. To bridge the gap, we design a cost optimization framework for tenant task placement in geo-distributed clouds, called TanGo. However, it is non-trivial to achieve an optimization framework while meeting all the tenant requirements. To this end, we first formulate the electricity cost minimization for task placement problem as a constrained mixed-integer non-linear programming problem. We then propose a near-optimal algorithm with a tight approximation ratio (1 − 1/e) using an effective submodular-based method. Results of in-depth simulations based on real-world datasets show the effectiveness of our algorithm as well as the overall 10%-30% reduction in electricity expenses compared to commonly-adopted alternatives. Luyao Luo, Gongming Zhao, Hongli Xu 0001, Zhuolong Yu, Liguang Xie |
INFOCOM | 1 |
| 2023 | Southbound Message Delivery With Virtual Network Topology Awareness in CloudsabstractSouthbound message delivery from the control plane to the data plane is one of the essential issues in multi-tenant clouds. A natural method of southbound message delivery is that the control plane directly communicates with compute nodes in the data plane. However, due to the large number of compute nodes, this method may result in massive control overhead. The Message Queue (MQ) model can solve this challenge by aggregating and distributing messages to queues. Existing MQ-based solutions often perform message aggregation based on the physical network topology, which do not align with the fundamental requirements of southbound message delivery, leading to high message redundancy on compute nodes. To address this issue, we design and implement VITA, the first-of-its-kind work on virtual network topology-aware southbound message delivery. However, it is intractable to optimally deliver southbound messages according to the virtual attributes of messages. Thus, we design two algorithms, submodular-based approximation algorithm and simulated annealing-based algorithm, to solve different scenarios of the problem. Both experiment and simulation results show that VITA can reduce the total traffic amount of redundant messages by 45%-75% and reduce the control overhead by 33%-80% compared with state-of-the-art solutions. Gongming Zhao, Luyao Luo, Hongli Xu 0001, Chun-Jen Chung, Liguang Xie |
IEEE/ACM Trans. Netw. | 2 |
| 2022 | SNIP: Southbound Message Delivery with In-network Pruning in CloudsabstractIn a hyper-scale cloud data center, a large number of control messages are distributed to hundreds of thousands of compute nodes from a logically centralized control plane. The delivery of these control messages, a.k.a. southbound messages, is critical to cloud infrastructure management, as it greatly affects customer experience. Existing works mainly deal with southbound message delivery by two methods: point-to-point transmission and Message Queue (MQ)-based solutions. With the point-to-point transmission method, each message is sent from the control plane to compute nodes directly, which may cause high control complexity and overhead in hyper-scale clouds. The MQ-based method can address the challenge of high complexity through message aggregation and subscribe/publish model. However, it usually brings in redundant messages, and further causes extra load on compute nodes. To solve the problem above, we design SNIP, which exploits the ability of programmable switches to perform in-network message pruning and to reduce message redundancy. Specffically, forwarding and processing information computed by the control plane is attached to the package header of every control message. Redundant messages can be identified and processed by programmable switches. In addition, we propose a rounding-based algorithm to prune messages with minimal redundancy. The simulation results show that SNIP can reduce the control overhead by 80%-85% and the total traffic of redundant messages by 35% compared with existing solutions. Gongming Zhao, Hongli Xu 0001, Huaqing Tu, Luyao Luo, Liguang Xie |
ICPADS | 5 |
| 2022 | VITA: Virtual Network Topology-aware Southbound Message Delivery in CloudsabstractSouthbound message delivery from the control plane to the data plane is one of the essential issues in multi-tenant clouds. A natural method of southbound message delivery is that the control plane directly communicates with compute nodes in the data plane. However, due to the large number of compute nodes, this method may result in massive control overhead. The Message Queue (MQ) model can solve this challenge by aggregating and distributing messages to queues. Existing MQ-based solutions often perform message aggregation based on the physical network topology, which do not align with the fundamental requirements of southbound message delivery, leading to high message redundancy on compute nodes. To address this issue, we design and implement VITA, the first-of-its-kind work on virtual network topology-aware southbound message delivery. However, it is intractable to optimally deliver southbound messages according to the virtual attributes of messages. Thus, we design two algorithms, submodular-based approximation algorithm and simulated annealing-based algorithm, to solve different scenarios of the problem. Both experiment and simulation results show that VITA can reduce the total traffic amount of redundant messages by 45%-75% and reduce the control overhead by 33%-80% compared with state-of-the-art solutions. Luyao Luo, Gongming Zhao, Hongli Xu 0001, Liguang Xie |
INFOCOM | 1 |
| 2021 | Robust Service Mapping in Multi-Tenant CloudsabstractIn a multi-tenant cloud, cloud vendors provide services (e.g., elastic load-balancing, virtual private networks) on service nodes for tenants. Thus, the mapping of tenants' traffic and service nodes is an important issue in multi-tenant clouds. In practice, unreliability of service nodes and uncertainty/dynamics of tenants' traffic are two critical challenges that affect the tenants' QoS. However, previous works often ignore the impact of these two challenges, leading to poor system robustness when encountering system accidents. To bridge the gap, this paper studies the problem of robust service mapping in multi-tenant clouds (RSMP). Due to traffic dynamics, we take a two-step approach: service node assignment and tenant traffic scheduling. For service node assignment, we prove its NP-Hardness and analyze its problem difficulty. Then, we propose an efficient algorithm with bounded approximation factors based on randomized rounding and knapsack. For tenant traffic scheduling, we design an approximation algorithm based on fully polynomial time approximation scheme (FPTAS). The proposed algorithm achieves the approximation factor of 2+ ε , where ε is an arbitrarily small value. Both small-scale experimental results and large-scale simulation results show the superior performance of our proposed algorithms compared with other alternatives. Jingzhou Wang, Gongming Zhao, Hongli Xu 0001, He Huang 0001, Luyao Luo, Yongqiang Yang |
INFOCOM | 5 |
| 2016 | TinySDM: Software Defined Measurement in Wireless Sensor NetworksabstractNetwork measurement, which provides detailed information about the behaviors of operational networks, is essential for network management in wireless sensor networks. In the literature, there have been many approaches focusing on measuring individual aspect of the network, e.g., per-packet routing path and per-hop delay. However, there lacks a general support for conducting different measurement tasks. When managing an operational network, a network operator often needs to switch the current measurement task to a different one, in order to diagnose the observed symptoms. In this paper, we propose TinySDM, a software-defined measurement architecture for WSNs. TinySDM provides a general support for conducting different measurement tasks. TinySDM defines a set of carefully selected hooks that allow the users to easily execute their own measurement tasks. In addition, TinySDM provides a C- like language called TinyCode Language (TCL) to enable easy customization of measurement tasks. By only transmitting the binary code of the measurement task, TinySDM significantly reduces the size of the disseminated data compared with existing reprogramming approaches. We implement TinySDM on the TinyOS/TelosB platform and evaluate its performance extensively in a testbed with 60 nodes. We also use TCL to implement four specific measurement tasks. Results show that TinySDM is flexible, efficient and easily programmable. Chenhong Cao, Luyao Luo, Yi Gao 0001, Wei Dong 0001, Chun Chen 0001 |
IPSN | 2 |
| 2016 | Dynamic Logging with Dylog in Networked Embedded SystemsabstractEvent logging is an important technique for networked embedded systems like wireless sensor networks. It can greatly help developers to understand complex system behaviors and diagnose program bugs. Existing logging facilities do not well satisfy three practical requirements: flexibility , efficiency , and high synchronization accuracy . To simultaneously satisfy these requirements, we present Dylog, a dynamic logging facility for networked embedded systems. Dylog employs several techniques. First, Dylog uses binary instrumentation for dynamically inserting or removing logging statements, enabling flexible and interactive debugging at runtime. Second, Dylog incorporates an efficient storage system and log collection protocol for recording and transferring the logging messages. Third, Dylog employs a lightweight data-driven approach for reconstructing the synchronized time of the logging messages. Dylog uses MAC-layer timestamping and drift compensation to achieve high synchronization accuracy . We implement Dylog on the TinyOS 2.1.1/TelosB platform. Results show the following: (1) Dylog incurs a small overhead. Indirections in Dylog incur an additional execution overhead of less than 1%. Dylog reduces the logging storage size by approximately 50% compared with the standard TinyOS radio printf library. Dylog reduces the patch size by more than 90%, compared with incremental reprogramming. (2) Dylog reduces the synchronization overhead by 78% in terms of transmission cost, compared with a traditional time synchronization protocol, FTSP, and it can achieve a high time synchronization accuracy of 5.4μs. (3) Dylog can help diagnose system problems effectively at the source-code level for three real-world scenarios. Wei Dong 0001, Luyao Luo, Chao Huang 0026 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2016 | Post-Deployment Anomaly Detection and Diagnosis in Networked Embedded Systems by Program Profiling and Symptom MiningabstractDetecting and diagnosing anomalies in networked embedded systems like sensor networks is a very difficult task, due to the variable workloads and severe resource constraints. In this paper, we focus on how to aid bug diagnosis after the system has been deployed. We notice that most node-level debugging tools can provide detailed program information inside the node but fail to detect when and where a problem occurs in the network. On the other hand, most network-level diagnosis tools can effectively detect a problem from the network but fail to narrow down the problem within the node because they lack detailed program information. To close the gap, we propose D2, a new method for post-deployment anomaly detection and diagnosis in networked embedded systems by combining program profiling and symptom mining. D2 employs binary instrumentation to perform lightweight function count profiling. Based on the statistics, D2 uses PCA (Principal Component Analysis) based approach for automatically detecting network anomalies. Compared with previous methods, D2 is able to point programmers closer to the most likely causes by a novel approach combining statistical tests and program call graph analysis. We implement our method based on TinyOS 2.1.1 and evaluate its effectiveness by case studies in the development of a working sensor network. Results show that our method can aid programmers to diagnose problems quickly in real-world sensor network systems, and at the same time, incurs an acceptable overhead to the running system. Wei Dong 0001, Luyao Luo, Chun Chen 0001, Jiajun Bu, Xue (Steve) Liu, Yunhao Liu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |