Xiaoxuan Luo

dblp:258/5419 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PaTGen: Temporal Similarity-Driven Proxy Benchmark Generation Method for Cloud Workloads
abstract
The rapid expansion of cloud computing has made precise performance evaluation a critical necessity. However, conventional cloud benchmarks often face significant limitations in simulation environments—necessary for scalable and cost-effective testing—due to the complexity of technology stacks and substantial runtime overheads. Proxy benchmarking has thus emerged as a practical alternative. Existing methods primarily focus on the global similarity of micro-architectural metrics between proxy benchmarks and real workloads but neglect their temporal similarity, leading to inaccurate performance evaluations, flawed cache behavior simulations, and misguided architectural optimization decisions. To address this, we present PaTGen , a phase-aware method for generating proxy benchmarks that accurately reflect both global and temporal similarity. By partitioning workloads into phases and formulating proxy generation as nonlinear optimization problems, PaTGen further refines intra-phase execution patterns via the delay-based temporal similarity optimization (DTSO) technique. Evaluations on 15 real-world workloads show PaTGen achieves over 97% global similarity in key metrics while significantly outperforming state-of-the-art methods in temporal similarity. Ablation studies confirm the efficacy of phase division and DTSO. Further experiments confirm its scalability and generalizability across architectures. Moreover, the effectiveness observed in downstream tasks provides empirical evidence that preserving temporal similarity is a fundamental requirement for proxy benchmarks to faithfully capture real workload behavior.
Haolang Yin, Weiwei Lin 0001, Huikang Huang, Xiaoxuan Luo, Haocheng Zhong, Keqin Li 0001
ACM Trans. Archit. Code Optim.4
2026 Dynamic Power Capping for Latency-Sensitive Cloud Applications: A Prediction-Based Power Efficiency Approach
abstract
Data centers often deploy servers that exceed the capacity of the power infrastructure to increase power utilization, i.e., power over-subscription. To address potential power overloads, power capping mechanisms are implemented for protection. However, the power efficiency disparities among heterogeneous servers, along with the adjustment intervals required for power capping at the cluster level, pose challenges to capping decisions, especially for latency-sensitive applications with higher quality of service requirements. To tackle this challenge, we propose a prediction-based, power efficiency-aware dynamic power capping framework (PPE-DPC), comprising two stages. First, we design a lightweight online prediction to capture requests generated within the power adjustment time intervals. Then, leveraging prior knowledge of the power efficiency of heterogeneous servers, we design a greedy strategy to fine-tune the capping power for better capping decisions. Extensive simulations and conducted in Testbed using Alibaba traces demonstrate that PPE-DPC outperforms existing solutions, optimizing request latency, power utilization, and mitigating the negative impact of power adjustment intervals. Finally, we also explore the effect of prediction error on PPE-DPC and find that only 3x the true prediction error is weaker than the existing optimal capping algorithm.
Huikang Huang, Weiwei Lin 0001, Xiaoxuan Luo, James Zijun Wang, Keqin Li 0001
IEEE Trans. Computers3
2026 BOTVPA: An SLO-Aware and Efficient Resource Scheduling Method for Microservice
Xiaoming Ye, Weiwei Lin 0001, Xiaoxuan Luo, Mengheng Li, Qi Mu
IEEE Trans. Serv. Comput.3
2025 Cold-Start Microservice Workload Prediction by Dynamically Annealed Graph-Regularized Matrix Factorization
Xiaoxuan Luo, Hong Shen 0001, Wei Ke 0001
PDCAT1
2025 MCG-Sched: Multi-Cluster GPU Scheduling for Resource Fragmentation Reduction and Load Balancing
abstract
Since the rapid development of deep learning (DL) technology, large-scale GPU clusters receive a large number of DL workloads daily. To speed up the completion time, the workloads usually occupy several GPUs on a server. However, workload scheduling inevitably generates resource fragmentation, which results in many scattered GPU resources being unavailable. Existing works address improving resource utilization by reducing GPU resource fragmentation, while they focus on resource scheduling for a single cluster and ignore multiple clusters. Multi-cluster scenarios, such as virtual clusters and geo-distributed clusters, require load balancing to avoid some clusters exhausting resources while some clusters are idle while improving resource utilization, which is not well addressed by existing works. In this paper, we propose MCG-Sched, a scheduling strategy to reduce resource fragmentation in multiple GPU clusters while maintaining load balancing among clusters. MCG-Sched measures the fragmented resources with the distribution of workload demands and uses a scheme that minimizes fragmentation in workload scheduling. Meanwhile, MCG-Sched achieves balanced load scheduling across clusters through the load balancing index. MCG-Sched senses the workload requests in the waiting queue, and prioritizes the workloads by combining fragmentation measurement and load balancing index to maximize resource utilization and load balancing during load peak. Our experiments show that MCG-Sched reduces unallocated GPUs up to 1.45× and workload waiting time by more than 40% compared to existing fragmentation-aware methods and achieves effective load balancing.
Haijie Wu, Xiaoxuan Luo, Wangbo Shen, Weiwei Lin 0001
IEEE Trans. Parallel Distributed Syst.3
2023 An Energy-Efficient Tuning Method for Cloud Servers Combining DVFS and Parameter Optimization
abstract
Emerging cloud computing applications place a growing demand on resources, leading to increasingly large data centers with significant energy consumption and carbon emissions. Various research conduct optimization methods to improve the energy efficiency of the server in the cloud data center. However, most existing optimization methods are designed for specific applications, thus making it difficult to handle complex cloud environments. In this paper, we propose a general parameter optimization method called MPOD to improve the energy efficiency of cloud servers in real time. MPOD considers issues in the cloud environment, such as SLA guarantee, user privacy, and dynamic workloads. We introduce energy efficiency curves to DVFS, implementing a low-overhead, fast response, and general frequency optimization strategy. Moreover, we design a workload classification framework and three prediction models based on machine learning algorithms to achieve accurate and adaptive Linux kernel parameters optimization. According to the experiment, MPOD can improve the energy efficiency of the server by an average of 30.5%, 20.1%, 10.8% in BenchSEE, SERT and TPC-H, respectively.
Weiwei Lin 0001, Xiaoxuan Luo, ChunKi Li, Jiechao Liang, Guokai Wu, Keqin Li 0001
IEEE Trans. Cloud Comput.2