Xue Ouyang 0003

dblp:165/1945-3 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0003-1196-0237ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 2 since 2021Computer networks · 6 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Can LLMs only talk? Experimental studies on task scheduling with Large Language Models
abstract
Large Language Models (LLMs) have emerged as a disruptive technology for Natural Language Processing (NLP), achieving success in NLP-related generative applications. However, the potential capability of LLMs in other domains remains largely unexplored. To explore the potential of task scheduling with LLMs, we model a typical task scheduling scenario in cloud computing and transfer scheduling problems as natural language prompts. Afterward, the knowledge and reasoning abilities of LLMs are enabled to generate scheduling decisions. Six well-known and open-source LLMs are integrated into our framework to perform experimental studies, and the results are evaluated from multiple perspectives and compared with each other. Besides, traditional heuristic algorithms and a basic Reinforcement Learning (RL) method are all performed for comparison. Our results demonstrate: 1) compared to most heuristic methods, the decisions made by LLMs achieve better scheduling performance; 2) compared to the basic RL method, LLMs exhibit better generalization on various workload patterns; 3) the larger parameter size of the LLMs has, the better scheduling performance it achieves. To the best of our knowledge, our experimental study is the first exploration to apply LLMs in task scheduling. Our findings highlight the promising potential of LLMs as a novel approach to task scheduling, offering new avenues for research and practice.
Mengjuan Li, Zhengguang Chen, Huan Zhou 0006, Yingwen Chen 0001, Baokang Zhao, Xue Ouyang 0003, Jinshu Su
ICCCN7
2023 An edge computing emulator incorporating moving devices and geospatial characteristics
abstract
No abstract available.
Guogui Yang, Baokang Zhao, Xue Ouyang 0003, Qin Xin 0001, Huan Zhou 0006
APNet5
2023 Robustness challenges in Reinforcement Learning based time-critical cloud resource scheduling: A Meta-Learning based solution
abstract
Cloud computing attracts increasing attention in processing dynamic computing tasks and automating the software development and operation pipeline. In many cases, the computing tasks have strict deadlines. The cloud resource manager (e.g., orchestrator) effectively manages the resources and provides tasks Quality of Service (QoS). Cloud task scheduling is tricky due to the dynamic nature of task workload and resource availability. Reinforcement Learning (RL) has attracted lots of research attention in scheduling. However, those RL-based approaches suffer from low scheduling performance robustness when the task workload and resource availability change, particularly when handling time-critical tasks. This paper focuses on both challenges of robustness and deadline guarantee among such RL, specifically Deep RL (DRL)-based scheduling approaches. We quantify the robustness measurements as the retraining time and investigate how to improve both robustness and deadline guarantee of DRL-based scheduling. We propose MLR-TC-DRLS, a practical, robust Meta Deep Reinforcement Learning-based scheduling solution to provide time-critical tasks deadline guarantee and fast adaptation under highly dynamic situations. We comprehensively evaluate MLR-TC-DRLS performance against RL-based and RL advanced variants-based scheduling approaches using real-world and synthetic data. The evaluations validate that our proposed approach improves the scheduling performance robustness of typical DRL variants scheduling approaches with 97%–98.5% deadline guarantees and 200%–500% faster adaptation.
Hongyun Liu, Peng Chen 0007, Xue Ouyang 0003, Hui Gao 0003, Bing Yan 0001, Paola Grosso, Zhiming Zhao
Future Gener. Comput. Syst.3
2022 The Extreme Counts: Modeling the Performance Uncertainty of Cloud Resources with Extreme Value Theory
Mengjuan Li, Jinshu Su, Hongyun Liu, Zhiming Zhao, Xue Ouyang 0003, Huan Zhou 0006
ICSOC5
2022 Efficient Interactive Global Cellular Signal Strength Visualization
abstract
Cellular Signal Strength (CSS), defined as the signal power received by mobile phones, is an important aspect of geographic information flow analysis, because the density of such information can reflect the urbanization variables such as population, gross domestic product, built-up area, electric power consumption, etc. Despite the importance, the real-time analysis of global CSS distribution remains a challenging problem due to the large data scale. In this article, a Display-driven Computing (DisDC) technique is designed and applied to provide efficient large scale interactive CSS visualization, generating results by calculating the value of each pixel that directly for display. Specifically, we present an efficient CSS measurement algorithm, which introduces spatial indexes and a corresponding query strategy; besides, an optimized parallel computing architecture is proposed to ensure the ability of real-time visualization. Experiments show that our approach obviously outperforms traditional methods and is capable of handling more than 40 million base stations in real-time. Moreover, an online demonstration is provided athttps://github.com/MemoryMmy/CSSMap.
Mengyu Ma, Xue Ouyang 0003, Jun Li 0020, Ning Jing
IEEE Trans. Big Data3
2021 Enforcing trustworthy cloud SLA with witnesses: A game theory-based model using smart contracts
abstract
There lacks trust between the cloud customer and provider to enforce traditional cloud SLA (Service Level Agreement) where the blockchain technique seems a promising solution. However, current explorations still face challenges to prove that the off-chain SLO (Service Level Objective) violations really happen before recorded into the on-chain transactions. In this paper, a witness model is proposed implemented with smart contracts to solve this trust issue. The introduced role, "Witness", gains rewards as an incentive for performing the SLO violation report, and the payoff function is carefully designed in a way that the witness has to tell the truth, for maximizing the rewards. This fact that the witness has to be honest is analyzed and proved using the Nash Equilibrium principle of game theory. For ensuring the chosen witnesses are random and independent, an unbiased selection algorithm is proposed to avoid possible collusions. An auditing mechanism is also introduced to detect potential malicious witnesses. Specifically, we define three types of malicious behaviors and propose quantitative indicators to audit and detect these behaviors. Moreover, experimental studies based on Ethereum blockchain demonstrate the proposed model is feasible, and indicate that the performance, ie, transaction fee, of each interface follows the design expectations.
Huan Zhou 0006, Xue Ouyang 0003, Jinshu Su, Cees T. A. M. de Laat, Zhiming Zhao
Concurr. Comput. Pract. Exp.2
2021 Building a blockchain-based decentralized ecosystem for cloud and edge computing: an ALLSTAR approach and empirical study
Huan Zhou 0006, Zeshun Shi, Xue Ouyang 0003, Zhiming Zhao
Peer-to-Peer Netw. Appl.3
2021 Jointgraph: A DAG-based efficient consensus algorithm for consortium blockchains
abstract
Summary The blockchain is a distributed ledger that records all transactions and operations in a shared manner. Public blockchains such as Bitcoin realize decentralization at the cost of mining overhead, which is not suitable for real‐life scenarios requiring high throughput. Techniques such as the consortium blockchain improve efficiency through partial decentralization. However, the consensus algorithms used in the existing state‐of‐the‐art consortium blockchains face many challenges when dealing with commercial applications. For example, the high communication overhead hinders the scalability of PBFT‐based consensus algorithms even though they are efficient at small scale. Hashgraph, one of the most popular Directed Acyclic Graph‐based (DAG‐based) consensus algorithms, achieves good performance in scalability; however, it does not allow users' dynamic participation. To deal with these challenges, we propose Jointgraph, a Byzantine fault‐tolerance consensus algorithm for consortium blockchains based on DAG. In Jointgraph, transactions are packed into events and validated by no less than 2/3 of all members. A supervisor is introduced in our design, who monitors member behaviors and improves consensus efficiency. Simulation results demonstrate that Jointgraph outperforms Hashgraph in both throughput and latency.
Xiang Fu 0002, Huaimin Wang 0001, Peichang Shi, Xue Ouyang 0003, Xunhui Zhang
Softw. Pract. Exp.4
2020 ALLSTAR: A Blockchain Based Decentralized Ecosystem for Cloud and Edge Computing
abstract
Last decades, Cloud computing has made significant impacts on traditional applications to change their development and operation methods. We witnessed ever more newly-built Clouds and data centers. However, the centralized management mechanism of current Clouds lacks the dispersion to satisfy the requirements of emerging collaborative applications, including AI, IoT, and autopilot. On the other hand, the Edge computing stays at the conceptual and experimental stage. Most organizations construct their own Edge nodes to operate applications. An efficient and incentive mechanism is missing to motivate the Edge and micro Cloud resource providers to join and constitute a more generalized and decentralized ecosystem. To address this issue, we propose ALLSTAR, a blockchain based architecture for equally combining all the Cloud and Edge resources to be seamlessly leveraged by the application in the DevOps (development and operations) lifecycle. The ALLSTAR architecture is a systematic solution to realize the "Cloud+Edge" management and contributes to constructing the corresponding ALLSTAR ecosystem. This paper describes the overall architecture of ALLSTAR, the related key techniques, and detailed application DevOps processes as well as the new business model.
Huan Zhou 0006, Xue Ouyang 0003, Zhiming Zhao
JCC2
2020 Tails in the cloud: a survey and taxonomy of straggler management within large-scale cloud data centres
Sukhpal Singh, Xue Ouyang 0003, Peter Garraghan
J. Supercomput.2
2019 A Blockchain based Witness Model for Trustworthy Cloud Service Level Agreement Enforcement
abstract
Traditional cloud Service Level Agreement (SLA) suffers from lacking a trustworthy platform for automatic enforcement. The emerging blockchain technique brings in an immutable solution for tracking transactions among business partners. However, it is still very challenging to prove the credibility of possible violations in the SLA before recording them onto the blockchain. To tackle this challenge, we propose a witness model using game theory and the smart contract techniques. The proposed model extends the existing service model with a new role called “witness” for detecting and reporting service violations. Witnesses gain revenue as an incentive for performing these duties, and the payoff function is carefully designed in a way that trustworthiness is guaranteed: in order to get the maximum profit, the witness has to always tell the truth. This is analyzed and proved through game theory using the Nash equilibrium principle. In addition, an unbiased sortition algorithm is proposed to ensure the randomness of the independent witnesses selection from the decentralized witness pool, to avoid possible unfairness or collusion. An auditing mechanism is also introduced in the paper to detect potential irrational or malicious witnesses. We have prototyped the system leveraging the smart contracts of Ethereum blockchain. Experimental results demonstrate the feasibility of the proposed model and indicate good performance in accordance with the design expectations.
Huan Zhou 0006, Xue Ouyang 0003, Zhijie Ren, Jinshu Su, Cees T. A. M. de Laat, Zhiming Zhao
INFOCOM2
2019 Mitigating stragglers to avoid QoS violation for time-critical applications through dynamic server blacklisting
Xue Ouyang 0003, Jie Xu 0007
Future Gener. Comput. Syst.1
2019 CloudsStorm: A framework for seamlessly programming and controlling virtual infrastructure functions during the DevOps lifecycle of cloud applications
abstract
Summary The infrastructure‐as‐a‐service (IaaS) model of cloud computing provides virtual infrastructure functions (VIFs), which allow application developers to flexibly provision suitable virtual machines' (VM) types and locations, and even configure the network connection for each VM. Because of the pay‐as‐you‐go business model, IaaS provides an elastic way to operate applications on demand. However, in current cloud applications DevOps (software development and operations) lifecycle, the VM provisioning steps mainly rely on manually leveraging these VIFs. Moreover, these functions cannot be programmatically embedded into the application logic to control the infrastructure at runtime. Especially, the vendor lock‐in issue, which different clouds provide different VIFs, also enlarges this gap between the cloud infrastructure management and application operation. To mitigate this gap, we designed and implemented a framework, CloudsStorm, which enables developers to easily leverage VIFs of different clouds and program them into their cloud applications. To be specific, CloudsStorm empowers applications with infrastructure programmability at design‐level, infrastructure‐level, and application‐level. CloudsStorm also provides two infrastructure controlling modes, ie, active and passive mode, for applications at runtime. Besides, case studies about operating task‐based and big data applications on clouds show that the monetary cost is significantly reduced through the seamless and on‐demand infrastructure management provided by CloudsStorm. Finally, the scaling and recovery operation evaluations of CloudsStorm are performed to show its controlling performance. Compared with other tools, ie, “jcloud” and “cloudinit.d”, the scaling and provisioning performance evaluations demonstrate that CloudsStorm can achieve at least 10% efficiency improvement in our experiment settings.
Huan Zhou 0006, Yang Hu 0013, Xue Ouyang 0003, Jinshu Su, Spiros Koulouzis, Cees T. A. M. de Laat, Zhiming Zhao
Softw. Pract. Exp.3
2019 Straggler Root-Cause and Impact Analysis for Massive-scale Virtualized Cloud Datacenters
abstract
Increased complexity and scale of virtualized distributed systems has resulted in the manifestation of emergent phenomena substantially affecting overall system performance. This phenomena is known as “Long Tail”, whereby a small proportion of task stragglers significantly impede job completion time. While work focuses on straggler detection and mitigation, there is limited work that empirically studies straggler root-cause and quantifies its impact upon system operation. Such analysis is critical to ascertain in-depth knowledge of straggler occurrence for focusing developmental and research efforts towards solving the Long Tail challenge. This paper provides an empirical analysis of straggler root-cause within virtualized Cloud datacenters; we analyze two large-scale production systems to quantify the frequency and impact stragglers impose, and propose a method for conducting root-cause analysis. Results demonstrate approximately 5 percent of task stragglers impact 50 percent of total jobs for batch processes, and 53 percent of stragglers occur due to high server resource utilization. We leverage these findings to propose a method for extreme straggler detection through a combination of offline execution patterns modeling and online analytic agents to monitor tasks at runtime. Experiments show the approach is capable of detecting stragglers less than 11 percent into their execution lifecycle with 95 percent accuracy for short duration jobs.
Peter Garraghan, Xue Ouyang 0003, Renyu Yang, David McKee 0001, Jie Xu 0007
IEEE Trans. Serv. Comput.2
2018 JCDTA: The Data Trading Archtecture Design in JointCloud Computing
abstract
JointCloud computing is a new generation cloud computing model based on collaboration among Cloud Service Providers, making resources from multiple clouds deeply integrated., and supporting customize cloud service. To achieve the data confirmation Right when doing data trading to prevent resell is one of the most important challenges faced by such JointCloud environment. In this paper, we propose the JointCloud Computing Data Trading Architecture (JCDTA), an optimized data trading architecture for cross-stakeholder. Firstly, the announced mechanism is used to extract the data resources summary and the owner information into blockchain. Secondly, the untampering of the blockchain is used to record the data resources information, such as data resources statement information, data resources transaction records and data resources operation records. Finally, the supervision mechanism confirms the declaration of the data sources while the encryption technology ensures the privacy of the data resources. JCDTA guarantees the security and the reliability of data trading in the JointCloud environment and ensures the value invariance of data resources.
Xikun Yue, Huaimin Wang 0001, Wei Li 0022, Peichang Shi, Xue Ouyang 0003
ICPADS6
2018 Adaptive Speculation for Efficient Internetware Application Execution in Clouds
abstract
Modern Cloud computing systems are massive in scale, featuring environments that can execute highly dynamic Internetware applications with huge numbers of interacting tasks. This has led to a substantial challenge—the straggler problem, whereby a small subset of slow tasks significantly impede parallel job completion. This problem results in longer service responses, degraded system performance, and late timing failures that can easily threaten Quality of Service (QoS) compliance. Speculative execution (or speculation) is the prominent method deployed in Clouds to tolerate stragglers by creating task replicas at runtime. The method detects stragglers by specifying a predefined threshold to calculate the difference between individual tasks and the average task progression within a job. However, such a static threshold debilitates speculation effectiveness as it fails to capture the intrinsic diversity of timing constraints in Internetware applications, as well as dynamic environmental factors, such as resource utilization. By considering such characteristics, different levels of strictness for replica creation can be imposed to adaptively achieve specified levels of QoS for different applications. In this article, we present an algorithm to improve the execution efficiency of Internetware applications by dynamically calculating the straggler threshold, considering key parameters including job QoS timing constraints, task execution progress, and optimal system resource utilization. We implement this dynamic straggler threshold into the YARN architecture to evaluate it’s effectiveness against existing state-of-the-art solutions. Results demonstrate that the proposed approach is capable of reducing parallel job response time by up to 20% compared to the static threshold, as well as a higher speculation success rate, achieving up to 66.67% against 16.67% in comparison to the static method.
Xue Ouyang 0003, Peter Garraghan, Bernhard Primas, David McKee 0001, Paul Townend, Jie Xu 0007
ACM Trans. Internet Techn.1
2017 ML-NA: A Machine Learning Based Node Performance Analyzer Utilizing Straggler Statistics
abstract
Current Cloud clusters often consist of heterogeneous machine nodes, which can trigger performance challenges such as the task straggler problem, whereby a small subset of parallel tasks running abnormally slower than the other sibling ones. The straggler problem leads to extended job response and deteriorates system throughput. Poor performance nodes are more likely to engender stragglers, and can undermine straggler mitigation effectiveness. For example, as the dominant mechanism for straggler alleviation, speculative execution functions by creating redundant task replicas on other machine nodes as soon as a straggler is detected. When speculative copies are assigned onto the poor performance nodes, it is hard for them to catch up with the stragglers compared to replicas run on fast nodes. And due to the fact that the performance heterogeneity is caused not only by static attribute variations such as physical capacity, but also dynamic characteristic uctuations such as contention level, analyzing node performance is important yet challenging. In this paper we develop ML-NA, a Machine Learning based Node performance Analyzer. By leveraging historical parallel tasks execution log data, ML-NA classies cluster nodes into different categories and predicts their performance in the near future as a scheduling guide to improve speculation effectiveness and minimize task straggler generation. We consider MapReduce as a representative framework to perform our analysis, and use the published OpenCloud trace as a case study to train and to evaluate our model. Results show that ML-NA can predict node performance categories with an average accuracy up to 92.86%.
Xue Ouyang 0003, Renyu Yang, Guogui Yang, Paul Townend, Jie Xu 0007
ICPADS1
2017 Mitigate data skew caused stragglers through ImKP partition in MapReduce
abstract
Speculative execution is the mechanism adopted by current MapReduce framework when dealing with the straggler problem, and it functions through creating redundant copies for identified stragglers. The result of the quicker task will be adopted to improve the overall job execution performance. Although proved to be effective for contention caused stragglers, speculative execution can easily meet its bottleneck when mitigating data skew caused stragglers due to its replication nature: the identical unbalanced input data will lead to a slow speculative task. The Map inputs are typically even in size according to the HDFS block configuration, therefore the skew caused stragglers happen mainly in the Reduce phase because of the unknown intermediate key distribution. In this paper, we focus on mitigating data skew caused Reduce stragglers, propose ImKP, an Intermediate Key Pre-processing framework that enables the even distributed partition for Reduce inputs. A group based ranking technique has been developed that dramatically decreases the pre-processing time, and ImKP manages to eliminate this timing overhead through parallelizing the pre-processing with the file uploading procedure (from local file system to HDFS). For jobs that take input directly from HDFS, ImKP minimizes the overhead by storing themapping result on every node within the cluster for reuse. Experiments are conducted on different datasets with various workloads. Results show that, compared to the popular hash partition, ImKP can dramatically decrease Reduce skew, achieving a 99.8% reduction in the coefficient of variation of the input sizes in average, and improve up to 29.37% job response performance.
Xue Ouyang 0003, Huan Zhou 0006, Stephen J. Clement, Paul Townend, Jie Xu 0007
IPCCC1
2016 Straggler Detection in Parallel Computing Systems through Dynamic Threshold Calculation
abstract
Cloud computing systems face the substantial challenge of the Long Tail problem: a small subset of straggling tasks significantly impede parallel jobs completion. This behavior results in longer service response times and degraded system utilization. Speculative execution, which create task replicas at runtime, is a typical method deployed in large-scale distributed systems to tolerate stragglers. This approach defines stragglers by specifying a static threshold value, which calculates the temporal difference between an individual task and the average task progression for a job. However, specifying static threshold debilitates speculation effectiveness as it fails to consider the intrinsic diversity of job timing constraints within modern day Cloud computing systems. Capturing such heterogeneity enables the ability to impose different levels of strictness for replica creation while achieving specified levels of QoS for different application types. Furthermore, a static threshold also fails to consider system environmental constraints in terms of replication overheads and optimal system resource usage. In this paper we present an algorithm for dynamically calculating a threshold value to identify task stragglers, considering key parameters including job QoS timing constraints, task execution characteristics, and optimal system resource utilization. We study and demonstrate the effectiveness of our algorithm through simulating a number of different operational scenarios based on real production cluster data against state-of-the-art solutions. Results demonstrate that our approach is capable of creating 58.62% less replicas under high resource utilization while reducing response time up to 17.86% for idle periods compared to a static threshold.
Xue Ouyang 0003, Peter Garraghan, David McKee 0001, Paul Townend, Jie Xu 0007
AINA1
2016 Tolerating Transient Late-Timing Faults in Cloud-Based Real-Time Stream Processing
abstract
Real-time stream processing is a frequently deployed application within Cloud datacenters that is required to provision high levels of performance and reliability. Numerous fault-tolerant approaches have been proposed to effectively achieve this objective in the presence of crash failures. However, such systems struggle with transient late-timing faults - a fault classification challenging to effectively tolerate - that manifests increasingly within large-scale distributed systems. Such faults represent a significant threat towards minimizing soft real-time execution of streaming applications in the presence of failures. This work proposes a fault-tolerant approach for QoS-aware data prediction to tolerate transient late-timing faults. The approach is capable of determining the most effective data prediction algorithm for imposed QoS constraints on a failed stream processor at run-time. We integrated our approach into Apache Storm with experiment results showing its ability to minimize stream processor end-to-end execution time by 61% compared to other fault-tolerant approaches. The approach incurs 12% additional CPU utilization while reducing network usage by 44%.
Peter Garraghan, Stuart Perks, Xue Ouyang 0003, David McKee 0001, Ismael Solís Moreno
ISORC3
2016 SEED: A Scalable Approach for Cyber-Physical System Simulation
abstract
Simulation is critical when studying real operational behavior of increasingly complex Cyber-Physical Systems, forecasting future behavior, and experimenting with hypothetical scenarios. A critical aspect of simulation is the ability to evaluate large-scale systems within a reasonable time frame while modeling complex interactions between millions of components. However, modern simulations face limitations in provisioning this functionality for CPSs in terms of balancing simulation complexity with performance, resulting in substantial operational costs required for completing simulation execution. Moreover, users are required to have expertise in modeling and configuring simulations to infrastructure which is time consuming. In this paper we present Simulation EnvironmEnt Distributor (SEED), a novel approach for simulating large-scale CPSs across a loosely-coupled distributed system requiring minimal user configuration. This is achieved through automated simulation partitioning and instantiation while enforcing tight event messaging across the system. SEED operates efficiently within both small and large-scale OTS hardware, agnostic of cluster heterogeneity and OS running, and is capable of simulating the full system and network stack of a CPS. Our approach is validated through experiments conducted in a cluster to simulate CPS operation. Results demonstrate that SEED is capable of simulating CPSs containing 2,000,000 tasks across 2,000 nodes with only 6.89× slow down relative to real time, and executes effectively across distributed infrastructure.
Peter Garraghan, David McKee 0001, Xue Ouyang 0003, David Webster, Jie Xu 0007
IEEE Trans. Serv. Comput.3
2015 Timely Long Tail Identification through Agent Based Monitoring and Analytics
abstract
The increasing complexity and scale of distributed systems has resulted in the manifestation of emergent behavior which substantially affects overall system performance. A significant emergent property is that of the "Long Tail", whereby a small proportion of task stragglers significantly impact job execution completion times. To mitigate such behavior, straggling tasks occurring within the system need to be accurately identified in a timely manner. However, current approaches focus on mitigation rather than identification, which typically identify stragglers too late in the execution lifecycle. This paper presents a method and tool to identify Long Tail behavior within distributed systems in a timely manner, through a combination of online and offline analytics. This is achieved through historical analysis to profile and model task execution patterns, which then inform online analytic agents that monitor task execution at runtime. Furthermore, we provide an empirical analysis of two large-scale production Cloud data enters that demonstrate the challenge of data skew within modern distributed systems, this analysis shows that approximately 5% of task stragglers caused by data skew impact 50% of the total jobs for batch processes. Our results demonstrate that our approach is capable of identifying task stragglers less than 11% into their execution lifecycle with 98% accuracy, signifying significant improvement over current state-of-the-art practice and enables far more effective mitigation strategies in large-scale distributed systems worldwide.
Peter Garraghan, Xue Ouyang 0003, Paul Townend, Jie Xu 0007
ISORC2