Xiaohui Gu

dblp:55/6848 · DBLP profile ↗
← Back
84ranked-venue papers
29as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 31 · 8 first-author · 2 since 2021Computer networks · 16 · 9 first-author · 11 since 2021Security and privacy · 9Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Robust Beamforming for Digital Twin-Enabled RIS Systems: A Hierarchical Meta-Learning Approach
Xiaohui Gu, Guoan Zhang
WCNC1
2025 You can be more trustworthy: A feature fusion reinforcement network for credible anti-noise fault diagnosis
Hongchong Peng, Mansong Rong, Xiaohui Gu, Xiangyan Chen
Adv. Eng. Informatics4
2025 Optimizing Federated Learning Performance: A Blockchain-Integrated Solution for Edge Networks
abstract
This paper proposes a blockchain-integrated federated learning (FL) framework tailored for secure, efficient, and energy-aware model training in edge computing environments. The framework follows an offload-train-aggregate paradigm where edge devices transmit local datasets to proximate servers for localized model updates. A reputation-driven RAFT consensus protocol is incorporated to achieve reliable, low-latency, and lightweight blockchain coordination while preserving privacy and accountability. To overcome the inherent mixed-integer nonlinear programming (MINLP) complexity, we develop a two-stage cross-layer optimization strategy. In the first stage, an alternating direction method of multipliers (ADMM)-based feedback control scheme jointly allocates bandwidth and computation resources under energy and delay constraints. In the second stage, server selection and sub-band assignment are modeled as a bipartite matching problem and solved via the Hungarian algorithm, guided by a convergence-aware performance bound. Extensive simulations demonstrate that our framework significantly improves learning accuracy, uplink throughput, and energy efficiency over state-of-the-art FL baselines. It also exhibits strong robustness to network fragmentation and resource heterogeneity, making it well suited for practical edge environments.
Xiaohui Gu, Guoan Zhang, Wei Duan 0001, Qiang Sun 0001, Miaowen Wen, Pin-Han Ho
IEEE Trans. Commun.1
2025 ARIS: Adaptive Beamforming Design Under Dynamic Environments
abstract
In the rapidly evolving field of wireless communications, the emergence of 5G and the progression towards 6G technologies highlight the demands for innovative frameworks capable of enhancing network performance under complex environmental factors. According to this trend, this work introduces a novel system model for aerial reconfigurable intelligent surface (ARIS)-assisted wireless communications, engineered to adeptly manage dynamic environmental influences such as fluctuating wind patterns and variable weather conditions. By treating these influences as random variables, we further introduce unpredictable variations in ARIS orientation (roll, yaw, and pitch), affecting network efficiency and reliability. To mitigate these environmental perturbations and uncertainties in the channel state information (CSI), we propose a robust beamforming strategy to reshape the original optimization problem into a form that is easier to analyze and solve, by employing stochastic optimization, successive convex approximation (SCA), and block coordinate descent (BCD) techniques to optimize active beamforming vectors and ARIS phase shifts. Comprehensive simulations rigorously evaluate the performance of our proposed algorithms across diverse conditions, including aerial disturbances, varying channel states, and different RIS element configurations. The results underscore the efficacy of our beamforming design in boosting the resilience and dependability of ARIS-enhanced wireless networks.
Xiaohui Gu, Guoan Zhang, Wei Duan 0001, Lei Zhang 0160, Miaowen Wen, Pin-Han Ho
IEEE Trans. Wirel. Commun.1
2024 A novel bearing intelligent fault diagnosis method based on spectrum sparse deep deconvolution
Huifang Shi, Yonghao Miao, Chenhui Li 0002, Xiaohui Gu
Eng. Appl. Artif. Intell.4
2024 Computing Offloading for RIS-Aided Internet of Everything: A Cybertwin Version
abstract
Cybertwin technology introduces a novel paradigm employing digital twins to model complex physical systems within a cyber environment, thus enhancing communication, collaboration, and decision-making capabilities. By harnessing advanced technologies, such as reconfigurable intelligent surfaces (RISs) and multiaccess edge computing (MEC), seamless interaction between physical and virtual entities is facilitated. In this article, we propose a cybertwin-driven edge computing framework that leverages RIS technology, complemented by an efficient computing offloading strategy to support large-scale Internet of Everything (IoE) applications. Specifically, the proposed strategy focuses on a multicell system where numerous randomly distributed end users have the option to offload delay-sensitive and computing-intensive tasks to edge computing nodes. The offloading channels are enhanced by RISs through passive beamforming, while cybertwin technology directs resource cooperation among multicells and allocates computing and communication resources. Our main objective is to optimize the system’s utility with respect to task completion latency and energy consumption reduction. To achieve this goal, we conduct the joint optimization of task offloading and resource allocation. Furthermore, we develop a joint task offloading and resource allocation (JTORA) algorithm to derive optimal solutions for passive beamforming design, computing offloading decisions, communication resource scheduling, and computing capacity allocation. The simulation results demonstrate the superiority of the proposed algorithm over benchmark schemes in terms of edge computing efficiency. Furthermore, the system utility can be further enhanced by increasing the number of embedded RIS elements.
Xiaohui Gu, Guoan Zhang, Wei Duan 0001, Shuping Dang, Miaowen Wen, Pin-Han Ho
IEEE Internet Things J.1
2024 Self-Supervised Machine Learning Framework for Online Container Security Attack Detection
abstract
Container security has received much research attention recently. Previous work has proposed to apply various machine learning techniques to detect security attacks in containerized applications. On one hand, supervised machine learning schemes require sufficient labeled training data to achieve good attack detection accuracy. On the other hand, unsupervised machine learning methods are more practical by avoiding training data labeling requirements, but they often suffer from high false alarm rates. In this article, we present a generic self-supervised hybrid learning (SHIL) framework for achieving efficient online security attack detection in containerized systems. SHIL can effectively combine both unsupervised and supervised learning algorithms but does not require any manual data labeling. We have implemented a prototype of SHIL and conducted experiments over 46 real-world security attacks in 29 commonly used server applications. Our experimental results show that SHIL can reduce false alarms by 33%–93% compared to existing supervised, unsupervised, or semi-supervised machine learning schemes while achieving a higher or similar detection rate.
Olufogorehan Tunde-Onadele, Yuhang Lin 0001, Xiaohui Gu, Jingzhu He, Hugo Latapie
ACM Trans. Auton. Adapt. Syst.3
2024 Sum-Rate Maximization for RIS-IoV: From Instantaneous to Statistical CSI
abstract
To fully exploit the potential of reconfigurable intelligent surface (RIS), the controllable channel state information (CSI) should be accurate for its future applications. Unfortunately, in vehicular communications, obtaining exact instantaneous CSI presents substantial challenges. Moreover, even with an instantaneous CSI acquisition, a processing latency for RIS phase shift adaption might occur before the vehicular system reacts to the instantaneous CSI information. To effectively introduce RIS into Internet of vehicle (IoV) networks, we employ a more realistic statistical CSI approach in designing RIS-assisted vehicular communication systems that are robust to the general characteristics of the channel, rather than its instantaneous fluctuations. We present a practical system framework, where a roadside unit employs an RIS to facilitate indirect wireless communications for vehicle-to-vehicle (V2V) communications. Particularly, the direct links between vehicles are susceptible to blockages caused by surrounding obstacles/vehicles. The deployment of RIS is to establish supplementary communication links between a multi-antenna vehicle source (VS) and multiple vehicular users (VUs) as they traverse areas with a poor service coverage. With the objective to maximize the time-averaged sum-rate of VUs, instead of instantaneous CSI, we rely on the delayed statistical CSI feedback to design active beamforming at the VS and passive beamforming at RIS. Moreover, we develop an efficient algorithm, named JAPBNB, which leverages the fractional programming technique to find a stationary solution for the formulated sum-of-logarithms-of-ratio problem. Specifically, a non-convex block coordinate descent (BCD) approach, collaborating with the alternating direction method of multipliers (ADMM), is applied for the joint optimization of active and passive beamforming. Finally, the complexity and convergence of the proposed JAPBNB algorithm are thoroughly discussed and validated. Simulation results demonstrate that the time-averaged sum-rate obtained by the proposed JAPBNB algorithm approaches that obtained by the instantaneous CSI scheme, when the delayed statistical CSI feedback interval is adequately small.
Wei Duan 0001, Xiaohui Gu, Guoan Zhang, Miaowen Wen, Zhiguo Ding 0001, Pin-Han Ho
IEEE Trans. Wirel. Commun.2
2023 Cooperative vehicular networks over Nakagami-m fading: Joint power control and spectrum scheduling
Guoan Zhang, Xiaohui Gu
Comput. Networks3
2023 A survey on UAV-assisted wireless communications: Recent advances and future trends
Xiaohui Gu, Guoan Zhang
Comput. Commun.1
2023 Partial-NOMA Based Physical Layer Security: Forwarding Design and Secrecy Analysis
abstract
Due to the inherent broadcast characteristics of wireless communications, the data transmission is difficult to be shielded from unintended recipients. Secure communications over wireless channel is regarded as an effective way to overcome this issue in the design of wireless networks. In this paper, a new cooperative relaying system based on partial non-orthogonal multiple access (P-NOMA) is proposed, where two relays help the communications between the source and destination nodes in the presence of an eavesdropper (Eve). Specifically, after receiving the P-NOMA signals transmitted from the base station, the relay nodes decode and forward the receptions under the attack of a passive Eve. To improve the physical layer security (PLS) of the proposed system, we design four cooperative schemes for signals decoded at the relays and destination, where maximum ratio combination (MRC) technique is adopted to both destination and Eve to construe the worst case for achievable secrecy rate. Without lose of generality, the channel conditions (weak or strong), new power allocations, decoding principles and Eve locations are considered in designing forwarding schemes. In particular, to further improve the achievable secrecy rate, we also propose that, if the Eve locating near to the weak relay node, the weak relay only forwards the signal with lower power allocation factor at the base station. The closed-form expressions of the achievable secrecy rates are derived for the proposed schemes over Rayleigh fading and Nakagami-$m$fading channels, which well match the simulation results. By means of the numerical result, it corroborates the superiority of the proposed P-NOMA schemes over the conventional NOMA schemes. It is also revealed that, with an increasing overlap ratio, the advantage in terms of the secrecy rate becomes more remarkable.
Biting Zhuo, Wei Duan 0001, Juping Gu, Xiaohui Gu, Guoan Zhang, Yancheng Ji, Miaowen Wen
IEEE Trans. Intell. Transp. Syst.4
2023 Performance Bug Analysis and Detection for Distributed Storage and Computing Systems
abstract
This article systematically studies 99 distributed performance bugs from five widely deployed distributed storage and computing systems (Cassandra, HBase, HDFS, Hadoop MapReduce and ZooKeeper). We present the TaxPerf database, which collectively organizes the analysis results as over 400 classification labels and over 2,500 lines of bug re-description. TaxPerf is classified into six bug categories (and 18 bug subcategories) by their root causes; resource, blocking, synchronization, optimization, configuration, and logic. TaxPerf can be used as a benchmark for performance bug studies and debug tool designs. Although it is impractical to automatically detect all categories of performance bugs in TaxPerf, we find that an important category of blocking bugs can be effectively solved by analysis tools. We analyze the cascading nature of blocking bugs and design an automatic detection tool called PCatch , which (i) performs program analysis to identify code regions whose execution time can potentially increase dramatically with the workload size; (ii) adapts the traditional happens-before model to reason about software resource contention and performance dependency relationship; and (iii) uses dynamic tracking to identify whether the slowdown propagation is contained in one job. Evaluation shows that PCatch can accurately detect blocking bugs of representative distributed storage and computing systems by observing system executions under small-scale workloads.
Yiming Zhang 0003, Shan Lu 0001, Haryadi S. Gunawi, Xiaohui Gu, Dongsheng Li 0001
ACM Trans. Storage5
2022 Understanding Software Security Vulnerabilities in Cloud Server Systems
abstract
Cloud systems have been widely adopted by many real world production applications. Thus, security vulnerabilities in those cloud systems can cause serious widespread impact. Although previous intrusion detection systems can detect security attacks, understanding the underlying software defects that cause those security vulnerabilities is little studied. In this paper, we conduct a systematic study over 110 software security vulnera-bilities in 13 popular cloud server systems. To understand the underlying vulnerabilities, we answer the following questions: 1) what are the root causes of those security vulnerabilities? 2) what threat impact do those vulnerable code have? 3) how do developers patch those vulnerable code? Our results show that the vulnerable code of the studied security vulnerabilities comprise five common categories: 1) improper execution restrictions, 2) improper permission checks, 3) improper resource path-name checks, 4) improper sensitive data handling, and 5) improper synchronization handling. We further extract principal vulnerable code patterns from those common vulnerability categories.
Olufogorehan Tunde-Onadele, Yuhang Lin 0001, Xiaohui Gu, Jingzhu He
IC2E3
2022 PerfSig: Extracting Performance Bug Signatures via Multi-modality Causal Analysis
abstract
Diagnosing a performance bug triggered in production cloud environments is notoriously challenging. Extracting performance bug signatures can help cloud operators quickly pinpoint the problem and avoid repeating manual efforts for diagnosing similar performance bugs. In this paper, we present PerfSig, a multi-modality performance bug signature extraction tool which can identify principal anomaly patterns and root cause functions for performance bugs. PerfSig performs fine-grained anomaly detection over various machine data such as system metrics, system logs, and function call traces. We then conduct causal analysis across different machine data using information theory method to pinpoint the root cause function of a performance bug. PerfSig generates bug signatures as the combination of the identified anomaly patterns and root cause functions. We have implemented a prototype of PerfSig and conducted evaluation using 20 real world performance bugs in six commonly used cloud systems. Our experimental results show that PerfSig captures various kinds of fine-grained anomaly patterns from different machine data and successfully identifies the root cause functions through multi-modality causal analysis for 19 out of 20 tested performance bugs.
Jingzhu He, Yuhang Lin 0001, Xiaohui Gu, Chin-Chia Michael Yeh, Zhongfang Zhuang
ICSE3
2022 A Motion Representation AI Based Video Conference Solution
abstract
Video conference services have been widely used today. Bandwidth cost is huge in HD video conference and thus reducing bandwidth is important for video conference services. In our paper, we proposed a real-time AI based cloud video conference solution on Intel Xeon which enables motion representation AI video reconstruction algorithms such as first order motion model (FOMM) to reduce bandwidth. We also proposed a background enhancement method to improve user experiences. The experimental results showed that our solution can achieve about 80% bandwidth reduction of same video quality compared with H.264. Our demo is available on https://github.com/Amygu1994/AIbasedVideoConference/raw/main/Demo_AI_based_Video_Conference_Solution.mp4.
Xiaohui Gu, Yanying Sun
MMSP1
2022 UAV-Aided Energy-Efficient Edge Computing Networks: Security Offloading Optimization
abstract
Unmanned aerial vehicles (UAVs) are widely applied for service provisioning in many domains, such as topographic mapping and traffic monitoring. These applications are complicated with huge computational resources and extremely low-latency requirements. However, the moderate computational capability and limited energy restrict the local data processing for the UAV. Fortunately, this impediment may be mitigated by utilizing wireless power transfer (WPT) and employing the multiaccess edge computing (MEC) paradigm for offloading demanding computational tasks from the UAV via wireless communications. Particularly, the offloaded information may become compromising by the eavesdropper (Eve) when UAVs offload the computational tasks to MEC servers. To address this issue, a UAV-MEC (UMEC) system with energy harvesting (EH) is studied, where the full-duplex protocol is considered to realize simultaneously receiving confidential data from the UAV and broadcasting the control instructions. It is worth noting that in our proposed scheme, these control instructions also serve as the artificial interference to confuse the Eve. To improve the energy efficiency for offloading, the computational communication resource allocation is optimized to minimize the energy consumption for UAV with the consumed and harvested energy. Specially, the worst case secrecy offloading rate and computation-latency constraint are considered, to further enhance the reliability and security of the proposed system. Since the objective optimization problem is nonconvex, we convert it into a convex one by analytical means. The semiclosed form expressions of the offloading time, offloading data size, and transmit power are, respectively, derived. Moreover, the conditions of nonoffloading, partial, and full offloading are also discussed from a physical perspective. With the specific conditions of activating the above-mentioned three offloading options, numerical results verify the performance of our proposed offloading strategy in various scenarios and show the superiority of our offloading strategy with the existing works in terms of the offloading capacity and energy efficiency.
Xiaohui Gu, Guoan Zhang, Wei Duan 0001, Miaowen Wen, Pin-Han Ho
IEEE Internet Things J.1
2022 Resource Management for Intelligent Vehicular Edge Computing Networks
abstract
To overcome the inherent defect of centralized data processing in cloud computing, the mobile edge computing (MEC) brings data storage and computing capacities, to the edge closer to end users. However, the uneven distribution of access vehicles, as well the volume of computing data, cause the workload diversity among various mobile edge computing servers (MECSs). In this paper, we propose a hierarchical model with quality of service (QoS)-aware and power-aware resource management for the cooperative edge-computing-based intelligent vehicular network (CEC-IoV), and the system latency and energy efficiency at MECSs are respectively optimized. Specifically, considering the changing response times versus MECSs’ workloads, the Minimum Latency with Migration Loads (MLML) scheme is developed for workload balance among multiple MECSs. By selecting the appropriate response time threshold and migration loads from overloading MECSs to idle MECSs simultaneously, the load-balancing problem can be efficiently solved for multiple MECSs with unbalanced workloads. On the other hand, through performing workload redistribution and dynamic reconfiguration of virtual machines (VMs) instantiated onto the parallel computing platform at one MECS, the energy-efficiency can be also optimized while guaranteeing the QoS requirement on the processing delay. With the latency constraint, the power minimization problem is formulated to be a convex one, and the semi-closed forms for optimal solutions of VMs’ workloads and processing rates are provided using KKT conditions. Compared with the performance obtained by benchmark schemes, numerical results exhibit that our resource management schemes gain lower system latency and higher energy efficiency.
Wei Duan 0001, Xiaohui Gu, Miaowen Wen, Yancheng Ji, Jianhua Ge, Guoan Zhang
IEEE Trans. Intell. Transp. Syst.2
2021 Time Series Prediction of Wind Speed Based on SARIMA and LSTM
Caiquan Xiong, Congcong Yu, Xiaohui Gu, Shiqiang Xu
CISIS3
2021 Offloading Optimization for Energy-Minimization Secure UAV-Edge-Computing Systems
abstract
This paper considers a mobile edge computing (MEC) system where an unmanned aerial vehicle (UAV) offloads demanding computational tasks to the access point (AP), and AP cooperatively sends jamming noise to protect the confidentiality and secrecy of data transmission. In this system, the computation from UAV is portioned into two parts with one executed locally on UAV and the other offloaded to MEC for processing. To minimize the total energy-consumption of UAV for computing and offloading, we jointly optimize computation and communication resource allocation, subject to secure offloading rate and computation-latency constraints. Although the formulated problem is non-convex, we transform it into a convex problem by analytical means. Numerical results verify the superiority of our proposed scheme in terms of energy-consumption and quantify the performance of our proposed offloading scheme in various scenarios.
Xiaohui Gu, Guoan Zhang, Jinyuan Gu
WCNC1
2021 Energy-efficient computation offloading for vehicular edge computing networks
Xiaohui Gu, Guoan Zhang
Comput. Commun.1
2021 UAV-Relaying Cooperation for Internet of Everything with CRT-Based NOMA
abstract
Due to the great potential of the combination of machine learning technology and unmanned aerial vehicle (UAV) enabled wireless communications, various optimization algorithms on resource allocation have been proposed for the Internet of Things. UAVs not only can perform the missions under the extreme conditions but also enhance the overall performance of the system as an aerial relay assisting transmission in the public and civil domains, which have been received extensive attentions. However, with the limited capacity and power constraints, they are difficult to support the transmission for the big data information users. In addition, the lack of spectrum resource poses challenges to satisfy the quality of service (QoS) of mobile users in wireless networks. To contribute to these urgent problems, this article first studies the potential and effective applications of UAVs, by introducing the Chinese remainder theorem (CRT) and nonorthogonal multiple access (NOMA) technologies into UAV relay networks. Two scenarios with/without direct transmissions between the source and destination nodes are investigated, following the decomposition and reconstruction mechanisms to satisfy the big data information transmission. Considering the user fairness, we further discuss the effect of the UAV numbers to the overall system capacity. To maximize the system capacity, the designs of transmission protocol and receiver are also discussed, in various channel conditions. Finally, a low complexity and efficient two‐stage power allocation scheme is established for the perspective of users and UAV relays.
Jinyuan Gu, Xiaohui Gu, Guoan Zhang, Wei Duan 0001
Wirel. Commun. Mob. Comput.2
2020 CDL: Classified Distributed Learning for Detecting Security Attacks in Containerized Applications
abstract
Containers have been widely adopted in production computing environments for its efficiency and low overhead of isolation. However, recent studies have shown that containerized applications are prone to various security attacks. Moreover, containerized applications are often highly dynamic and short-lived, which further exacerbates the problem. In this paper, we present CDL, a classified distributed learning framework to achieve efficient security attack detection for containerized applications. CDL integrates online application classification and anomaly detection to overcome the challenge of lacking sufficient training data for dynamic short-lived containers while considering diversified normal behaviors in different applications. We have implemented a prototype of CDL and evaluated it over 33 real world vulnerability attacks in 24 commonly used server applications. Our experimental results show that CDL can reduce the false positive rate from over 12% to 0.24% compared to traditional anomaly detection schemes without aggregating training data. By introducing application classification into container behavior learning, CDL can improve the detection rate from catching 20 attacks to 31 attacks before those attacks succeed. CDL is light-weight, which can complete application classification and anomaly detection for each data sample within a few milliseconds.
Yuhang Lin 0001, Olufogorehan Tunde-Onadele, Xiaohui Gu
ACSAC3
2020 HangFix: automatically fixing software hang bugs for production cloud systems
abstract
Software hang bugs are notoriously difficult to debug, which often cause serious service outages in cloud systems. In this paper, we present HangFix, a software hang bug fixing framework which can automatically fix a hang bug that is triggered and detected in production cloud environments. HangFix first leverages stack trace analysis to localize the hang function and then performs root cause pattern matching to classify hang bugs into different types based on likely root causes. Next, HangFix generates effective code patches based on the identified root cause patterns. We have implemented a prototype of HangFix and evaluated the system on 42 real-world software hang bugs in 10 commonly used cloud server applications. Our results show that HangFix can successfully fix 40 out of 42 hang bugs in seconds.
Jingzhu He, Xiaohui Gu, Guoliang Jin
SoCC3
2019 FabZK: Supporting Privacy-Preserving, Auditable Smart Contracts in Hyperledger Fabric
abstract
On a Blockchain network, transaction data are exposed to all participants. To preserve privacy and confidentiality in transactions, while still maintaining data immutability, we design and implement FabZK. FabZK conceals transaction details on a shared ledger by storing only encrypted data from each transaction (e.g., payment amount), and by anonymizing the transactional relationship (e.g., payer and payee) between members in a Blockchain network. It achieves both privacy and auditability by supporting verifiable Pedersen commitments and constructing zero-knowledge proofs. FabZK is implemented as an extension to the open source Hyperledger Fabric. It provides APIs to easily enable data privacy in both client code and chaincode. It also supports on-demand, automated auditing based on encrypted data. Our evaluation shows that FabZK offers strong privacy-preserving capabilities, while delivering reasonable performance for the applications developed based on its framework.
Nerla Jean-Louis, Shu Tao, Xiaohui Gu
DSN5
2019 A Study on Container Vulnerability Exploit Detection
abstract
Containers have become increasingly popular for deploying applications in cloud computing infrastructures. However, recent studies have shown that containers are prone to various security attacks. In this paper, we conduct a study on the effectiveness of various vulnerability detection schemes for containers. Specifically, we implement and evaluate a set of static and dynamic vulnerability attack detection schemes using 28 real world vulnerability exploits that widely exist in docker images. Our results show that the static vulnerability scanning scheme only detects 3 out of 28 tested vulnerabilities and dynamic anomaly detection schemes detect 22 vulnerability exploits. Combining static and dynamic schemes can further improve the detection rate to 86% (i.e., 24 out of 28 exploits). We also observe that the dynamic anomaly detection scheme can achieve more than 20 seconds lead time (i.e., a time window before attacks succeed) for a group of commonly seen attacks in containers that try to gain a shell and execute arbitrary code.
Olufogorehan Tunde-Onadele, Jingzhu He, Xiaohui Gu
IC2E4
2019 TFix: Automatic Timeout Bug Fixing in Production Server Systems
abstract
Timeout is widely used to handle unexpected failures in distributed systems. However, improper use of timeout schemes can cause serious availability and performance issues, which is often difficult to fix due to lack of diagnostic information. In this paper, we present TFix, an automatic timeout bug fixing system for correcting misused timeout bugs in production systems. TFix adopts a drill-down bug analysis protocol that can narrow down the root cause of a misused timeout bug and producing recommendations for correcting the root cause. TFix first employs a system call frequent episode mining scheme to check whether a timeout bug is caused by a misused timeout variable. TFix then employs application tracing to identify timeout affected functions. Next, TFix uses taint analysis to localize the misused timeout variable. Last, TFix produces recommendations for proper timeout variable values based on the tracing results during normal runs. We have implemented a prototype of TFix and conducted extensive experiments using 13 real world server timeout bugs. Our experimental results show that TFix can correctly localize the misused timeout variables and suggest proper timeout values for fixing those bugs.
Jingzhu He, Xiaohui Gu
ICDCS3
2019 Hytrace: A Hybrid Approach to Performance Bug Diagnosis in Production Cloud Infrastructures
abstract
Server applications running inside production cloud infrastructures are prone to various performance problems (e.g., software hang, performance slowdown). When those problems occur, developers often have little clue to diagnose those problems. In this paper, we present Hytrace, a novel hybrid approach to diagnosing performance problems in production cloud infrastructures. Hytrace combines rule-based static analysis and runtime inference techniques to achieve higher bug localization accuracy than pure-static and pure-dynamic approaches for performance bugs. Hytrace does not require source code and can be applied to both compiled and interpreted programs such as C/C++ and Java. We conduct experiments using real performance bugs from seven commonly used server applications in production cloud infrastructures. The results show that our approach can significantly improve the performance bug diagnosis accuracy compared to existing diagnosis techniques.
Daniel Joseph Dean, Xiaohui Gu, Shan Lu 0001
IEEE Trans. Parallel Distributed Syst.4
2018 DScope: Detecting Real-World Data Corruption Hang Bugs in Cloud Server Systems
abstract
Cloud server systems such as Hadoop and Cassandra have enabled many real-world data-intensive applications running inside computing clouds. However, those systems present many data-corruption and performance problems which are notoriously difficult to debug due to the lack of diagnosis information. In this paper, we present DScope, a tool that statically detects data-corruption related software hang bugs in cloud server systems. DScope statically analyzes I/O operations and loops in a software package, and identifies loops whose exit conditions can be affected by I/O operations through returned data, returned error code, or I/O exception handling. After identifying those loops which are prone to hang problems under data corruption, DScope conducts loop bound and loop stride analysis to prune out false positives. We have implemented DScope and evaluated it using 9 common cloud server systems. Our results show that DScope can detect 42 real software hang bugs including 29 newly discovered software hang bugs. In contrast, existing bug detection tools miss detecting most of those bugs.
Jingzhu He, Xiaohui Gu, Shan Lu 0001
SoCC3
2018 Pcatch: automatically detecting performance cascading bugs in cloud systems
abstract
Distributed systems have become the backbone of modern clouds. Users often expect high scalability and performance isolation from distributed systems. Unfortunately, a type of poor software design, which we refer to as performance cascading bugs (PCbugs), can often cause the slowdown of non-scalable code in one job to propagate, causing global performance degradation and even threatening system availability.
Shan Lu 0001, Yiming Zhang 0003, Haryadi S. Gunawi, Xiaohui Gu, Xicheng Lu, Dongsheng Li 0001
EuroSys7
2018 Understanding Real-World Timeout Problems in Cloud Server Systems
abstract
Timeouts are commonly used to handle unexpected failures in distributed systems. In this paper, we conduct a comprehensive study to characterize real-world timeout problems in 11 commonly used cloud server systems (e.g., Hadoop, HDSF, Spark, Cassandra, etc.). Our study reveals timeout problems are widespread among cloud server systems. We categorize those timeout problems in three aspects: 1) what are the root causes of those timeout problems? 2) what impact can timeout problems impose to cloud systems? 3) how are timeout problems currently diagnosed or misdiagnosed? Our results show that root causes of timeout problems include misused timeout, missing timeout, improper timeout handling, unnecessary timeout, and clock drifting. We further find timeout bugs impose serious impact (e.g., system hang or crash, job failure, performance degradation, data loss) to both applications and systems. Our study also shows that 60% of the bugs do not produce any error messages and 12% bugs produce misleading error messages, which makes it difficult to diagnose those timeout bugs.
Jingzhu He, Xiaohui Gu, Shan Lu 0001
IC2E3
2017 Hytrace: a hybrid approach to performance bug diagnosis in production cloud infrastructures
abstract
Server applications running inside production cloud infrastructures are prone to various performance problems (e.g., software hang, performance slow down). When those problems occur, developers often have little clue to diagnose those problems. We present HyTrace, a novel hybrid approach to diagnosing performance problems in production cloud infrastructures. HyTrace combines rule-based static analysis and runtime inference techniques to achieve higher bug localization accuracy than pure-static and pure-dynamic approaches for performance bugs. HyTrace does not require source code and can be applied to both compiled and interpreted programs such as C/C++ and Java. We conduct experiments using real performance bugs from seven commonly used server applications. The results show that our approach can significantly improve the performance bug diagnosis accuracy compared to existing diagnosis techniques.
Daniel Dean, Xiaohui Gu, Shan Lu 0001
SoCC4
2017 A Study of Security Vulnerabilities on Docker Hub
abstract
Docker containers have recently become a popular approach to provision multiple applications over shared physical hosts in a more lightweight fashion than traditional virtual machines. This popularity has led to the creation of the Docker Hub registry, which distributes a large number of official and community images. In this paper, we study the state of security vulnerabilities in Docker Hub images. We create a scalable Docker image vulnerability analysis (DIVA) framework that automatically discovers, downloads, and analyzes both official and community images on Docker Hub. Using our framework, we have studied 356,218 images and made the following findings: (1) both official and community images contain more than 180 vulnerabilities on average when considering all versions; (2) many images have not been updated for hundreds of days; and (3) vulnerabilities commonly propagate from parent images to child images. These findings demonstrate a strong need for more automated and systematic methods of applying security updates to Docker images and our current Docker image analysis framework provides a good foundation for such automatic security update.
Xiaohui Gu, William Enck
CODASPY2
2016 Performance Analysis of a Multi-tenant In-Memory Data Grid
abstract
Distributed key-value stores have become indispensable for large scale low latency applications. Many cloud services have deployed in-memory data grids for their enterprise infrastructures and support multi-tenancy services. But it is still difficult to provide consistent performance to all tenants for fluctuating workloads that need to scale out. Many popular key-value stores suffer from performance problems at scale and different tenant requirements. To this front, we present our study with Hazelcast, a popular open source data grid, and provide insights to contention and performance bottlenecks. Through experimental analysis, this paper uncovers scenarios of performance degradation followed by optimized performance via end-point multiplexing. Our study suggests that processing increasing number of client requests spawning fewer number of threads help improve performance.
Anwesha Das 0001, Frank Mueller 0001, Xiaohui Gu, Arun Iyengar
CLOUD3
2016 RDE: Replay DEbugging for Diagnosing Production Site Failures
abstract
Online service failures in production computing environments are notoriously difficult to debug. One of the key challenges is to allow the developer to replay the failure execution within an interactive debugging tool such as GDB. Previous work has proposed in-situ approaches to inferring the production-run failure path within the production environment. However, those tools may sometimes suggest failure execution paths that are infeasible to reach by any program inputs. Moreover, production site often does not record or provide failure-triggering inputs due to the user privacy concern. In this paper, we present RDE, a Replay DEbug system that can replay a production-site failure at the development site within an interactive debugging environment without requiring user inputs. RDE takes an inferred production failure path as input and performs execution synthesis using a new guided symbolic execution technique. RDE can tolerate imprecise or inaccurate failure path information by navigating the symbolic execution along a set of selected paths. RDE synthesizes an input from the selected symbolic execution path which can be fed to a debugging tool to replay the failure. We have implemented an initial prototype of RDE and tested it with a set of coreutils bugs. The results show that RDE can successfully replay all the tested bugs within GDB.
Hiep Nguyen, Xiaohui Gu, Shan Lu 0001
SRDS3
2016 PerfCompass: Online Performance Anomaly Fault Localization and Inference in Infrastructure-as-a-Service Clouds
abstract
Infrastructure-as-a-service clouds are becoming widely adopted. However, resource sharing and multi-tenancy have made performance anomalies a top concern for users. Timely debugging those anomalies is paramount for minimizing the performance penalty for users. Unfortunately, this debugging often takes a long time due to the inherent complexity and sharing nature of cloud infrastructures. When an application experiences a performance anomaly, it is important to distinguish between faults with a global impact and faults with a local impact as the diagnosis and recovery steps forfaults with a global impact or local impact are quite different. In this paper, we present PerfCompass, an online performance anomaly fault debugging tool that can quantify whether a production-run performance anomaly has a global impact or local impact. PerfCompass can use this information to suggest the root cause as either an external fault (e.g., environment-based) or an internal fault (e.g., software bugs). Furthermore, PerfCompass can identify top affected system calls to provide useful diagnostic hints for detailed performance debugging. PerfCompass does not require source code or runtime application instrumentation, which makes it practical for production systems. We have tested PerfCompass by running five common open source systems (e.g., Apache, MySQL, Tomcat, Hadoop, Cassandra) inside a virtualized cloud testbed. Our experiments use a range of common infrastructure sharing issues and real software bugs. The results show that PerfCompass accurately classifies 23 out of the 24 tested cases without calibration and achieves 100 percent accuracy with calibration. PerfCompass provides useful diagnosis hints within several minutes and imposes negligible runtime overhead to the production system during normal execution time.
Daniel Joseph Dean, Hiep Nguyen, Xiaohui Gu, Anca Sailer, Andrzej Kochut
IEEE Trans. Parallel Distributed Syst.4
2015 Understanding Real World Data Corruptions in Cloud Systems
abstract
Big data processing is one of the killer applications for cloud systems. MapReduce systems such as Hadoop are the most popular big data processing platforms used in the cloud system. Data corruption is one of the most critical problems in cloud data processing, which not only has serious impact on the integrity of individual application results but also affects the performance and availability of the whole data processing system. In this paper, we present a comprehensive study on 138 real world data corruption incidents reported in Hadoop bug repositories. We characterize those data corruption problems in four aspects: 1) what impact can data corruption have on the application and system? 2) how is data corruption detected? 3) what are the causes of the data corruption? and 4) what problems can occur while attempting to handle data corruption? Our study has made the following findings: 1) the impact of data corruption is not limited to data integrity, 2) existing data corruption detection schemes are quite insufficient: only 25% of data corruption problems are correctly reported, 42% are silent data corruption without any error message, and 21% receive imprecise error report. We also found the detection system raised 12% false alarms, 3) there are various causes of data corruption such as improper runtime checking, race conditions, inconsistent block states, improper network failure handling, and improper node crash handling, and 4) existing data corruption handling mechanisms (i.e., data replication, replica deletion, simple re-execution) make frequent mistakes including replicating corrupted data blocks, deleting uncorrupted data blocks, or causing undesirable resource hogging.
Daniel Joseph Dean, Xiaohui Gu
IC2E3
2014 PerfScope: Practical Online Server Performance Bug Inference in Production Cloud Computing Infrastructures
abstract
Performance bugs which manifest in a production cloud computing infrastructure are notoriously difficult to diagnose because of both the difficulty of reproducing those bugs and the lack of debugging information. In this paper, we present PerfScope, a practical online performance bug inference tool to help the developer understand how a performance bug happened during the production run. PerfScope achieves online bug inference to obviate the need for offline bug reproduction. PerfScope does not require application source code or any runtime instrumentation to the production system. PerfScope is application-agnostic, which can support both interpreted and compiled programs running inside a cloud infrastructure.
Daniel Joseph Dean, Hiep Nguyen, Xiaohui Gu, Hui Zhang 0002, Junghwan Rhee, Nipun Arora, Geoff Jiang
SoCC3
2014 PREC: practical root exploit containment for android devices
abstract
Application markets such as the Google Play Store and the Apple App Store have become the de facto method of distributing software to mobile devices. While official markets dedicate significant resources to detecting malware, state-of-the-art malware detection can be easily circumvented using logic bombs or checks for an emulated environment. We present a Practical Root Exploit Containment (PREC) framework that protects users from such conditional malicious behavior. PREC can dynamically identify system calls from high-risk components (e.g., third-party native libraries) and execute those system calls within isolated threads. Hence, PREC can detect and stop root exploits with high accuracy while imposing low interference to benign applications. We have implemented PREC and evaluated our methodology on 140 most popular benign applications and 10 root exploit malicious applications. Our results show that PREC can successfully detect and stop all the tested malware while reducing the false alarm rates by more than one order of magnitude over traditional malware detection algorithms. PREC is light-weight, which makes it practical for runtime on-device root exploit detection and containment.
Tsung-Hsuan Ho, Daniel Joseph Dean, Xiaohui Gu, William Enck
CODASPY3
2014 Insight: In-situ Online Service Failure Path Inference in Production Computing Infrastructures
Hiep Nguyen, Daniel Joseph Dean, Kamal Kc, Xiaohui Gu
USENIX ATC4
2014 Scalable Distributed Service Integrity Attestation for Software-as-a-Service Clouds
abstract
Software-as-a-service (SaaS) cloud systems enable application service providers to deliver their applications via massive cloud computing infrastructures. However, due to their sharing nature, SaaS clouds are vulnerable to malicious attacks. In this paper, we present IntTest, a scalable and effective service integrity attestation framework for SaaS clouds. IntTest provides a novel integrated attestation graph analysis scheme that can provide stronger attacker pinpointing power than previous schemes. Moreover, IntTest can automatically enhance result quality by replacing bad results produced by malicious attackers with good results produced by benign service providers. We have implemented a prototype of the IntTest system and tested it on a production cloud computing infrastructure using IBM System S stream processing applications. Our experimental results show that IntTest can achieve higher attacker pinpointing accuracy than existing approaches. IntTest does not require any special hardware or secure kernel support and imposes little performance impact to the application, which makes it practical for large-scale cloud systems.
Juan Du 0006, Daniel Joseph Dean, Yongmin Tan, Xiaohui Gu, Ting Yu 0001
IEEE Trans. Parallel Distributed Syst.4
2013 FChain: Toward Black-Box Online Fault Localization for Cloud Systems
abstract
Distributed applications running inside cloud systems are prone to performance anomalies due to various reasons such as resource contentions, software bugs, and hardware failures. One big challenge for diagnosing an abnormal distributed application is to pinpoint the faulty components. In this paper, we present a black-box online fault localization system called FChain that can pinpoint faulty components immediately after a performance anomaly is detected. FChain first discovers the onset time of abnormal behaviors at different components by distinguishing the abnormal change point from many change points caused by normal workload fluctuations. Faulty components are then pinpointed based on the abnormal change propagation patterns and inter-component dependency relationships. FChain performs runtime validation to further filter out false alarms. We have implemented FChain on top of the Xen platform and tested it using several benchmark applications (RUBiS, Hadoop, and IBM System S). Our experimental results show that FChain can quickly pinpoint the faulty components with high accuracy within a few seconds. FChain can achieve up to 90% higher precision and 20% higher recall than existing schemes. FChain is non-intrusive and light-weight, which imposes less than 1% overhead to the cloud system.
Hiep Nguyen, Zhiming Shen, Yongmin Tan, Xiaohui Gu
ICDCS4
2013 Resilient Self-Compressive Monitoring for Large-Scale Hosting Infrastructures
abstract
Large-scale hosting infrastructures have become the fundamental platforms for many real-world systems such as cloud computing infrastructures, enterprise data centers, and massive data processing systems. However, it is a challenging task to achieve both scalability and high precision while monitoring a large number of intranode and internode attributes (e.g., CPU usage, free memory, free disk, internode network delay). In this paper, we present the design and implementation of a Resilient self-Compressive Monitoring (RCM) system for large-scale hosting infrastructures. RCM achieves scalable distributed monitoring by performing online data compression to reduce remote data collection cost. RCM provides failure resilience to achieve robust monitoring for dynamic distributed systems where host and network failures are common. We have conducted extensive experiments using a set of real monitoring data from NCSU's virtual computing lab (VCL), PlanetLab, a Google cluster, and real Internet traffic matrices. The experimental results show that RCM can achieve up to 200 percent higher compression ratio and several orders of magnitude less overhead than the existing approaches.
Yongmin Tan, Vinay Venkatesh, Xiaohui Gu
IEEE Trans. Parallel Distributed Syst.3
2012 PREPARE: Predictive Performance Anomaly Prevention for Virtualized Cloud Systems
abstract
Virtualized cloud systems are prone to performance anomalies due to various reasons such as resource contentions, software bugs, and hardware failures. In this paper, we present a novel Predictive Performance Anomaly Prevention (PREPARE) system that provides automatic performance anomaly prevention for virtualized cloud computing infrastructures. PREPARE integrates online anomaly prediction, learning-based cause inference, and predictive prevention actuation to minimize the performance anomaly penalty without human intervention. We have implemented PREPARE on top of the Xen platform and tested it on the NCSU's Virtual Computing Lab using a commercial data stream processing system (IBM System S) and an online auction benchmark (RUBiS). The experimental results show that PREPARE can effectively prevent performance anomalies while imposing low overhead to the cloud infrastructure.
Yongmin Tan, Hiep Nguyen, Zhiming Shen, Xiaohui Gu, Chitra Venkatramani, Deepak Rajan
ICDCS4
2011 CloudScale: elastic resource scaling for multi-tenant cloud systems
abstract
Elastic resource scaling lets cloud systems meet application service level objectives (SLOs) with minimum resource provisioning costs. In this paper, we present CloudScale, a system that automates fine-grained elastic resource scaling for multi-tenant cloud computing infrastructures. CloudScale employs online resource demand prediction and prediction error handling to achieve adaptive resource allocation without assuming any prior knowledge about the applications running inside the cloud. CloudScale can resolve scaling conflicts between applications using migration, and integrates dynamic CPU voltage/frequency scaling to achieve energy savings with minimal effect on application SLOs. We have implemented CloudScale on top of Xen and conducted extensive experiments using a set of CPU and memory intensive applications (RUBiS, Hadoop, IBM System S). The results show that CloudScale can achieve significantly higher SLO conformance than other alternatives with low resource and energy cost. CloudScale is non-intrusive and light-weight, and imposes negligible overhead (< 2% CPU in Domain 0) to the virtualized computing cluster.
Zhiming Shen, Sethuraman Subbiah, Xiaohui Gu, John Wilkes
SoCC3
2011 Adaptive data-driven service integrity attestation for multi-tenant cloud systems
abstract
Cloud systems provide a cost-effective service hosting infrastructure for application service providers (ASPs). However, cloud systems are often shared by multiple tenants from different security domains, which makes them vulnerable to various malicious attacks. Moreover, cloud systems often host long-running applications such as massive data processing, which provides more opportunities for attackers to exploit the system vulnerability and perform strategic attacks. In this paper, we present AdapTest, a novel adaptive data-driven runtime service integrity attestation framework for multi-tenant cloud systems. AdapTest can significantly reduce attestation overhead and shorten detection delay by adaptively selecting attested nodes based on dynamically derived trust scores. Our scheme treats attested services as black-boxes and does not impose any special hardware or software requirements on the cloud system or ASPs. We have implemented AdapTest on top of the IBM System S stream processing system and tested it within a virtualized computing cluster. Our experimental results show that AdapTest can reduce attestation overhead by up to 60% and shorten the detection delay by up to 40% compared to previous approaches.
Juan Du 0006, Xiaohui Gu, Nidhi Shah
IWQoS2
2011 OLIC: OnLine Information Compression for scalable hosting infrastructure monitoring
abstract
Quality-of-service (QoS) management often requires a continuous monitoring service to provide updated information about different hosts and network links in the managed system. However, it is a challenging task to achieve both scalability and precision for monitoring various intra-node and inter-node metrics (e.g., CPU, memory, disk, network delay) in a large-scale hosting infrastructure. In this paper, we present a novel OnLine Information Compression (OLIC) system to achieve scalable fine-grained hosting infrastructure monitoring. OLIC models continuous snapshots of a hosting infrastructure as a sequence of images and performs online monitoring data compression to significantly reduce the monitoring cost. We have implemented a prototype of the OLIC system and deployed it on the PlanetLab and NCSU's virtual computing lab (VCL). We have conducted extensive experiments using a set of real monitoring data from VCL, Planetlab, and a Google cluster as well as a real Internet traffic matrix trace. The experimental results show that OLIC can achieve much higher compression ratios with several orders of magnitude less overhead than previous approaches.
Yongmin Tan, Xiaohui Gu, Vinay Venkatesh
IWQoS2
2011 ELT: Efficient Log-based Troubleshooting System for Cloud Computing Infrastructures
abstract
We present an Efficient Log-based Troubleshooting(ELT) system for cloud computing infrastructures. ELT adopts a novel hybrid log mining approach that combines coarse-grained and fine-grained log features to achieve both high accuracy and low overhead. Moreover, ELT can automatically extract key log messages and perform invariant checking to greatly simplify the troubleshooting task for the system administrator. We have implemented a prototype of the ELT system and conducted an extensive experimental study using real management console logs of a production cloud system and a Hadoop cluster. Our experimental results show that ELT can achieve more efficient and powerful troubleshooting support than existing schemes. More importantly, ELT can find software bugs that cannot be detected by current cloud system management practice.
Kamal Kc, Xiaohui Gu
SRDS2
2010 On verifying stateful dataflow processing services in large-scale cloud systems
abstract
Cloud computing needs to provide integrity assurance in order to support security sensitive application services such as critical dataflow processing. In this paper, we present a novel RObust Service Integrity Attestation (ROSIA) framework that can efficiently verify the integrity of stateful dataflow processing services and pinpoint malicious service providers within a large-scale cloud system. ROSIA achieves robustness by supporting stateful dataflow services such as windowed stream operators, and performing integrated consistency check to detect colluding attacks. We have implemented ROSIA on top of the IBM System S dataflow processing system and tested it on the NCSU virtual computing lab. Our experimental results show that our scheme is feasible and efficient for large-scale cloud systems.
Juan Du 0006, Xiaohui Gu, Ting Yu 0001
CCS2
2010 RunTest: assuring integrity of dataflow processing in cloud computing infrastructures
abstract
Cloud computing has emerged as a multi-tenant resource sharing platform, which allows different service providers to deliver software as services in an economical way. However, for many security sensitive applications such as critical data processing, we must provide necessary security protection for migrating those critical application services into shared open cloud infrastructures. In this paper, we present RunTest, a scalable runtime integrity attestation framework to assure the integrity of dataflow processing in cloud infrastructures. RunTest provides light-weight application-level attestation methods to dynamically verify the integrity of data processing results and pinpoint malicious service providers when inconsistent results are detected. We have implemented RunTest within IBM System S dataflow processing system and tested it on NCSU virtual computing lab. Our experimental results show that our scheme is effective and imposes low performance impact for dataflow processing in the cloud infrastructure.
Juan Du 0006, Xiaohui Gu, Ting Yu 0001
AsiaCCS3
2010 PRESS: PRedictive Elastic ReSource Scaling for cloud systems
abstract
Cloud systems require elastic resource allocation to minimize resource provisioning costs while meeting service level objectives (SLOs). In this paper, we present a novel PRedictive Elastic reSource Scaling (PRESS) scheme for cloud systems. PRESS unobtrusively extracts fine-grained dynamic patterns in application resource demands and adjust their resource allocations automatically. Our approach leverages light-weight signal processing and statistical learning algorithms to achieve online predictions of dynamic application resource requirements. We have implemented the PRESS system on Xen and tested it using RUBiS and an application load trace from Google. Our experiments show that we can achieve good resource prediction accuracy with less than 5% over-estimation error and near zero under-estimation error, and elastic resource scaling can both significantly reduce resource waste and SLO violations.
Zhenhuan Gong, Xiaohui Gu, John Wilkes
CNSM2
2010 Highly available component sharing in large-scale multi-tenant cloud systems
abstract
A multi-tenant cloud system allows multiple users to share a common physical computing infrastructure in a cost-effective way. Component sharing is highly desired in such a shared computing infrastructure, where different tenants can leverage each other's information and expertise to fulfill their own tasks. However, it is challenging to maintain the availability of sharable component resources in a large-scale cloud infrastructure, as cloud tenants are fully autonomous and highly dynamic. In this paper, we present a novel highly available component sharing system for large-scale multi-tenant cloud systems. We describe a component availability prediction scheme to identify endangered components (i.e., components at risk of extinction) within the infrastructure. The system then performs predictive replication based on the availability prediction results to preserve those endangered components. Thus, our system can preserve the availability of all component resources with low cost. Theoretical analysis and large-scale simulation are used to quantify the accuracy of our component availability prediction, and the efficiency of predictive replication. Experimental results show that our scheme can predict endangered components with high accuracy, and achieve up to 99% availability with about 15% of the full replication cost.
Juan Du 0006, Xiaohui Gu, Douglas S. Reeves
HPDC2
2010 PAC: Pattern-driven Application Consolidation for Efficient Cloud Computing
abstract
To reduce cloud system resource cost, application consolidation is a must. In this paper, we present a novel pattern driven application consolidation (PAC) system to achieve efficient resource sharing in virtualized cloud computing infrastructures. PAC employs signal processing techniques to dynamically discover significant patterns called signatures of different applications and hosts. PAC then performs dynamic application consolidation based on the extracted signatures. We have implemented a prototype of the PAC system on top of the Xen virtual machine platform and tested it on the NCSU Virtual Computing Lab. We have tested our system using RUBiS benchmarks, Hadoop data processing systems, and IBM System S stream processing system. Our experiments show that 1) PAC can efficiently discover repeating resource usage patterns in the tested applications; 2) Signatures can reduce resource prediction errors by 50-90% compared to traditional coarse-grained schemes; 3) PAC can improve application performance by up to 50% when running a large number of applications on a shared cluster.
Zhenhuan Gong, Xiaohui Gu
MASCOTS2
2010 On Predictability of System Anomalies in Real World
abstract
As computer systems become increasingly complex, system anomalies have become major concerns in system management. In this paper, we present a comprehensive measurement study to quantify the predictability of different system anomalies. Online anomaly prediction allows the system to foresee impending anomalies so as to take proper actions to mitigate anomaly impact. Our anomaly prediction approach combines feature value prediction with statistical classification methods. We conduct extensive measurement study to investigate anomalous behavior of three systems in the real world: PlanetLab, SMART hard drive data, and IBM System S. We observe that real world system anomalies do exhibit predictability, which can be predicted with high accuracy and significant lead time.
Yongmin Tan, Xiaohui Gu
MASCOTS2
2010 Adaptive system anomaly prediction for large-scale hosting infrastructures
abstract
Large-scale hosting infrastructures require automatic system anomaly management to achieve continuous system operation. In this paper, we present a novel adaptive runtime anomaly prediction system, called ALERT, to achieve robust hosting infrastructures. In contrast to traditional anomaly detection schemes, ALERT aims at raising advance anomaly alerts to achieve just-in-time anomaly prevention. We propose a novel context-aware anomaly prediction scheme to improve prediction accuracy in dynamic hosting infrastructures. We have implemented the ALERT system and deployed it on several production hosting infrastructures such as IBM System S stream processing cluster and PlanetLab. Our experiments show that ALERT can achieve high prediction accuracy for a range of system anomalies and impose low overhead to the hosting infrastructure.
Yongmin Tan, Xiaohui Gu, Haixun Wang
PODC2
2009 SecureMR: A Service Integrity Assurance Framework for MapReduce
abstract
MapReduce has become increasingly popular as a powerful parallel data processing model. To deploy MapReduce as a data processing service over open systems such as service oriented architecture, cloud computing, and volunteer computing, we must provide necessary security mechanisms to protect the integrity of MapReduce data processing services. In this paper, we present SecureMR, a practical service integrity assurance framework for MapReduce. SecureMR consists of five security components, which provide a set of practical security mechanisms that not only ensure MapReduce service integrity as well as to prevent replay and denial of service (DoS) attacks, but also preserve the simplicity, applicability and scalability of MapReduce. We have implemented a prototype of SecureMR based on Hadoop, an open source MapReduce implementation. Our analytical study and experimental results show that SecureMR can ensure data processing service integrity while imposing low performance overhead.
Juan Du 0006, Ting Yu 0001, Xiaohui Gu
ACSAC4
2009 Online Anomaly Prediction for Robust Cluster Systems
abstract
In this paper, we present a stream-based mining algorithm for online anomaly prediction. Many real-world applications such as data stream analysis requires continuous cluster operation. Unfortunately, today's large-scale cluster systems are still vulnerable to various software and hardware problems. System administrators are often overwhelmed by the tasks of correcting various system anomalies such as processing bottlenecks (i.e., full stream buffers), resource hot spots, and service level objective (SLO) violations. Our anomaly prediction scheme raises early alerts for impending system anomalies and suggests possible anomaly causes. Specifically, we employ Bayesian classification methods to capture different anomaly symptoms and infer anomaly causes. Markov models are introduced to capture the changing patterns of different measurement metrics. More importantly, our scheme combines Markov models and Bayesian classification methods to predict when a system anomaly will appear in the foreseeable future and what are the possible anomaly causes. To the best of our knowledge, our work provides the first stream-based mining algorithm for predicting system anomalies. We have implemented our approach within the IBM System S distributed stream processing cluster, and conducted case study experiments using fully implemented distributed data analysis applications processing real application workloads. Our experiments show that our approach efficiently predicts and diagnoses several bottleneck anomalies with high accuracy while imposing low overhead to the cluster system.
Xiaohui Gu, Haixun Wang
ICDE1
2009 SigLM: Signature-driven load management for cloud computing infrastructures
abstract
Cloud computing has emerged as a promising platform that grants users with direct yet shared access to computing resources and services without worrying about the internal complex infrastructure. Unlike traditional batch service model, cloud service model adopts a pay-as-you-go form, which demands explicit and precise resource control. In this paper, we present SigLM, a novel Signature-driven Load Management system to achieve quality-aware service delivery in shared cloud computing infrastructures. SigLM dynamically captures fine-grained signatures of different application tasks and cloud nodes using time series patterns, and performs precise resource metering and allocation based on the extracted signatures. SigLM employs dynamic time warping algorithm and multi-dimensional time series indexing to achieve efficient signature pattern matching. Our experiments using real load traces collected on the PlanetLab show that SigLM can improve resource provisioning performance by 30–80% compared to existing approaches. SigLM is scalable and efficient, which imposes less than 1% overhead to the system and can perform signature matching within tens of milliseconds.
Zhenhuan Gong, Prakash Ramaswamy, Xiaohui Gu, Xiaosong Ma
IWQoS3
2009 QoS-Aware Shared Component Composition for Distributed Stream Processing Systems
abstract
Many emerging online data analysis applications require applying continuous query operations such as correlation, aggregation, and filtering to data streams in real time. Distributed stream processing systems allow in-network stream processing to achieve better scalability and quality-of-service (QoS) provision. In this paper, we present Synergy, a novel distributed stream processing middleware that provides automatic sharing-aware component composition capability. Synergy enables efficient reuse of both result streams and processing components, while composing distributed stream processing applications with QoS demands. It provides a set of fully distributed algorithms to discover and evaluate the reusability of available result streams and processing components when instantiating new stream applications. Specifically, Synergy performs QoS impact projection to examine whether the shared processing can cause QoS violations on currently running applications. The QoS impact projection algorithm can handle different types of streams including both regular traffic and bursty traffic. If no existing processing components can be reused, Synergy dynamically deploys new components at strategic locations to satisfy new application requests. We have implemented a prototype of the Synergy middleware and evaluated its performance on both PlanetLab and simulation testbeds. The experimental results show that Synergy can achieve much better resource utilization and QoS provisioning than previously proposed schemes, by judiciously sharing streams and components during application composition.
Thomas Repantis, Xiaohui Gu, Vana Kalogeraki
IEEE Trans. Parallel Distributed Syst.2
2008 Toward Predictive Failure Management for Distributed Stream Processing Systems
abstract
Distributed stream processing systems (DSPSs) have many important applications such as sensor data analysis, network security, and business intelligence. Failure management is essential for DSPSs that often require highly-available system operations. In this paper, we explore a new predictive failure management approach that employs online failure prediction to achieve more efficient failure management than previous reactive or proactive failure management approaches. We employ light-weight stream-based classification methods to perform online failure forecast. Based on the prediction results, the system can take differentiated failure preventions on abnormal components only. Our failure prediction model is tunable, which can achieve a desired tradeoff between failure penalty reduction and prevention cost based on a user-defined reward function. To achieve low-overhead online learning, we propose adaptive data stream sampling schemes to adaptively adjust measurement sampling rates based on the states of monitored components, and maintain a limited size of historical training data using reservoir sampling. We have implemented an initial prototype of the predictive failure management framework within the IBM System S distributed stream processing system. Experiment results show that our system can achieve more efficient failure management than conventional reactive and proactive approaches, while imposing low overhead to the DSPS.
Xiaohui Gu, Spiros Papadimitriou, Philip S. Yu, Shu-Ping Chang
ICDCS1
2008 Online Failure Forecast for Fault-Tolerant Data Stream Processing
abstract
In this paper, we present a new online failure forecast system to achieve predictive failure management for fault-tolerant data stream processing. Different from previous reactive or proactive approaches, predictive failure management employs failure forecast to perform informed and just-in-time preventive actions on abnormal components only. We employ stream-based online learning methods to continuously classify runtime operator state into normal, alert, or failure, based on collected feature streams. We have implemented the online failure forecast system as part of the IBM System S stream processing system. Our experiments show that the on-line failure forecast system can achieve good prediction accuracy for a range of stream processing software failures, while imposing low overhead to the stream system.
Xiaohui Gu, Spiros Papadimitriou, Philip S. Yu, Shu-Ping Chang
ICDE1
2008 peerTalk: A Peer-to-Peer Multiparty Voice-over-IP System
abstract
Multiparty voice-over-IP (MVolP) services allow a group of people to freely communicate with each other via the Internet, which have many important applications such as online gaming and teleconferencing. In this paper, we present a peer-to-peer MVolP system called peerTalk. Compared to traditional approaches such as server-based mixing, peerTalk achieves better scalability and failure resilience by dynamically distributing the stream processing workload among different peers. Particularly, peerTalk decouples the MVolP service delivery into two phases: mixing phase and distribution phase. The decoupled model allows us to explore the asymmetric property of MVolP services (for example, distinct speaking/listening activities and unequal inbound/outbound bandwidths) so that the system can better adapt to distinct stream mixing and distribution requirements. To overcome arbitrary peer departures/ failures, peerTalk provides lightweight backup schemes to achieve fast failure recovery. We have implemented a prototype of the peerTalk system and evaluated its performance using both a large-scale simulation testbed and a real Internet environment. Our initial implementation demonstrates the feasibility of our approach and shows promising results: peerTalk can outperform existing approaches such as P2P overlay multicast and coupled distributed processing for providing MVolP services.
Xiaohui Gu, Philip S. Yu, Zon-Yin Shae
IEEE Trans. Parallel Distributed Syst.1
2007 Adaptive Load Diffusion for Multiway Windowed Stream Joins
abstract
In this paper, we present an adaptive load diffusion operator to enable scalable processing of multiway windowed stream joins (MWSJs) using a cluster system. The load diffusion is achieved by a set of novel semantics-pre serving tuple routing algorithms. Different from previous work, the load diffusion operator can (1) preserve the MWSJ semantics while spreading tuples to different hosts for parallel join processing; (2) achieve fine-grained load balancing among distributed hosts; and (3) perform semantics-preserving online adaptations to maintain optimal performance in dynamic stream environments. We have implemented a prototype of the distributed MWSJ framework on top of the System S distributed stream processing system. Our experiment results based on both real data streams and synthetic workloads show that the load diffusion algorithms can efficiently scale-up the performance of MWSJ processing with low overhead.
Xiaohui Gu, Philip S. Yu, Haixun Wang
ICDE1
2007 Toward Self-Managed Media Stream Processing Service Overlays
abstract
On-demand media stream processing service provisioning on top of a service overlay network (SON) has emerged as a promising approach to providing quality-aware and failure-resilient media streaming services. Although previous work has addressed different problems in media streaming, it is still an open problem to manage SON under dynamic stream environments. In this paper, we propose a novel model-based adaptive SON management framework that can achieve better QoS-aware media stream processing services than unmanaged SONs. We first describe a statistical SON model that can capture the popularity of different service functions and the interaction patterns between different services. We then present the model-based SON management framework and use the SON topology manager as a case study to illustrate the main idea. We are implementing a prototype of the self-managed SON framework and our initial experimental evaluation shows promising results of our approach.
Xiaohui Gu, Philip S. Yu
ICME1
2007 BridgeNet: An Adaptive Multi-Source Stream Dissemination Overlay Network
abstract
Emerging stream processing applications such as on-line data analysis often need to collect streaming information from geographically dispersed locations (e.g., different sensor networks). Different from conventional discrete data (e.g., messages), streaming data aretime-varyingandlong-lived, which provides both new challenges and opportunities for optimizing wide-area continuous data dissemination. In this paper, we presentBridgeNet, a novel biology-inspired stream dissemination overlay network that can dynamically learn stream patterns to achieve efficient multi-source stream dissemination. We propose a new distributed cell tree structure that can adaptively expand or contract itself in response to workload changes. BridgeNet performspattern-basedadaptations to deliver efficient stream disseminations without losing system stability. We have implemented a prototype of BridgeNet and conducted extensive experiments using both simulations and Planetlab deployment. The experimental results based on both synthetic workload and real data streams show that BridgeNet outperforms existing schemes for efficient multi-source stream dissemination.
Xiaohui Gu, Philip S. Yu
INFOCOM1
2007 Self-Configuring Information Management for Large-Scale Service Overlays
abstract
Service overlay networks (SON) provide important infrastructure support for many emerging distributed applications such as web service composition, distributed stream processing, and workflow management. Quality-sensitive distributed applications such as multimedia services and on-line data analysis often desire the SON to provide up-to-date dynamic information about different overlay nodes and overlay links. However, it is a challenging task to provide scalable and efficient information management for large-scale SONs, where both system conditions and application requirements can change over time. In this paper, we present InfoEye, a model-based self-configuring distributed information management system that consists of a set of monitoring sensors deployed on different overlay nodes. InfoEye can dynamically configure the operations of different sensors based on current statistical application query patterns and system attribute distributions. Thus, InfoEye can greatly improve the scalability of SON by answering information queries with minimum monitoring overhead. We have implemented a prototype of InfoEye and evaluated its performance using both extensive simulations and micro-benchmark experiments on PlanetLab. The experimental results show that InfoEye can significantly reduce the information management overhead compared with existing approaches. In addition, InfoEye can quickly reconfigure itself in response to application requirement and system information pattern changes.
Xiaohui Gu, Klara Nahrstedt
INFOCOM2
2007 Towards Graph Containment Search and Indexing
Chen Chen 0005, Xifeng Yan, Philip S. Yu, Jiawei Han 0001, Dong-Qing Zhang, Xiaohui Gu
VLDB6
2007 Challenges and Experience in Prototyping a Multi-Modal Stream Analytic and Monitoring Application on System S
Kun-Lung Wu, Philip S. Yu, Bugra Gedik, Kirsten Hildrum, Charu C. Aggarwal, Eric Bouillet, Wei Fan 0001, Xiaohui Gu, Gang Luo 0001, Haixun Wang
VLDB9
2006 Synergy: Sharing-Aware Component Composition for Distributed Stream Processing Systems
Thomas Repantis, Xiaohui Gu, Vana Kalogeraki
Middleware2
2006 ViCo: an adaptive distributed video correlation system
abstract
Many emerging applications such as video sensor monitoring can benefit from an on-line video correlation system, which can be used to discover linkages between different video streams in realtime. However, on-line video correlations are often resource-intensive where a single host can be easily overloaded. We present a novel adaptive distributed on-line video correlation system called ViCo. Unlike single stream processing, correlations between different video streams require a distributed execution system to observe a new correlation constraint that any two correlated data must be distributed to the same host. ViCo achieves three unique features: (1) correlation-awareness that ViCo can guarantee the correlation accuracy while spreading excessive workload on multiple hosts; (2) adaptability that the system can adjust algorithm behaviors and switch between different algorithms to adapt to dynamic stream environments; and (3) fine-granularity that the workload of one resource-intensive correlation request can be divided and distributed among multiple hosts. We have implemented and deployed a prototype of ViCo on a commercial cluster system. Our experiment results using both real videos and synthetic workloads show that ViCo outperforms existing techniques for scaling-up the performance of video correlations.
Xiaohui Gu, Ching-Yung Lin, Philip S. Yu
ACM Multimedia1
2006 Distributed multimedia service composition with statistical QoS assurances
abstract
Service composition allows multimedia services to be automatically composed from atomic service components based on dynamic service requirements. Previous work falls short for distributed multimedia service composition in terms of scalability, flexibility and quality-of-service (QoS) management. In this paper, we present a fully decentralized service composition framework, called SpiderNet, to address the challenges. SpiderNet provides statistical multiconstrained QoS assurances and load balancing for service composition. Moreover, SpiderNet supports directed acyclic graph composition topologies and exchangeable composition orders. We have implemented a prototype of SpiderNet and conducted experiments on both wide-area networks and a simulation testbed. Our experimental results show the feasibility and efficiency of the SpiderNet service composition framework.
Xiaohui Gu, Klara Nahrstedt
IEEE Trans. Multim.1
2006 On Composing Stream Applications in Peer-to-Peer Environments
abstract
Stream processing has become increasingly important as many emerging applications call for continuous real-time processing over data streams, such as voice-over-IP telephony, security surveillance, and sensor data analysis. In this paper, we propose a composable stream processing system for cooperative peer-to-peer environments. The system can dynamically select and compose stream processing elements located on different peers into user desired applications. We investigate multiple alternative approaches to composing stream applications: 1) global-state-based centralized versus local-state-based distributed algorithms for initially composing stream applications at setup phase. The centralized algorithm performs periodical global state maintenance while the distributed algorithm performs on-demand state collection. 2) Reactive versus proactive failure recovery schemes for maintaining composed stream applications during runtime. The reactive failure recovery algorithm dynamically recomposes a new stream application upon failures while the proactive approach maintains a number of backup compositions for failure recovery. We conduct both theoretical analysis and experimental evaluations to study the properties of different approaches. Our study illustrates the performance and overhead trade-offs among different design alternatives, which can provide important guidance for selecting proper algorithms to compose stream applications in cooperative peer-to-peer environments.
Xiaohui Gu, Klara Nahrstedt
IEEE Trans. Parallel Distributed Syst.1
2005 Optimal Component Composition for Scalable Stream Processing
abstract
Stream processing has become increasingly important with emergence of stream applications such as audio/video surveillance, stock price tracing, and sensor data analysis. A challenging problem is to provide optimal component composition in a distributed stream processing environment. The goal of optimal component composition is to achieve load balancing subject to multiple function, resource, and quality-of-service (QoS) constraints while composing stream applications. In this paper, we present an adaptive composition probing (ACP) approach to the problem. Different from previous work, ACP provides a new hybrid approach that combines distributed composition probing with coarse-grain global state management. Guided by the coarse-grain global state information, ACP selectively probes a subset of candidate components to discover an approximately optimal component composition. Further, ACP is self-tuning, which can adoptively adjust the number of probes to maintain a specified composition performance target (i.e., composition success rate) in a dynamic stream environment. While the optimal component composition problem is NP-hard, our ACP approach provides an adaptive polynomial approximation solution. We have conducted extensive simulation experiments to show the efficiency, scalability, and adaptability of the ACP approach by comparing with other alternative solutions
Xiaohui Gu, Philip S. Yu, Klara Nahrstedt
ICDCS1
2005 Adaptive Load Diffusion for Stream Joins
Xiaohui Gu, Philip S. Yu
Middleware1
2005 Supporting multi-party voice-over-IP services with peer-to-peer stream processing
abstract
Multi-party voice-over-IP (MVoIP) services provide economical and natural group communication mechanisms for many emerging applications such as on-line gaming, distance collaboration, and tele-immersion. In this paper, we present a novel peer-to-peer (P2P) stream processing system called peerTalk to provide resource-efficient and failure-resilient MVoIP services. Different from previous work, our solution is fully distributed and self-organizing without requiring specialized servers or IP multicast support. Particularly, we decouple the stream processing in MVoIP services into two phases: (1) aggregation phase that mixes audio streams from active speakers into a single stream; and (2) distribution phase that distributes the mixed audio stream to all listeners. The decoupled model allows us to optimize and adapt the P2P stream mixing and distribution processes separately. Specifically, we can adaptively spread stream mixing workload among resource-constrained peer hosts according to current speaking activities. We have implemented a prototype of the peerTalk system and conducted experiments in real-world wide-area networks. The results show that peerTalk can achieve lower resource contention and better service quality than previous common solution.
Xiaohui Gu, Philip S. Yu, Zon-Yin Shae
ACM Multimedia1
2004 SpiderNet: An Integrated Peer-to-Peer Service Composition Framework
Xiaohui Gu, Klara Nahrstedt, Bin Yu 0010
HPDC1
2004 An overlay based QoS-aware voice-over-IP conferencing system
abstract
Ubiquitous IP telephony has become a feasible Internet service, and it is expected to meet the quality standards of traditional telephone services. The work presents a distributed voice-over-IP (VoIP) conferencing system called Venus that is implemented as a composable application-level service overlay network. Compared to the traditional centralized approach, Venus achieves better scalability and resource utilization by efficiently aggregating resources across distributed voice mixers. Moreover, Venus provides multi-constrained quality-of-service (QoS) provisioning by establishing each conferencing session based on multiple QoS constraints (e.g., delay, loss rate) and resource requirements (e.g., bandwidth, audio channels). Venus provides a failure resilient VoIP conferencing service by leveraging the fast failure recovery capability of the application-level service overlay network. Large-scale simulation results illustrate the efficiency of the Venus system.
Xiaohui Gu, Klara Nahrstedt, Rong Chang 0001, Zon-Yin Shae
ICME1
2003 QoS-Assured Service Composition in Managed Service Overlay Networks
abstract
Many value-added and content delivery services are being offered via service level agreements (SLAs). These services can be interconnected to form a service overlay network (SON) over the Internet. Service composition in SON has emerged as a cost-effective approach to quickly creating new services. Previous research has addressed the reliability, adaptability, and compatibility issues for composed services. However little has been done to manage generic quality-of-service (QoS) provisioning for composed services, based on the SLA contracts of individual services. In this paper we present QUEST a QoS assUred composEable Service infrasTructure, to address the problem. QUEST framework provides: (1) initial service composition, which can compose a qualified service path under multiple QoS constraints (e.g., response time, availability). If multiple qualified service paths exist, QUEST chooses the best one according to the load balancing metric; and (2) dynamic service composition, which can dynamically recompose the service path to quickly recover from service outages and QoS violations. Different from the previous work, QUEST can simultaneously achieve QoS assurances and good load balancing in SON.
Xiaohui Gu, Klara Nahrstedt, Rong Chang 0001, Christopher Ward
ICDCS1
2003 Adaptive Offloading Inference for Delivering Applications in Pervasive Computing Environments
abstract
Pervasive computing allows a user to access an application on heterogeneous devices continuously and consistently. However it is challenging to deliver complex applications on resource-constrained mobile devices, such as cellular telephones and PDA. Different approaches, such as application-based or system-based adaptations, have been proposed to address the problem. However existing solutions often require degrading application fidelity. We believe that this problem can be overcome by dynamically partitioning the application and offloading part of the application execution to a powerful nearby surrogate. This will enable pervasive application delivery to be realized without significant fidelity degradation or expensive application rewriting. Because pervasive computing environments are highly dynamic, the runtime offloading system needs to adapt to both application execution patterns and resource fluctuations. Using the fuzzy control model, we have developed an offloading inference engine to adaptively solve two key decision-making problems during runtime offloading: (1) timely triggering of adaptive offloading, and (2) intelligent selection of an application partitioning policy. Extensive trace-driven evaluations show the effectiveness of the offloading inference engine.
Xiaohui Gu, Klara Nahrstedt, Alan Messer, Ira Greenberg 0002, Dejan S. Milojicic
PerCom1
2002 A Scalable QoS-Aware Service Aggregation Model for Peer-to-Peer Computing Grids
abstract
Peer-to-peer (P2P) computing grids consist of peer nodes that communicate directly among themselves through wide-area networks and can act as both clients and servers. These systems have drawn much research attention since they promote Internet-scale resource and service sharing without any administration cost or centralized infrastructure support. However aggregating different application services into a high-performance distributed application delivery in such systems is challenging due to the presence of dynamic performance information, arbitrary peer arrivals/departures, and systems' scalability requirement. In this paper we propose a scalable QoS-aware service aggregation model to address the challenges. The model includes two tiers: (1) on-demand service composition tier which is responsible for choosing and composing different application services into a service path satisfying the user's quality requirements; and (2) dynamic peer selection tier, which decides the specific peers where the chosen services are actually instantiated based on the dynamic, composite and distributed performance information. The model is designed and implemented in a fully distributed and self-organizing fashion. Conducting extensive simulations of a large-scale P2P system (10/sup 4/ peers), we show that our proposed model and algorithms achieve better performance than several common heuristic algorithms.
Xiaohui Gu, Klara Nahrstedt
HPDC1
2002 Dynamic QoS-Aware Multimedia Service Configuration in Ubiquitous Computing Environments
abstract
Ubiquitous computing promotes the proliferation of various stationary, embedded and mobile devices interconnected by heterogeneous networks. It leads to a highly dynamic distributed system with many devices and services coming and going frequently. Many emerging distributed multimedia applications are being deployed in such a computing environment. In order to make the experience for a user truly seamless and to provide soft performance guarantees, we must meet the following challenges: (1) users should be able to perform tasks continuously, despite changes of resources, devices and locations; (2) users should be able to efficiently utilize all accessible resources within runtime environments to receive the best possible Quality-of-Service (QoS). In this paper, we propose an integrated QoS-aware service configuration model to address the above problems. The configuration model includes two tiers: (1) service composition tier, which is responsible for choosing and composing current available service components appropriately and coordinating arbitrary interactions between them to achieve the user's objectives; and (2) service distribution tier which is responsible for dividing an application into several partitions and distributing them to different available devices appropriately. Our initial experimental results based on both prototype and simulations show the soundness of our model and algorithms.
Xiaohui Gu, Klara Nahrstedt
ICDCS1
2002 Towards a Distributed Platform for Resource-Constrained Devices
abstract
Many visions of the future predict a world with pervasive computing, where computing services and resources permeate the environment. In these visions, people will want to execute a service on any available device without worrying about whether the service has been tailored for the device. We believe that it will be difficult to create services that can execute well on the wide variety of devices that are being developed because of problems with diversity and resource constraints. We believe that these problems can be greatly reduced by using an ad-hoc distributed platform to transparently off-load portions of a service from a resource-constrained device to a nearby server. We implemented a preliminary prototype and emulator to study this approach. Our experiments show the beneficial use of nearby resources to relieve both memory and processing constraints, when it is appropriate to do so. We believe that this approach will reduce the burden on developers by masking more device details.
Alan Messer, Ira Greenberg 0002, Philippe Bernadat, Dejan S. Milojicic, DeQing Chen, Thomas J. Giuli, Xiaohui Gu
ICDCS7
2002 A programming framework for quality-aware ubiquitous multimedia applications
abstract
Ubiquitous computing promises a computing environment that seamlessly and pervasively delivers applications to the user, despite changes of resources, devices, and locations. However, few ubiquitous multimedia applications (UMAs) exist up-to-date. One of the main reasons lies in the fact that it is difficult and error-prone to build a UMA which is mobile and deployable in different ubiquitous environments, and still provides acceptable application-specific Quality-of-Service (QoS) guarantees. In this paper, we present the design and implementation of a novel programming framework, called 'QCompiler" to address the challenges. The framework includes (1) a high-level application specification for the application developer to easily write a UMA with specific quality, mobility, and ubiquity supports, (2) a meta-data compilation, which provides automated consistency checks, translations, and substitutions, to relieve the application developer from dealing with complex programming related to quality, mobility, and ubiquity, (3) a binding, which prepares a quality-aware specification to be executable, in a specific deployment environment, and (4)a run-time meta-data execution, utilizing the meta-data compilation's results, to manage and control a quality-aware multimedia application. As a case study, we apply the programming framework to build a mobile Video-on-Demand (VoD) application. The experimental results show tradeoffs between easiness and flexibility to develop and deploy UMA, and overheads during UMA instantiation and adaptation.
Duangdao Wichadakul, Xiaohui Gu, Klara Nahrstedt
ACM Multimedia2
2001 Visual QoS Programming Environment for Ubiquitous Multimedia Services
abstract
The provision of distributed multimedia services is becoming mobile and ubiquitous. Different multimedia services require application-specific Quality of Service (QoS). In this paper, we present QoSTalk, a unified component-based programming environment that allows application developers to specify different application-specific QoS requirements easily. In QoSTalk, we adopt a hierarchical approach to model application configuration graphs for different distributed multimedia services. We design and implement the XML-based Hierarchical QoS Markup Language, called HQML, to describe the hierarchical configuration graph as well as other application-specific QoS requirements and policies. QoSTalk promotes the separation of concerns in developing QoS-aware ubiquitous multimedia applications and thus enables easy programming of QoS-aware applications, running on top of a unified QoS-aware middleware framework. We have prototyped the QoSTalk in Java and CORBA. Our case studies with several multimedia applications show that QoSTalk effectively fills the gap for application developers between the very general facilities provided by the QoS-aware middleware and different kinds of distributed multimedia applications.
Xiaohui Gu, Duangdao Wichadakul, Klara Nahrstedt
ICME1
2001 2K: An Integrated Approach of QoS Compilation and Reconfigurable, Component-Based Run-Time Middleware for the Unified QoS Management Framework
Duangdao Wichadakul, Klara Nahrstedt, Xiaohui Gu, Dongyan Xu
Middleware3