EDBT 2026 Demo / reviewers in the wild / expert
Qingyang Wang 0001
dblp:75/1008-1
· DBLP profile ↗
54ranked-venue papers
8as first author
14since 2021 · last 2025
0000-0002-5729-2898ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-authorSecurity and privacy · 7 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 7 · 2 since 2021Computer networks · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CryptMove: Moving Stealthily through Legitimate and Encrypted Communication ChannelsabstractTo move laterally inside an enterprise environment, Advanced Persistent Threat (APT) attacks have used multiple techniques. Due to the arms race between the attacks and the defenses, such techniques have evolved over time, with the latest one capable of reusing existing network connections for stealthy lateral movement. However, this technique has limited impact because it cannot reuse encrypted connections that are becoming the norm. In this paper, we present CryptMove, a novel technique that can abuse existing and encrypted channels for lateral movement. CryptMove secretly accesses the memory of the target process to duplicate the security context that is used by the target process to perform encryption/decryption; it also secretly duplicates sockets owned by the target process and injects encrypted malicious commands through these sockets into the encrypted communication channels. Since the location of the security context is specific to the target application, CryptMove employs automated analysis of the target application's binary code, in order to learn a path to reach the security context via a sequence of memory accesses. To demonstrate the feasibility of CryptMove, we built PoC attack tools (on both Windows and Linux) that successfully attacked popular applications (e.g., OpenSSH, PuTTY, WinSCP and WinRM) under 63 different cipher-protocol combinations. We also confirmed that the CryptMove PoC is not detectable by several popular Antivirus and Endpoint Detection and Response systems. Md Rabbi Alam, Jinpeng Wei, Qingyang Wang 0001 |
CODASPY | 3 |
| 2025 | PathFence: Reducing Cross-Path Dependencies in MicroservicesabstractMaintaining low and consistent response times is crucial for mission-critical, user-facing applications (e.g., e-commerce sites and social media platforms) that are built on microservices architectures. However, through extensive benchmarking of microservice applications in cloud environments, we find that response time stability in microservices applications is fragile, with delays ranging from milliseconds to seconds, even under moderate CPU utilization level (e.g., 60%). An important cause of this instability is cross-path dependency, where multiple execution paths, triggered by different user requests, share certain common component microservices. As a result, a slowdown in one execution path (e.g., due to a transient bottleneck) can propagate and degrade the performance of other execution paths that share the same microservices. To address this challenge, we propose PathFence, a novel approach that reduces the impact of cross-path dependencies in microservices applications by isolating workloads from different execution paths at shared microservices. By dynamically allocating software resources (e.g., thread and connection pools) with the awareness of execution path and optimizing concurrency level on shared microservices, PathFence significantly improves response time stability. Based on three real-world workload traces and three representative microservices benchmarks, we show that PathFence reduces the 99th percentile response time by up to 80% and decreases the number of dropped requests by over 90%. Xuhang Gu, Qingyang Wang 0001 |
HPDC | 2 |
| 2024 | Sync-Millibottleneck Attack on Microservices Cloud ArchitectureabstractThe modern web services landscape is characterized by numerous fine-grained, loosely coupled microservices with increasingly stringent low-latency requirements. However, this architecture also brings new performance vulnerabilities. In this paper, we introduce a novel low-volume application layer DDoS attack called the Sync-Millibottleneck (SyncM) attack, specifically targeting microservices. The goal of this attack is to cause a long-tail latency problem that violates the service-level agreement (SLA) while evading state-of-the-art DDoS detection/defense mechanisms. The SyncM attack exploits two unique features of microservices architecture: (1) the shared frontend gateway that directs user requests to mid-tier/backend microservices, and (2) the co-existence of multiple logically independent execution paths, each with its own bottleneck resource. By creating synchronized millibottlenecks (i.e., sub-second duration bottlenecks) on multiple independent execution paths, SyncM attack can cause the queuing effect in each execution path to be propagated and superimposed in the shared frontend gateway. As a result, SyncM triggers surprisingly high latency spikes in the system, even when all system resources are far from saturation, making it challenging to trace the cause of performance instability. Xuhang Gu, Qingyang Wang 0001, Qiben Yan 0001, Jianshu Liu, Calton Pu |
AsiaCCS | 2 |
| 2024 | Grunt Attack: Exploiting Execution Dependencies in MicroservicesabstractLoosely-coupled and lightweight microservices running in containers are likely to form complex execution dependencies inside the system. The execution dependency arises when two execution paths partially share component microservices, resulting in potential runtime blocking effects. In this paper, we present Grunt Attack - a novel low-volume DDoS attack that takes advantage of the execution dependencies of microservice applications. Grunt Attack utilizes legitimate HTTP requests to accurately profile the internal pairwise dependencies of all supported execution paths in the target system. By grouping and characterizing all the execution paths based on their pairwise dependencies, the Grunt attacker can target only a few execution paths to launch a low-volume DDoS attack that achieves large performance damage to the entire system. To increase the attack stealthiness, the Grunt attacker avoids creating a persistent bottleneck by alternating the target execution paths within their dependency group. We validate the effectiveness of Grunt attack through experiments of open-source microservices benchmark applications on real clouds (e.g., EC2, Azure) equipped with state-of-the-art IDS/IPS systems and live attack scenarios. Our results show that Grunt attack consumes less than 20% additional CPU resource of the target system while increasing its average response time by over 10x. Xuhang Gu, Qingyang Wang 0001, Jianshu Liu, Jinpeng Wei |
DSN | 2 |
| 2024 | Totoro: A Scalable Federated Learning Engine for the EdgeabstractFederated Learning (FL) is an emerging distributed machine learning (ML) technique that enables in-situ model training and inference on decentralized edge devices. We propose Totoro, a novel scalable FL engine, that enables massive FL applications to run simultaneously on edge networks. The key insight is to explore a distributed hash table (DHT)-based peer-to-peer (P2P) model to re-architect the centralized FL system design into a fully decentralized one. In contrast to previous studies where many FL applications shared one centralized parameter server, Totoro assigns a dedicated parameter server to each individual application. Any edge node can act as any application's coordinator, aggregator, client selector, worker (participant device), or any combination of the above, thereby radically improving scalability and adaptivity. Totoro introduces three innovations to realize its design: a locality-aware P2P multi-ring structure, a publish/subscribe-based forest abstraction, and a bandit-based exploitation-exploration path planning model. Real-world experiments on 500 Amazon EC2 servers show that Totoro scales gracefully with the number of FL applications and N edge nodes, speeds up the total training time by 1.2 × -14.0×, achieves O (logN) hops for model dissemination and gradient aggregation with millions of nodes, and efficiently adapts to the practical edge networks and churns. Cheng-Wei Ching, Xin Chen 0084, Taehwan Kim 0012, Bo Ji 0001, Qingyang Wang 0001, Dilma Da Silva, Liting Hu |
EuroSys | 5 |
| 2023 | μConAdapter: Reinforcement Learning-based Fast Concurrency Adaptation for Microservices in CloudabstractModern web-facing applications such as e-commerce comprise tens or hundreds of distributed and loosely coupled microservices that promise to facilitate high scalability. While hardware resource scaling approaches [28] have been proposed to address response time fluctuations in critical microservices, little attention has been given to the scaling of soft resources (e.g., threads or database connections), which control hardware resource concurrency. This paper demonstrates that optimal soft resource allocation for critical microservices significantly impacts overall system performance, particularly response time. This suggests the need for fast and intelligent runtime reallocation of soft resources as part of microservices scaling management. We introduce μConAdapter, an intelligent and efficient framework for managing concurrency adaptation. It quickly identifies optimal soft resource allocations for critical microservices and adjusts them to mitigate violations of service-level objectives (SLOs). μConAdapter utilizes fine-grained online monitoring metrics from both the system and application levels and a Deep Q-Network (DQN) to quickly and adaptively provide optimal concurrency settings for critical microservices. Using six realistic bursty workload traces and two representative microservices-based benchmarks (SockShop and SocialNetwork), our experimental results show that μConAdapter can effectively mitigate large response time fluctuation and reduce the tail latency at the 99th percentile by 3× on average when compared to the hardware-only scaling strategies like Kubernetes Autoscaling and FIRM [28], and by 1.6× to the state-of-the-art concurrency-aware system scaling strategy like ConScale [21]. Jianshu Liu, Shungeng Zhang, Qingyang Wang 0001 |
SoCC | 3 |
| 2023 | Scalable Federated Learning with System HeterogeneityabstractFederated learning (FL) is a distributed learning framework that inherently provides data privacy and parallel computation capability over a set of participating devices (clients). In real-life applications, these clients can have a great variety in terms of resources (storage, RAM, CPU/GPU speed, network speed, etc.). However, most previous FL studies do not consider this scenario with system heterogeneity and assume that all clients can operate on the same full-size deep neural network (DNN) model. In this work, we demonstrate a scalable FL approach, ScaleFL, which tackles system heterogeneity through hierarchically downscaling the DNN model for clients with limited resources. ScaleFL utilizes early exits to form multi-exit DNN models by injecting early exit networks into the given DNN. During FL, the model is adaptively split along depth (exits) and width (hidden dimensions) based on the resource budget of each participating client. A proof-of-concept demonstration is provided with interactive features, demonstrating the system flow on image classification and NLP benchmark workloads. Fatih Ilhan, Gong Su, Qingyang Wang 0001, Ling Liu 0001 |
ICDCS | 3 |
| 2023 | Sora: A Latency Sensitive Approach for Microservice Soft Resource AdaptationabstractFast response time for modern web services that include numerous distributed and lightweight microservices becomes increasingly important due to its business impact. While hardware-only resource scaling approaches (e.g., FIRM [47] and PARSLO [40]) have been proposed to mitigate response time fluctuations on critical microservices, the re-adaptation of soft resources (e.g., threads or connections) that control the concurrency of hardware resource usage has been largely ignored. This paper shows that the soft resource adaptation of critical microservices has a significant impact on system scalability because either under- or over-allocation of soft resources can lead to inefficient usage of underlying hardware resources. We present Sora, an intelligent, fast soft resource adaptation management framework for quickly identifying and adjusting the optimal concurrency level of critical microservices to mitigate service-level objective (SLO) violations. Sora leverages online fine-grained system metrics and the propagated deadline along the critical path of request execution to quickly and accurately provide optimal concurrency setting for critical microservices. Based on six real-world bursty workload traces and two representative microservices benchmarks (Sock Shop and Social Network), our experimental results show that Sora can effectively mitigate large response time fluctuations and reduce the 99th percentile latency by up to 2.5× compared to the hardware-only scaling strategy FIRM [47] and 1.5× to the state-of-the-art concurrency-aware system scaling strategy ConScale. Jianshu Liu, Qingyang Wang 0001, Shungeng Zhang, Liting Hu, Dilma Da Silva |
Middleware | 2 |
| 2022 | Decentralized Allocation of Geo-distributed Edge Resources using Smart ContractsabstractIn the Internet of Things (loT) era, edge computing is a promising paradigm to improve the quality of service for latency sensitive applications by filling gaps between the loT devices and the cloud infrastructure. Highly geo-distributed edge computing resources that are managed by independent and competing service providers pose new challenges in terms of resource allocation and effective resource sharing to achieve a globally efficient resource allocation. In this paper, we propose a novel blockchain-based model for allocating computing resources in an edge computing platform that allows service providers to establish resource sharing contracts with edge infrastructure providers apriori using smart contracts in Ethereum. The smart contract in the proposed model acts as the auctioneer and replaces the trusted third-party to handle the auction. The blockchain-based auctioning protocol increases the transparency of the auction-based resource allocation for the participating edge service and infrastructure providers. The design of sealed bids and bid revealing methods in the proposed protocol make it possible for the participating bidders to place their bids without revealing their true valuation of the goods. The truthful auction design and the utility-aware bidding strategies incorporated in the proposed model enables the edge service providers and edge infrastructure providers to maximize their utilities. We implement a prototype of the model on a real blockchain test bed and our extensive experiments demonstrate the effectiveness, scalability and performance efficiency of the proposed approach. Jinlai Xu, Balaji Palanisamy, Qingyang Wang 0001, Heiko Ludwig, Sandeep Gopisetty |
CCGRID | 3 |
| 2022 | ShadowSync: latency long tail caused by hidden synchronization in real-time LSM-tree based stream processing systemsabstractMission-critical, real-time, continuous stream processing applications that interact with the real world have stringent latency requirements. For example, e-commerce websites like Amazon improve their marketing strategy by performing real-time advertising based on customers' behavior, and latency long tail can cause significant revenue loss. Recent work [39] showed a positive correlation between latency long tail and variance in the execution time of synchronous invocation chains (critical paths) in microservices benchmarks. This paper shows that asynchronous, very short but intense resource demands (called millibottlenecks) outside of critical paths can also cause significant latency long tail. Shungeng Zhang, Qingyang Wang 0001, Yasuhiko Kanemasa, Julius Michaelis, Jianshu Liu, Calton Pu |
Middleware | 2 |
| 2022 | Amnis: Optimized stream processing for edge computing
Jinlai Xu, Balaji Palanisamy, Qingyang Wang 0001, Heiko Ludwig, Sandeep Gopisetty |
J. Parallel Distributed Comput. | 3 |
| 2022 | Coordinating Fast Concurrency Adapting With Autoscaling for SLO-Oriented Web ApplicationsabstractCloud providers tend to support dynamic computing resources reallocation (e.g., Autoscaling) to handle the bursty workload for web applications (e.g., e-commerce) in the cloud environment. Nevertheless, we demonstrate that directly scaling a bottleneck server without quickly adjusting its soft resources (e.g., server threads and database connections) can cause significant response time fluctuations of the target web application. Since soft resources determine the request processing concurrency of each server in the system, simply scaling out/in the bottleneck service can unintentionally change the concurrency level of related services, inducing either under- or over-utilization of the critical hardware resource. In this paper, we propose the Scatter-Concurrency-Throughput (SCT) model, which can rapidly identify the near-optimal soft resource allocation of each server in the system using the measurement of each server’s real-time throughput and concurrency. Furthermore, we implement a Concurrency-aware autoScaling (ConScale) framework that integrates the SCT model to quickly reallocate the soft resources of the key servers in the system to best utilize the new hardware resource capacity after the system scaling. Based on extensive experimental comparisons with two widely used hardware-only scaling mechanisms for web applications: EC2-AutoScaling (VM-based autoscaler) and Kubernetes HPA (container-based autoscaler), we show that ConScale can successfully mitigate the response time fluctuations over the system scaling phase in both VM-based and container-based environments. Jianshu Liu, Shungeng Zhang, Qingyang Wang 0001, Jinpeng Wei |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Resilient Stream Processing in Edge ComputingabstractThe proliferation of Internet-of-Things (IoT) devices is rapidly increasing the demands for efficient processing of low latency stream data generated close to the edge of the network. A large number of IoT applications require continuous processing of data streams in real-time. Examples include virtual reality applications, connected autonomous vehicles and smart city applications. Although current distributed stream processing systems offer various forms of fault tolerance, existing schemes do not understand the dynamic characteristics of edge computing infrastructures and the unique requirements of edge computing applications. Optimizing fault tolerance techniques to meet latency requirements while minimizing resource usage becomes a critical dimension of resource allocation and scheduling when dealing with latency-sensitive IoT applications in edge computing. In this paper, we present a novel resilient stream processing framework that achieves system-wide fault tolerance while meeting the latency requirements for edge-based applications. The proposed approach employs a novel resilient physical plan generation for stream queries and optimizes the placement of operators to minimize the processing latency during recovery and reduces the overhead of checkpointing. We implement a prototype of the proposed techniques in Apache Storm and evaluate it in a real testbed. Our results demonstrate that the proposed approach is highly effective and scalable while ensuring low latency and low-cost recovery for edge-based stream processing applications. Jinlai Xu, Balaji Palanisamy, Qingyang Wang 0001 |
CCGRID | 3 |
| 2021 | Blockumulus: A Scalable Framework for Smart Contracts on the CloudabstractPublic blockchains have spurred the growing popularity of decentralized transactions and smart contracts, especially on the financial market. However, public blockchains exhibit their limitations on the transaction throughput, storage availability, and compute capacity. To avoid transaction gridlock, public blockchains impose large fees and per-block resource limits, making it difficult to accommodate the ever-growing high transaction demand. Previous research endeavors to improve the scalability and performance of blockchain through various technologies, such as side-chaining, sharding, secured off-chain computation, communication network optimizations, and efficient consensus protocols. However, these approaches have not attained a widespread adoption due to their inability in delivering a cloud-like performance, in terms of the scalability in transaction throughput, storage, and compute capacity. In this work, we determine that the major obstacle to public blockchain scalability is their underlying unstructured P2P networks. We further show that a centralized network can support the deployment of decentralized smart contracts. We propose a novel approach for achieving scalable decentralization: instead of trying to make blockchain scalable, we deliver decentralization to already scalable cloud by using an Ethereum smart contract. We introduce Blockumulus, a framework that can deploy decentralized cloud smart contract environments using a novel technique called overlay consensus. Through experiments, we demonstrate that Blockumulus is scalable in all three dimensions: computation, data storage, and transaction throughput. Besides eliminating the current code execution and storage restrictions, Blockumulus delivers a transaction latency between 2 and 5 seconds under normal load. Moreover, the stress test of our prototype reveals the ability to execute 20,000 simultaneous transactions under 26 seconds, which is on par with the average throughput of worldwide credit card transactions. Qiben Yan 0001, Qingyang Wang 0001 |
ICDCS | 3 |
| 2020 | FP4S: Fragment-based Parallel State Recovery for Stateful Stream ApplicationsabstractStreaming computations are by nature long-running. They run in highly dynamic distributed environments where many stream operators may leave or fail at the same time. Most of them are stateful, in which stream operators need to store and maintain large-sized state in memory, resulting in expensive time and space costs to recover them. The state-of-the-art stream processing systems offer failure recovery mainly through three approaches: replication recovery, checkpointing recovery, and DStream-based lineage recovery, which are either slow, resource-expensive or fail to handle many simultaneous failures.We present FP4S, a novel fragment-based parallel state recovery mechanism that can handle many simultaneous failures for a large number of concurrently running stream applications. The novelty of FP4S is that we organize all the application's operators into a distributed hash table (DHT) based consistent ring to associate each operator with a unique set of neighbors. Then we divide each operator's in-memory state into many fragments and periodically save them in each node's neighbors, ensuring that different sets of available fragments can reconstruct lost state in parallel. This approach makes this failure recovery mechanism extremely scalable, and allows it to tolerate many simultaneous operator failures. We apply FP4S on Apache Storm and evaluate it using large-scale real-world experiments, which demonstrate its scalability, efficiency, and fast failure recovery features. When compared to the state-of-the-art solutions (Apache Storm), FP4S reduces 37.8% latency of state recovery and saves more than half of the hardware costs. It can scale to many simultaneous failures and successfully recover the states when up to 66.6% of states fail or get lost. Pinchao Liu, Hailu Xu, Dilma Da Silva, Qingyang Wang 0001, Sarker Tanzir Ahmed, Liting Hu |
IPDPS | 4 |
| 2020 | Mitigating Large Response Time Fluctuations through Fast Concurrency Adapting in CloudsabstractDynamically reallocating computing resources to handle bursty workloads is a common practice for web applications (e.g., e-commerce) in clouds. However, our empirical analysis on a standard n-tier benchmark application (RUBBoS) shows that simply scaling an n-tier application by reallocating hardware resources without fast adapting soft resources (e.g., server threads, connections) may lead to large response time fluctuations. This is because soft resources control the workload concurrency of component servers in the system: adding or removing hardware resources such as Virtual Machines (VMs) can implicitly change the workload concurrency of dependent servers, causing either under- or over-utilization of the critical hardware resource in the system. To quickly identify the optimal soft resource allocation of each server in the system and stabilize response time fluctuation, we propose a novel Scatter-Concurrency-Throughput (SCT) model based on the monitoring of each server's real-time concurrency and throughput. We then implement a Concurrency-aware system Scaling (ConScale) framework which integrates the SCT model to fast adapt the soft resource allocations of key servers during the system scaling process. Our experiments using six realistic bursty workload traces show that ConScale can effectively mitigate the response time fluctuations of the target web application compared to the state-of-the-art cloud scaling strategies such as EC2-AutoScaling. Jianshu Liu, Shungeng Zhang, Qingyang Wang 0001, Jinpeng Wei |
IPDPS | 3 |
| 2020 | DoubleFaceAD: A New Datastore Driver Architecture to Optimize Fanout Query PerformanceabstractThe broad adoption of fanout queries on distributed datastores has made asynchronous event-driven datastore drivers a natural choice due to reduced multithreading overhead. However, through extensive experiments using the latest datastore drivers (e.g., MongoDB, HBase, DynamoDB) and YCSB benchmark, we show that an asynchronous datastore driver can cause unexpected performance degradation especially in fanout-query scenarios. For example, the default MongoDB asynchronous driver adopts the latest Java asynchronous I/O library, which uses a hidden on-demand JVM level thread pool to process fanout query responses, causing a surprising multithreading overhead when the query response size is large. A second instance is the traditional wisdom of modular design of an application server and the embedded asynchronous datastore driver can cause an im-balanced workload between the two components due to lack of coordination, incurring frequent unnecessary system calls. To address the revealed problems, we introduce DoubleFaceAD--a new asynchronous datastore driver architecture that integrates the management of both upstream and downstream workload traffic through a few shared reactor threads, with fanout-query-aware priority-based scheduling to reduce the overall query waiting time. Our experimental results on two representative application scenarios (YCSB and DBLP) show DoubleFaceAD outperforms all other types of datastore drivers up to 34% on throughput and 1.9× faster on 99th percentile response time. Shungeng Zhang, Qingyang Wang 0001, Yasuhiko Kanemasa, Jianshu Liu, Calton Pu |
Middleware | 2 |
| 2020 | ShadowMove: A Stealthy Lateral Movement Strategy
Amirreza Niakanlahiji, Jinpeng Wei, Md Rabbi Alam, Qingyang Wang 0001, Bei-tseng Chu |
USENIX Security Symposium | 4 |
| 2020 | The Impact of Event Processing Flow on Asynchronous Server EfficiencyabstractAsynchronous event-driven server architecture has been considered as a superior alternative to the thread-based counterpart due to reduced multithreading overhead. In this paper, we conduct empirical research on the efficiency of asynchronous Internet servers, showing that an asynchronous server may perform significantly worse than a thread-based one due to two design deficiencies. The first one is the widely adopted one-event-one-handler event processing model in current asynchronous Internet servers, which could generate frequent unnecessary context switches between event handlers, leading to significant CPU overhead of the server. The second one is a write-spin problem (i.e., repeatedly making unnecessary I/O system calls) in asynchronous servers due to some specific runtime workload and network conditions (e.g., large response size and non-trivial network latency). To address these two design deficiencies, we present a hybrid solution by exploiting the merits of different asynchronous architectures so that the server is able to adapt to dynamic runtime workload and network conditions in the cloud. Concretely, our hybrid solution applies a lightweight runtime request checking and seeks for the most efficient path to process each request from clients. Our results show that the hybrid solution can achieve from 10 to 90 percent higher throughput than all the other types of servers under the various realistic workload and network conditions in the cloud. Shungeng Zhang, Qingyang Wang 0001, Yasuhiko Kanemasa, Huasong Shan, Liting Hu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Tail Amplification in n-Tier Systems: A Study of Transient Cross-Resource Contention AttacksabstractFast response time becomes increasingly important for modern web applications (e.g., e-commerce) due to intense competitive pressure. In this paper, we present a new type of Denial of Service (DoS) Attacks in the cloud, MemCA, with the goal of causing performance uncertainty (the long-tail response time problem) of the target n-tier web application while keeping stealthy. MemCA exploits the sharing nature of public cloud computing platforms by co-locating the adversary VMs with the target VMs that host the target web application, and causing intermittent and short-lived cross-resource contentions on the target VMs. We show that these short-lived cross-resource contentions can cause transient performance interferences that lead to large response time fluctuations of the target web application, due to complex resource dependencies in the system. We further model the attack scenario in n-tier systems based on queuing network theory, and analyze cross-tier queue overflow and tail response time amplification under our attacks. Through extensive benchmark experiments in both private and public clouds (e.g., Amazon EC2), we confirm that MemCA can cause significant performance uncertainty of the target n-tier system while keeping stealthy. Specifically, we show that MemCA not only bypasses the cloud elastic scaling mechanisms, but also the state-of-the-art cloud performance interference detection mechanisms. Shungeng Zhang, Huasong Shan, Qingyang Wang 0001, Jianshu Liu, Qiben Yan 0001, Jinpeng Wei |
ICDCS | 3 |
| 2019 | Mitigating Tail Response Time of n-Tier Applications: The Impact of Asynchronous InvocationsabstractConsistent low response time is essential for e-commerce due to intense competitive pressure. However, practitioners of web applications have often encountered the long-tail response time problem in cloud data centers as the system utilization reaches moderate levels (e.g., 50%). Our fine-grained measurements of an open source n-tier benchmark application (RUBBoS) show such long response times are often caused by Cross-tier Queue Overflow (CTQO). Our experiments reveal the CTQO is primarily created by the synchronous nature of RPC-style call/response inter-tier communications, which create strong inter-tier dependencies due to the request processing chain of classic n-tier applications composed of synchronous RPC/thread-based servers. We remove gradually the dependencies in n-tier applications by replacing the classic synchronous servers (e.g., Apache, Tomcat, and MySQL) with their corresponding event-driven asynchronous version (e.g., Nginx, XTomcat, and XMySQL) one-by-one. Our measurements with two application scenarios (virtual machine co-location and background monitoring interference) show that replacing a subset of asynchronous servers will shift the CTQO, without significant improvements in long-tail response time. Only when all the servers become asynchronous the CTQO is resolved. In synchronous n-tier applications, long-tail response times resulting from CTQO arise at utilization as low as 43%. On the other hand, the completely asynchronous n-tier system can disrupt CTQO and remove the long tail latency at utilization as high as 83%. Qingyang Wang 0001, Shungeng Zhang, Yasuhiko Kanemasa, Calton Pu |
ACM Trans. Internet Techn. | 1 |
| 2019 | Integrating Concurrency Control in n-Tier Application Scaling Management in the CloudabstractScaling complex distributed systems such as e-commerce is an importance practice to simultaneously achieve high performance and high resource efficiency in the cloud. Most previous research focuses on hardware resource scaling to handle runtime workload variation. Through extensive experiments using a representative n-tier web application benchmark (RUBBoS), we demonstrate that scaling an n-tier system by adding or removing VMs without appropriately re-allocating soft resources (e.g., server threads and connections) may lead to significant performance degradation resulting from implicit change of request processing concurrency in the system, causing either over- or under-utilization of the critical hardware resource in the system. We build a concurrency-aware model that determines a near optimal soft resource allocation of each tier by combining some operational queuing laws and the fine-grained online measurement data of the system. We then develop a dynamic concurrency management (DCM) framework that integrates the concurrency-aware model to intelligently reallocate soft resources in the system during the system scaling process. We compare DCM with Amazon EC2-AutoScale, the state-of-the-art hardware only scaling management solution using six real-world bursty workload traces. The experimental results show that DCM achieves significantly shorter tail latency and higher throughput compared to Amazon EC2-AutoScale under all the workload traces. Qingyang Wang 0001, Shungeng Zhang, Liting Hu, Balaji Palanisamy |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | Lightweight Indexing and Querying Services for Big Spatial DataabstractWith the widespread use of GPS-equipped smartphones and Internet of Things devices, a huge amount of data with location information is being generated at an unprecedented rate. To gain a deeper insight into such a plethora of spatial data, scientists and engineers are widely using spatial queries for their big data applications. However, because of not only the massive spatial data size but also the complexity of spatial query processing, they are struggling to efficiently process the spatial queries. In this paper, we propose lightweight and scalable indexing and querying services for big spatial data stored in distributed storage systems or graph-based systems. Our spatial services have several advantages over existing approaches. First, our services can be easily applied to existing storage systems or graph-based models without modifying the internal implementation of existing systems/models. Second, our services achieve high pruning power by efficiently selecting only relevant spatial objects based on a simple yet effective filter. Third, our services support a customizable and easy-to-use control of index data size by adjusting the precision of indexed geometries. Lastly, our services support efficient updates of spatial data. Our experimental results using real-world datasets validate the effectiveness and efficiency of our spatial services. Kisung Lee, Ling Liu 0001, Raghu K. Ganti, Mudhakar Srivatsa, Qi Zhang 0009, Yang Zhou 0001, Qingyang Wang 0001 |
IEEE Trans. Serv. Comput. | 7 |
| 2018 | A Toolset for Detecting Containerized Application's Dependencies in CaaS CloudsabstractThere has been a dramatic increase in the popularity of Container as a Service (CaaS) clouds. The CaaS multi-tier applications could be optimized by using network topology, link or server load knowledge to choose the best endpoints to run in CaaS cloud. However, it is difficult to apply those optimizations to the public datacenter shared by multi-tenants. This is because of the opacity between the tenants and the datacenter providers: Providers have no insight into tenant's container workloads and dependencies, while tenants have no clue about the underlying network topology, link, and load. As a result, containers might be booted at wrong physical nodes that lead to performance degradation due to bi-section bandwidth bottleneck or co-located container interference. We propose 'DocMan', a toolset that adopts a black-box approach to discover container ensembles and collect information about intra-ensemble container interactions. It uses a combination of techniques such as distance identification and hierarchical clustering. The experimental results demonstrate that DocMan enables optimized containers placement to reduce the stress on bi-section bandwidth of the datacenter's network. The method can detect container ensembles at low cost and with 92% accuracy and significantly improve performance for multi-tier applications under the best of circumstances. Pinchao Liu, Liting Hu, Hailu Xu, Jason Liu 0001, Qingyang Wang 0001, Jai Dayal, Yuzhe Tang |
IEEE CLOUD | 6 |
| 2018 | Oases: An Online Scalable Spam Detection System for Social NetworksabstractWeb-based social networks enable new community-based opportunities for participants to engage, share their thoughts, and interact with each other. Theses related activities such as searching and advertising are threatened by spammers, content polluters, and malware disseminators. We propose a scalable spam detection system, termed Oases, for uncovering social spam in social networks using an online and scalable approach. The novelty of our design lies in two key components: (1) a decentralized DHT-based tree overlay deployment for harvesting and uncovering deceptive spam from social communities; and (2) a progressive aggregation tree for aggregating the properties of these spam posts for creating new spam classifiers to actively filter out new spam. We design and implement the prototype of Oases and discuss the design considerations of the proposed approach. Our large-scale experiments using real-world Twitter data demonstrate scalability, attractive load-balancing, and graceful efficiency in online spam detection for social networks. Hailu Xu, Liting Hu, Pinchao Liu, Jai Dayal, Qingyang Wang 0001, Yuzhe Tang |
IEEE CLOUD | 7 |
| 2018 | To Sell or Not To Sell: Trading Your Reserved Instances in Amazon EC2 MarketplaceabstractRecently, Amazon EC2 offers a reserved instance marketplace, where cloud users can sell their idle reserved instances varying in contract lengths and pricing options for avoiding the waste of their unused reservations. However, without knowing the future demands, it is hard for users to determine how to sell instances optimally, for it would incur more cost if new demands arrive after selling their reserved instances. For dealing with this problem, in this paper we first propose three online selling algorithms to guide cloud users in making decisions whether or not to sell their reservations in Amazon EC2 marketplace while guaranteeing competitive ratios. We prove theoretically that the three proposed online algorithms can guarantee bounded competitive ratios, whose values are specific to the type of reserved instances under consideration. Specifically, for all standard instances (Linux, US East) for 1-year terms in Amazon EC2, compared with a benchmark optimal offline algorithm, our algorithm A3T/4 can achieve a ratio of 2-α-a/4 in managing instance purchasing cost, where α is the entitled discount due to reservation and a is the selling discount specified by the user who sells its reservations. Finally, through extensive experiments based on workload data collected from real-world applications, we validate the effectiveness of our online instance selling algorithms by showing that it can bring significant cost savings to cloud users compared with always keeping their reservations in Amazon EC2 reserved instance marketplace. Shengsong Yang, Li Pan 0001, Qingyang Wang 0001, Shijun Liu |
ICDCS | 3 |
| 2018 | Improving Asynchronous Invocation Performance in Client-Server SystemsabstractIn this paper, we conduct an experimental study of asynchronous invocation on the performance of client-server systems. Through extensive measurements of both realistic macro-and micro-benchmarks, we show that servers with the asynchronous event-driven architecture may perform significantly worse than the thread-based version resulting from two nontrivial reasons. First, the traditional wisdom of one-event-one-handler event processing flow can create large amounts of intermediate context switches that significantly degrade the performance of an asynchronous server. Second, some runtime workload (e.g., response size) and network conditions (e.g., network latency) may cause significant negative performance impact on the asynchronous event-driven servers, but not on threadbased ones. We provide a hybrid solution by taking advantage of different asynchronous architectures to adapt to varying workload and network conditions. Our hybrid solution searches for the most efficient execution path for each client request based on the runtime request profiling and type checking. Our experimental results show that the hybrid solution outperforms all the other types of servers up to 19%~90% on throughput, depending on specific workload and network conditions. Shungeng Zhang, Qingyang Wang 0001, Yasuhiko Kanemasa |
ICDCS | 2 |
| 2018 | Subscription or Pay-as-You-Go: Optimally Purchasing IaaS Instances in Public CloudsabstractIn public clouds such as Amazon EC2, there are two main pricing models in purchasing Infrastructure-as-a-Service (IaaS) instances: the pay-as-you-go model and the subscription model. For these two options, users can dynamically combine them to provide services for demands to save their instance acquisition costs. Making optimal decisions toward the purchase of IaaS instances generally requires prior knowledge of future demands; however, it is difficult for users to predict all future workloads accurately. To deal with this problem, online reservation algorithms have been proposed to guide users in reserving instances. However, existing online algorithms do not conform to the pricing rules currently used in public cloud platforms. Therefore, we put forward a new online reserving algorithm for instance in accordance with the pricing policies used in most public IaaS offerings. Specifically, in this study, we use Amazon EC2 as an example to illustrate our algorithm. Through theoretical analysis, we prove that the cost of the proposed algorithm Aβin this paper is not greater than 2-1/β times of the optimal offline algorithm, where β>1 is a critical point in the online reservation algorithm proposed in this paper. Via extensive experimental simulations using both synthetic and actual workload datasets, we demonstrated that the online algorithm Aβis much more cost effective for cloud users than always paying-as-you-go in public IaaS markets. Shengsong Yang, Li Pan 0001, Qingyang Wang 0001, Shijun Liu |
ICWS | 3 |
| 2018 | Privacy-Preserving Publishing of Multilevel Utility-Controlled Graph DatasetsabstractConventional private data publication schemes are targeted at publication of sensitive datasets either after the k -anonymization process or through differential privacy constraints. Typically these schemes are designed with the objective of retaining as much utility as possible for the aggregate queries while ensuring the privacy of the individual records. Such an approach, though suitable for publishing aggregate information as public datasets, is inapplicable when users have different levels of access to the same data. We argue that existing schemes either result in increased disclosure of private information or lead to reduced utility when some users have more access privileges than the others. In this article, we present an anonymization framework for publishing large datasets with the goals of providing different levels of utility to the users based on their access privilege levels. We design and implement our proposed multilevel utility-controlled anonymization schemes in the context of large association graphs considering three levels of user utility, namely, (1) users having access to only the graph structure, (2) users having access to the graph structure and aggregate query results, and (3) users having access to the graph structure, aggregate query results, and individual associations. Our experiments on real large association graphs show that the proposed techniques are effective and scalable and yield the required level of privacy and utility for each user privacy and access privilege level. Balaji Palanisamy, Ling Liu 0001, Yang Zhou 0001, Qingyang Wang 0001 |
ACM Trans. Internet Techn. | 4 |
| 2017 | Minimal Coflow Routing and Scheduling in OpenFlow-Based Cloud Storage Area NetworksabstractResearches affirm that coflow scheduling/routing substantially shortens the average application inner communication time in data center networks(DCNs). The commonly desirable critical features of existing coflow scheduling/routing framework includes (1) coflow scheduling, (2) coflow routing, and (3) per-flow rate-limiting. However, to provide the 3 features, existing frameworks require customized computing frameworks, customized operating systems, or specific external commercial monitoring frameworks on software-defined networking(SDN) switches. These requirements defer or even prohibit the deployment of coflow scheduling/routing in production DCNs. In this paper, we design a coflow scheduling and routing framework, MinCOF which has minimal requirements on hosts and switches for cloud storage area networks(SANs) based on OpenFlow SDN. MinCOF accommodates all critical features of coflow scheduling/routing from previous works. The deployability in production environment is especially taken into consideration. The OpenFlow architecture is capable of processing the traffic load in a cloud SAN. Not necessary requirements for hosts from existing frameworks are migrated to the mature commodity OpenFlow 1.3 Switch and our coflow scheduler. Transfer applications on hosts only need slight enhancements on their existing connection establishment and progress reporting functions. Evaluations reveal that MinCOF decreases the average coflow completion time (CCT) by 12.94% compared to the latest OpenFlow-based coflow scheduling and routing framework. Chui-Hui Chiu, Dipak Kumar Singh, Qingyang Wang 0001, Kisung Lee, Seung-Jong Park |
CLOUD | 3 |
| 2017 | Coflourish: An SDN-Assisted Coflow Scheduling Framework for CloudsabstractExisting coflow scheduling frameworks effectively shorten communication time and completion time of cluster applications. However, existing frameworks only consider available bandwidth on hosts and overlook congestion in the network when making scheduling decisions. Through extensive simulations using the realistic workload probability distribution from Facebook, we observe the performance degradation of the state-of-the-art coflow scheduling framework, Varys, in the cloud environment on a shared data center network (DCN) because of the lack of network congestion information. We propose Coflourish, the first coflow scheduling framework that exploits the congestion feedback assistances from the software-defined-networking(SDN)-enabled switches in the networks for available bandwidth estimation. Our simulation results demonstrate that Coflourish outperforms Varys by up to 75.5% in terms of average coflow completion time under various workload conditions. The proposed work also reveals the potentials of integration with traffic engineering mechanisms in lower levels for further performance optimization. Chui-Hui Chiu, Dipak Kumar Singh, Qingyang Wang 0001, Seung-Jong Park |
CLOUD | 3 |
| 2017 | An Experimental Study of the Impact of vCPU Provisioning on the Performance of a 2-Tier Application Running in CloudabstractLeveraging Virtual Machine (VM) technologies to host multiple Web applications on the same physical machine can improve the resource utilization and thus save a cloud provider's provisioning cost. By allocating and scheduling virtual CPU (vCPU) resources for running VMs, a hosted Web application may achieve varying performances. Thus, when facing an end user with a specific SLA (Service Level Agreement) requirement, a cloud provider needs to decide how many vCPUs to provision for the target SaaS application to meet the user's end-to-end performance requirement while saving cost. However, it is a non-trivial task to economically determine an optimal resource configuration to meet an end user's SLA requirement. Accurate performance analytic models based on traditional modeling techniques such as queuing systems are difficult to construct for web applications. In this paper, we describe our experience in studying the impact of vCPU provisioning on the performance of 2-tier Web applications, through benchmarking a 2-tier web application in the context of provisioning vCPUs to SaaS applications in a cloud environment. From a cloud provider's perspective, we focus on a generic approach for SaaS benchmarking, which can help to study the impact of vCPU allocations on a Web application's performance and make the optimal vCPU allocation decisions to meet end users' QoS requirements while saving provisioning costs. Besides, based on our benchmark experimental results, we also propose an adaptive controller with a vCPU allocation optimization algorithm which can automatically adjust the vCPU allocations to meet the end users' workload requirements. Li Pan 0001, Qingyang Wang 0001, Shijun Liu, Dahui Chen |
CLOUD | 2 |
| 2017 | Tail Attacks on Web ApplicationsabstractAs the extension of Distributed Denial-of-Service (DDoS) attacks to application layer in recent years, researchers pay much interest in these new variants due to a low-volume and intermittent pattern with a higher level of stealthiness, invaliding the state-of-the-art DDoS detection/defense mechanisms. We describe a new type of low-volume application layer DDoS attack--Tail Attacks on Web Applications. Such attack exploits a newly identified system vulnerability of n-tier web applications (millibottlenecks with sub-second duration and resource contention with strong dependencies among distributed nodes) with the goal of causing the long-tail latency problem of the target web application (e.g., 95th percentile response time > 1 second) and damaging the long-term business of the service provider, while all the system resources are far from saturation, making it difficult to trace the cause of performance degradation. Huasong Shan, Qingyang Wang 0001, Calton Pu |
CCS | 2 |
| 2017 | DCM: Dynamic Concurrency Management for Scaling n-Tier Applications in CloudabstractScaling web applications such as e-commerce in cloud by adding or removing servers in the system is an important practice to handle workload variations, with the goal of achieving both high quality of service (QoS) and high resource efficiency. Through extensive scaling experiments of an n-tier application benchmark (RUBBoS), we have observed that scaling only hardware resources without appropriate adaptation of soft resource allocations (e.g., thread or connection pool size) of each server would cause significant performance degradation of the overall system by either under- or over-utilizing the bottleneck resource in the system. We develop a dynamic concurrency management (DCM) framework which integrates soft resource allocations into the system scaling management. DCM introduces a model which determines a near-optimal concurrency setting to each tier of the system based on a combination of operational queuing laws and online analysis of fine-grained measurement data. We implement DCM as a two-level actuator which scales both hardware and soft resources in an n-tier system on the fly without interrupting the runtime system performance. Our experimental results demonstrate that DCM can achieve significantly more stable performance and higher resource efficiency compared to the state-of-the-art hardware-only scaling solutions (e.g., Amazon EC2-AutoScale) under realistic bursty workload traces. Qingyang Wang 0001, Balaji Palanisamy, PengCheng Xiong |
ICDCS | 2 |
| 2017 | milliScope: A Fine-Grained Monitoring Framework for Performance Debugging of n-Tier Web ServicesabstractModern distributed systems are often considered to be black boxes that greatly limit the potential to understand behaviors at the level of detail necessary to diagnose some of the most important types of performance problems. Recently researchers have found abnormal response time delays, one to two orders of magnitude longer than the average response time, that exist in short periods and cause economic loss for service providers. These very short bottlenecks are hard to detect due to their short life spans and their variety of possible reasons. In this paper, we propose milliScope (mScope), the first millisecond-granularity software-based resource and event monitoring for distributed systems that achieves both performance, low overhead at high frequency, and high accuracy matched with other firmware monitoring tool. More specifically, milliScope is a fine-grained monitoring framework to collaborate multiple mScopeMonitors for event and resource monitoring to reconstruct the flow of each client request and profile execution performance in a distributed system. We utilize the resource mScopeMonitors for system resource monitoring, and we develop our own event mScopeMonitors to identify the execution boundary in a lightweight, precise and systematic methodology. The semantic and syntactic of these monitoring logs with arbitrary formats are enriched by our multistage data transformation tool, mScopeDataTransformer, which unifies the diverse monitoring logs into a dynamic data warehouse, mScopeDB, for advanced analysis. We conduct several illustrative scenarios in which milliScope successfully diagnoses the response time anomalies caused by very short bottlenecks using a representative web application benchmark (RUBBoS). Chien-An Lai, Josh Kimball, Qingyang Wang 0001, Calton Pu |
ICDCS | 4 |
| 2017 | The Millibottleneck Theory of Performance Bugs, and Its Experimental VerificationabstractThe performance of n-tier web-facing applications often suffer from response time long-tail problem. With relatively low resource utilization (less than 50%) and the majority of requests returning within a few milliseconds, a non-negligible num-ber of normally short requests may take seconds to return. We propose the millibottleneck theory of performance bugs (that lead to long-tail problems). Several case studies have confirmed the millibottlenecks (that last a few tens to hundreds of milliseconds) as causal agents of long requests. A concrete example (garbage collection) illustrates the experimental verification of millibottlenecks. An open source fine-grain monitoring toolkit is being devel-oped to facilitate the experimental research on millibottlenecks. Calton Pu, Josh Kimball, Chien-An Lai, Jack Li 0001, Junhee Park, Qingyang Wang 0001, Deepal Jayasinghe, PengCheng Xiong, Simon Malkowski, Qinyi Wu, Gueyoung Jung, Younggyun Koh, Galen S. Swint |
ICDCS | 7 |
| 2017 | A Study of Long-Tail Latency in n-Tier Systems: RPC vs. Asynchronous InvocationsabstractLong-tail latency of web-facing applications continues to be a serious problem. Most of the previously published research addresses two classes of long latency problems: uneven workloads such as web search, and resource saturation in single nodes. We describe an experimental study of a third class of long tail latency problems that are specific to distributed systems: Cross-Tier Queue Overflow (CTQO) due to a combination of millibottlenecks (with sub-second duration) and tightly-coupled servers in n-tier systems (e.g., Apache, Tomcat, and MySQL) using RPC-style request-response communications. Our experiments show that the appearance of millibottlenecks (e.g., created by short workload bursts) in one server often causes another server (which has no saturated resources) in the synchronous invocation chain to fill up its queues (CTQO) and drop packets, creating very long response time queries. CTQO can be reduced or avoided by replacing the server dropping packets with an asynchronous server. In synchronous n-tier system experiments, long tail latency due to CTQO can be reproduced consistently at utilization as low as 43%. In contrast, when all n-tier servers are replaced by asynchronous versions, CTQO and consequent dropped packets remain absent at utilization levels as high as 83%, despite the same millibottlenecks. Qingyang Wang 0001, Chien-An Lai, Yasuhiko Kanemasa, Shungeng Zhang, Calton Pu |
ICDCS | 1 |
| 2017 | Limitations of Load Balancing Mechanisms for N-Tier Systems in the Presence of MillibottlenecksabstractThe scalability of n-tier systems relies on effective load balancing to distribute load among the servers of the same tier. We found that load balancing mechanisms (and some policies) in servers used in typical n-tier systems (e.g., Apache and Tomcat) have issues of instability when very long response time (VLRT) requests appear due to millibottlenecks, very short bottlenecks that last only tens to hundreds of milliseconds. Experiments with standard n-tier benchmarks show that during millibottlenecks, some load balancing policy/mechanism combinations make the mistake of sending new requests to the node(s) suffering from millibottlenecks, instead of the idle nodes as load balancers are supposed to do. Several of these mistakes are due to the implicit assumptions made by load balancing policies and mechanisms on the stability of system state. Our study shows that appropriate remedies at policy and mechanism levels can avoid these mistakes during millibottlenecks and remove the VLRT requests, thus improving the average response time by a factor of 12. Jack Li 0001, Josh Kimball, Junhee Park, Chien-An Lai, Calton Pu, Qingyang Wang 0001 |
ICDCS | 7 |
| 2017 | An Experimental Study of a Biosequence Big Data Analysis ServiceabstractWith the development of next-generation sequencing (NGS), DNA/RNA sequencing has become cheaper and more efficient. Today, a whole human genome can be sequenced under $1,000, providing opportunities for large-scale bioinformatic analysis on big datasets. However, most of existing bioinformatic analysis tools are programmed for single server based computing platform and not suitable to process such big datasets. As Hadoop MapReduce and Spark are gaining popularity as cluster computing based big data processing platform, more and more bioinformatic applications start to explore cluster computing platform for large scale data analysis. In this paper we present an in-depth experimental study on deploying Spark clusters for high performance bioinformatic short sequence reconstruction. Our experimental results enable us to answer a number of challenging and yet most frequently asked questions regarding efficient management of bioinformatic data analysis services on Spark systems. Example questions include how to best split big dataset into multiple partitions, and how to distribute data partitions and bioinformatic analysis tasks on a Spark cluster for carrying out a high performance distributed analysis job? What types of memory models are effective for bioinformatic data analysis services on a Spark cluster? Why do different bioinformatic data analysis operations exhibit different throughput performance on the same Spark cluster? We conjecture that this experimental study not only demonstrates the feasibility of high performance bioinformatic data analysis on Spark platform, but also will help bioinformatic application developers to make more informed decisions on both design and configuration of Spark Cluster, managing and tuning parameters of Spark runtime system for enhancing the performance of large scale big data analytics. Wei Zhou 0011, Ling Liu 0001, Calton Pu, Qingyang Wang 0001, Wenkun Xiang, Shaowen Yao 0001 |
ICWS | 5 |
| 2017 | Very Short Intermittent DDoS Attacks in an Unsaturated System
Huasong Shan, Qingyang Wang 0001, Qiben Yan 0001 |
SecureComm | 2 |
| 2016 | Performance Interference of Memory Thrashing in Virtualized Cloud Environments: A Study of Consolidated n-Tier ApplicationsabstractModern datacenters employ server virtualization and consolidation to reduce the cost of operation and to maximize profit. However, interference among consolidated virtual machines (VMs) has barred mission-critical applications due to unpredictable performance. Through extensive measurements of RUBBoS n-tier benchmark, we found a major source of performance unpredictability: the memory thrashing caused by VM consolidation can reduce the system throughput by 46% although memory was not over-committed. On a physical host with 4 consolidated VMs, we observed two distinct operational modes during a typical RUBBoS benchmark experiment. Over the first half of run-time session we found frequent CPU IOwait causing very long response time requests even though the system is under read-only CPU intensive workload, however, the latter half showed no such CPU abnormalities (IOwait). Using ElbaLens - a lightweight tracing tool, we conducted fine-grain analyses at time granularities as short as 50ms and found that the abnormal IOwait is caused by transient memory thrashing among consolidated VMs. The abnormal IOwait induces queue overflows that propagate through the entire n-tier system, resulting in very long response time requests due to frequent TCP retransmissions. We provide three practical techniques such as VM migration, memory reallocation, soft resource reallocation and show that they can mitigate the effects of performance interference among consolidated VMs. Junhee Park, Qingyang Wang 0001, Jack Li 0001, Chien-An Lai, Calton Pu |
CLOUD | 2 |
| 2015 | A Hybrid Cloud Framework for Scientific ComputingabstractCloud services are transforming many computing tasks, but the unique requirements of scientific computing have caused it to lag behind in cloud adoption because of the performance variation of cloud resources. Based on our experience with the Organic Grid, we propose a framework for a hybrid cloud that will intelligently distribute work to appropriate computing resources to mitigate the impact of performance variation. We describe a cloud framework that integrates with specialized hardware and distributes work intelligently among heterogeneous computing resources. Our approach is to organize a set of computing nodes in an overlay network, to allow each node as an individual agent to position itself within the network to maximize its productivity. An application finds the resources and decides which task to run on which cloud nodes. Our simulations demonstrate that our methods can significantly reduce communication burdens of the most overworked nodes, especially on networks with the highest task to node ratios. Brian Peterson, Gerald Baumgartner, Qingyang Wang 0001 |
CLOUD | 3 |
| 2014 | IO Performance Interference among Consolidated n-Tier Applications: Sharing Is Better Than Isolation for DisksabstractThe performance unpredictability associated with migrating applications into cloud computing infrastructures has impeded this migration. For example, CPU contention between co-located applications has been shown to exhibit counter-intuitive behavior. In this paper, we investigate IO performance interference through the experimental study of consolidated n-tier applications leveraging the same disk. Surprisingly, we found that specifying a specific disk allocation, e.g., limiting the number of Input/Output Operations Per Second (IOPs) per VM, results in significantly lower performance than fully sharing disk across VMs. Moreover, we observe severe performance interference among VMs can not be totally eliminated even with a sharing strategy (e.g., response times for constant workloads still increase over 1,100%). By using a micro-benchmark (Filebench) and an n-tier application benchmark systems (RUBBoS), we demonstrate the existence of disk contention in consolidated environments, and how performance loss occurs when co-located database systems in order to maintain database consistency flush their logs from memory to disk. Potential solutions to these isolation issues are (1) to increase the log buffer size to amortize the disk IO cost (2) to decrease the number of write threads to alleviate disk contention. We validate these methods experimentally and find a 64% and 57% reduction in response time (or more generally, a reduction in performance interference) for constant and increasing workloads respectively. Chien-An Lai, Qingyang Wang 0001, Josh Kimball, Jack Li 0001, Junhee Park, Calton Pu |
IEEE CLOUD | 2 |
| 2014 | The Impact of Software Resource Allocation on Consolidated n-Tier ApplicationsabstractConsolidating several under-utilized user applications together to achieve higher utilization of hardware resources is important for cloud vendors to reduce cost and maximize profit. In this paper, we study the impact of tuning software resources (e.g., server thread pool size or connection pool size) on n-tier web application performance in a consolidated cloud environment. By measuring CPU utilizations and performance of two consolidated n-tier web application benchmark systems running RUBBoS, we found significant differences depending on the amount of soft resources allocated. When the two systems have different soft resource allocations and are fully utilized, the application with more software resources may steal up to 8% CPU from the co-resident application. Further analysis shows that the CPU stealing is due to more threads being scheduled for the system with higher software resources. By limiting the number of runnable active threads for the consolidated VMs, we were able to mitigate the performance interference. More generally, our results show that careful software resource allocation is a significant factor when deploying and tuning n-tier application performance in clouds. Jack Li 0001, Qingyang Wang 0001, Chien-An Lai, Junhee Park, Daisaku Yokoyama, Calton Pu |
IEEE CLOUD | 2 |
| 2014 | Variations in Performance and Scalability: An Experimental Study in IaaS Clouds Using Multi-Tier WorkloadsabstractThe increasing popularity of clouds drives researchers to find answers to a large variety of new and challenging questions. Through extensive experimental measurements, we show variance in performance and scalability of clouds for two non-trivial scenarios. In the first scenario, we target the public Infrastructure as a Service (IaaS) clouds, and study the case when a multi-tier application is migrated from a traditional datacenter to one of the three IaaS clouds. To validate our findings in the first scenario, we conduct similar study with three private clouds built using three mainstream hypervisors. We used the RUBBoS benchmark application and compared its performance and scalability when hosted in Amazon EC2, Open Cirrus, and Emulab. Our results show that a best-performing configuration in one cloud can become the worst-performing configuration in another cloud. Subsequently, we identified several system level bottlenecks such as high context switching and network driver processing overheads that degraded the performance. We experimentally evaluate concrete alternative approaches as practical solutions to address these problems. We then built the three private clouds using a commercial hypervisor (CVM), Xen, and KVM respectively and evaluated performance characteristics using both RUBBoS and Cloudstone benchmark applications. The three clouds show significant performance variations; for instance, Xen outperforms CVM by 75 percent on the read-write RUBBoS workload and CVM outperforms Xen by over 10 percent on the Cloudstone workload. These observed problems were confirmed at a finer granularity through micro-benchmark experiments that measure component performance directly. Deepal Jayasinghe, Simon Malkowski, Jack Li 0001, Qingyang Wang 0001, Zhikui Wang, Calton Pu |
IEEE Trans. Serv. Comput. | 4 |
| 2013 | An Experimental Study of Rapidly Alternating Bottlenecks in n-Tier ApplicationsabstractIdentifying the location of performance bottlenecks is a non-trivial challenge when scaling n-tier applications in computing clouds. Specifically, we observed that an n-tier application may experience significant performance loss when bottlenecks alternate rapidly between component servers. Such rapidly alternating bottlenecks arise naturally and often from resource dependencies in an n-tier system and bursty workloads. These rapidly alternating bottlenecks are difficult to detect because the saturation in each participating server may have a very short lifespan (e.g., milliseconds) compared to current system monitoring tools and practices with sampling at intervals of seconds or minutes. Using passive network tracing at fine-granularity (e.g., aggregate at every 50ms), we are able to correlate throughput (i.e., request service rate) and load (i.e., number of concurrent requests) in each server of an n-tier system. Our experimental results show conclusive evidence of rapidly alternating bottlenecks caused by system software (JVM garbage collection) and middleware (VM collocation). Qingyang Wang 0001, Yasuhiko Kanemasa, Jack Li 0001, Deepal Jayasinghe, Toshihiro Shimizu, Masazumi Matsubara, Motoyuki Kawaba, Calton Pu |
IEEE CLOUD | 1 |
| 2013 | Detecting Transient Bottlenecks in n-Tier Applications through Fine-Grained AnalysisabstractIdentifying the location of performance bottlenecks is a non-trivial challenge when scaling n-tier applications in computing clouds. Specifically, we observed that an n-tier application may experience significant performance loss when there are transient bottlenecks in component servers. Such transient bottlenecks arise frequently at high resource utilization and often result from transient events (e.g., JVM garbage collection) in an n-tier system and bursty workloads. Because of their short lifespan (e.g., milliseconds), these transient bottlenecks are difficult to detect using current system monitoring tools with sampling at intervals of seconds or minutes. We describe a novel transient bottleneck detection method that correlates throughput (i.e., request service rate) and load (i.e., number of concurrent requests) of each server in an n-tier system at fine time granularity. Both throughput and load can be measured through passive network tracing at millisecond-level time granularity. Using correlation analysis, we can identify the transient bottlenecks at time granularities as short as 50ms. We validate our method experimentally through two case studies on transient bottlenecks caused by factors at the system software layer (e.g., JVM garbage collection) and architecture layer (e.g., Intel SpeedStep). Qingyang Wang 0001, Yasuhiko Kanemasa, Jack Li 0001, Deepal Jayasinghe, Toshihiro Shimizu, Masazumi Matsubara, Motoyuki Kawaba, Calton Pu |
ICDCS | 1 |
| 2012 | Expertus: A Generator Approach to Automate Performance Testing in IaaS CloudsabstractCloud computing is an emerging technology paradigm that revolutionizes the computing landscape by providing on-demand delivery of software, platform, and infrastructure over the Internet. Yet, architecting, deploying, and configuring enterprise applications to run well on modern clouds remains a challenge due to associated complexities and non-trivial implications. The natural and presumably unbiased approach to these questions is thorough testing before moving applications to production settings. However, thorough testing of enterprise applications on modern clouds is cumbersome and error-prone due to a large number of relevant scenarios and difficulties in testing process. We address some of these challenges through Expertus---a flexible code generation framework for automated performance testing of distributed applications in Infrastructure as a Service (IaaS) clouds. Expertus uses a multi-pass compiler approach and leverages template-driven code generation to modularly incorporate different software applications on IaaS clouds. Expertus automatically handles complex configuration dependencies of software applications and significantly reduces human errors associated with manual approaches for software configuration and testing. To date, Expertus has been used to study three distributed applications on five IaaS clouds with over 10,000 different hardware, software, and virtualization configurations. The flexibility and extensibility of Expertus and our own experience on using it shows that new clouds, applications, and software packages can easily be incorporated. Deepal Jayasinghe, Galen S. Swint, Simon Malkowski, Jack Li 0001, Qingyang Wang 0001, Junhee Park, Calton Pu |
IEEE CLOUD | 5 |
| 2012 | Challenges and Opportunities in Consolidation at High Resource Utilization: Non-monotonic Response Time Variations in n-Tier ApplicationsabstractA central goal of cloud computing is high resource utilization through hardware sharing; however, utilization often remains modest in practice due to the challenges in predicting consolidated application performance accurately. We present a thorough experimental study of consolidated n-tier application performance at high utilization to address this issue through reproducible measurements. Our experimental method illustrates opportunities for increasing operational efficiency by making consolidated application performance more predictable in high utilization scenarios. The main focus of this paper are non-trivial dependencies between SLA-critical response time degradation effects and software configurations (i.e., readily available tuning knobs). Methodologically, we directly measure and analyze the resource utilizations, request rates, and performance of two consolidated n-tier application benchmark systems (RUBBoS) in an enterprise-level computer virtualization environment. We find that monotonically increasing the workload of an n-tier application system may unexpectedly spike the overall response time of another co-located system by 300 percent despite stable throughput. Based on these findings, we derive a software configuration best-practice to mitigate such non-monotonic response time variations by enabling higher request-processing concurrency (e.g., more threads) in all tiers. More generally, this experimental study increases our quantitative understanding of the challenges and opportunities in the widely used (but seldom supported, quantified, or even mentioned) hypothesis that applications consolidate with linear performance in cloud environments. Simon Malkowski, Yasuhiko Kanemasa, Hanwei Chen, Masao Yamamoto, Qingyang Wang 0001, Deepal Jayasinghe, Calton Pu, Motoyuki Kawaba |
IEEE CLOUD | 5 |
| 2012 | Response Time Reliability in Cloud Environments: An Empirical Study of n-Tier Applications at High Resource UtilizationabstractWhen running mission-critical web-facing applications (e.g., electronic commerce) in cloud environments, predictable response time, e.g., specified as service level agreements (SLA), is a major performance reliability requirement. Through extensive measurements of n-tier application benchmarks in a cloud environment, we study three factors that significantly impact the application response time predictability: bursty workloads (typical of web-facing applications), soft resource management strategies (e.g., global thread pool or local thread pool), and bursts in system software consumption of hardware resources (e.g., Java Virtual Machine garbage collection). Using a set of profit-based performance criteria derived from typical SLAs, we show that response time reliability is brittle, with large response time variations (order of several seconds) depending on each one of those factors. For example, for the same workload and hardware platform, modest increases in workload burstiness may result in profit drops of more than 50%. Our results show that profitbased performance criteria may contribute significantly to the successful delimitation of performance unreliability boundaries and thus support effective management of clouds. Qingyang Wang 0001, Yasuhiko Kanemasa, Jack Li 0001, Deepal Jayasinghe, Motoyuki Kawaba, Calton Pu |
SRDS | 1 |
| 2011 | Variations in Performance and Scalability When Migrating n-Tier Applications to Different CloudsabstractThe increasing popularity of computing clouds continues to drive both industry and research to provide answers to a large variety of new and challenging questions. We aim to answer some of these questions by evaluating performance and scalability when an n-tier application is migrated from a traditional datacenter environment to an IaaS cloud. We used a representative n-tier macro-benchmark (RUBBoS) and compared its performance and scalability in three different test beds: Amazon EC2, Open Cirrus (an open scientific research cloud), and Emulab (academic research test bed). Interestingly, we found that the best-performing configuration in Emulab can become the worst-performing configuration in EC2. Subsequently, we identified the bottleneck components, high context switch overhead and network driver processing overhead, to be at the system level. These overhead problems were confirmed at a finer granularity through micro-benchmark experiments that measure component performance directly. We describe concrete alternative approaches as practical solutions for resolving these problems. Deepal Jayasinghe, Simon Malkowski, Qingyang Wang 0001, Jack Li 0001, PengCheng Xiong, Calton Pu |
IEEE CLOUD | 3 |
| 2011 | Economical and Robust Provisioning of N-Tier Cloud Workloads: A Multi-level Control ApproachabstractResource provisioning for N-tier web applications in Clouds is non-trivial due to at least two reasons. First, there is an inherent optimization conflict between cost of resources and Service Level Agreement (SLA) compliance. Second, the resource demands of the multiple tiers can be different from each other, and varying along with the time. Resources have to be allocated to multiple (virtual) containers to minimize the total amount of resources while meeting the end-to-end performance requirements for the application. In this paper we address these two challenges through the combination of the resource controllers on both application and container levels. On the application level, a decision maker (i.e., an adaptive feedback controller) determines the total budget of the resources that are required for the application to meet SLA requirements as the workload varies. On the container level, a second controller partitions the total resource budget among the components of the applications to optimize the application performance (i.e., to minimize the round trip time). We evaluated our method with three different workload models -- open, closed, and semi-open - that were implemented in the RUBiS web application benchmark. Our evaluation indicates two major advantages of our method in comparison to previous approaches. First, fewer resources are provisioned to the applications to achieve the same performance. Second, our approach is robust enough to address various types of workloads with time-varying resource demand without reconfiguration. PengCheng Xiong, Zhikui Wang, Simon Malkowski, Qingyang Wang 0001, Deepal Jayasinghe, Calton Pu |
ICDCS | 4 |
| 2011 | The Impact of Soft Resource Allocation on n-Tier Application ScalabilityabstractGood performance and efficiency, in terms of high quality of service and resource utilization for example, are important goals in a cloud environment. Through extensive measurements of an n-tier application benchmark (RUBBoS), we show that overall system performance is surprisingly sensitive to appropriate allocation of soft resources (e.g., server thread pool size). Inappropriate soft resource allocation can quickly degrade overall application performance significantly. Concretely, both under-allocation and over-allocation of thread pool can lead to bottlenecks in other resources because of non-trivial dependencies. We have observed some non-obvious phenomena due to these correlated bottlenecks. For instance, the number of threads in the Apache web server can limit the total useful throughput, causing the CPU utilization of the C-JDBC clustering middleware to decrease as the workload increases. We provide a practical iterative solution approach to this challenge through an algorithmic combination of operational queuing laws and measurement data. Our results show that soft resource allocation plays a central role in the performance scalability of complex systems such as n-tier applications in cloud environments. Qingyang Wang 0001, Simon Malkowski, Deepal Jayasinghe, PengCheng Xiong, Calton Pu, Yasuhiko Kanemasa, Motoyuki Kawaba, Lilian Harada |
IPDPS | 1 |
| 2006 | A New Architecture of Data Access Middleware under Grid EnvironmentabstractData sharing is one of the most important research areas in data grid. Distributed data resource and heterogeneous data schema bring difficulties to data access and sharing. This article mainly focuses on how to deal with the heterogeneity of data schema and put forward a blueprint to solve the data access and sharing problem of heterogeneous physical data resources in the grid. To solve this problem we propose the SDB resource model which extracts three data layers from the physical data resource in order to facilitate the access and sharing of data resource Qingyang Wang 0001, Jingshu Chen, Xibin Gao, Wei Zhou 0011, Baoping Yan |
APSCC | 1 |