VLDB 2026 Research / reviewers in the wild / expert
Long Wang 0003
dblp:68/4459-3
· DBLP profile ↗
24ranked-venue papers
11as first author
3since 2021 · last 2025
0000-0001-6073-1225ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 9 · 2 first-author · 3 since 2021Security and privacy · 8 · 6 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-authorSystems, architecture and hardware · 6 · 4 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Tracing Service Request Processing in CloudabstractABSTRACT Nowadays, more and more IT services are being hosted on cloud systems, which render cloud systems to grow into a huge complex with millions of physical servers, multi‐layer software stacks and the processing of cloud service requests across many servers and software layers. It is highly demanded for cloud service providers to have the capability of getting the knowledge on cloud service behaviour directly from the service execution instead of from people's expertise. This paper studies the problem of tracing cloud service's processing of requests across components in cloud environments and proposes cloud tracing mechanisms for this purpose. We also developed model‐based studies of our proposed mechanisms for analysing certain designs of the mechanisms. The implementation of the proposed cloud tracing is deployed onto the environments of OpenStack, Kubernetes and Hadoop, and the experiments on these environments demonstrate that our mechanisms effectively trace cloud service behaviour and generate a single complete request execution path, while without our mechanisms the cloud tracing either fails to work or results in thousands of path segments. Our mechanisms have a low performance overhead (2.3%) in the experiments. Yinqin Zhao, Long Wang 0003, Xuanqing Shi, Yong Yang 0011, Ying Li 0012, Zhengang Wang, Dongdong Shangguan |
Softw. Test. Verification Reliab. | 4 |
| 2023 | Capturing Request Execution Path for Understanding Service Behavior and Detecting Anomalies Without Code InstrumentationabstractWith the increasing scale and complexity of cloud platforms and big-data analytics platforms, it is becoming more and more challenging to understand and diagnose the processing of a service request across multi-layer software stacks of such platforms. One way that helps to deal with this problem is to accurately capture the complete end-to-end execution path of service requests among all involved components. This paper presents REPTrace, a generic methodology for capturing such execution paths in a transparent fashion. Moreover, this paper demonstrates the effectiveness of REPTrace by presenting how REPTrace can be leveraged for knowledge extraction and anomaly detection on the platforms’ request processing. Our experimental results show that, REPTrace enables capturing a holistic view of the request processing across multiple layers of the platforms (which is missing in official documentation) and discovering important undocumented features of the platforms. Fault injection experiments show execution anomalies are detected with 93% precision and 96% recall with aid of REPTrace. Yong Yang 0011, Long Wang 0003, Ying Li 0012 |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | Tracing Processing of Service Requests in Cloud EnvironmentsabstractCloud computing is growingly popular for hosting IT services, and is also growing into a huge complex with millions of physical servers, multi-layer software stacks and the processing of cloud service requests across many servers and software layers. It is highly demanded for cloud service providers to have the capability of getting the knowledge on cloud service behavior directly from the service execution instead of from people's expertise. This paper studies the problem of tracing cloud services' processing of requests across components in cloud environments, and proposes cloud tracing mechanisms for this purpose. The implementation of the proposed cloud tracing is deployed onto an OpenStack cloud environment, and the experiments performed on the cloud environment shows that our mechanisms effectively trace cloud service behavior and generate a single complete request execution path, while without our mechanisms the cloud tracing either could not work or results in thousands of path segments. Our mechanisms' performance overhead is low (2.3%). Yinqin Zhao, Long Wang 0003, Xuanqing Shi, Yong Yang 0011, Ying Li 0012, Zhengang Wang, Dongdong Shangguan |
PRDC | 4 |
| 2020 | Scheduling Physical Machine Maintenance on Qualified Clouds: What if Migration is not Allowed?abstractIncreasingly, cloud systems host IT services for critical domains such as public health and electric power. Such cloud systems need to be qualified and validated in accordance with compliance policies or government regulations. Configurations of such qualified cloud systems, including VM placement and software deployment, are limited to a small set of configurations which passed the qualification and validation process. So, on practical qualified cloud systems, on-demand VM migration or replication is not allowed. We address the problem of optimally scheduling physical machine maintenance on qualified cloud systems where VM placement cannot change dynamically. Particularly, we try to minimize the number of maintenance waves, and hence, minimize the qualified cloud systems' exposure to loss of high availability. Moreover, we study how the initial placement of service component replicas can help optimize the maintenance schedule while minimize the hardware cost, which is much higher on qualified cloud systems than regular ones. Evaluations in simulations and our real-world qualified cloud system demonstrate the effectiveness of our algorithms and justify our algorithm selection in the real qualified system. In particular, our approach eliminated impacts on service workload and availability during maintenance, while simultaneously largely reduced maintenance time. Long Wang 0003, HariGovind V. Ramasamy, Richard E. Harper |
CLOUD | 1 |
| 2020 | How Far Have We Come in Detecting Anomalies in Distributed Systems? An Empirical Study with a Statement-level Fault Injection MethodabstractAnomaly detection in distributed systems has been a fertile research area, and a range of anomaly detectors have been proposed for distributed systems. Unfortunately, there is no systematic quantitative study of the efficacy of different anomaly detectors, which is of great importance to reveal the deficiencies of existing anomaly detectors and shed light on future research directions. In this paper, we investigate how various anomaly detectors behave on anomalies of different types and the reasons for the same, by extensively injecting software faults into three widely-used distributed systems. We use a statementlevel fault injection method to observe the anomalies, characterize these anomalies, and analyze the detection results from anomaly detectors of three categories. We find that: (1) the distributed systems' own error reporting mechanisms are able to report most of the anomalies (from 82.1% to 92.8%) but they incur a high false alarm rate of 26.6%. (2) State-of-the-art anomaly detectors are able to detect the existence of anomalies with 99.08% precision and 90.60% recall, but there is still a long way to go to pinpoint the accurate location of the detected anomalies, and (3) Log-based anomaly detection techniques outperform other anomaly detection techniques, but not for all anomaly types. Yong Yang 0011, Yifan Wu 0002, Karthik Pattabiraman, Long Wang 0003, Ying Li 0012 |
ISSRE | 4 |
| 2019 | LADRA: Log-based abnormal task detection and root-cause analysis in big data processing with Spark
Siyang Lu, Wei Xiang 0007, BingBing Rao, Byung-Chul Tak, Long Wang 0003, Liqiang Wang 0001 |
Future Gener. Comput. Syst. | 5 |
| 2018 | Transparently Capturing Execution Path of Service/Job Request Processing
Yong Yang 0011, Long Wang 0003, Ying Li 0012 |
ICSOC | 2 |
| 2017 | Predicting Misconfiguration-Induced Unsuccessful Executions of Jobs in Big Data SystemabstractAs the complex workload scheduling and resource allocating mechanism in big data system, programmers' configuration error is one of the most typical root causes of unsuccessful termination of jobs, which can result in performance deterioration, availability degradation, resource inefficiency and user unsatisfactory. In this paper, we propose an approach called SD-Predictor, to predict misconfiguration-induced unsuccessful executions of jobs combining static job configurations and dynamic runtime system state before scheduling and execution, so as to save computing resource and scheduling overheads in big data system. We implement and incorporate SD-Predictor with a popular scheduling framework YARN to optimize job scheduling so as to avoid negative impacts by misconfigured jobs. Moreover, we explore correlations between configurations and termination status of jobs and provide some recommendations for configuration optimization. The experiment results show that our approach performs at 78% of precision, 52% of recall and 2% of false positive rate in unsuccessful job prediction, with significantly better recall and false positive rate than related works. Hongyan Tang, Ying Li 0012, Long Wang 0003, Zhonghai Wu |
COMPSAC (1) | 3 |
| 2017 | Log-based Abnormal Task Detection and Root Cause Analysis for SparkabstractApplication delays caused by abnormal tasks arecommon problems in big data computing frameworks. Anabnormal task in Spark, which may run slowly withouterror or warning logs, not only reduces its resident node'sperformance, but also affects other nodes' efficiency.Spark log files report neither root causes of abnormal tasks,nor where and when abnormal scenarios happen. AlthoughSpark provides a “speculation” mechanism to detect stragglertasks, it can only detect tailed stragglers in each stage. Sincethe root causes of abnormal happening are complicated, thereare no effective ways to detect root causes.This paper proposes an approach to detect abnormality andanalyzes root causes using Spark log files. Unlike commononline monitoring or analysis tools, our approach is a pureoff-line method that can analyze abnormality accurately. Ourapproach consists of four steps. First, a parser preprocessesraw log files to generate structured log data. Second, ineach stage of Spark application, we choose features relatedto execution time and data locality of each task, as well asmemory usage and garbage collection of each node. Third,based on the selected features, we detect where and whenabnormalities happen. Finally, we analyze the problems usingweighted factors to decide the probability of root causes. In thispaper, we consider four potential root causes of abnormalities,which include CPU, memory, network, and disk. The proposedmethod has been tested on real-world Spark benchmarks.To simulate various scenario of root causes, we conductedinterference injections related to CPU, memory, network,and Disk. Our experimental results show that the proposedapproach is accurate on detecting abnormal tasks as well asfinding the root causes Siyang Lu, BingBing Rao, Wei Xiang 0007, Byung-Chul Tak, Long Wang 0003, Liqiang Wang 0001 |
ICWS | 5 |
| 2017 | Failure Diagnosis for Distributed Systems Using Targeted Fault InjectionabstractThis paper introduces a novel approach to automating failure diagnostics in distributed systems by combining fault injection and data analytics. We use fault injection to populate the database of failures for a target distributed system. When a failure is reported from production environment, the database is queried to find “matched” failures generated by fault injections. Relying on the assumption that similar faults generate similar failures, we use information from the matched failures as hints to locate the actual root cause of the reported failures. In order to implement this approach, we introduce techniques for (i) reconstructing end-to-end execution flows of distributed software components, (ii) computing the similarity of the reconstructed flows, and (iii) performing precise fault injection at pre-specified executing points in distributed systems. We have evaluated our approach using an OpenStack cloud platform, a popular cloud infrastructure management system. Our experimental results showed that this approach is effective in determining the root causes, e.g., fault types and affected components, for 71-100 percent of tested failures. Furthermore, it can provide fault locations close to actual ones and can easily be used to find and fix actual root causes. We have also validated this technique by localizing real bugs that occurred in OpenStack. Cuong Pham 0003, Long Wang 0003, Byung-Chul Tak, Salman Baset, Chunqiang Tang, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Auto-tuning Performance of MPI Parallel Programs Using Resource Management in Container-Based Virtual CloudabstractLoad imbalance problem is one of the major obstacles to achieving optimal performance of High Performance Computing applications. The approach of trying to distribute the problem pieces to each node with the hope of balancing execution time has limits since the performance depends not only on data size but also on many other dynamic factors. This paper describes an approach that uses adaptive resource management enabled by the container-based virtualization to solve the load imbalance problem of MPI programs running in the cloud. Our techniques dynamically adjust CPU resource allocation to MPI processes running as container instances according to the current program execution state and system resource status. The resource allocation among MPI processes is adjusted in two ways: the intra-host level, which dynamically adjusts resources within a host, and the inter-host level, which migrates containers together with MPI processes from one host to another host. We have implemented and evaluated our approach on Amazon EC2 platform using real-world scientific benchmarks and applications, which demonstrates that the performance can be improved up to 31% (with an average of 15%) when compared with the baseline. Hongyi Ma, Liqiang Wang 0001, Byung-Chul Tak, Long Wang 0003, Chunqiang Tang |
CLOUD | 4 |
| 2016 | Disaster Recovery for Cloud-Hosted Enterprise ApplicationsabstractWe describe disaster protection and recovery of cloud-hosted enterprise applications both at the cloud infrastructure level and at the application level. We explore scenarios which favor one option over the other, and scenarios where a combination of both are required for effective and end-to-end protection. Through case studies grounded in the experience of implementing disaster recovery for IBM's Cloud Managed Services (CMS) platform, we highlight the complexities of protecting enterprise applications on the cloud. For recovery planning and execution, we present a scheduling algorithm that recovers machines hosted on cloud by taking into account application-level logical dependencies and the business criticalities of the applications. Long Wang 0003, Richard E. Harper, Ruchi Mahindru, HariGovind V. Ramasamy |
CLOUD | 1 |
| 2015 | Experiences with Building Disaster Recovery for Enterprise-Class CloudsabstractThe ability to recover from disasters is an important requirement for many enterprises. With enterprise-class workloads increasingly hosted on the cloud, cloud customers have come to expect disaster recovery (DR) as a necessary feature from cloud platforms. This paper identifies key challenges in providing DR as a service on enterprise cloud platforms, and portrays DR solutions for a managed cloud platform. In particular, we present the reference architecture for DR solutions, and describe our practical experiences in providing a portfolio of DR solutions for the cloud platform. The solutions cover diverse target recovery sites, such as an equivalent cloud site, a dedicated recovery site, and a customer-owned site. From the experiences, we provide insights into and lessons on implementing DR for enterprise-class clouds. Long Wang 0003, HariGovind V. Ramasamy, Richard E. Harper, Mahesh Viswanathan 0002, Edmond Plattier |
DSN | 1 |
| 2015 | VM-μCheckpoint: Design, Modeling, and Assessment of Lightweight In-Memory VM CheckpointingabstractCheckpointing and rollback techniques enhance reliability and availability of virtual machines and their hosted IT services. This paper proposes VM-μCheckpoint, a light-weight pure-software mechanism for high-frequency checkpointing and rapid recovery for VMs. Compared with existing techniques of VM checkpointing, VM-μCheckpoint tries to minimize checkpoint overhead and speed up recovery by means of copy-on-write, dirty-page prediction and in-place recovery, as well as saving incremental checkpoints in volatile memory. Moreover, VM-μCheckpoint deals with the issue that latency in error detection potentially results in corrupted checkpoints, particularly when checkpointing frequency is high. We also constructed Markov models to study the availability improvements provided by VM-μCheckpoint (from 99 to 99.98 percent on reasonably reliable hypervisors). We designed and implemented VM-μCheckpoint in the Xen VMM. The evaluation results demonstrate that VM-μCheckpoint incurs an average of 6.3 percent overhead (in terms of program execution time) for 50 ms checkpoint intervals when executing the SPEC CINT 2006 benchmark. Error injection experiments demonstrate that VM-μCheckpoint, combined with error detection techniques in RMK, provides high coverage of recovery. Long Wang 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Arun Iyengar |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2013 | CAP3: A Cloud Auto-Provisioning Framework for Parallel Processing Using On-Demand and Spot InstancesabstractCloud computing has drawn increasing attention from the scientific computing community due to its ease of use, elasticity, and relatively low cost. Because a high-performance computing (HPC) application is usually resource demanding, without careful planning, it can incur a high monetary expense even in Cloud. We design a tool called CAP3 (Cloud Auto-Provisioning framework for Parallel Processing) to help a user minimize the expense of running an HPC application in Cloud, while meeting the user-specified job deadline. Given an HPC application, CAP3 automatically profiles the application, builds a model to predict its performance, and infers a proper cluster size that can finish the job within its deadline while minimizing the total cost. To further reduce the cost, CAP3 intelligently chooses the Cloud's reliable on-demand instances or low-cost spot instances, depending on whether the remaining time is tight in meeting the application's deadline. Experiments on Amazon EC2 show that the execution strategy given by CAP3 is cost-effective, by choosing a proper cluster size and a proper instance type (on-demand or spot). Liqiang Wang 0001, Byung-Chul Tak, Long Wang 0003, Chunqiang Tang |
IEEE CLOUD | 4 |
| 2013 | PseudoApp: Performance prediction for application migration to cloud
Byung-Chul Tak, Chunqiang Tang, Long Wang 0003 |
IM | 4 |
| 2012 | Remediating Overload in Over-Subscribed Computing EnvironmentsabstractResource over subscription brings the risk of resource overload. This paper proposes a mechanism to remediate overload without assuming there is always resource available for migration. A work value notion is introduced to compare importance of VMs, and the overload remediation problem is formulated as a variant of Removable Online Multi-Knapsack Problem. An algorithm is proposed to solve this optimization problem. The mechanism is implemented in a large commercial Cloud environment. Experiments and model-based studies demonstrate the effectiveness of the proposed mechanism in remediating overload and its performance in maximizing work values provided by computing environments (27% higher work values than the baseline algorithm in our study). Long Wang 0003, Rafah Hosn, Chunqiang Tang |
IEEE CLOUD | 1 |
| 2010 | Checkpointing virtual machines against transient errorsabstractThis paper proposes VM-μCheckpoint, a lightweight software mechanism for high-frequency checkpointing and rapid recovery of virtual machines. VM-μCheckpoint minimizes checkpoint overhead and speeds up recovery by saving incremental checkpoints in volatile memory and by employing copy-on-write, dirty-page prediction, and in-place recovery. In our approach, knowledge of fault/error latency is used to explicitly address checkpoint corruption, a critical problem, especially when checkpoint frequency is high. We designed and implemented VM-μCheckpoint in the Xen VMM. The evaluation results demonstrate that VM-μCheckpoint incurs an average of 6.3% execution-time overhead for 50ms checkpoint intervals when executing the SPEC CINT 2006 benchmark. Long Wang 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Arun Iyengar |
IOLTS | 1 |
| 2008 | Formalizing System Behavior for Evaluating a System Hang DetectorabstractThis paper presents an approach to formally verify the detection capability of a system hang detector. To achieve this goal, an abstract formal model of a typical Linux system is created to thoroughly exercise all execution scenarios that may lead to hangs. The goal is to expose cases (i.e., hang scenarios) that escape detection. Our system model abstracts the basic hardware (e.g., timer, hardware counter) and software (e.g., processes/threads) components present in the Linux system. The model enables: (i) capturing behavior of these components so as to depict execution scenarios that lead to hangs, and (ii) evaluating hang detection coverage. Explicit-state model checking is applied to reason about system behavior and uncover hang scenarios that escape detection. The results indicate that the proposed framework allows identification of corner cases of hang scenarios that escape detection and provides valuable insight to developers for enhancing detection mechanisms. Long Wang 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SRDS | 1 |
| 2007 | Reliability MicroKernel: Providing Application-Aware Reliability in the OSabstractThis paper describes the reliability MicroKernel (RMK) framework, a loadable kernel module (or a device driver) for providing application-aware reliability, and dynamically configuring reliability mechanisms. Characteristics of application/system execution are exploited transparently through application-aware reliability techniques to achieve low-latency detection, and low-overhead checkpointing. The RMK prototype is implemented in both Linux, and Windows; and it supports detection of application/OS failures, and transparent application checkpointing. Experiment results show that the system hang detection and application hang detection, which exploit characteristics of application, and system behavior, can achieve high coverage (100% observed in our experiments) with a low false positive rate. Moreover, the performance overhead of RMK, and its detection/checkpointing mechanisms, is small: 0.6% for application hang detection, and 0.1% for transparent application checkpointing in the experiments. Long Wang 0003, Zbigniew T. Kalbarczyk, Weining Gu, Ravishankar K. Iyer |
IEEE Trans. Reliab. | 1 |
| 2006 | An OS-level Framework for Providing Application-Aware ReliabilityabstractThe paper describes the reliability microkernel framework (RMK), a loadable kernel module for providing application-aware reliability and dynamically configuring reliability mechanisms installed in RMK. The RMK prototype is implemented in Linux and supports detection of application/OS failures and transparent application checkpointing. Experiment results show that the OS hang detection, which exploits characteristics of application and system behavior, can achieve high coverage (100% in our experiments) and low false positive rate. Moreover, the performance overhead is negligible because instruction counting is performed in hardware Long Wang 0003, Zbigniew T. Kalbarczyk, Weining Gu, Ravishankar K. Iyer |
PRDC | 1 |
| 2005 | Modeling Coordinated Checkpointing for Large-Scale SupercomputersabstractCurrent supercomputing systems consisting of thousands of nodes cannot meet the demands of emerging high-performance scientific applications. As a result, a new generation of supercomputing systems consisting of hundreds of thousands of nodes is being proposed. However, these systems are likely to experience far more frequent failures than today's systems, and such failures must be tackled effectively. Coordinated checkpointing is a common technique to deal with failures in supercomputers. This paper presents a model of a coordinated checkpointing protocol for large-scale supercomputers, and studies its scalability by considering both the coordination overhead and the effect of failures. Unlike most of the existing checkpointing models, the proposed model takes into account failures during checkpointing and recovery, as well as correlated failures. Stochastic activity networks (SANs) are used to model the system, and the model is simulated to study the scalability, reliability, and performance of the system. Long Wang 0003, Karthik Pattabiraman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Lawrence G. Votta, Christopher A. Vick, Alan Wood |
DSN | 1 |
| 2004 | Checkpointing of Control Structures in Main Memory Database SystemsabstractThis paper proposes an application-transparent, low-overhead checkpointing strategy for maintaining consistency of control structures in a commercial main memory database (MMDB) system, based on the ARMOR (adaptive reconfigurable mobile object of reliability) infrastructure. Performance measurements and availability estimates show that the proposed checkpointing scheme significantly enhances database availability (an extra nine in improvement compared with major-recovery-based solutions) while incurring only a small performance overhead (less than 2% in a typical workload of real applications). Long Wang 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, H. Vora, T. Chahande |
DSN | 1 |
| 2003 | Group Communication Protocols under ErrorsabstractGroup communication protocols constitute a basic building block for highly dependable distributed applications. Designing and correctly implementing a group communication system (GCS) is a difficult task. While many theoretical algorithms have been formalized and proved for correctness, only few research projects have experimentally assessed the dependability of GCS implementations under complex error scenarios. This paper describes a thorough error-injection experimental campaign conducted on Ensemble, a popular GCS. By employing synthetic benchmark applications, we stress selected components of the GCS $the group membership service, the FIFO-ordered reliable multicast - under various error models, including errors in the memory (text and heap segments) and in the network messages. The data show that about 5-6% of the failures are due to an error escaping Ensemble's error-containment mechanism and manifesting as a fail silence violation. This constitutes an impediment to achieving high dependability, the natural objective of GCSs. Our results are derived for a particular system (Ensemble), and more investigation involving other GCSs is required to generalize the conclusions. Nevertheless, through an accurate analysis of the failure causes and the error propagation patterns, this paper offers insights into the design and the implementation of robust GCSs. Claudio Basile, Long Wang 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SRDS | 2 |