Zbigniew T. Kalbarczyk

dblp:18/5985 · also Zbigniew Kalbarczyk · DBLP profile ↗
← Back
143ranked-venue papers
2as first author
18since 2021 · last 2026
0009-0002-6040-6865ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 77 · 4 since 2021Systems, architecture and hardware · 74 · 1 first-author · 10 since 2021Software engineering, systems software and programming languages · 23 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 3 since 2021Artificial intelligence and machine learning · 8 · 2 since 2021Computer networks · 4Databases, data management, data science and information retrieval · 3Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Detection of Compromised Entities in Cross-Domain Communication Systems
abstract
Cross-domain communication systems are important in multiple areas such as smart factories and automatic driving. Derived from real-life scenarios, each domain has a local network and accesses the public network via a central server, which makes it easy to manage the local network and protect entities in the domain. Based on this, a cross-domain communication system follows the “ device-server-server-device ” pattern. For each communication session, there are two steps: authentication and communication. The authentication phase helps two parties to build a secure communication channel. If the server is compromised, the authentication request can be sent to the attacker and then the attacker can impersonate the original receiver. As a result, it is necessary to effectively solve this server-compromise case. Further, anomaly detection is required to monitor the behaviors of entities in domains. Current deep learning solutions deploy autoencoder-based models to reconstruct input data and then detect anomalies. However, they feed raw input data into the models, which makes it hard to reconstruct evolving data streams. In this work, we focus on the detection of compromised entities for cross-domain communication systems. Specifically, we split the problem into two parts: the server-compromise case in the authentication phase and anomaly detection for the whole system. Toward the server-compromise case, we propose a blockchain-based “double verification” scheme to prevent the server from making decisions on its own. Specifically, nodes in the blockchain network evaluate submitted records using the public key infrastructure. Meanwhile, the distributed nature of blockchain allows us to deploy more machines as verifiers to monitor the behavior of servers. For anomaly detection, we propose to feed the degree of change of items instead of raw item values into deep learning models, which is more adaptive to evolving data streams. We propose to use the linear combination of historical data and recent data to compute the degree of change. Finally, we analyze the security properties of the proposed system and evaluate the proposed anomaly detection method using real datasets. We build a simple cross-domain communication system using the Fabric framework to simulate the “double verification” scheme and the proposed anomaly detection method achieves around 0.11 accuracy improvement on average.
Shiqing Li, Utku Tefek, Ertem Esiner, Zbigniew T. Kalbarczyk, Deming Chen
ACM Trans. Cyber Phys. Syst.4
2026 Resilient Path Tracking of Autonomous Driving under Few-shot Action Space Attacks
abstract
Modern autonomous vehicles face growing cybersecurity risks, especially from action space attacks that directly target vehicle actuators. This article systematically evaluates the resilience of three representative Autonomous Driving (AD) architectures, including modular, end-to-end, and feature-fused agents, against few-shot action space attacks crafted via deep reinforcement learning under a black-box setting. The adversary perturbs the vehicle’s lateral control only during safety-critical moments, using either a camera or an inertial measurement unit. Our results reveal distinct vulnerabilities and behavioral patterns across AD architectures, which underscore the necessity for adaptive and robust defense strategies. However, existing adversarial training defense methods show limitations of overfitting and reliance on attack knowledge. To address these limitations, we propose a learning-based Path Correction System (PCS) that integrates traditional feedback control with an adversarially trained correction loop. The correction loop is selectively activated by a kinematic model-based attack detector to counteract abnormal control deviations. Evaluation experiments show that PCS reduces path-tracking deviation by 78% when the system is under attack.
Xin Lou 0005, Rui Tan 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
ACM Trans. Cyber Phys. Syst.5
2025 Page Migration for Hardware Memory Disaggregation Across a Network
abstract
Hardware memory disaggregation (HMD) is an emerging technology that enables access to remote memory, thereby creating expansive memory pools and reducing memory underutilization in datacenters.However, a significant challenge arises when accessing remote memory over a network: increased contention that can lead to severe application performance degradation.To reduce the performance penalty of using remote memory, the operating system uses page migration to promote frequently accessed pages closer to the processor.However, previously proposed page migration mechanisms do not achieve the best performance in HMD systems because of obliviousness to variable page transfer costs that occur due to network contention.To address these limitations, we present INDIGO: a network-aware page migration framework that uses novel page telemetry and a learning-based approach for network adaptation.We implemented INDIGO in the Linux kernel and evaluated it with common cloud and HPC applications on a real disaggregated memory system prototype.Our evaluation shows that INDIGO offers up to 50-70% improvement in application performance compared to other state-of-the-art page migration policies and reduces network traffic up to 2×.
Archit Patke, Christian Pinto, Saurabh Jha, Haoran Qiu, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
ICS5
2025 Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
abstract
This study characterizes GPU resilience in Delta, a large-scale AI system that consists of 1,056 A100 and H100 GPUs, with over 1,300 petaflops of peak throughput. We used 2.5 years of operational data (11.7 million GPU hours) on GPU errors. Our major findings include: (i) H100 GPU memory resilience is worse than A100 GPU memory, with 3.2x lower per-GPU MTBE for memory errors, (ii) The GPU memory error-recovery mechanisms on H100 GPUs are insufficient to handle the increased memory capacity, (iii) H100 GPUs demonstrate significantly improved GPU hardware resilience over A100 GPUs with respect to critical hardware components, (iv) GPU errors on both A100 and H100 GPUs frequently result in job failures due to the lack of robust recovery mechanisms at the application level, and (v) We project the impact of GPU node availability on larger-scales and find that significant overprovisioning of 5% is necessary to handle GPU failures.
Shengkun Cui, Archit Patke, Aditya Ranjan, Ziheng Chen 0006, Phuong Cao, Gregory H. Bauer, Brett M. Bode, Catello Di Martino, Saurabh Jha, Chandrasekhar Narayanaswami 0001, Daby M. Sow, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SC13
2024 On Practicality of Using ARM TrustZone Trusted Execution Environment for Securing Programmable Logic Controllers
abstract
Programmable logic controllers (PLCs) are crucial devices for implementing automated control in various industrial control systems (ICS), such as smart power grids, water treatment systems, manufacturing, and transportation systems. Owing to their importance, PLCs are often the target of cyber attackers that are aiming at disrupting the operation of ICS, including the nation's critical infrastructure, by compromising the integrity of control logic execution. While a wide range of cybersecurity solutions for ICS have been proposed, they cannot counter strong adversaries with a foothold on the PLC devices, which could manipulate memory, I/O interface, or PLC logic itself. These days, many ICS devices in the market, including PLCs, run on ARM-based processors, and there is a promising security technology called ARM TrustZone, to offer a Trusted Execution Environment (TEE) on embedded devices. Envisioning that such a hardware-assisted security feature becomes available for ICS devices in the near future, this paper investigates the application of the ARM TrustZone TEE technology for enhancing the security of PLC. Our aim is to evaluate the feasibility and practicality of the TEE-based PLCs through the proof-of-concept design and implementation using open-source software such as OP-TEE and OpenPLC. Our evaluation assesses the performance and resource consumption in real-world ICS configurations, and based on the results, we discuss bottlenecks in the OP-TEE secure OS towards a large-scale ICS and desired changes for its application on ICS devices. Our implementation is made available to public for further study and research.
Daisuke Mashima, Wen Shei Ong, Ertem Esiner, Zbigniew T. Kalbarczyk, Ee-Chien Chang
AsiaCCS5
2024 Queue Management for SLO-Oriented Large Language Model Serving
abstract
Large language model (LLM) serving is becoming an increasingly critical workload for cloud providers. Existing LLM serving systems focus on interactive requests, such as chatbots and coding assistants, with tight latency SLO requirements. However, when such systems execute batch requests that have relaxed SLOs along with interactive requests, it leads to poor multiplexing and inefficient resource utilization. To address these challenges, we propose QLM, a queue management system for LLM serving. QLM maintains batch and interactive requests across different models and SLOs in a request queue. Optimal ordering of the request queue is critical to maintain SLOs while ensuring high resource utilization. To generate this optimal ordering, QLM uses a Request Waiting Time (RWT) Estimator that estimates the waiting times for requests in the request queue. These estimates are used by a global scheduler to orchestrate LLM Serving Operations (LSOs) such as request pulling, request eviction, load balancing, and model swapping. Evaluation on heterogeneous GPU devices and models with real-world LLM serving dataset shows that QLM improves SLO attainment by 40-90% and throughput by 20-400% while maintaining or improving device utilization compared to other state-of-the-art LLM serving systems. QLM's evaluation is based on the production requirements of a cloud provider. QLM is publicly available at https://www.github.com/QLM-project/QLM.
Archit Patke, Dhemath Reddy, Saurabh Jha, Haoran Qiu, Christian Pinto, Chandrasekhar Narayanaswami 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SoCC7
2024 Mutiny! How Does Kubernetes Fail, and What Can We Do About It?
abstract
In this paper, we i) analyze and classify real-world failures of Kubernetes (the most popular container orchestration system), ii) develop a framework to perform a fault/error injection campaign targeting the data store preserving the cluster state, and iii) compare results of our fault/error injection experiments with real-world failures, showing that our fault/error injections can recreate many real-world failure patterns. The paper aims to address the lack of studies on systematic analyses of Kubernetes failures to date. Our results show that even a single fault/error (e.g., a bit-flip) in the data stored can propagate, causing cluster-wide failures (3% of injections), service networking issues (4%), and service under/overprovisioning (24%). Errors in the fields tracking dependencies between object caused 51% of such cluster-wide failures. We argue that controlled fault/error injection-based testing should be employed to proactively assess Kubernetes' resiliency and guide the design of failure mitigation strategies.
Marco Barletta, Marcello Cinque, Catello Di Martino, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN4
2024 Power-aware Deep Learning Model Serving with μ-Serve
Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Tamer Basar, Ravishankar K. Iyer
USENIX ATC8
2024 True Attacks, Attack Attempts, or Benign Triggers? An Empirical Measurement of Network Alerts in a Security Operations Center
Zhi Chen 0028, Chenkai Wang 0001, Zhenning Zhang, Sushruth Booma, Phuong Cao, Constantin Adam, Alexander Withers, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Gang Wang 0011
USENIX Security Symposium9
2023 Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task Similarity
abstract
Multi-agent reinforcement learning (MARL) has primarily focused on solving a single task in isolation, while in practice the environment is often evolving, leaving many related tasks to be solved. In this paper, we investigate the benefits of meta-learning in solving multiple MARL tasks collectively. We establish the first line of theoretical results for meta-learning in a wide range of fundamental MARL settings, including learning Nash equilibria in two-player zero-sum Markov games and Markov potential games, as well as learning coarse correlated equilibria in general-sum Markov games. Under natural notions of task similarity, we show that meta-learning achieves provable sharper convergence to various game-theoretical solution concepts than learning each task separately. As an important intermediate step, we develop multiple MARL algorithms with initialization-dependent convergence guarantees. Such algorithms integrate optimistic policy mirror descents with stage-based value updates, and their refined convergence guarantees (nearly) recover the best known results even when a good initialization is unknown. To our best knowledge, such results are also new and might be of independent interest. We further provide numerical simulations to corroborate our theoretical findings.
Weichao Mao, Haoran Qiu, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Tamer Basar
NeurIPS5
2023 AWARE: Automate Workload Autoscaling with Reinforcement Learning in Production Cloud Systems
Haoran Qiu, Weichao Mao, Chen Wang 0039, Hubertus Franke, Alaa Youssef, Zbigniew T. Kalbarczyk, Tamer Basar, Ravishankar K. Iyer
USENIX ATC6
2023 CyberSAGE: The cyber security argument graph evaluation tool
William G. Temple, Carmen Cheh, Binbin Chen 0001, Zbigniew T. Kalbarczyk, William H. Sanders, David M. Nicol
Empir. Softw. Eng.6
2023 Message Authentication and Provenance Verification for Industrial Control Systems
abstract
Successful attacks against industrial control systems (ICSs) often exploit insufficient checking mechanisms. While firewalls, intrusion detection systems, and similar appliances introduce essential checks, their efficacy depends on the attackers’ ability to bypass such middleboxes. We propose a provenance solution to enable the verification of an end-to-end message delivery path and the actions performed on a message. Fast and flexible provenance verification (F2-Pro) provides cryptographically verifiable evidence that a message has originated from a legitimate source and gone through the necessary checks before reaching its destination. F2-Prorelies on lightweight cryptographic primitives and flexibly supports various communication settings and protocols encountered in ICS thanks to its transparent, bump-in-the-wire design. We provide formal definitions and cryptographically prove F2-Pro’s security. For human interaction with ICS via a field service device, F2-Profeatures a multi-factor authentication mechanism that starts the provenance chain from a human user issuing commands. We compatibility tested F2-Proon a smart power grid testbed and reported a sub-millisecond latency overhead per communication hop using a modest ARM Cortex-A15 processor.
Ertem Esiner, Utku Tefek, Daisuke Mashima, Binbin Chen 0001, Zbigniew T. Kalbarczyk, David M. Nicol
ACM Trans. Cyber Phys. Syst.5
2022 SIMPPO: a scalable and incremental online learning framework for serverless resource management
abstract
Serverless Function-as-a-Service (FaaS) offers improved programmability for customers, yet it is not server-"less" and comes at the cost of more complex infrastructure management (e.g., resource provisioning and scheduling) for cloud providers. To maintain service-level objectives (SLOs) and improve resource utilization efficiency, recent research has been focused on applying online learning algorithms such as reinforcement learning (RL) to manage resources. Despite the initial success of applying RL, we first show in this paper that the state-of-the-art single-agent RL algorithm (S-RL) suffers up to 4.8x higher p99 function latency degradation on multi-tenant serverless FaaS platforms compared to isolated environments and is unable to converge during training. We then design and implement a scalable and incremental multi-agent RL framework based on Proximal Policy Optimization (SIMPPO). Our experiments demonstrate that in multi-tenant environments, SIMPPO enables each RL agent to efficiently converge during training and provides online function latency performance comparable to that of S-RL trained in isolation with minor degradation (<9.2%). In addition, SIMPPO reduces the p99 function latency by 4.5x compared to S-RL in multi-tenant cases.
Haoran Qiu, Weichao Mao, Archit Patke, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Tamer Basar, Ravishankar K. Iyer
SoCC6
2022 Exploiting Temporal Data Diversity for Detecting Safety-critical Faults in AV Compute Systems
abstract
Silent data corruption caused by random hardware faults in autonomous vehicle (AV) computational elements is a significant threat to vehicle safety. Previous research has explored design diversity, data diversity, and duplication techniques to detect such faults in other safety-critical domains. However, these are challenging to use for AVs in practice due to significant resource overhead and design complexity. We propose, DiverseAV, a low-cost data-diversity-based redundancy technique for detecting safety-critical random hardware faults in computational elements. DiverseAV introduces data-diversity between the redundant agents by exploiting the temporal semantic consistency available in the AV sensor data. DiverseAV is a black-box technique that offers a plug-and-play solution as it requires no knowledge of the internals of the AI agent responsible for executing driving decisions, requiring little to no modification to the agent itself for achieving high coverage of transient and permanent hardware faults. It is commercially viable because it avoids software modifications to agents that are costly in terms of development and testing time. Specifically, DiverseAV distributes the sensor data between the two software agents in a round-robin manner. As a result, the sensor data for two consecutive time steps are semantically similar in terms of their worldview but significantly different at the bit level, thus ensuring the state and data diversity between the two agents necessary for detecting faults. We demonstrate DiverseAV using an open-source self-driving AI agent which is controlling a car in an open-source world simulator.
Saurabh Jha, Shengkun Cui, Timothy Tsai 0002, Siva Kumar Sastry Hari, Michael B. Sullivan 0001, Zbigniew T. Kalbarczyk, Stephen W. Keckler, Ravishankar K. Iyer
DSN6
2022 A Mean-Field Game Approach to Cloud Resource Management with Function Approximation
abstract
Reinforcement learning (RL) has gained increasing popularity for resource management in cloud services such as serverless computing. As self-interested users compete for shared resources in a cluster, the multi-tenancy nature of serverless platforms necessitates multi-agent reinforcement learning (MARL) solutions, which often suffer from severe scalability issues. In this paper, we propose a mean-field game (MFG) approach to cloud resource management that is scalable to a large number of users and applications and incorporates function approximation to deal with the large state-action spaces in real-world serverless platforms. Specifically, we present an online natural actor-critic algorithm for learning in MFGs compatible with various forms of function approximation. We theoretically establish its finite-time convergence to the regularized Nash equilibrium under linear function approximation and softmax parameterization. We further implement our algorithm using both linear and neural-network function approximations, and evaluate our solution on an open-source serverless platform, OpenWhisk, with real-world workloads from production traces. Experimental results demonstrate that our approach is scalable to a large number of users and significantly outperforms various baselines in terms of function latency and resource utilization efficiency.
Weichao Mao, Haoran Qiu, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Tamer Basar
NeurIPS5
2021 BayesPerf: minimizing performance monitoring errors using Bayesian statistics
abstract
Hardware performance counters (HPCs) that measure low-level architectural and microarchitectural events provide dynamic contextual information about the state of the system. However, HPC measurements are error-prone due to non determinism (e.g., undercounting due to event multiplexing, or OS interrupt-handling behaviors). In this paper, we present BayesPerf, a system for quantifying uncertainty in HPC measurements by using a domain-driven Bayesian model that captures microarchitectural relationships between HPCs to jointly infer their values as probability distributions. We provide the design and implementation of an accelerator that allows for low-latency and low-power inference of the BayesPerf model for x86 and ppc64 CPUs. BayesPerf reduces the average error in HPC measurements from 40.1% to 7.6% when events are being multiplexed. The value of BayesPerf in real-time decision-making is illustrated with a simple example of scheduling of PCIe transfers.
Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
ASPLOS3
2021 Delay sensitivity-driven congestion mitigation for HPC systems
abstract
Modern high-performance computing (HPC) systems concurrently execute multiple distributed applications that contend for the high-speed network leading to congestion. Consequently, application runtime variability and suboptimal system utilization are observed in production systems. To address these problems, we propose Netscope, a congestion mitigation framework based on a novel delay sensitivity metric. Delay sensitivity of an application is used to quantify the impact of congestion on its runtime. Netscope uses delay sensitivity estimates to drive a congestion mitigation mechanism to selectively throttle applications that are less susceptible to congestion. We evaluate Netscope on two Cray Aries systems, including a production supercomputer, on common scientific applications. Our evaluation shows that Netscope has a low training cost and accurately estimates the impact of congestion on application runtime with a correlation between 0.7 and 0.9. Moreover, Netscope reduces application tail runtime increase by up to 16.3x while improving the median system utility by 12%.
Archit Patke, Saurabh Jha, Haoran Qiu, Jim M. Brandt, Ann C. Gentile, Joe Greenseid, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
ICS7
2020 Identifying Failing Point Machines from Sensor-Free Train System Logs
abstract
A great many train systems worldwide are legacy systems, without modern sensors whose data can be mined to detect and predict failures. In this paper, we show how to support failure identification in a legacy system with no sensors, using alarm and natural-language described event logs as the only data sources. With too few failures in a mass of log data to train a traditional machine learning model, we propose a new approach called SA-HMM (Survival Analysis-Hidden Markov Model). After enriching the event logs with Word2vec, SA-HMM uses HMMs and survival analysis to identify failure trends in individual assets and failure tendencies in types of assets, respectively, then combines the two part in a weighted sum that indicates the priority of each asset for preventative maintenance. Our evaluation of SA-HMM with a large amount of urban train data shows that SA-HMM greatly outperforms naive method, HMM, and one-class SVM methods in terms of precision and recall in identifying failing assets, while also offering a tunable balance between those two aspects of performance.
Xin Lou 0005, Binbin Chen 0001, Marianne Winslett, Zbigniew T. Kalbarczyk
IEEE BigData5
2020 ML-Driven Malware that Targets AV Safety
abstract
Ensuring the safety of autonomous vehicles (AVs) is critical for their mass deployment and public adoption. However, security attacks that violate safety constraints and cause accidents are a significant deterrent to achieving public trust in AVs, and that hinders a vendor's ability to deploy AVs. Creating a security hazard that results in a severe safety compromise (for example, an accident) is compelling from an attacker's perspective. In this paper, we introduce an attack model, a method to deploy the attack in the form of smart malware, and an experimental evaluation of its impact on production-grade autonomous driving software. We find that determining the time interval during which to launch the attack is{ critically} important for causing safety hazards (such as collisions) with a high degree of success. For example, the smart malware caused 33X more forced emergency braking than random attacks did, and accidents in 52.6% of the driving simulations.
Saurabh Jha, Shengkun Cui, Subho S. Banerjee, James Cyriac, Timothy Tsai 0002, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN6
2020 Inductive-bias-driven Reinforcement Learning For Efficient Schedules in Heterogeneous Clusters
abstract
The problem of scheduling of workloads onto heterogeneous processors (e.g., CPUs, GPUs, FPGAs) is of fundamental importance in modern data centers. Current system schedulers rely on application/system-specific heuristics that have to be built on a case-by-case basis. Recent work has demonstrated ML techniques for automating the heuristic search by using black-box approaches which require significant training data and time, which make them challenging to use in practice. This paper presents Symphony, a scheduling framework that addresses the challenge in two ways: (i) a domain-driven Bayesian reinforcement learning (RL) model for scheduling, which inherently models the resource dependencies identified from the system architecture; and (ii) a sampling-based technique to compute the gradients of a Bayesian model without performing full probabilistic inference. Together, these techniques reduce both the amount of training data and the time required to produce scheduling policies that significantly outperform black-box approaches by up to 2.2{\texttimes}.
Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
ICML3
2020 AV-FUZZER: Finding Safety Violations in Autonomous Driving Systems
abstract
This paper proposes AV-FUZZER, a testing framework, to find the safety violations of an autonomous vehicle (AV) in the presence of an evolving traffic environment. We perturb the driving maneuvers of traffic participants to create situations in which an AV can run into safety violations. To optimally search for the perturbations to be introduced, we leverage domain knowledge of vehicle dynamics and genetic algorithm to minimize the safety potential of an AV over its projected trajectory. The values of the perturbation determined by this process provide parameters that define participants' trajectories. To improve the efficiency of the search, we design a local fuzzer that increases the exploitation of local optima in the areas where highly likely safety-hazardous situations are observed. By repeating the optimization with significantly different starting points in the search space, AV-FUZZER determines several diverse AV safety violations. We demonstrate AV-FUZZER on an industrial-grade AV platform, Baidu Apollo, and find five distinct types of safety violations in a short period of time. In comparison, other existing techniques can find at most two. We analyze the safety violations found in Apollo and discuss their overarching causes.
Guanpeng Li, Saurabh Jha, Timothy Tsai 0002, Michael B. Sullivan 0001, Siva Kumar Sastry Hari, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
ISSRE7
2020 Measuring Congestion in High-Performance Datacenter Interconnects
Saurabh Jha, Archit Patke, Jim M. Brandt, Ann C. Gentile, Benjamin Lim, Michael T. Showerman, Gregory H. Bauer, Larry Kaplan, Zbigniew T. Kalbarczyk, William T. Kramer, Ravishankar K. Iyer
NSDI9
2020 FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented Microservices
Haoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
OSDI4
2020 Live forensics for HPC systems: a case study on distributed storage systems
abstract
Large-scale high-performance computing systems frequently experience a wide range of failure modes, such as reliability failures (e.g., hang or crash), and resource overload-related failures (e.g., congestion collapse), impacting systems and applications. Despite the adverse effects of these failures, current systems do not provide methodologies for proactively detecting, localizing, and diagnosing failures. We present Kaleidoscope, a near real-time failure detection and diagnosis framework, consisting of of hierarchical domain-guided machine learning models that identify the failing components, the corresponding failure mode, and point to the most likely cause indicative of the failure in near real-time (within one minute of failure occurrence). Kaleidoscope has been deployed on Blue Waters supercomputer and evaluated with more than two years of production telemetry data. Our evaluation shows that Kaleidoscope successfully localized 99.3% and pinpointed the root causes of 95.8% of 843 real-world production issues, with less than 0.01% runtime overhead.
Saurabh Jha, Shengkun Cui, Subho S. Banerjee, Tianyin Xu, Jeremy Enos, Michael T. Showerman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SC7
2020 Assessing and Mitigating Impact of Time Delay Attack: Case Studies for Power Grid Controls
abstract
Due to recent cyber attacks on various cyber-physical systems (CPSes), traditional isolation based security schemes in the critical systems are insufficient to deal with the smart adversaries in CPSes with advanced information and communication technologies (ICTs). In this paper, we develop real-time assessment and mitigation of an attack's impact as a system's built-in mechanisms. We study a general class of attacks, which we call time delay attack, that delays the transmissions of control data packets in the CPS control loops. Based on a joint stability-safety criterion, we propose the attack impact assessment consisting of (i) a machine learning (ML) based safety classification, and (ii) a tandem stability-safety classification that exploits a basic relationship between stability and safety, namely that an unstable system must be unsafe whereas a stable system may not be safe. In this assessment approach, the ML addresses a state explosion problem in the safety classification, whereas the tandem structure reduces false negatives in detecting unsafety arising from imperfect ML. We apply our approach to assess the impact of the attack on power grid automatic generation control, and accordingly develop a two-tiered mitigation that tunes the control gain automatically to restore safety where necessary and shed load only if the tuning is insufficient. We also apply our attack impact assessment approach to a thermal power plant control system consisting of two PID control loops. A mitigation approach by tuning the PID controller is also proposed. Extensive simulations based on a 37-bus system model and a thermal power plant control system are conducted to evaluate the effectiveness of our assessment and mitigation approaches.
Xin Lou 0005, Cuong Tran 0006, Rui Tan 0001, David K. Y. Yau, Zbigniew T. Kalbarczyk, Ambarish Kumar Banerjee, Prakhar Ganesh
IEEE J. Sel. Areas Commun.5
2020 Data Integrity Threats and Countermeasures in Railway Spot Transmission Systems
abstract
Modern trains rely on balises (communication beacons) located on the track to provide location information as they traverse a rail network. Balises, such as those conforming to the Eurobalise standard, were not designed with security in mind and are thus vulnerable to cyber attacks targeting data availability, integrity, or authenticity. In this work, we discuss data integrity threats to balise transmission modules and use high-fidelity simulation to study the risks posed by data integrity attacks. To mitigate such risk, we propose a practical two-layer solution: At the device level, we design a lightweight and low-cost cryptographic solution to protect the integrity of the location information; at the system layer, we devise a secure hybrid train speed controller to mitigate the impact under various attacks. Our simulation results demonstrate the effectiveness of our proposed solutions.
Hoon Wei Lim, William G. Temple, Bao Anh N. Tran, Binbin Chen 0001, Zbigniew T. Kalbarczyk, Jianying Zhou 0001
ACM Trans. Cyber Phys. Syst.5
2019 AcMC 2 : Accelerating Markov Chain Monte Carlo Algorithms for Probabilistic Models
abstract
Probabilistic models (PMs) are ubiquitously used across a variety of machine learning applications. They have been shown to successfully integrate structural prior information about data and effectively quantify uncertainty to enable the development of more powerful, interpretable, and efficient learning algorithms. This paper presents AcMC2, a compiler that transforms PMs into optimized hardware accelerators (for use in FPGAs or ASICs) that utilize Markov chain Monte Carlo methods to infer and query a distribution of posterior samples from the model. The compiler analyzes statistical dependencies in the PM to drive several optimizations to maximally exploit the parallelism and data locality available in the problem. We demonstrate the use of AcMC2 to implement several learning and inference tasks on a Xilinx Virtex-7 FPGA. AcMC2-generated accelerators provide a 47-100× improvement in runtime performance over a 6-core IBM Power8 CPU and a 8-18× improvement over an NVIDIA K80 GPU. This corresponds to a 753-1600× improvement over the CPU and 248-463× over the GPU in performance-per-watt terms.
Subho S. Banerjee, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
ASPLOS2
2019 ML-Based Fault Injection for Autonomous Vehicles: A Case for Bayesian Fault Injection
abstract
The safety and resilience of fully autonomous vehicles (AVs) are of significant concern, as exemplified by several headline-making accidents. While AV development today involves verification, validation, and testing, end-to-end assessment of AV systems under accidental faults in realistic driving scenarios has been largely unexplored. This paper presents DriveFI, a machine learning-based fault injection engine, which can mine situations and faults that maximally impact AV safety, as demonstrated on two industry-grade AV technology stacks (from NVIDIA and Baidu). For example, DriveFI found 561 safety-critical faults in less than 4 hours. In comparison, random injection experiments executed over several weeks could not find any safety-critical faults.
Saurabh Jha, Subho S. Banerjee, Timothy Tsai 0002, Siva Kumar Sastry Hari, Michael B. Sullivan 0001, Zbigniew T. Kalbarczyk, Stephen W. Keckler, Ravishankar K. Iyer
DSN6
2019 CAUDIT: Continuous Auditing of SSH Servers To Mitigate Brute-Force Attacks
Phuong Cao, Yuming Wu, Subho S. Banerjee, Justin Azoff, Alexander Withers, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
NSDI6
2019 Smart Malware that Uses Leaked Control Data of Robotic Applications: The Case of Raven-II Surgical Robots
Key-whan Chung, Peicheng Tang, Zeran Zhu, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Thenkurussi Kesavadas
RAID5
2019 ASAP: Accelerated Short-Read Alignment on Programmable Hardware
abstract
The proliferation of high-throughput sequencing machines ensures rapid generation of up to billions of short nucleotide fragments in a short period of time. This massive amount of sequence data can quickly overwhelm today's storage and compute infrastructure. This paper explores the use of hardware acceleration to significantly improve the runtime of short-read alignment, a crucial step in preprocessing sequenced genomes. We focus on the Levenshtein distance (edit-distance) computation kernel and propose the ASAP accelerator, which utilizes the intrinsic delay of circuits for edit-distance computation elements as a proxy for computation. Our design is implemented on an Xilinx Virtex 7 FPGA in an IBM POWER8 system that uses the CAPI interface for cache coherence across the CPU and FPGA. Our design is$200\times$faster than an equivalent Smith-Waterman-C implementation of the kernel running on the host processor,$40-60\times$faster than an equivalent Landau-Vishkin-C++ implementation of the kernel running on the IBM Power8 host processor, and$2\times$faster for an end-to-end alignment tool for 120–150 base-pair short-read sequences. Further the design represents a$3760\times$improvement over the CPU in performance/Watt terms.
Subho S. Banerjee, Mohamed El-Hadedy 0001, Jong Bin Lim, Zbigniew T. Kalbarczyk, Deming Chen, Steven S. Lumetta, Ravishankar K. Iyer
IEEE Trans. Computers4
2018 A ML-based Runtime System for Executing Dataflow Graphs on Heterogeneous Processors
abstract
No abstract available.
Subho S. Banerjee, Arjun P. Athreya, Zbigniew T. Kalbarczyk, Steven S. Lumetta, Ravishankar K. Iyer
SoCC3
2018 Characterizing Supercomputer Traffic Networks Through Link-Level Analysis
abstract
We present techniques for characterizing bandwidth and congestion characteristics of supercomputer High-Speed Networks (HSN). By utilizing a link-level perspective, we gain generality over analyses which are tied to specific topologies. We illustrate these techniques using five months of a Blue Waters production dataset consisting of network utilization and congestion counters. We find that: i) execution time of the communication-heavy applications is highly correlated to network stalls observed in the network topology and increase in application runtime can be as high as 1.7x with nominal increase in stalls, ii) heterogeneity in the available link bandwidth in the network can lead to backpressure and congestion even when the network is not underprovisioned, and (iii) links connected to I/O nodes are no more likely to observe congestion during operational hours than any other link in the system.
Saurabh Jha, Jim M. Brandt, Ann C. Gentile, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
CLUSTER4
2018 Hands Off the Wheel in Autonomous Vehicles?: A Systems Perspective on over a Million Miles of Field Data
abstract
Autonomous vehicle (AV) technology is rapidly becoming a reality on U.S. roads, offering the promise of improvements in traffic management, safety, and the comfort and efficiency of vehicular travel. The California Department of Motor Vehicles (DMV) reports that between 2014 and 2017, manufacturers tested 144 AVs, driving a cumulative 1,116,605 autonomous miles, and reported 5,328 disengagements and 42 accidents involving AVs on public roads. This paper investigates the causes, dynamics, and impacts of such AV failures by analyzing disengagement and accident reports obtained from public DMV databases. We draw several conclusions. For example, we find that autonomous vehicles are 15 - 4000× worse than human drivers for accidents per cumulative mile driven; that drivers of AVs need to be as alert as drivers of non-AVs; and that the AVs' machine-learning-based systems for perception and decision-and-control are the primary cause of 64% of all disengagements.
Subho S. Banerjee, Saurabh Jha, James Cyriac, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN4
2018 Resiliency of HPC Interconnects: A Case Study of Interconnect Failures and Recovery in Blue Waters
abstract
Availability of the interconnection network in high-performance computing (HPC) systems is fundamental to sustaining the continuous execution of applications at scale. When failures occur, interconnect recovery mechanisms orchestrate complex operations to recover network connectivity between the nodes. As the scale and design complexity of HPC systems increase, so does the system's susceptibility to failures during execution of interconnect-recovery procedures. This study characterizes the recovery procedures of the Gemini interconnect network, the largest Gemini network built by Cray, on Blue Waters, a 13.3 petaflop supercomputer at the National Center for Supercomputing Applications (NCSA). We propose a propagation model that captures interconnect failures and recovery procedures to help understand types of failures and their propagation in both the system and applications during recovery. The measurements show that recovery procedures occur very frequently and that the unsuccessful execution of recovery procedures, when additional failures occur during recovery, causes system-wide outages (SWOs, 28 out of 101) and application failures (3.4 percent of all running applications).
Saurabh Jha, Valerio Formicola, Catello Di Martino, Mark Dalton, William T. Kramer, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
IEEE Trans. Dependable Secur. Comput.6
2017 Holistic Measurement-Driven System Assessment
abstract
In high-performance computing systems, application performance and throughput are dependent on a complex interplay of hardware and software subsystems and variable workloads with competing resource demands. Data-driven insights into the potentially widespread scope and propagationof impact of events, such as faults and contention for shared resources, can be used to drive more effective use of resources, for improved root cause diagnosis, and for predicting performance impacts. We present work developing integrated capabilities for holistic monitoring and analysis to understand and characterize propagation of performance-degrading events. These characterizations can be used to determine and invoke mitigating responses by system administrators, applications, and system software.
Saurabh Jha, Jim M. Brandt, Ann C. Gentile, Zbigniew T. Kalbarczyk, Gregory H. Bauer, Jeremy Enos, Michael T. Showerman, Larry Kaplan, Brett M. Bode, Annette Greiner, Amanda Bonnie, Mike Mason, Ravishankar K. Iyer, William T. Kramer
CLUSTER4
2017 Smart Maintenance via Dynamic Fault Tree Analysis: A Case Study on Singapore MRT System
abstract
Urban railway systems, as the most heavily used systems in daily life, suffer from frequent service disruptions resulting millions of affected passengers and huge economic losses. Maintenance of the systems is done by maintaining individual devices in fixed cycles. It is time consuming, yet not effective. Thus, to reduce service failures through smart maintenance is becoming one of the top priorities of the system operators. In this paper, we propose a data driven approach that is to decide maintenance cycle based on estimating the mean time to failure of the system. There are two challenges: 1) as a cyber physical system, hardwares of cyber components (like signalling devices) fail more frequently than physical components (like power plants), 2) as a system of systems, functional dependency exists not only between components within a sub-system but also between different sub-systems, for example, a train relies on traction power system to operate. To meet the challenges, a Dynamic Fault Tree (DFT) based approach is adopted for the expressiveness of the modelling formalism and an efficient tool support by DFTCalc. Our case study shows interesting results that the Singapore Massive Rapid Train (MRT) system is likely to fail in 20 days from the full functioning status based on the manufacture data.
Zbigniew T. Kalbarczyk
DSN3
2017 ASAP: Accelerated Short Read Alignment on Programmable Hardware (Abstract Only)
Subho S. Banerjee, Mohamed El-Hadedy 0001, Jong Bin Lim, Daniel Chen 0001, Zbigniew T. Kalbarczyk, Deming Chen, Ravishankar K. Iyer
FPGA5
2017 On accelerating pair-HMM computations in programmable hardware
abstract
This paper explores hardware acceleration to significantly improve the runtime of computing the forward algorithm on Pair-HMM models, a crucial step in analyzing mutations in sequenced genomes. We describe 1) the design and evaluation of a novel accelerator architecture that can efficiently process real sequence data without performing wasteful work; and 2) aggressive memoization techniques that can significantly reduce the number of invocations of, and the amount of data transferred to the accelerator. We describe our demonstration of the design on a Xilinx Virtex 7 FPGA in an IBM Power8 system. Our design achieves a 14.85× higher throughput than an 8-core CPU baseline (that uses SIMD and multi-threading) and a 147.49 × improvement in throughput per unit of energy expended on the NA12878 sample.
Subho S. Banerjee, Mohamed El-Hadedy 0001, Ching Y. Tan, Zbigniew T. Kalbarczyk, Steven S. Lumetta, Ravishankar K. Iyer
FPL4
2017 Trustworthy Services Built on Event-Based Probing for Layered Defense
abstract
Numerous event-based probing methods exist for cloud computing environments allowing a hypervisor to gain insight into guest activities. Such event-based probing has been shown to be useful for detecting attacks, system hangs through watchdogs, and for inserting exploit detectors before a system can be patched, among others. Here, we illustrate how to use such probing for trustworthy logging and highlight some of the challenges that existing event-based probing mechanisms do not address. Challenges include ensuring a probe inserted at given address is trustworthy despite the lack of attestation available for probes that have been inserted dynamically. We show how probes can be inserted to ensure proper logging of every invocation of a probed instruction. When combined with attested boot of the hypervisor and guest machines, we can ensure the output stream of monitored events is trustworthy. Using these techniques we build a trustworthy log of certain guest-system-call events. The log powers a cloud-tuned Intrusion Detection System (IDS). New event types are identified that must be added to existing probing systems to ensure attempts to circumvent probes within the guest appear in the log. We highlight the overhead penalties paid by guests to increase guarantees of log completeness when faced with attacks on the guest kernel. Promising results (less that 10% for guests) are shown when a guest relaxes the trade-off between log completeness and overhead. Our demonstrative IDS detects common attack scenarios with simple policies built using our guest behavior recording system.
Read Sprabery, Zachary Estrada, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Rakesh Bobba, Roy H. Campbell
IC2E3
2017 Attack Induced Common-Mode Failures on PLC-Based Safety System in a Nuclear Power Plant: Practical Experience Report
abstract
This paper demonstrates attack induced common-mode failures on an industrial-grade (Tricon) Triple-Modular-Redundant PLC (programmable logic controller) and its impact in a Nuclear Power Plant settings. The attack exploits the fact that during the configuration phase the same control logic is downloaded to all three redundant modules. We describe how an attacker can exploit this vulnerability to embed malicious control logic and how to trigger the attack. The feasibility and the attack impact are evaluated on a testbed, which includes the Tricon PLC as part of a safety protection system in a simulated nuclear power plant.
Bernard Lim, Daniel Chen 0001, Yongkyu An, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
PRDC4
2017 On Train Automatic Stop Control Using Balises: Attacks and a Software-Only Countermeasure
abstract
The components and systems involved in railway operation are subject to stringent reliability and safety requirements, but up until now the cyber security of those same systems has been largely under-explored. In this work, we examine a widely-used railway technology, track beacons or balises, which provide a train with its position on the track and often assist with accurate stopping at stations. Balises have been identified as one potential weak link in train signalling systems. We evaluate an automatic train stop controller that is used in real deployment and show that attackers who can compromise the availability or integrity of the balises' data can cause the trains to stop dozens of meters away from the right position, disrupting train service. To address this risk, we have developed a novel countermeasure that ensures the correct stopping of the trains in the presence of attacks, with only a small extra stopping delay.
William G. Temple, Bao Anh N. Tran, Binbin Chen 0001, Zbigniew T. Kalbarczyk, William H. Sanders
PRDC4
2017 Using OS Design Patterns to Provide Reliability and Security as-a-Service for VM-based Clouds
abstract
This paper extends the concepts behind cloud services to offer hypervisor-based reliability and security monitors for cloud virtual machines. Cloud VMs can be heterogeneous and as such guest OS parameters needed for monitoring can vary across different VMs and must be obtained in some way. Past work involves running code inside the VM, which is unacceptable for a cloud environment. We solve this problem by recognizing that there are common OS design patterns that can be used to infer monitoring parameters from the guest OS. We extract information about the cloud user's guest OS with the user's existing VM image and knowledge of OS design patterns as the only inputs to analysis. To demonstrate the range of monitoring functionality possible with this technique, we implemented four sample monitors: a guest OS process tracer, an OS hang detector, a return-to-user attack detector, and a process-based keylogger detector.
Zachary Estrada, Read Sprabery, Lok K. Yan, Zhongzhi Yu, Roy H. Campbell, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
VEE6
2017 Modeling and Mitigating Impact of False Data Injection Attacks on Automatic Generation Control
abstract
This paper studies the impact of false data injection (FDI) attacks on automatic generation control (AGC), a fundamental control system used in all power grids to maintain the grid frequency at a nominal value. Attacks on the sensor measurements for AGC can cause frequency excursion that triggers remedial actions, such as disconnecting customer loads or generators, leading to blackouts, and potentially costly equipment damage. We derive an attack impact model and analyze an optimal attack, consisting of a series of FDIs that minimizes the remaining time until the onset of disruptive remedial actions, leaving the shortest time for the grid to counteract. We show that, based on eavesdropped sensor data and a few feasible-to-obtain system constants, the attacker can learn the attack impact model and achieve the optimal attack in practice. This paper provides essential understanding on the limits of physical impact of the FDIs on power grids, and provides an analysis framework to guide the protection of sensor data links. For countermeasures, we develop efficient algorithms to detect the attack, estimate which sensor data links are under attack, and mitigate attack impact. Our analysis and algorithms are validated by experiments on a physical 16-bus power system test bed and extensive simulations based on a 37-bus power system model.
Rui Tan 0001, Hoang Hai Nguyen, Yi Shyh Eddy Foo, David K. Y. Yau, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Hoay Beng Gooi
IEEE Trans. Inf. Forensics Secur.5
2017 Failure Diagnosis for Distributed Systems Using Targeted Fault Injection
abstract
This paper introduces a novel approach to automating failure diagnostics in distributed systems by combining fault injection and data analytics. We use fault injection to populate the database of failures for a target distributed system. When a failure is reported from production environment, the database is queried to find “matched” failures generated by fault injections. Relying on the assumption that similar faults generate similar failures, we use information from the matched failures as hints to locate the actual root cause of the reported failures. In order to implement this approach, we introduce techniques for (i) reconstructing end-to-end execution flows of distributed software components, (ii) computing the similarity of the reconstructed flows, and (iii) performing precise fault injection at pre-specified executing points in distributed systems. We have evaluated our approach using an OpenStack cloud platform, a popular cloud infrastructure management system. Our experimental results showed that this approach is effective in determining the root causes, e.g., fault types and affected components, for 71-100 percent of tested failures. Furthermore, it can provide fault locations close to actual ones and can easily be used to find and fix actual root causes. We have also validated this technique by localizing real bugs that occurred in OpenStack.
Cuong Pham 0003, Long Wang 0003, Byung-Chul Tak, Salman Baset, Chunqiang Tang, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
IEEE Trans. Parallel Distributed Syst.6
2017 Analysis and Diagnosis of SLA Violations in a Production SaaS Cloud
abstract
A software-as-a-service (SaaS) needs to provide its intended service as per its stated service-level agreements (SLAs). While SLA violations in a SaaS platform have been reported, not much work has been done to empirically characterize failures of SaaS. In this paper, we study SLA violations of a production SaaS platform, diagnose the causes, unearth several critical failure modes, and then, suggest various solution approaches to increase the availability of the platform as perceived by the end user. Our approach combines field failure data analysis (FFDA) and fault injection. Our study is based on 283 days of operational logs of the platform. During this time, the platform received business workload from 42 customers spread over 22 countries. We have first developed a set of home-grown FFDA tools to analyze the log, and second implemented a fault injector to automatically inject several runtime errors in the application code written in .NET/C#, and then, collate the injection results. We summarize our finding as: first, system failures have caused 93% of all SLA violations; second, our fault injector has been able to recreate a few cases of bursts of SLA violations that could not be diagnosed from the logs; and third, the fault injection mechanism could recreate several error propagation paths leading to data corruptions that the failure data analysis could not reveal. Finally, the paper presents some system-level implication of this study and how the joint use of fault injection and log analysis may help in improving the reliability of the measured platform.
Catello Di Martino, Santonu Sarkar, Rajeshwari Ganesan, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
IEEE Trans. Reliab.4
2016 Towards longitudinal analysis of a population's electronic health records using factor graphs
abstract
In this feasibility study, we demonstrate the use of a factor-graph-based probabilistic graphical model approach to process longitudinal data derived from a population's electronic health records (EHR). Processing of EHR allows for fore-casting patient-specific health complications and inference of population-level statistics on several epidemiological factors. As a case-study, we provide preliminary results and demonstrate feasibility of our approach by processing the EHR of a diabetic cohort in Singapore. Our model passes the feasibility test as we are able to forecast a series of health complications of a new patient based on the factor functions inferred from EHR of 100 diabetic patients spanning 10-years. This forecast gives both the caregivers and the patient a better view of the patient's health in the coming years and increases patient's motivation to stay healthy and conform to medication plan. Furthermore, our approach informs commonly occurring health complications in the population that warrant hospital readmissions, which helps a physician/clinician in decide when to intervene to avoid complications in order to improve the patient's quality of life and minimize the cost of care.
Arjun P. Athreya, Kee Yuan Ngiam, Zhaojing Luo, E. Shyong Tai, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
BDCAT5
2016 Unsupervised single-cell analysis in triple-negative breast cancer: A case study
abstract
This paper demonstrates an unsupervised learning approach to identify genes with significant differential expression across single-cell subpopulations induced by therapeutic treatment. Identifying this set of genes makes it possible to use well-established bioinformatics approaches such as pathway analysis to establish their biological relevance. Then, a biologist can use his/her prior knowledge to investigate in the laboratory, a few particular candidates among the subset of genes overlapping with relevant pathways. Due to the large size of the human genome and limitations in cost and skilled resources, biologists benefit from analytical methods combined with pathway analysis to design laboratory experiments focusing on only a few significant genes. As an example, we show how model-based unsupervised methods can identify a small set of genes (1% of the genome) that have significant differential expression in single-cells and are also highly correlated to pathways (p-value < 1E − 7) with anticancer effects driven by the antidiabetic drug metformin. Further analysis of genes on these relevant pathways reveal three candidate genes previously implicated in several anticancer mechanisms in other cancers, not driven by metformin. Identification of these genes can help biologists and clinicians design laboratory experiments to establish the molecular mechanisms of metformin in triple-negative breast cancer. In a domain where there is no prior knowledge of small biologically significant data, we demonstrate that careful data-driven methods can infer such significant small data to explain biological mechanisms.
Arjun P. Athreya, Alan J. Gaglio, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Junmei Cairns, Krishna Rani Kalari, Richard M. Weinshilboum, Liewei Wang
BIBM3
2016 Targeted Attacks on Teleoperated Surgical Robots: Dynamic Model-Based Detection and Mitigation
abstract
This paper demonstrates targeted cyber-physical attacks on teleoperated surgical robots. These attacks exploit vulnerabilities in the robot's control system to infer a critical time during surgery to drive injection of malicious control commands to the robot. We show that these attacks can evade the safety checks of the robot, lead to catastrophic consequences in the physical system (e.g., sudden jumps of robotic arms or system's transition to an unwanted halt state), and cause patient injury, robot damage, or system unavailability in the middle of a surgery. We present a model-based analysis framework that can estimate the consequences of control commands through real-time computation of robot's dynamics. Our experiments on the RAVEN II robot demonstrate that this framework can detect and mitigate the malicious commands before they manifest in the physical system with an average accuracy of 90%.
Homa Alemzadeh, Daniel Chen 0001, Thenkurussi Kesavadas, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN5
2016 A hardware-in-the-loop simulator for safety training in robotic surgery
abstract
This paper presents a simulation-based safety training simulator for robot assisted surgery. While adverse events occur rarely during training, they could be fatal to the patients if they happen during real surgical procedures and are not handled properly by the surgical team. In this work we propose a hardware-in-the-loop robotic surgery simulator with high fidelity of the robot motion in a simulated environment, which is capable of reproducing adverse events during surgery. The proposed simulator is built upon the Raven-II open source surgical robot, integrated with a simulated surgeon console and a safety hazard injection engine, which automatically injects faults into modules of the robot control software. We simulate representative safety hazards seen in the adverse events, related to da Vinci™ robot, reported to the FDA MAUDE database. A novel haptic feedback strategy is provided to the operator when the underlying dynamics differ from the real robot states.
Homa Alemzadeh, Daniel Chen 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Thenkurussi Kesavadas
IROS4
2016 A Data-Driven Approach to Soil Moisture Collection and Prediction
abstract
Agriculture has been one of the most under-investigated areas in technology, and the development of Precision Agriculture (PA) is still in its early stages. This paper proposes a data-driven methodology on building PA solutions for collection and data modeling systems. Soil moisture, a key factor in the crop growth cycle, is selected as an example to demonstrate the effectiveness of our data-driven approach. On the collection side, a reactive wireless sensor node is developed that aims to capture the dynamics of soil moisture using MicaZ mote and VH400 soil moisture sensor. The prototyped device is tested on field soil to demonstrate its functionality and the responsiveness of the sensors. On the data analysis side, a unique, site-specific soil moisture prediction framework is built on top of models generated by the machine learning techniques SVM (support vector machine) and RVM (relevance vector machine). The framework predicts soil moisture n days ahead based on the same soil and environmental attributes that can be collected by our sensor node. Due to the large data size required by the machine learning algorithms, our framework is evaluated under the Illinois historical data, not field collected sensor data. It achieves low error rates (15%) and high correlations (95%) between predicted values and actual values across 9 different sites when forecasting soil moisture about 2 weeks ahead. Also, it is shown that the prediction outputs can remain accurate over a long period of time (one year) when reliable data are fed to the model every 45 days.
Zhihao Hong, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SMARTCOMP2
2015 Measuring and Understanding Extreme-Scale Application Resilience: A Field Study of 5, 000, 000 HPC Application Runs
abstract
This paper presents an in-depth characterization of the resiliency of more than 5 million HPC application runs completed during the first 518 production days of Blue Waters, a 13.1 petaflop Cray hybrid supercomputer. Unlike past work, we measure the impact of system errors and failures on user applications, i.e., the compiled programs launched by user jobs that can execute across one or more XE (CPU) or XK (CPU+GPU) nodes. The characterization is performed by means of a joint analysis of several data sources, which include workload and error/failure logs. In order to relate system errors and failures to the executed applications, we developed LogDiver, a tool to automate the data pre-processing and metric computation. Some of the lessons learned in this study include: i) while about 1.53% of applications fail due to system problems, the failed applications contribute to about 9% of the production node hours executed in the measured period, i.e., the system consumes computing resources, and system-related issues represent a potentially significant energy cost for the work lost, ii) there is a dramatic increase in the application failure probability when executing full-scale applications: 20x (from 0.008 to 0.162) when scaling XE applications from 10,000 to 22,000 nodes, and 6x (from 0.02 to 0.129) when scaling GPU/hybrid applications from 2000 to 4224 nodes, and iii) the resiliency of hybrid applications is impaired by the lack of adequate error detection capabilities in hybrid nodes.
Catello Di Martino, William T. Kramer, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN3
2015 Model-Based Cybersecurity Assessment with NESCOR Smart Grid Failure Scenarios
abstract
The transformation of traditional power systems to smart grids brings significant benefits, but also exposes the grids to various cyber threats. The recent effort led by US National Electric Sector Cybersecurity Organization Resource (NESCOR) Technical Working Group 1 to compile failure scenarios is an important initiative to document typical cybersecurity threats to smart grids. While these scenarios are an invaluable thought-aid, companies still face challenges in systematically and efficiently applying the failure scenarios to assess security risks for their specific infrastructure. In this work, we develop a model-based process for assessing the security risks from NESCOR failure scenarios. We extend our cybersecurity assessment tool, Cyber-SAGE, to support this process, and use it to analyze 25 failure scenarios. Our results show that CyberSAGE can generate precise and structured security argument graphs to quantitatively reason about the risk of each failure scenario. Further, CyberSAGE can significantly reduce the assessment effort by allowing the reuse of models across different failure scenarios, systems, and attacker profiles to perform "what if?" analysis.
Sumeet Jauhar, Binbin Chen 0001, William G. Temple, Xinshu Dong, Zbigniew T. Kalbarczyk, William H. Sanders, David M. Nicol
PRDC5
2015 Systems-Theoretic Safety Assessment of Robotic Telesurgical Systems
Homa Alemzadeh, Daniel Chen 0001, Zbigniew T. Kalbarczyk, Jaishankar Raman, Nancy G. Leveson, Ravishankar K. Iyer
SAFECOMP4
2015 VM-μCheckpoint: Design, Modeling, and Assessment of Lightweight In-Memory VM Checkpointing
abstract
Checkpointing and rollback techniques enhance reliability and availability of virtual machines and their hosted IT services. This paper proposes VM-μCheckpoint, a light-weight pure-software mechanism for high-frequency checkpointing and rapid recovery for VMs. Compared with existing techniques of VM checkpointing, VM-μCheckpoint tries to minimize checkpoint overhead and speed up recovery by means of copy-on-write, dirty-page prediction and in-place recovery, as well as saving incremental checkpoints in volatile memory. Moreover, VM-μCheckpoint deals with the issue that latency in error detection potentially results in corrupted checkpoints, particularly when checkpointing frequency is high. We also constructed Markov models to study the availability improvements provided by VM-μCheckpoint (from 99 to 99.98 percent on reasonably reliable hypervisors). We designed and implemented VM-μCheckpoint in the Xen VMM. The evaluation results demonstrate that VM-μCheckpoint incurs an average of 6.3 percent overhead (in terms of program execution time) for 50 ms checkpoint intervals when executing the SPEC CINT 2006 benchmark. Error injection experiments demonstrate that VM-μCheckpoint, combined with error detection techniques in RMK, provides high coverage of recovery.
Long Wang 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Arun Iyengar
IEEE Trans. Dependable Secur. Comput.2
2015 Integrity Attacks on Real-Time Pricing in Electric Power Grids
abstract
Modern information and communication technologies used by electric power grids are subject to cyber-security threats. This article studies the impact of integrity attacks on real-time pricing (RTP), an emerging feature of advanced power grids that can improve system efficiency. Recent studies have shown that RTP creates a closed loop formed by the mutually dependent real-time price signals and price-taking demand. Such a closed loop can be exploited by an adversary whose objective is to destabilize the pricing system. Specifically, small malicious modifications to the price signals can be iteratively amplified by the closed loop, causing highly volatile prices, fluctuating power demand, and increased system operating cost. This article adopts a control-theoretic approach to deriving the fundamental conditions of RTP stability under basic demand, supply, and RTP models that characterize the essential behaviors of consumers, suppliers, and system operators, as well as two broad classes of integrity attacks, namely, the scaling and delay attacks. We show that, under an approximated linear time-invariant formulation, the RTP system is at risk of being destabilized only if the adversary can compromise the price signals advertised to consumers, by either reducing their values in the scaling attack or providing old prices to over half of all consumers in the delay attack. The results provide useful guidelines for system operators to analyze the impact of various attack parameters on system stability so that they may take adequate measures to secure RTP systems.
Rui Tan 0001, Varun Badrinath Krishna, David K. Y. Yau, Zbigniew T. Kalbarczyk
ACM Trans. Inf. Syst. Secur.4
2014 Automated Classification of Computer-Based Medical Device Recalls: An Application of Natural Language Processing and Statistical Learning
abstract
This paper presents MedSafe, a framework for automated classification of computer-based medical device recalls. The data is collected from the U.S. Food and Drug Administration (FDA) recalls database. We combined techniques in natural language processing and statistical learning to automatically identify the computer-related recalls, by interpreting the natural language semantics of recall descriptions. We evaluated MedSafe on over 16K recall records submitted to the FDA between years 2007-2013.
Homa Alemzadeh, Raymond Hoagland, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
CBMS3
2014 A Performance Evaluation of Sequence Alignment Software in Virtualized Environments
abstract
The prospect of simpler infrastructure management and affordability has garnered interest in cloud computing from bioinformaticians. However, the performance cost of adopting such an infrastructure model for bioinformatics is not fully known. In an effort to help quantify this performance cost, we ran synthetic benchmarks and measured the runtimes of two short-read alignment applications on cloud-like virtualization environments. The environments were implemented utilizing the KVM hypervisor, the Xen hypervisor, and Linux Containers. We compare the runtime in each environment against a physical server and offer discussion and insights. Though the applications perform similar operations, we observe that their performance characteristics differ, as do their performance in the different virtualized environments. We attribute the differences to the way that these programs utilize system resources. We find that the more CPU-bound Novo align is much less sensitive to virtualization environments than BWA is, and has near-physical server performance even when virtualized. Additionally, we find that static CPU pinning can improve performance, and we demonstrate that Linux Containers offer performance comparable to that of a physical server.
Zachary Estrada, Zachary Stephens, Cuong Manh Pham, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
CCGRID4
2014 AHEMS: Asynchronous Hardware-Enforced Memory Safety
abstract
This paper presents AHEMS (Asynchronous Hardware-Enforced Memory Safety), an architectural support for enforcing spatial and temporal memory safety to protect against memory corruption attacks. We integrated AHEMS with the Leon3 open-source processor and prototype on an FPGA. In an evaluation of the detection coverage using 677 security test cases (including spatial and temporal memory errors), selected from the Juliet Test Suite, AHEMS detected all but one memory safety violation. The missed test case involves overflow of a sub-object in a data structure whose detection is not supported by the current prototype. Performance assessment using the Olden benchmarks shows an average 10.6% overhead, and negligible impact on the processor-critical path (0.06% overhead) and power consumption (0.5% overhead).
Kuan-Yu Tseng, Dao Lu, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSD3
2014 Lessons Learned from the Analysis of System Failures at Petascale: The Case of Blue Waters
abstract
This paper provides an analysis of failures and their impact for Blue Waters, the Cray hybrid (CPU/GPU) supercomputer at the University of Illinois at Urbana-Champaign. The analysis is based on both manual failure reports and automatically generated event logs collected over 261 days. Results include i) a characterization of the root causes of single-node failures, ii) a direct assessment of the effectiveness of system-level fail over as well as memory, processor, network, GPU accelerator, and file system error resiliency, and iii) an analysis of system-wide outages. The major findings of this study are as follows. Hardware is not the main cause of system downtime. This is notwithstanding the fact that hardware-related failures are 42% of all failures. Failures caused by hardware were responsible for only 23% of the total repair time. These results are partially due to the fact that processor and memory protection mechanisms (x8 and x4 Chip kill, ECC, and parity) are able to handle a sustained rate of errors as high as 250 errors/h while providing a coverage of 99.997% out of a set of more than 1.5 million of analyzed errors. Only 28 multiple-bit errors bypassed the employed protection mechanisms. Software, on the other hand, was the largest contributor to the node repair hours (53%), despite being the cause of only 20% of the total number of failures. A total of 29 out of 39 system-wide outages involved the Lustre file system with 42% of them caused by the inadequacy of the automated fail over procedures.
Catello Di Martino, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Fabio Baccanico, Joseph Fullop, William T. Kramer
DSN2
2014 Reliability and Security Monitoring of Virtual Machines Using Hardware Architectural Invariants
abstract
This paper presents a solution that simultaneously addresses both reliability and security (RnS) in a monitoring framework. We identify the commonalities between reliability and security to guide the design of Hyper Tap, a hyper visor-level framework that efficiently supports both types of monitoring in virtualization environments. In Hyper Tap, the logging of system events and states is common across monitors and constitutes the core of the framework. The audit phase of each monitor is implemented and operated independently. In addition, Hyper Tap relies on hardware invariants to provide a strongly isolated root of trust. Hyper Tap uses active monitoring, which can be adapted to enforce a wide spectrum of RnS policies. We validate Hyper Tap by introducing three example monitors: Guest OS Hang Detection (GOSHD), Hidden Root Kit Detection (HRKD), and Privilege Escalation Detection (PED). Our experiments with fault injection and real root kits/exploits demonstrate that Hyper Tap provides robust monitoring with low performance overhead.
Cuong Manh Pham, Zachary Estrada, Phuong Cao, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN4
2014 Analysis and Diagnosis of SLA Violations in a Production SaaS Cloud
abstract
This paper investigates SLA violations of a production SaaS platform by means of joint use of field failure data analysis (FFDA) and fault injection. The objective of this study is to diagnose the causes of SLA violations, pinpoint critical failure modes under realistic error assumptions and identify potential means to increase the user perceived availability of the platform and assurance of SLA requirements. We base our study on 283 days of logs obtained during the production time of the platform, while it was employed to process business data received by 42 customers in 22 countries. In this paper, we develop a set of tools that include i) a FFDA toolset used to analyze the data extracted from the platform and by the operating system event logs and ii) a. NET/C++ injector able to automate the injection of specific runtime errors in the production code and the collection of results. Major findings include i) 93% of all service level agreement (SLA) violations were due to system failures, ii) there were a few cases of bursts of SLA violations that could not be diagnosed from the logs and were revealed from the performed injections, and iii) the error injection revealed several error propagation paths leading to data corruptions that could not be detected from the analysis of failure data.
Catello Di Martino, Daniel Chen 0001, Geetika Goel, Rajeshwari Ganesan, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
ISSRE5
2014 Automatic Generation of Security Argument Graphs
abstract
Graph-based assessment formalisms have proven to be useful in the safety, dependability, and security communities to help stakeholders manage risk and maintain appropriate documentation throughout the system lifecycle. In this paper, we propose a set of methods to automatically construct security argument graphs, a graphical formalism that integrates various security-related information to argue about the security level of a system. Our approach is to generate the graph in a progressive manner by exploiting logical relationships among pieces of diverse input information. Using those emergent argument patterns as a starting point, we define a set of extension templates that can be applied iteratively to grow a security argument graph. Using a scenario from the electric power sector, we demonstrate the graph generation process and highlight its application for system security evaluation in our prototype software tool, Cyber SAGE.
Nils Ole Tippenhauer, William G. Temple, An Hoa Vu, Binbin Chen 0001, David M. Nicol, Zbigniew T. Kalbarczyk, William H. Sanders
PRDC6
2014 Semantic Security Analysis of SCADA Networks to Detect Malicious Control Commands in Power Grids (Poster)
abstract
In this poster, we present a semantic analysis framework based on a collaborative network of intrusion detection systems (IDSes) that we proposed in [3] to detect control-related attacks in power systems. The framework combines system knowledge of both cyber and physical infrastructure in power grids to help the IDS to estimate execution consequences of control commands. We demonstrate the implementation based on Bro IDS [11] and the experimental results on the performance overhead of the semantic analysis framework.
Hui Lin 0005, Adam J. Slagell, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SIN3
2014 An evaluation of zookeeper for high availability in system S
abstract
ZooKeeper provides scalable, highly available coordination services for distributed applications. In this paper, we evaluate the use of ZooKeeper in a distributed stream computing system called System S to provide a resilient name service, dynamic configuration management, and system state management. The evaluation shed light on the advantages of using ZooKeeper in these contexts as well as its limitations. We also describe design changes we made to handle named objects in System S to overcome the limitations. We present detailed experimental results, which we believe will be beneficial to the community.
Cuong Manh Pham, Victor Dogaru, Rohit Wagle, Chitra Venkatramani, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
ICPE5
2013 Impact of integrity attacks on real-time pricing in smart grids
abstract
Modern information and communication technologies used by smart grids are subject to cybersecurity threats. This paper studies the impact of integrity attacks on real-time pricing (RTP), a key feature of smart grids that uses such technologies to improve system efficiency. Recent studies have shown that RTP creates a closed loop formed by the mutually dependent real-time price signals and price-taking demand. Such a closed loop can be exploited by an adversary whose objective is to destabilize the pricing system. Specifically, small malicious modifications to the price signals can be iteratively amplified by the closed loop, causing inefficiency and even severe failures such as blackouts. This paper adopts a control-theoretic approach to deriving the fundamental conditions of RTP stability under two broad classes of integrity attacks, namely, the scaling and delay attacks. We show that the RTP system is at risk of being destabilized only if the adversary can compromise the price signals advertised to smart meters by reducing their values in the scaling attack, or by providing old prices to over half of all consumers in the delay attack. The results provide useful guidelines for system operators to analyze the impact of various attack parameters on system stability, so that they may take adequate measures to secure RTP systems.
Rui Tan 0001, Varun Badrinath Krishna, David K. Y. Yau, Zbigniew T. Kalbarczyk
CCS4
2013 Reliability analysis reloaded: how will we survive?
abstract
In safety related applications and in products with long lifetimes reliability is a must. Moreover, facing future technology nodes of integrated circuit device level reliability may decrease, i.e., counter-measures have to be taken to ensure product level reliability. But assessing the reliability of a large system is not a trivial task. This paper revisits the state-of-the-art in reliability evaluation starting from the physical device level, to the software system level, all the way up to the product level. Relevant standards and future trends are discussed.
Robert C. Aitken, Görschwin Fey, Zbigniew T. Kalbarczyk, Frank Reichenbach, Matteo Sonza Reorda
DATE3
2013 Pluggable Watchdog: Transparent Failure Detection for MPI Programs
abstract
This paper presents a framework and its techniques that can detect various types of runtime errors and failures in MPI programs. The presented framework offloads its detection techniques to an external device (e.g., extension card). By developing intelligence on the normal behavioral and semantic execution patterns of monitored parallel threads, the presented external error detectors can accurately and quickly detect errors and failures. This architecture allows us to use powerful detectors without directly using the computing power of the monitored system. The separation of hardware of the monitored and monitoring systems offers an extra advantage in terms of system reliability. We have prototyped our system on a parallel computer system by using an FPGA-based PCI extension card as a monitoring device. We have conducted a fault injection experiment to evaluate the presented techniques using eight MPI-based parallel programs. The techniques cover ~98.5% of faults, on average. The average performance overhead is 1.8% for techniques that detect crash and hang failures and 6.6% for techniques that detect SDC failures.
Keun Soo Yim, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
IPDPS2
2013 Go with the flow: toward workflow-oriented security assessment
abstract
In this paper we advocate the use of workflow---describing how a system provides its intended functionality---as a pillar of cybersecurity analysis and propose a holistic workflow-oriented assessment framework. While workflow models are currently used in the area of performance and reliability assessment, these approaches are designed neither to assess a system in the presence of an active attacker, nor to assess security aspects such as confidentiality. On the other hand, existing security assessment methods typically focus on modeling the active attacker (e.g., attack graphs), but many rely on restrictive models that are not readily applicable to complex (e.g., cyber-physical or cyber-human) systems.
Binbin Chen 0001, Zbigniew T. Kalbarczyk, David M. Nicol, William H. Sanders, Rui Tan 0001, William G. Temple, Nils Ole Tippenhauer, An Hoa Vu, David K. Y. Yau
NSPW2
2013 SymPLFIED: Symbolic Program-Level Fault Injection and Error Detection Framework
abstract
This paper introduces SymPLFIED, a program-level framework that allows specification of arbitrary error detectors and the verification of their efficacy against hardware errors. SymPLFIED comprehensively enumerates all transient hardware errors in registers, memory, and computation (expressed symbolically as value errors) that potentially evade detection and cause program failure. The framework uses symbolic execution to abstract the state of erroneous values in the program and model checking to comprehensively find all errors that evade detection. We demonstrate the use of SymPLFIED on a widely deployed aircraft collision avoidance application, tcas. Our results show that the SymPLFIED framework can be used to uncover hard-to-detect catastrophic cases caused by transient errors in programs that may not be exposed by random fault injection-based validation. Further, the errors exposed by the framework help us formulate a set of error detectors for the application to avoid the catastrophic case and other incorrect outcomes.
Karthik Pattabiraman, Nithin Nakka, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
IEEE Trans. Computers3
2012 Characterization of the error resiliency of power grid substation devices
abstract
With the advent of modern technologies, microprocessor-based devices are used to monitor and control critical infrastructures, e.g., electric power grids, oil and gas distribution. However, the security and reliability of these microprocessor-based systems is a significant issue, since they are more susceptible to transient errors and malicious attacks. An error in one of these systems could have a cascading and catastrophic impact on the whole infrastructure. This paper explores the error resiliency of power grid substation devices. A software-implemented fault injection technique is used to induce errors/faults inside devices used in power grid substations. The goal is to test the ability of these systems to compute through errors/faults. Our results demonstrate that a single error in a substation device may render the operator in the control center unable to control the operation of a relay in the substation.
Kuan-Yu Tseng, Daniel Chen 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN3
2011 Modeling stream processing applications for dependability evaluation
abstract
This paper describes a modeling framework for evaluating the impact of faults on the output of streaming applications. Our model is based on three abstractions: stream operators, stream connections, and tuples. By composing these abstractions within a Stochastic Activity Network, we allow the modeling of complete applications. We consider faults that lead to data loss and to silent data corruption (SDC). Our framework captures how faults originating in one operator propagate to other operators down the stream processing graph. We demonstrate the extensibility of our framework by evaluating three different fault tolerance techniques: checkpointing, partial graph replication, and full graph replication. We show that under crashes that lead to data loss, partial graph replication has a great advantage in maintaining the accuracy of the application output when compared to checkpointing. We also show that SDC can break the no data duplication guarantees of a full graph replication-based fault tolerance technique.
Gabriela Jacques-Silva, Zbigniew T. Kalbarczyk, Bugra Gedik, Henrique Andrade, Kun-Lung Wu, Ravishankar K. Iyer
DSN2
2011 Improving Log-based Field Failure Data Analysis of multi-node computing systems
abstract
Log-based Field Failure Data Analysis (FFDA) is a widely-adopted methodology to assess dependability properties of an operational system. A key step in FFDA is filtering out entries that are not useful and redundant error entries from the log. The latter is challenging: a fault, once triggered, can generate multiple errors that propagate within the system. Grouping the error entries related to the same fault manifestation is crucial to obtain realistic measurements. This paper deals with the issues of the tuple heuristic, used to group the error entries in the log, in multi-node computing systems. We demonstrate that the tuple heuristic can group entries incorrectly; thus, an improved heuristic that adopts statistical indicators is proposed. We assess the impact of inaccurate grouping on dependability measurements by comparing the results obtained with both the heuristics. The analysis encompasses the log of the Mercury cluster at the National Center for Supercomputing Applications.
Antonio Pecchia, Domenico Cotroneo, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN3
2011 CloudVal: A framework for validation of virtualization environment in cloud infrastructure
abstract
We present CloudVal, a framework to validate the reliability of virtualization environment in Cloud Computing infrastructure. A case study, based on injecting faults in the KVM hypervisor and Xen hypervisor, was conducted to show the viability of the framework. The study shows that due to the architectural differences between KVM and Xen, a direct comparison of the two virtualization systems is not feasible. In order to confidently weigh error resiliency of virtualization systems, more comprehensive studies are required. We believe, however, that the fault injection approach and the fault models proposed in this paper are a good starting point towards designing and implementing a benchmark which would enable the assessment of different virtualization infrastructures in a common manner.
Cuong Manh Pham, Daniel Chen 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN3
2011 Analysis of security data from a large computing organization
abstract
This paper presents an in-depth study of the forensic data on security incidents that have occurred over a period of 5 years at the National Center for Supercomputing Applications at the University of Illinois. The proposed methodology combines automated analysis of data from security monitors and system logs with human expertise to extract and process relevant data in order to: (i) determine the progression of an attack, (ii) establish incident categories and characterize their severity, (iii) associate alerts with incidents, and (iv) identify incidents missed by the monitoring tools and examine the reasons for the escapes. The analysis conducted provides the basis for incident modeling and design of new techniques for security monitoring.
Aashish Sharma, Zbigniew T. Kalbarczyk, James Barlow, Ravishankar K. Iyer
DSN2
2011 Hauberk: Lightweight Silent Data Corruption Error Detector for GPGPU
abstract
High performance and relatively low cost of GPU-based platforms provide an attractive alternative for general purpose high performance computing (HPC). However, the emerging HPC applications have usually stricter output correctness requirements than typical GPU applications (i.e., 3D graphics). This paper first analyzes the error resiliency of GPGPU platforms using a fault injection tool we have developed for commodity GPU devices. On average, 16-33% of injected faults cause silent data corruption (SDC) errors in the HPC programs executing on GPU. This SDC ratio is significantly higher than that measured in CPU programs (<;2.3%). In order to tolerate SDC errors, customized error detectors are strategically placed in the source code of target GPU programs so as to minimize performance impact and error propagation and maximize recoverability. The presented Hauberk technique is deployed in seven HPC benchmark programs and evaluated using a fault injection. The results show a high average error detection coverage (~87%) with a small performance overhead (~15%).
Keun Soo Yim, Cuong Manh Pham, Mushfiq Saleheen, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
IPDPS4
2011 Identifying Compromised Users in Shared Computing Infrastructures: A Data-Driven Bayesian Network Approach
abstract
The growing demand for processing and storage capabilities has led to the deployment of high-performance computing infrastructures. Users log into the computing infrastructure remotely, by providing their credentials (e.g., username and password), through the public network and using well-established authentication protocols, e.g., SSH. However, user credentials can be stolen and an attacker (using a stolen credential) can masquerade as the legitimate user and penetrate the system as an insider. This paper deals with security incidents initiated by using stolen credentials and occurred during the last three years at the National Center for Supercomputing Applications (NCSA) at the University of Illinois. We analyze the key characteristics of the security data produced by the monitoring tools during the incidents and use a Bayesian network approach to correlate (i) data provided by different security tools (e.g., IDS and Net Flows) and (ii) information related to the users' profiles to identify compromised users, i.e., the users whose credentials have been stolen. The technique is validated with the real incident data. The experimental results demonstrate that the proposed approach is effective in detecting compromised users, while allows eliminating around 80% of false positives (i.e., not compromised user being declared compromised).
Antonio Pecchia, Aashish Sharma, Zbigniew T. Kalbarczyk, Domenico Cotroneo, Ravishankar K. Iyer
SRDS3
2011 Automated Derivation of Application-Aware Error Detectors Using Static Analysis: The Trusted Illiac Approach
abstract
This paper presents a technique to derive and implement error detectors to protect an application from data errors. The error detectors are derived automatically using compiler-based static analysis from the backward program slice of critical variables in the program. Critical variables are defined as those that are highly sensitive to errors, and deriving error detectors for these variables provides high coverage for errors in any data value used in the program. The error detectors take the form of checking expressions and are optimized for each control-flow path followed at runtime. The derived detectors are implemented using a combination of hardware and software and continuously monitor the application at runtime. If an error is detected at runtime, the application is stopped so as to prevent error propagation and enable a clean recovery. Experiments show that the derived detectors achieve low-overhead error detection while providing high coverage for errors that matter to the application.
Karthik Pattabiraman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
IEEE Trans. Dependable Secur. Comput.2
2011 Automated Derivation of Application-Specific Error Detectors Using Dynamic Analysis
abstract
This paper proposes a novel technique for preventing a wide range of data errors from corrupting the execution of applications. The proposed technique enables automated derivation of fine-grained, application-specific error detectors based on dynamic traces of application execution. The technique derives a set of error detectors using rule-based templates to maximize the error detection coverage for the application. A probability model is developed to guide the choice of the templates and their parameters for error-detection. The paper also presents an automatic framework for synthesizing the set of detectors in hardware to enable low-overhead, runtime checking of the application. The coverage of the derived detectors is evaluated using fault-injection experiments, while the performance and area overheads of the detectors are evaluated by synthesizing them on reconfigurable hardware.
Karthik Pattabiraman, Giacinto Paolo Saggese, Daniel Chen 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
IEEE Trans. Dependable Secur. Comput.4
2010 A Soldier Health Monitoring System for Military Applications
abstract
With recent advances in technology, various wearable sensors have been developed for the monitoring of human physiological parameters. A Body Sensor Network (BSN) consisting of such physiological and biomedical sensor nodes placed on, near or within a human body can be used for real-time health monitoring. In this paper, we describe an on-going effort to develop a system consisting of interconnected BSNs for real-time health monitoring of soldiers. We discuss the background and an application scenario for this project. We describe the preliminary prototype of the system and present a blast source localization application.
Hock-Beng Lim, Di Ma 0001, Bang Wang 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Kenneth L. Watkin
BSN4
2010 Measurement-based analysis of fault and error sensitivities of dynamic memory
abstract
This paper presents a measurement-based analysis of the fault and error sensitivities of dynamic memory. We extend a software-implemented fault injector to support data-type-aware fault injection into dynamic memory. The results indicate that dynamic memory exhibits about 18 times higher fault sensitivity than static memory, mainly because of the higher activation rate. Furthermore, we show that errors in a large portion of static and dynamic memory space are recoverable by simple software techniques (e.g., reloading data from a disk). The recoverable data include pages filled with identical values (e.g., `0') and pages loaded from files unmodified during the computation. Consequently, the selection of targets for protection should be based on knowledge of recoverability rather than on error sensitivity alone.
Keun Soo Yim, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN2
2010 Checkpointing virtual machines against transient errors
abstract
This paper proposes VM-μCheckpoint, a lightweight software mechanism for high-frequency checkpointing and rapid recovery of virtual machines. VM-μCheckpoint minimizes checkpoint overhead and speeds up recovery by saving incremental checkpoints in volatile memory and by employing copy-on-write, dirty-page prediction, and in-place recovery. In our approach, knowledge of fault/error latency is used to explicitly address checkpoint corruption, a critical problem, especially when checkpoint frequency is high. We designed and implemented VM-μCheckpoint in the Xen VMM. The evaluation results demonstrate that VM-μCheckpoint incurs an average of 6.3% execution-time overhead for 50ms checkpoint intervals when executing the SPEC CINT 2006 benchmark.
Long Wang 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Arun Iyengar
IOLTS2
2010 Analysis of Credential Stealing Attacks in an Open Networked Environment
abstract
This paper analyses the forensic data on credential stealing incidents over a period of 5 years across 5000 machines monitored at the National Center for Supercomputing Applications at the University of Illinois. The analysis conducted is the first attempt in an open operational environment (i) to evaluate the intricacies of carrying out SSH-based credential stealing attacks, (ii) to highlight and quantify key characteristics of such attacks, and (iii) to provide the system level characterization of such incidents in terms of distribution of alerts and incident consequences.
Aashish Sharma, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, James Barlow
NSS2
2009 An end-to-end approach for the automatic derivation of application-aware error detectors
abstract
Critical Variable Recomputation (CVR) based error detection provides high coverage for data critical to an application while reducing the performance overhead associated with detecting benign errors. However, when implemented exclusively in software, the performance penalty associated with CVR based detection is unsuitably high. This paper addresses this limitation by providing a hybrid hardware/software tool chain which allows for the design of efficient error detectors while minimizing additional hardware. Detection mechanisms are automatically derived during compilation and mapped onto hardware where they are executed in parallel with the original task at runtime. When tested using an FPGA platform, results show that our approach incurs an area overhead of 53% while increasing execution time by 27% on average.
Galen Lyle, Shelley Cheny, Karthik Pattabiraman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN4
2009 Effectiveness of machine checks for error diagnostics
abstract
Machine Check Architecture (MCA) is a processor internal architecture subsystem that detects and logs correctable and uncorrectable errors in the data or control paths in each CPU core and the Northbridge. These errors include parity errors associated with caches, TLBs, ECC errors associated with caches and DRAM, and system bus errors. This paper reports on an experimental study on: (i) monitoring a computing cluster for machine checks and using this data to identify patterns that can be employed for error diagnostics and (ii) introducing faults into the machine to understand the resulting machine checks and correlate this data with relevant performance metrics.
Nikhil Pandit, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN2
2009 Second workshop on Compiler and Architectural Techniques for Application Reliability and Security (CATARS)
abstract
The second workshop on Compiler and Architectural Techniques for Application Reliability and Security (CATARS) aims to bring together compiler designers and computer architects with dependability researchers and practitioners. The goal is to provide a forum for research on application dependability using techniques drawn from compilers and computer architecture. While there has been a mushrooming of interest in the areas of dependability-centric compiler-and architecture research, there has been no unified forum for researchers in these areas to publish their work. The CATARS workshop aims to provide such a venue.
Karthik Pattabiraman, Zbigniew T. Kalbarczyk
DSN2
2009 Pervasive embedded systems for detection of traumatic brain injury
abstract
Transient explosions on the battlefield result in blast injuries that are polytrauma in nature. That is, along with physical wounds additional injury can include cognitive and communication impairments related to traumatic brain injury (TBI). The design of a multisensor system embedded in an advanced combat helmet is presented that is capable of real time tracking of physiological signals (EEG, blast pressure, head acceleration, oxygen saturation and heart rate) and facilitating a reliable and dependable decision making process that provide alerts for potential traumatic brain injury. The cyperphysical system focuses on the use of heterogeneous sensors within the helmet pads and a processing element for real time algorithmic processing.
Ajay M. Cheriyan, Albert O. Jarvi, Zbigniew T. Kalbarczyk, Tanya M. Gallagher, Ravishankar K. Iyer, Kenneth L. Watkin
ICME3
2009 Quantitative Analysis of Long-Latency Failures in System Software
abstract
This paper presents a study on long latency failures using accelerated fault injection. The data collected from the experiments are used to analyze the significance, causes, and characteristics of long latency failures caused by soft errors in the processor and the memory. The results indicate that a non-negligible portion of soft errors in the code and data memory lead to long latency failures. The long latency failures are caused by errors with long fault activation times and errors causing failures only under certain runtime conditions. On the other hand, less than 0.5% of soft errors in the processor registers used in kernel mode lead to a failure with latency longer than a thousand seconds. This is due to a strong temporal locality of the register values. The study shows also that the obtained insight can be used to guide design and placement (in the application code and/or system) of application-specific error detectors.
Keun Soo Yim, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
PRDC2
2009 Discovering Application-Level Insider Attacks Using Symbolic Execution
Karthik Pattabiraman, Nithin Nakka, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SEC3
2008 Workshop on compiler and architectural techniques for application reliability and security (CATARS)
abstract
Compiler and architectural techniques can play a vital role in dependability enhancement of applications. These techniques have traditionally been focused on performance enhancement. However this trend is changing as reliability and security are becoming first-class constraints for application design. The goal of the workshop is to provide a common platform for researchers in the dependability community to interact with researchers in the compiler and computer architecture communities so that effective cross-pollination of ideas can occur between these areas.
Karthik Pattabiraman, Shuo Chen 0001, Zbigniew T. Kalbarczyk
DSN3
2008 SymPLFIED: Symbolic program-level fault injection and error detection framework
abstract
This paper introduces SymPLFIED, a program-level framework that allows specification of arbitrary error detectors and the verification of their efficacy against hardware errors. SymPLFIED comprehensively enumerates all transient hardware errors in registers, memory, and computation (expressed as value errors) that potentially evade detection and cause program failure. The framework uses symbolic execution to abstract the state of erroneous values in the program and model checking to comprehensively find all errors that evade detection. We demonstrate the use of SymPLFIED on a widely deployed aircraft collision avoidance application, tcas. Our results show that the SymPLFIED framework can be used to uncover hard-to-detect corner cases caused by transient errors in programs that may not be exposed by random fault-injection based validation.
Karthik Pattabiraman, Nithin Nakka, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN3
2008 Error Behavior Comparison of Multiple Computing Systems: A Case Study Using Linux on Pentium, Solaris on SPARC, and AIX on POWER
abstract
This paper presents an approach to conducting experimental studies for the characterization and comparison of the error behavior in different computing systems. The proposed approach is applied to characterize and compare the error behavior of three commercial systems (Linux 2.6 on Pentium 4, Solaris 10 on UltraSPARC IIIi, and AIX 5.3 on POWER 5) under hardware transient faults. The data is obtained by conducting extensive fault injection into kernel code, kernel stack, and system registers with the NFTAPE framework while running the Apache Web server as a workload. The error behavior comparison shows that the Linux system has the highest average crash latency, the Solaris system has the highest hang rate, and the AIX system has the lowest error sensitivity and the least amount of crashes in the more severe categories.
Daniel Chen 0001, Gabriela Jacques-Silva, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Bruce G. Mealey
PRDC3
2008 Formalizing System Behavior for Evaluating a System Hang Detector
abstract
This paper presents an approach to formally verify the detection capability of a system hang detector. To achieve this goal, an abstract formal model of a typical Linux system is created to thoroughly exercise all execution scenarios that may lead to hangs. The goal is to expose cases (i.e., hang scenarios) that escape detection. Our system model abstracts the basic hardware (e.g., timer, hardware counter) and software (e.g., processes/threads) components present in the Linux system. The model enables: (i) capturing behavior of these components so as to depict execution scenarios that lead to hangs, and (ii) evaluating hang detection coverage. Explicit-state model checking is applied to reason about system behavior and uncover hang scenarios that escape detection. The results indicate that the proposed framework allows identification of corner cases of hang scenarios that escape detection and provides valuable insight to developers for enhancing detection mechanisms.
Long Wang 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SRDS2
2007 How Do Mobile Phones Fail? A Failure Data Analysis of Symbian OS Smart Phones
abstract
While the new generation of hand-held devices, e.g., smart phones, support a rich set of applications, growing complexity of the hardware and runtime environment makes the devices susceptible to accidental errors and malicious attacks. Despite these concerns, very few studies have looked into the dependability of mobile phones. This paper presents measurement-based failure characterization of mobile phones. The analysis starts with a high level failure characterization of mobile phones based on data from publicly available web forums, where users post information on their experiences in using hand-held devices. This initial analysis is then used to guide the development of a failure data logger for collecting failure-related information on SymbianOS-based smart phones. Failure data is collected from 25 phones (in Italy and USA) over the period of 14 months. Key findings indicate that: (i) the majority of kernel exceptions are due to memory access violation errors (56%) and heap management problems (18%), and (ii) on average users experience a failure (freeze or self shutdown) every 11 days. While the study provide valuable insight into the failure sensitivity of smart-phones, more data and further analysis are needed before generalizing the results.
Marcello Cinque, Domenico Cotroneo, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN3
2007 Automated Derivation of Application-aware Error Detectors using Static Analysis
abstract
This paper presents a technique to derive and implement error detectors to protect an application from data errors. The error detectors are derived automatically using compiler-based static analysis from the backward program slice of critical variables in the program. Critical variables are defined as those that are highly sensitive to errors, and deriving error detectors for these variables provides high coverage for errors in any data value used in the program. The error detectors take the form of checking expressions and are optimized for each control flow path followed at runtime. The derived detectors are implemented using a combination of hardware and software. Experiments show that the derived detectors incur low performance overheads while achieving high detection coverage for errors that impact the application.
Karthik Pattabiraman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
IOLTS2
2007 Inner-Circle Consistency for Wireless Ad Hoc Networks
abstract
This paper proposes and evaluates strategies to build reliable and secure wireless ad hoc networks. Our contribution is based on the notion of inner-circle consistency, where local node interaction is used to neutralize errors/attacks at the source, both preventing errors/attacks from propagating in the network and improving the fidelity of the propagated information. We achieve this goal by combining statistical (a proposed fault-tolerant duster algorithm) and security (threshold cryptography) techniques with application-aware checks to exploit the data/computation that is partially and naturally replicated in wireless applications. We have prototyped an inner-circle framework and used it to demonstrate the idea of inner-circle consistency in two significant wireless scenarios: 1) the neutralization of black hole attacks in AODV networks and 2) the neutralization of sensor errors in a target detection/ localization application executed over a wireless sensor network
Claudio Basile, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
IEEE Trans. Mob. Comput.2
2007 Reliability MicroKernel: Providing Application-Aware Reliability in the OS
abstract
This paper describes the reliability MicroKernel (RMK) framework, a loadable kernel module (or a device driver) for providing application-aware reliability, and dynamically configuring reliability mechanisms. Characteristics of application/system execution are exploited transparently through application-aware reliability techniques to achieve low-latency detection, and low-overhead checkpointing. The RMK prototype is implemented in both Linux, and Windows; and it supports detection of application/OS failures, and transparent application checkpointing. Experiment results show that the system hang detection and application hang detection, which exploit characteristics of application, and system behavior, can achieve high coverage (100% observed in our experiments) with a low false positive rate. Moreover, the performance overhead of RMK, and its detection/checkpointing mechanisms, is small: 0.6% for application hang detection, and 0.1% for transparent application checkpointing in the experiments.
Long Wang 0003, Zbigniew T. Kalbarczyk, Weining Gu, Ravishankar K. Iyer
IEEE Trans. Reliab.2
2006 Workshop on Applied Software Reliability (WASR)
abstract
Research on software reliability has been active for several decades now and has produced massive amount of literature to explore new ideas, and system prototypes to experiment with the proposed ideas. Notwithstanding this proficiency, only in a few instances has research work found its way into industrial applications. This apparent uncoordination is exacerbated by the increasing need for quality and dependability guarantees in the more and more computerized modern world.
Adnan Agbaria, Claudio Basile, Zbigniew T. Kalbarczyk
DSN3
2006 An Approach for Detecting and Distinguishing Errors versus Attacks in Sensor Networks
abstract
Distributed sensor networks are highly prone to accidental errors and malicious activities, owing to their limited resources and tight interaction with the environment. Yet only a few studies have analyzed and coped with the effects of corrupted sensor data. This paper contributes with the proposal of an on-the-fly statistical technique that can detect and distinguish faulty data from malicious data in a distributed sensor network. Detecting faults and attacks is essential to ensure the correct semantic of the network, while distinguishing faults from attacks is necessary to initiate a correct recovery action. The approach uses hidden Markov models (HMMs) to capture the error/attack-free dynamics of the environment and the dynamics of error/attack data. It then performs a structural analysis of these HMMs to determine the type of error/attack affecting sensor observations. The methodology is demonstrated with real data traces collected over one month of observation from motes deployed on the Great Duck Island
Claudio Basile, Meeta Gupta, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN3
2006 An OS-level Framework for Providing Application-Aware Reliability
abstract
The paper describes the reliability microkernel framework (RMK), a loadable kernel module for providing application-aware reliability and dynamically configuring reliability mechanisms installed in RMK. The RMK prototype is implemented in Linux and supports detection of application/OS failures and transparent application checkpointing. Experiment results show that the OS hang detection, which exploits characteristics of application and system behavior, can achieve high coverage (100% in our experiments) and low false positive rate. Moreover, the performance overhead is negligible because instruction counting is performed in hardware
Long Wang 0003, Zbigniew T. Kalbarczyk, Weining Gu, Ravishankar K. Iyer
PRDC2
2006 Security Vulnerabilities: From Analysis to Detection and Masking Techniques
abstract
This paper presents a study that uses extensive analysis of real security vulnerabilities to drive the development of: 1) runtime techniques for detection/masking of security attacks and 2) formal source code analysis methods to enable identification and removal of potential security vulnerabilities. A finite-state machine (FSM) approach is employed to decompose programs into multiple elementary activities, making it possible to extract simple predicates to be ensured for security. The FSM analysis pinpoints common characteristics among a broad range of security vulnerabilities: predictable memory layout, unprotected control data, and pointer taintedness. We propose memory layout randomization and control data randomization to mask the vulnerabilities at runtime. We also propose a static analysis approach to detect potential security vulnerabilities using the notion of pointer taintedness.
Shuo Chen 0001, Jun Xu 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
Proc. IEEE3
2006 Active Replication of Multithreaded Applications
abstract
Software-based active replication is expensive in terms of performance overhead. Multithreading can help improve performance; however, thread scheduling is a source of nondeterminism in replica behavior. To achieve strong replica consistency in multithreaded environments, this paper proposes intercepting mutex lock/unlock operations performed by threads on accessing the shared data and contributes with two algorithmic solutions: 1) a loose synchronization algorithm (LSA), which captures the natural concurrency in a leader replica and projects it on follower replicas through interreplica communication, and 2) a preemptive deterministic scheduler (PDS) algorithm, which removes the need for interreplica communication through the notion of round and by suspending threads when it is unable (yet) to schedule them deterministically. Failure behavior and performance of LSA and PDS implementations are evaluated in a triplicated system and compared with existing solutions. A performance evaluation indicates that LSA and PDS outperform existing solutions, with PDS offering lower throughput than LSA. A fault-injection campaign shows that PDS is more robust to errors due to the absence of interreplica communication. Hence, LSA and PDS represent a trade-off between performance and dependability. Finally, LSA and PDS are demonstrated in replicating the Apache Web server, a substantial real-world application.
Claudio Basile, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
IEEE Trans. Parallel Distributed Syst.2
2005 Neutralization of Errors and Attacks in Wireless Ad Hoc Networks
abstract
This paper proposes and evaluates strategies to build reliable and secure wireless ad hoc networks. Our contribution is based on the notion of inner-circle consistency, where local node interaction is used to neutralize errors/attacks at the source, both preventing errors/attacks from propagating in the network and improving the fidelity of the propagated information. We achieve this goal by combining statistical (a proposed fault-tolerant cluster algorithm) and security (threshold cryptography) techniques with application-aware checks to exploit the data/computation that is partially and naturally replicated in wireless applications. We have prototyped an inner-circle framework with the ns-2 network simulator, and we use it to demonstrate the idea of inner-circle consistency in two significant wireless scenarios: (1) the neutralization of black hole attacks in AODV networks and (2) the neutralization of sensor errors in a target detection/localization application executed over a wireless sensor network.
Claudio Basile, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN2
2005 Defeating Memory Corruption Attacks via Pointer Taintedness Detection
abstract
Most malicious attacks compromise system security through memory corruption exploits. Recently proposed techniques attempt to defeat these attacks by protecting program control data. We have constructed a new class of attacks that can compromise network applications without tampering with any control data. These non-control data attacks represent a new challenge to system security. In this paper, we propose an architectural technique to defeat both control data and non-control data attacks based on the notion of pointer taintedness. A pointer is said to be tainted if user input can be used as the pointer value. A security attack is detected whenever a tainted value is dereferenced during program execution. The proposed architecture is implemented on the SimpleScalar processor simulator and is evaluated using synthetic programs as well as real-world network applications. Our technique can effectively detect both control data and non-control data attacks, and it offers better security coverage than current methods. The proposed architecture is transparent to existing programs.
Shuo Chen 0001, Jun Xu 0003, Nithin Nakka, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN4
2005 Microprocessor Sensitivity to Failures: Control vs Execution and Combinational vs Sequential Logic
abstract
The goal of this study is to characterize the impact of soft errors on embedded processors. We focus on control versus speculation logic on one hand, and combinational versus sequential logic on the other. The target system is a gate-level implementation of a DLX-like processor. The synthesized design is simulated, and transients are injected to stress the processor while it is executing selected applications. Analysis of the collected data shows that fault sensitivity of the combinational logic (4.2% for a fault duration of one clock cycle) is not negligible, even though it is smaller than the fault sensitivity of flip-flops (10.4%). Detailed study of the error impact, measured at the application level, reveals that errors in speculation and control blocks collectively contribute to about 34% of crashes, 34% of fail-silent violations and 69% of application incomplete executions. These figures indicate the increasing need for processor-level detection techniques over generic methods, such as ECC and parity, to prevent such errors from propagating beyond the processor boundaries.
Giacinto Paolo Saggese, Anoop Vetteth, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN3
2005 Modeling Coordinated Checkpointing for Large-Scale Supercomputers
abstract
Current supercomputing systems consisting of thousands of nodes cannot meet the demands of emerging high-performance scientific applications. As a result, a new generation of supercomputing systems consisting of hundreds of thousands of nodes is being proposed. However, these systems are likely to experience far more frequent failures than today's systems, and such failures must be tackled effectively. Coordinated checkpointing is a common technique to deal with failures in supercomputers. This paper presents a model of a coordinated checkpointing protocol for large-scale supercomputers, and studies its scalability by considering both the coordination overhead and the effect of failures. Unlike most of the existing checkpointing models, the proposed model takes into account failures during checkpointing and recovery, as well as correlated failures. Stochastic activity networks (SANs) are used to model the system, and the model is simulated to study the scalability, reliability, and performance of the system.
Long Wang 0003, Karthik Pattabiraman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Lawrence G. Votta, Christopher A. Vick, Alan Wood
DSN3
2005 Assessing the Crash-Failure Assumption of Group Communication Protocols
abstract
Designing and correctly implementing group communication systems (GCSs) is notoriously difficult. Assuming that processes fail only by crashing provides a powerful means to simplify the theoretical development of these systems. When making this assumption, however, one should not forget that clean crash failures provide only a coarse approximation of the effects that errors can have in distributed systems. Ignoring such a discrepancy can lead to complex GCS-based applications that pay a large price in terms of performance overhead yet fail to deliver the promised level of dependability. This paper provides a thorough study of error effects in real systems by demonstrating an error-injection-driven design methodology, where error injection is integrated in the core steps of the design process of a robust fault-tolerant system. The methodology is demonstrated for the Fortika toolkit, a Java-based GCS. Error injection enables us to uncover subtle reliability bottlenecks both in the design of Fortika and in the implementation of Java. Based on the obtained insights, we enhance Fortika's design to reduce the identified bottlenecks. Finally, a comparison of the results obtained for Fortika with the results obtained for the OCAML-based Ensemble system in a previous work, allows us to investigate the reliability implications that the choice of the development platform (Java versus OCAML) can have
Sergio Mena, Claudio Basile, Zbigniew T. Kalbarczyk, André Schiper, Ravishankar K. Iyer
ISSRE3
2005 Application-Based Metrics for Strategic Placement of Detectors
abstract
The goal of this paper is to provide low-latency detection and prevent error propagation due to value errors. This paper introduces metrics to guide the strategic placement of detectors and evaluates (using fault injection) the coverage provided by ideal detectors embedded at program locations selected using the computed metrics. The computation is represented in the form of a dynamic dependence graph (DDG), a directed-acyclic graph that captures the dynamic dependencies among the values produced during the course of program execution. The DDG is employed to model error propagation in the program and to derive metrics (e.g., value fanout or lifetime) for detector placement. The coverage of the detectors placed is evaluated using fault injections in real programs, including two large SPEC95 integer benchmarks fgcc and perl). Results show that a small number of detectors, strategically placed, can achieve a high degree of detection coverage.
Karthik Pattabiraman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
PRDC2
2004 Error Sensitivity of the Linux Kernel Executing on PowerPC G4 and Pentium 4 Processors
abstract
The goals of this study are: (i) to compare Linux kernel (2.4.22) behavior under a broad range of errors on two target processors - the Intel Pentium 4 (P4) running RedHat Linux 9.0 and the Motorola PowerPC (G4) running YellowDog Linux 3.0 - and (ii) to understand how architectural characteristics of the target processors impact the error sensitivity of the operating system. Extensive error injection experiments involving over 115,000 faults/errors are conducted targeting the kernel code, data, stack, and CPU system registers. Analysis of the obtained data indicates significant differences between the two platforms in how errors manifest and how they are detected in the hardware and the operating system. In addition to quantifying the observed differences and similarities, the paper provides several examples to support the insights gained from this research.
Weining Gu, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN2
2004 An Architectural Framework for Providing Reliability and Security Support
abstract
This paper explores hardware-implemented error-detection and security mechanisms embedded as modules in a hardware-level framework called the reliability and security engine (RSE), which is implemented as an integral part of a modern microprocessor. The RSE interacts with the processor through an input/output interface. The CHECK instruction, a special extension of the instruction set architecture of the processor, is the interface of the application with the RSE. The detection mechanisms described here in detail are: (I) the memory layout randomization (MLR) module, which randomizes the memory layout of a process in order to foil attackers who assume a fixed system layout, (2) the data dependency tracking (DDT) module, which tracks the dependencies among threads of a process and maintains checkpoints of shared memory pages in order to rollback the threads when an offending (potentially malicious) thread is terminated, and (3) the instruction checker module (ICM), which checks an instruction for its validity or the control-flow of the program just as the instruction enters the pipeline for execution. Performance simulations for the studied modules indicate low overhead of the proposed solutions.
Nithin Nakka, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Jun Xu 0003
DSN2
2004 Checkpointing of Control Structures in Main Memory Database Systems
abstract
This paper proposes an application-transparent, low-overhead checkpointing strategy for maintaining consistency of control structures in a commercial main memory database (MMDB) system, based on the ARMOR (adaptive reconfigurable mobile object of reliability) infrastructure. Performance measurements and availability estimates show that the proposed checkpointing scheme significantly enhances database availability (an extra nine in improvement compared with major-recovery-based solutions) while incurring only a small performance overhead (less than 2% in a typical workload of real applications).
Long Wang 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, H. Vora, T. Chahande
DSN2
2004 Formal Reasoning of Various Categories of Widely Exploited Security Vulnerabilities by Pointer Taintedness Semantics
abstract
This paper is motivated by a low level analysis of various categories of severe security vulnerabilities, which indicates that a common characteristic of many classes of vulnerabilities is pointer taintedness. A pointer is said to be tainted if a user input can directly or indirectly be used as a pointer value. In order to reason about pointer taintedness, a memory model is needed. The main contribution of this paper is the formal definition of a memory model using equational logic, which is used to reason about pointer taintedness. The reasoning is applied to several library functions to extract security preconditions, which must be satisfied to eliminate the possibility of pointer taintedness. The results show that pointer taintedness analysis can expose different classes of security vulnerabilities, such as format string, heap corruption and buffer overflow vulnerabilities, leading us to believe that pointer taintedness provides a unifying perspective for reasoning about security vulnerabilities.
Shuo Chen 0001, Karthik Pattabiraman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SEC3
2004 Hardware Support for High Performance, Intrusion- and Fault-Tolerant Systems
abstract
The paper proposes a combined hardware/software approach for realizing high performance, intrusion- and fault-tolerant services. The approach is demonstrated for (yet not limited to) an attribute authority server, which provides a compelling application due to its stringent performance and security requirements. The key element of the proposed architecture is an FPGA-based, parallel crypto-engine providing (1) optimally dimensioned RSA Processors for efficient execution of computationally intensive RSA signatures and (2) a KeyStore facility used as tamper-resistant storage for preserving secret keys. To achieve linear speed-up (with the number of RSA Processors) and deadlock-free execution in spite of resource-sharing and scheduling/synchronization issues, we have resorted to a number of performance enhancing techniques (e.g., use of different clock domains, optimal balance between internal and external parallelism) and have formally modeled and mechanically proved our crypto-engine with the Spin model checker. At the software level, the architecture combines active replication and threshold cryptography, but in contrast with previous work, the code of our replicas is multithreaded so it can efficiently use an attached parallel crypto-engine to compute an attribute authority partial signature (as required by threshold cryptography). Resulting replicated systems that exhibit nondeterministic behavior, which cannot be handled with conventional replication approaches. Our architecture is based on a preemptive deterministic scheduling algorithm to govern scheduling of replica threads and guarantee strong replica consistency.
Giacinto Paolo Saggese, Claudio Basile, Luigi Romano, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SRDS4
2004 Modeling and evaluating the security threats of transient errors in firewall software
Shuo Chen 0001, Jun Xu 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Keith Whisnant
Perform. Evaluation3
2004 Dependable Systems and Networks-Performance and Dependability Symposium (DSN-PDS) 2002: Selected Papers
Sachin Garg, Zbigniew T. Kalbarczyk
Perform. Evaluation2
2004 Reflections on Industry Trends and Experimental Research in Dependability
abstract
Experimental research in dependability has evolved over the past 30 years accompanied by dramatic changes in the computing industry. To understand the magnitude and nature of this evolution, this paper analyzes industrial trends, namely: 1) shifting error sources, 2) explosive complexity, and 3) global volume. Under each-of these trends, the paper explores research technologies that are applicable either to the finished product or artifact, and the processes that are used to produce products. The study gives a framework to not only reflect on the research of the past, but also project the needs of the future.
Daniel P. Siewiorek, Ram Chillarege, Zbigniew T. Kalbarczyk
IEEE Trans. Dependable Secur. Comput.3
2004 The Effects of an ARMOR-Based SIFT Environment on the Performance and Dependability of User Applications
abstract
Few, distributed software-implemented fault tolerance (SIFT) environments have been experimentally evaluated using substantial applications to show that they protect both themselves and the applications from errors. We present an experimental evaluation of a SIFT environment used to oversee spaceborne applications as part of the Remote Exploration and Experimentation (REE) program at the Jet Propulsion Laboratory. The SIFT environment is built around a set of self-checking ARMOR processes running on different machines that provide error detection and recovery services to themselves and to the REE applications. An evaluation methodology is presented in which over 28,000 errors were injected into both the SIFT processes and two representative REE applications. The experiments were split into three groups of error injections, with each group successively stressing the SIFT error detection and recovery more than the previous group. The results show that the SIFT environment added negligible overhead to the application's execution time during failure-free runs. Correlated failures affecting a SIFT process and application process are possible, but the division of detection and recovery responsibilities in the SIFT environment allows it to recover from these multiple failure scenarios. Only 28 cases were observed in which either the application failed to start or the SIFT environment failed to recognize that the application had completed. Further investigations showed that assertions within the SIFT processes-coupled with object-based incremental checkpointing-were effective in preventing system failures by protecting dynamic data within the SIFT processes.
Keith Whisnant, Ravishankar K. Iyer, Zbigniew T. Kalbarczyk, Phillip H. Jones, David A. Rennels, Raphael R. Some
IEEE Trans. Software Eng.3
2003 A Preemptive Deterministic Scheduling Algorithm for Multithreaded Replicas
abstract
Software-based active replication is expensive in terms of performance overhead. Multithreading can help improve performance; however, thread scheduling is a source of nondeterminism in replica behavior. This paper presents a Preemptive Deterministic Scheduling (PDS) algorithm for ensuring deterministic replica behavior while preserving concurrency. Threads are synchronized only on updates to the shared state. A replica execution is broken into a sequence of rounds and in a round each thread can acquire up to two mutexes. If a thread cannot acquire a mutex it requests, then it checks if all other threads are suspended. If so, the thread fires a new round; otherwise, the thread is suspended. When a new round fires, all threads' mutex requests are known; thus, it is possible to form a deterministic scheduling of mutex acquisitions in the round. No inter-replica communication is required. The algorithm is formally specified, and the proposed formalism is used to prove its correctness. Failure
Claudio Basile, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN2
2003 A Data-Driven Finite State Machine Model for Analyzing Security Vulnerabilities
abstract
This paper combines an analysis of data on security vulnerabilities (published in Bugtraq database) and a focused source-code examination to develop a finite state machine (FSM) model to depict and reason about security vulnerabilities. An in-depth analysis of the vulnerability reports and the corresponding source code of the applications leads to three observations: (i) exploits must pass through multiple elementary activities, (ii) multiple vulnerable operations on several objects are involved in exploiting a vulnerability, and (iii) the vulnerability data and corresponding code inspections allow us to derive a predicate for each elementary activity. Each predicate is represented as a primitive FSM (pFSM). Multiple pFSMs are then combined to create an FSM model of vulnerable operations and possible exploits. The proposed FSM methodology is exemplified by analyzing several types of vulnerabilities reported in the data: stack buffer overflow, integer overflow, heap overflow, input validation vulnerabilities, and format string vulnerabilities. For the studied vulnerabilities, we identify three types of pFSMs, which can be used to analyze operations involved in exploiting vulnerabilities and to identify the security checks to be performed at the elementary activity level. A demonstration of the practical usefulness of the FSM modeling approach was the discovery of a new heap overflow vulnerability now published in Bugtraq. Key words: security vulnerabilities, data analysis, finite state machine modeling. 1.
Shuo Chen 0001, Zbigniew T. Kalbarczyk, Jun Xu 0003, Ravishankar K. Iyer
DSN2
2003 Characterization of Linux Kernel Behavior under Errors
abstract
This paper describes an experimental study of Linux kernel behavior in the presence of errors that impact the instruction stream of the kernel code. Extensive error injection experiments including over 35,000 errors are conducted targeting the most fre- quently used functions in the selected kernel subsystems. Three types of faults/errors injection campaigns are conducted: (1) ran- dom non-branch instruction, (2) random conditional branch, and (3) valid but incorrect branch. The analysis of the obtained data shows: (i) 95% of the crashes are due to four major causes, namely, unable to handle kernel NULL pointer, unable to handle kernel paging request, invalid opcode, and general protection fault, (ii) less than 10% of the crashes are associated with fault propagation and nearly 40% of crash latencies are within 10 cycles, (iii) errors in the kernel can result in crashes that require reformatting the file system to restore system operation; the process of bringing up the system can take nearly an hour. Subsequently, over 35,000 faults/errors are injected into the kernel functions within four subsystems: architecture- dependent code (arch), virtual file system interface (fs), cen- tral section of the kernel (kernel), and memory management (mm). Three types of fault/error injection campaigns are con- ducted: random non-branch, random conditional branch, and valid but incorrect conditional branch. The data is analyzed to quantify the response of the OS as a whole based on the sub- system and to determine which functions are responsible for error sensitivity. The analysis provides a detailed insight into the OS behavior under faults/errors. The major findings in- clude: • Most crashes (95%) are due to four major causes: unable to handle kernel NULL pointer, unable to handle kernel paging request, invalid opcode, and general protection fault. • Nine errors in the kernel result in crashes (most severe crash category), which require reformatting the file system. The process of bringing up the system can take nearly an hour. • Less than 10% of the crashes are associated with fault propagation, and nearly 40% of crash latencies are within 10 cycles. The closer analysis of the propagation patterns indicates that it is feasible to identify strategic locations for embedding additional assertions in the source code of a given subsystem to detect errors and, hence, to prevent er- ror propagation.
Weining Gu, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Zhen-Yu Yang
DSN2
2003 Error-Injection-Based Failure Characterization of the IEEE 1394 Bus
abstract
This paper investigates the behavior of the IEEE 1394 bus in the presence of transient errors in the hardware layers of the protocol. Software-implemented error injection is used to introduce errors into the internals of the 1394 bus hardware chipset. Results from this study indicate that the IEEE 1394 bus protocol provides robust network communication in the presence of single-bit errors in the chipset.
D. J. Beauregard, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Savio N. Chau, Leon Alkalai
IOLTS2
2003 Group Communication Protocols under Errors
abstract
Group communication protocols constitute a basic building block for highly dependable distributed applications. Designing and correctly implementing a group communication system (GCS) is a difficult task. While many theoretical algorithms have been formalized and proved for correctness, only few research projects have experimentally assessed the dependability of GCS implementations under complex error scenarios. This paper describes a thorough error-injection experimental campaign conducted on Ensemble, a popular GCS. By employing synthetic benchmark applications, we stress selected components of the GCS $the group membership service, the FIFO-ordered reliable multicast - under various error models, including errors in the memory (text and heap segments) and in the network messages. The data show that about 5-6% of the failures are due to an error escaping Ensemble's error-containment mechanism and manifesting as a fail silence violation. This constitutes an impediment to achieving high dependability, the natural objective of GCSs. Our results are derived for a particular system (Ensemble), and more investigation involving other GCSs is required to generalize the conclusions. Nevertheless, through an accurate analysis of the failure causes and the error propagation patterns, this paper offers insights into the design and the implementation of robust GCSs.
Claudio Basile, Long Wang 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SRDS3
2003 Transparent Runtime Randomization for Security
abstract
A large class of security attacks exploit software implementation vulnerabilities such as unchecked buffers. This paper proposes transparent runtime randomization (TRR), a generalized approach for protecting against a wide range of security attacks. TRR dynamically and randomly relocates a program's stack, heap, shared libraries, and parts of its runtime control data structures inside the application memory address space. Making a program's memory layout different each time it runs foils the attacker's assumptions about the memory layout of the vulnerable program and makes the determination of critical address values difficult if not impossible. TRR is implemented by changing the Linux dynamic program loader, hence it is transparent to applications. We demonstrate that TRR is effective in defeating real security attacks, including malloc-based heap overflow, integer overflow, and double-free attacks, for which effective prevention mechanisms are yet to emerge. Furthermore, TRR incurs less than 9% program startup overhead and no runtime overhead.
Jun Xu 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SRDS2
2002 An Adaptive Architecture for Monitoring and Failure Analysis of High-Speed Networks
abstract
Describes the design of a reconfigurable device using an FPGA (field programmable gate array) whose primary function is high-speed (several Gb/s) network data monitoring and run-time adaptive fault injection and statistics gathering for failure analysis. The device is designed for two types of media: Myrinet SAN and Fibre Channel, and failure analysis can be performed simultaneously over both of these networks. Although the device intercepts and retransmits signals on the network, no impact on the data transfer rate is observed and the latency caused by inserting the device in the network is negligible. The fault injection capabilities are demonstrated on a Myrinet LAN. Fault injection experiments are conducted on data transmitted across the network, including control packets previously inaccessible to software-based techniques.
Benjamin Floering, B. Brothers, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN3
2002 Joint Panel - IPDS and Workshop on Dependability Benchmarking
Ravishankar K. Iyer, Zbigniew T. Kalbarczyk, Philip Koopman, Henrique Madeira, Gunter Heiner, Karama Kanoun, Haim Levendel, Brendan Murphy, Lawrence G. Votta, Don Wilson
DSN2
2002 NFTAPE: Networked Fault Tolerance and Performance Evaluator
abstract
The NFTAPE is a software implemented, highly flexible fault injection environment for conducting automated fault/error injection-based dependability characterization. NFTAPE: (1) enables a user: (i) to specify a fault/error injection plan, (ii) to carry out injection experiments, and (iii) to collect the experimental results for analysis; (2) targets assessment of a broad set of dependability metrics, e.g., availability, reliability, coverage; (3) operates in a distributed environment; (4) can be configured to implement a variety of fault/error injection strategies and thus to serve multiple users and target systems; (5) imposes minimal disturbance of target systems.
David T. Stott, Phillip H. Jones, M. Hamman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN4
2002 Loose Synchronization of Multithreaded Replicas
abstract
Although multithreading can improve performance, it is a source of nondeterminism in application behavior. Existing approaches to replicating multithreaded applications either synchronize replicas at the interrupt level, at the expense of performance, or use a nonpreemptive deterministic scheduler at the expense of concurrency. This paper presents a loose synchronization algorithm for ensuring deterministic replica behavior while preserving concurrency. The algorithm synchronizes replica threads only on state updates by enforcing an equivalent order of mutex acquisitions across replicas.
Claudio Basile, Keith Whisnant, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SRDS3
2001 A Framework for Database Audit and Control Flow Checking for a Wireless Telephone Network Controller
abstract
The paper presents the design and implementation of a dependability framework for a call-processing environment in a digital mobile telephone network controller. The framework contains a data audit subsystem to maintain the structural and semantic integrity of the database and a preemptive control flow checking technique, PECOS, to protect call-processing clients. Evaluation of the dependability-enhanced system is performed (using NFTAPE, a software-implemented error injection environment). The evaluation shows that for control flow errors in the client, the combination of PECOS and data audit eliminates fail-silence violations, reduces the incidence of client crashes, and eliminates client hangs. For database injections, data audit detects 85% of the errors and reduces the incidence of escaped errors. Evaluation of combined use of data and control checking (with error injection targeting the database and the client) shows coverage increase from 35% to 80% and indicates data flow errors as a key reason for error escapes.
Saurabh Bagchi, Keith Whisnant, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Ytzhak H. Levendel, Lawrence G. Votta
DSN4
2001 An Experimental Study of Security Vulnerabilities Caused by Errors
abstract
The paper presents an experimental study which shows that, for the Intel x86 architecture, single-bit control flow errors in the authentication sections of targeted applications can result in significant security vulnerabilities. The experiment targets two well-known Internet server applications: FTP and SSH (secure shell), injecting single-bit control flow errors into user authentication sections of the applications. The injected sections constitute approximately 2-8% of the text segment of the target applications. The results show that out of all activated errors: (a) 1-2% comprised system security (create a permanent window of vulnerability); (b) 43-62% resulted in crash failures (about 8.5% of these errors create a transient window of vulnerability); and (c) 7-12% resulted in fail silence violations. A key reason for the measured security vulnerabilities is that, in the x86 architecture, conditional branch instructions are a minimum of one Hamming distance apart. The design and evaluation of a new encoding scheme that reduces or eliminates this problem is presented.
Jun Xu 0003, Shuo Chen 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
DSN3
2001 Comparing Fail-Sailence Provided by Process Duplication versus Internal Error Detection for DHCP Server
abstract
This paper uses fault injection to compare the ability of two fault-tolerant software architectures to protect an application from faults. These two architectures are Voltan, which uses process duplication, and Chameleon ARMORs, which use self-checking. The target application is a Dynamic Host Configuration Protocol (DHCP) server, a widely used application for managing IP addresses. NFTAPE, a software-based fault injection environment, is used to inject three classes of faults, namely random memory bit-flip, control-flow and high-level target specific faults, into each software architecture and into baseline Solaris and Linux versions.
David T. Stott, Neil A. Speirs, Zbigniew T. Kalbarczyk, Saurabh Bagchi, Jun Xu 0003, Ravishankar K. Iyer
IPDPS3
2000 Hierarchical Error Detection in a Software Implemented Fault Tolerance (SIFT) Environment
abstract
Proposes a hierarchical error detection framework for a software-implemented fault tolerance (SIFT) layer of a distributed system. A four-level error detection hierarchy is proposed in the context of Chameleon, a software environment for providing adaptive fault tolerance in an environment of commercial off-the-shelf (COTS) system components and software. The design and implementation of a software-based distributed signature monitoring scheme, which is central to the proposed four-level hierarchy, is described. Both intra-level and inter-level optimizations that minimize the overhead of detection and are capable of adapting to runtime requirements are proposed. The paper presents results from a prototype implementation of two levels of the error detection hierarchy and results of a detailed simulation of the overall environment. The results indicate a substantial increase in availability due to the detection framework and help in understanding the tradeoffs between overhead and coverage for different combinations of techniques.
Saurabh Bagchi, Balaji Srinivasan, Keith Whisnant, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
IEEE Trans. Knowl. Data Eng.4
1999 Networked Windows NT System Field Failure Data Analysis
abstract
This paper presents a measurement-based dependability study of a Networked Windows NT system based on field data collected from NT System Logs from 503 servers running in a production environment over a four-month period. The event logs at hand contains only system reboot information. We study individual server failures and domain behavior in order to characterize failure behavior and explore error propagation between servers. The key observations from this study are: (1) system software and hardware failures are the two major contributors to the total system downtime (22% and 10%), (2) recovery from application software failures are usually quick, (3) in many cases, more than one reboots are required to recover from a failure, (4) the average availability of an individual server is over 99%, (5) there is a strong indication of error dependency or error propagation across the network, (6) most (58%) reboots are unclassified indicating the need for better logging techniques, (7) maintenance and configuration contribute to 24% of system downtime.
Jun Xu 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
PRDC2
1999 Failure Data Analysis of a LAN of Windows NT based Computers
abstract
This paper presents results of a failure data analysis of a LAN of Windows NT machines. Data for the study was obtained from event logs collected over a six-month period from the mail routing network of a commercial organization. The study focuses on characterizing causes of machine reboots. The key observations from this study are: 1) most of the problems that lead to reboots are software related; 2) rebooting the machine does not always solve the problem; 3) there are indications of propagated or correlated failures; and 4) though the average availability evaluates to over 99%, the machine downtime lasts (on average) two hours. Since the machines are dedicated mail servers, bringing down one or more of them can potentially disrupt storage, forwarding, reception and delivery of mail. This suggests that the average availability is not a good measure to characterize this type of network service.
M. Kalyanakrishnam, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SRDS2
1999 A Software Multilevel Fault Injection Mechanism: Case Study Evaluating the Virtual Interface Architecture
abstract
The characteristics of failures occurring in networked computing systems are still poorly understood. As a consequence, this is a rich area for exploration, especially with the arrival of new network interface standards, such as the Virtual Interface Architecture (VIA) adopted by Microsoft, Intel and Compaq. The goal of VIA is to improve the performance of distributed applications by reducing the latency associated with the exchange of critical message between processes in Windows NT-based systems. In this paper, we propose the SMiFI (Software Multilevel Fault Injection) mechanism to evaluate the failure characteristics of networked systems, specifically VIA. The mechanism covers all software protocol layers of the host interface and corrupts both the messages and the computation engines that manipulate the messages.
Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SRDS2
1999 Stress-Based and Path-Based Fault Injection
abstract
The objective of fault injection is to mimic the existence of faults and to force the exercise of the fault tolerance mechanisms of the target system. To maximize the efficacy of each injection, the locations, timing, and conditions for faults being injected must be carefully chosen. Faults should be injected with a high probability of being accessed. This paper presents two fault injection methodologies-stress-based injection and path-based injection; both are based on resource activity analysis to ensure that injections cause fault tolerance activity and, thus, the resulting exercise of fault tolerance mechanisms. The difference between these two methods is that stress-based injection validates the system dependability by monitoring the run-time workload activity at the system level to select faults that coincide with the locations and times of greatest workload activity, while path-based injection validates the system from the application perspective by using an analysis of the program flow and resource usage at the application program level to select faults during the program execution. These two injection methodologies focus separately on the system and process viewpoints to facilitate the testing of system dependability. Details of these two injection methodologies are discussed in this paper, along with their implementations, experimental results, and advantages and disadvantages.
Timothy K. Tsai, Mei-Chen Hsueh, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
IEEE Trans. Computers4
1999 Chameleon: A Software Infrastructure for Adaptive Fault Tolerance
abstract
This paper presents Chameleon, an adaptive infrastructure, which allows different levels of availability requirements to be simultaneously supported in a networked environment. Chameleon provides dependability through the use of special ARMORs-Adaptive. Reconfigurable, and Mobile Objects for Reliability-that control all operations in the Chameleon environment. Three broad classes of ARMORs are defined: 1) Managers oversee other ARMORs and recover from failures in their subordinates. 2) Daemons provide communication gateways to the ARMORs at the host node. They also make available a host's resources to the Chameleon environment. 3) Common ARMORs implement specific techniques for providing application-required dependability. Employing ARMORs, Chameleon makes available different fault-tolerant configurations and maintains run-time adaptation to changes in the availability requirements of an application. Flexible ARMOR architecture allows their composition to be reconfigured at run-time, i.e., the ARMORs may dynamically adapt to changing application requirements. In this paper, we describe ARMOR architecture, including ARMOR class hierarchy, basic building blocks, ARMOR composition, and use of ARMOR factories. We present how ARMORs can be reconfigured and reengineered and demonstrate how the architecture serves our objective of providing an adaptive software infrastructure. To our knowledge, Chameleon is one of the few real implementations which enables multiple fault tolerance strategies to exist in the same environment and supports fault-tolerant execution of substantially off-the-shelf applications via a software infrastructure only. Chameleon provides fault tolerance from the application's point of view as well as from the software infrastructure's point of view. To demonstrate the Chameleon capabilities, we have implemented a prototype infrastructure which provides set of ARMORs to initialize the environment and to support the dual and TMR application execution modes. Through this testbed environment, we measure the execution overhead and recovery times from failures in the user application, the Chameleon ARMORs, the hardware, and the operating system.
Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Saurabh Bagchi, Keith Whisnant
IEEE Trans. Parallel Distributed Syst.1
1999 Hierarchical Simulation Approach to Accurate Fault Modeling for System Dependability Evaluation
abstract
This paper presents a hierarchical simulation methodology that enables accurate system evaluation under realistic faults and conditions. In this methodology, effects of low-level (i.e., transistor or circuit level) faults are propagated to higher levels (i.e., system level) using fault dictionaries. The primary fault models are obtained via simulation of the transistor-level effect of a radiation particle penetrating a device. The resulting current bursts constitute the first-level fault dictionary and are used in the circuit-level simulation to determine the impact on circuit latches and flip-flops. The latched outputs constitute the next level fault dictionary in the hierarchy and are applied in conducting fault injection simulation at the chip-level under selected workloads or application programs. Faults injected at the chip-level result in memory corruptions, which are used to form the next level fault dictionary for the system-level simulation of an application running on simulated hardware. When an application terminates, either normally or abnormally, the overall fault impact on the software behavior is quantified and analyzed. The system in this sense can be a single workstation or a network. The simulation method is demonstrated and validated in the case study of Myrinet (a commercial, high-speed network) based network system.
Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Gregory L. Ries, Jaqdish U. Patel, Myeong S. Lee, Yuxiao Xiao
IEEE Trans. Software Eng.1
1998 The Chameleon Infrastructure for Adaptive, Software Implemented Fault Tolerance
abstract
This paper presents Chameleon, an adaptive software infrastructure for supporting different levels of availability requirements in a heterogeneous networked environment. Chameleon provides dependability through the use of ARMORs-Adaptive, Reconfigurable, and Mobile Objects for Reliability. Three broad classes of ARMORs are defined: Managers, Daemons, and Common ARMORs. Key concepts that support adaptive fault tolerance include the construction of fault tolerance execution strategies from a comprehensive set of ARMORs, the creation of ARMORs from a library of reusable basic building blocks, the dynamic adaptation to changing fault tolerance requirements, and the ability to detect and recover from errors in applications and in ARMORs.
Saurabh Bagchi, Keith Whisnant, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SRDS3
1998 Dependability Analysis of a Cache-Based RAID System via Fast Distributed Simulation
abstract
We propose a new speculation-based, distributed simulation method for dependability analysis of complex systems in which a detailed functional simulation of a system component is essential to obtain an accurate overall result. Our target example is a networked cluster with compute nodes and a single I/O node. Accurate system dependability characterization is achieved via a combination of detailed simulation of the I/O subsystem behavior in the presence of faults and more abstract simulation of the compute nodes and the switching network. Dependability measures like error coverage, error detection latency and performance measures such as delivery time in the presence of faults are obtained. The approach is implemented on a network of workstations, and experimental results show significant improvements over a Time Warp simulator for the same model.
Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SRDS2
1997 Chameleon: Adaptive Fault Tolerance Using Reliable, Mobile Agents
abstract
In networked computing systems, a broad range of commercial and scientific applications that need varying degrees of availability must coexist. It is not cost-effective to develop a reliable platform in each case. It is more efficient to build an infrastructure that provides the required level of dependability for each application's needs. It is also essential that the proposed alternatives should leverage off-the-shelf components. There have been exhaustive studies on fault tolerance strategies capable of providing efficient mechanisms to deal with system operational failures. Most of this work has focused on specific application needs and thus provided only piecemeal solutions. Little work has been done in addressing how to build a reliable networked computing system out of unreliable computation nodes. As a result, there is no comprehensive solution for providing a wide range of fault-tolerant services in a single networked environment. The most feasible way of understanding how such a software environment would fit on top of existing layers (the operating system, the network interfaces, etc.) is to implement an infrastructure for providing a range of reliable services. Fundamental components of the envisioned infrastructure (Chameleon) have been designed so that none of them is a single point of failure. Each of the components is active for a certain period, e.g. during the setting up the system configuration. If a component fails during its active phase, there is a provision for recovery, either by switching to a backup or by regenerating the component.
Ravishankar K. Iyer, Zbigniew T. Kalbarczyk, Saurabh Bagchi
SRDS2
1995 An Attempt to Evaluate Functional Diversity Employed in a Reactor Protection System
Jörgen Christmansson, Zbigniew T. Kalbarczyk, Jan Torin
SAFECOMP2
1994 Dependable flight control system by data diversity and self-checking components
Jörgen Christmansson, Zbigniew T. Kalbarczyk, Jan Torin
Microprocess. Microprogramming2