EDBT 2026 Demo / reviewers in the wild / expert
Ravishankar K. Iyer
dblp:i/RavishankarKIyer · also Ravi K. Iyer, Ravishankar Krishnan Iyer
· DBLP profile ↗
216ranked-venue papers
18as first author
21since 2021 · last 2026
0000-0003-2245-3038ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 122 · 8 first-author · 13 since 2021Security and privacy · 79 · 6 first-author · 5 since 2021Software engineering, systems software and programming languages · 39 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 4 since 2021Artificial intelligence and machine learning · 10 · 3 since 2021Computer networks · 5Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Praxis: Integrating Program Analysis with Observability for Root-Cause AnalysisabstractUnresolved production cloud incidents cost an average of over $2M per hour. This paper introduces PRAXIS, an orchestrator that manages and deploys an agentic workflow for diagnosing code- and configuration-caused cloud incidents. PRAXIS employs an LLM-driven structured traversal over two types of graph: (1) a service dependency graph (SDG) that captures microservice-level dependencies; and (2) a hammock-block program dependence graph (PDG) that captures code-level dependencies for each microservice. Compared to state-of-the-art ReAct baselines, PRAXIS improves RCA accuracy by up to 6.3x while reducing token consumption by 5.3x. PRAXIS is demonstrated on a set of 30 comprehensive real-world incidents that is being compiled into an RCA benchmark. Shengkun Cui, Rahul Krishna, Saurabh Jha, Ravishankar K. Iyer |
DSN | 4 |
| 2026 | Resilient Path Tracking of Autonomous Driving under Few-shot Action Space AttacksabstractModern autonomous vehicles face growing cybersecurity risks, especially from action space attacks that directly target vehicle actuators. This article systematically evaluates the resilience of three representative Autonomous Driving (AD) architectures, including modular, end-to-end, and feature-fused agents, against few-shot action space attacks crafted via deep reinforcement learning under a black-box setting. The adversary perturbs the vehicle’s lateral control only during safety-critical moments, using either a camera or an inertial measurement unit. Our results reveal distinct vulnerabilities and behavioral patterns across AD architectures, which underscore the necessity for adaptive and robust defense strategies. However, existing adversarial training defense methods show limitations of overfitting and reliance on attack knowledge. To address these limitations, we propose a learning-based Path Correction System (PCS) that integrates traditional feedback control with an adversarially trained correction loop. The correction loop is selectively activated by a kinematic model-based attack detector to counteract abnormal control deviations. Evaluation experiments show that PCS reduces path-tracking deviation by 78% when the system is under attack. Xin Lou 0005, Rui Tan 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ACM Trans. Cyber Phys. Syst. | 6 |
| 2025 | Page Migration for Hardware Memory Disaggregation Across a NetworkabstractHardware memory disaggregation (HMD) is an emerging technology that enables access to remote memory, thereby creating expansive memory pools and reducing memory underutilization in datacenters.However, a significant challenge arises when accessing remote memory over a network: increased contention that can lead to severe application performance degradation.To reduce the performance penalty of using remote memory, the operating system uses page migration to promote frequently accessed pages closer to the processor.However, previously proposed page migration mechanisms do not achieve the best performance in HMD systems because of obliviousness to variable page transfer costs that occur due to network contention.To address these limitations, we present INDIGO: a network-aware page migration framework that uses novel page telemetry and a learning-based approach for network adaptation.We implemented INDIGO in the Linux kernel and evaluated it with common cloud and HPC applications on a real disaggregated memory system prototype.Our evaluation shows that INDIGO offers up to 50-70% improvement in application performance compared to other state-of-the-art page migration policies and reduces network traffic up to 2×. Archit Patke, Christian Pinto, Saurabh Jha, Haoran Qiu, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ICS | 6 |
| 2025 | Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUsabstractThis study characterizes GPU resilience in Delta, a large-scale AI system that consists of 1,056 A100 and H100 GPUs, with over 1,300 petaflops of peak throughput. We used 2.5 years of operational data (11.7 million GPU hours) on GPU errors. Our major findings include: (i) H100 GPU memory resilience is worse than A100 GPU memory, with 3.2x lower per-GPU MTBE for memory errors, (ii) The GPU memory error-recovery mechanisms on H100 GPUs are insufficient to handle the increased memory capacity, (iii) H100 GPUs demonstrate significantly improved GPU hardware resilience over A100 GPUs with respect to critical hardware components, (iv) GPU errors on both A100 and H100 GPUs frequently result in job failures due to the lack of robust recovery mechanisms at the application level, and (v) We project the impact of GPU node availability on larger-scales and find that significant overprovisioning of 5% is necessary to handle GPU failures. Shengkun Cui, Archit Patke, Aditya Ranjan, Ziheng Chen 0006, Phuong Cao, Gregory H. Bauer, Brett M. Bode, Catello Di Martino, Saurabh Jha, Chandrasekhar Narayanaswami 0001, Daby M. Sow, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SC | 14 |
| 2024 | Queue Management for SLO-Oriented Large Language Model ServingabstractLarge language model (LLM) serving is becoming an increasingly critical workload for cloud providers. Existing LLM serving systems focus on interactive requests, such as chatbots and coding assistants, with tight latency SLO requirements. However, when such systems execute batch requests that have relaxed SLOs along with interactive requests, it leads to poor multiplexing and inefficient resource utilization. To address these challenges, we propose QLM, a queue management system for LLM serving. QLM maintains batch and interactive requests across different models and SLOs in a request queue. Optimal ordering of the request queue is critical to maintain SLOs while ensuring high resource utilization. To generate this optimal ordering, QLM uses a Request Waiting Time (RWT) Estimator that estimates the waiting times for requests in the request queue. These estimates are used by a global scheduler to orchestrate LLM Serving Operations (LSOs) such as request pulling, request eviction, load balancing, and model swapping. Evaluation on heterogeneous GPU devices and models with real-world LLM serving dataset shows that QLM improves SLO attainment by 40-90% and throughput by 20-400% while maintaining or improving device utilization compared to other state-of-the-art LLM serving systems. QLM's evaluation is based on the production requirements of a cloud provider. QLM is publicly available at https://www.github.com/QLM-project/QLM. Archit Patke, Dhemath Reddy, Saurabh Jha, Haoran Qiu, Christian Pinto, Chandrasekhar Narayanaswami 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SoCC | 8 |
| 2024 | Mutiny! How Does Kubernetes Fail, and What Can We Do About It?abstractIn this paper, we i) analyze and classify real-world failures of Kubernetes (the most popular container orchestration system), ii) develop a framework to perform a fault/error injection campaign targeting the data store preserving the cluster state, and iii) compare results of our fault/error injection experiments with real-world failures, showing that our fault/error injections can recreate many real-world failure patterns. The paper aims to address the lack of studies on systematic analyses of Kubernetes failures to date. Our results show that even a single fault/error (e.g., a bit-flip) in the data stored can propagate, causing cluster-wide failures (3% of injections), service networking issues (4%), and service under/overprovisioning (24%). Errors in the fields tracking dependencies between object caused 51% of such cluster-wide failures. We argue that controlled fault/error injection-based testing should be employed to proactively assess Kubernetes' resiliency and guide the design of failure mitigation strategies. Marco Barletta, Marcello Cinque, Catello Di Martino, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 5 |
| 2024 | iPrism: Characterize and Mitigate Risk by Quantifying Change in Escape Routes
Shengkun Cui, Saurabh Jha, Ziheng Chen 0006, Zbigniew T. Kalbarczvk, Ravishankar K. Iyer |
DSN | 5 |
| 2024 | Power-aware Deep Learning Model Serving with μ-Serve
Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Tamer Basar, Ravishankar K. Iyer |
USENIX ATC | 10 |
| 2024 | True Attacks, Attack Attempts, or Benign Triggers? An Empirical Measurement of Network Alerts in a Security Operations Center
Zhi Chen 0028, Chenkai Wang 0001, Zhenning Zhang, Sushruth Booma, Phuong Cao, Constantin Adam, Alexander Withers, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Gang Wang 0011 |
USENIX Security Symposium | 10 |
| 2023 | Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task SimilarityabstractMulti-agent reinforcement learning (MARL) has primarily focused on solving a single task in isolation, while in practice the environment is often evolving, leaving many related tasks to be solved. In this paper, we investigate the benefits of meta-learning in solving multiple MARL tasks collectively. We establish the first line of theoretical results for meta-learning in a wide range of fundamental MARL settings, including learning Nash equilibria in two-player zero-sum Markov games and Markov potential games, as well as learning coarse correlated equilibria in general-sum Markov games. Under natural notions of task similarity, we show that meta-learning achieves provable sharper convergence to various game-theoretical solution concepts than learning each task separately. As an important intermediate step, we develop multiple MARL algorithms with initialization-dependent convergence guarantees. Such algorithms integrate optimistic policy mirror descents with stage-based value updates, and their refined convergence guarantees (nearly) recover the best known results even when a good initialization is unknown. To our best knowledge, such results are also new and might be of independent interest. We further provide numerical simulations to corroborate our theoretical findings. Weichao Mao, Haoran Qiu, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Tamer Basar |
NeurIPS | 6 |
| 2023 | AWARE: Automate Workload Autoscaling with Reinforcement Learning in Production Cloud Systems
Haoran Qiu, Weichao Mao, Chen Wang 0039, Hubertus Franke, Alaa Youssef, Zbigniew T. Kalbarczyk, Tamer Basar, Ravishankar K. Iyer |
USENIX ATC | 8 |
| 2022 | Individualized seizure cluster prediction using machine learning and ambulatory intracranial EEGabstractSeizure clusters, i.e., seizures that occur within a short duration of each other, occur in several epilepsy patients and are associated with increased disease severity. Understanding the characteristics of seizure clusters and predicting whether a given seizure will cluster or not is valuable both from a patient’s and clinician’s perspective. We propose a novel methodology for studying seizure clusters based on bivariate intracranial EEG (iEEG) features and develop one of the first individualized seizure cluster prediction models by combining machine learning with relative entropy (a bivariate feature). Relative entropy was used to quantify interactions between brain regions and capture potential differences in interactions underlying isolated and cluster seizures. We evaluated our methodology using one of the largest ambulatory iEEG datasets, consisting of data from 15 patients with up to 2 years of recordings each. This provided us a sufficient number of seizures in each patient to enable individualized analyses and prediction. On data of 3710 seizures consisting of 3341 cluster seizures (from 427 clusters) and 369 isolated seizures, machine learning models based on relative entropy predicted seizure clusters with up to 73.6% F1-score and outperformed baseline predictors. Our results are beneficial in addressing the clinical burden of clusters. Krishnakant V. Saboo, Yurui Cao, Václav Kremen, Vladimir Sladky, Nicholas M. Gregg, Paul M. Arnold, Philippa J. Karoly, Dean R. Freestone, Mark J. Cook, Gregory A. Worrell, Ravishankar K. Iyer |
BIBM | 11 |
| 2022 | SIMPPO: a scalable and incremental online learning framework for serverless resource managementabstractServerless Function-as-a-Service (FaaS) offers improved programmability for customers, yet it is not server-"less" and comes at the cost of more complex infrastructure management (e.g., resource provisioning and scheduling) for cloud providers. To maintain service-level objectives (SLOs) and improve resource utilization efficiency, recent research has been focused on applying online learning algorithms such as reinforcement learning (RL) to manage resources. Despite the initial success of applying RL, we first show in this paper that the state-of-the-art single-agent RL algorithm (S-RL) suffers up to 4.8x higher p99 function latency degradation on multi-tenant serverless FaaS platforms compared to isolated environments and is unable to converge during training. We then design and implement a scalable and incremental multi-agent RL framework based on Proximal Policy Optimization (SIMPPO). Our experiments demonstrate that in multi-tenant environments, SIMPPO enables each RL agent to efficiently converge during training and provides online function latency performance comparable to that of S-RL trained in isolation with minor degradation (<9.2%). In addition, SIMPPO reduces the p99 function latency by 4.5x compared to S-RL in multi-tenant cases. Haoran Qiu, Weichao Mao, Archit Patke, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Tamer Basar, Ravishankar K. Iyer |
SoCC | 8 |
| 2022 | Exploiting Temporal Data Diversity for Detecting Safety-critical Faults in AV Compute SystemsabstractSilent data corruption caused by random hardware faults in autonomous vehicle (AV) computational elements is a significant threat to vehicle safety. Previous research has explored design diversity, data diversity, and duplication techniques to detect such faults in other safety-critical domains. However, these are challenging to use for AVs in practice due to significant resource overhead and design complexity. We propose, DiverseAV, a low-cost data-diversity-based redundancy technique for detecting safety-critical random hardware faults in computational elements. DiverseAV introduces data-diversity between the redundant agents by exploiting the temporal semantic consistency available in the AV sensor data. DiverseAV is a black-box technique that offers a plug-and-play solution as it requires no knowledge of the internals of the AI agent responsible for executing driving decisions, requiring little to no modification to the agent itself for achieving high coverage of transient and permanent hardware faults. It is commercially viable because it avoids software modifications to agents that are costly in terms of development and testing time. Specifically, DiverseAV distributes the sensor data between the two software agents in a round-robin manner. As a result, the sensor data for two consecutive time steps are semantically similar in terms of their worldview but significantly different at the bit level, thus ensuring the state and data diversity between the two agents necessary for detecting faults. We demonstrate DiverseAV using an open-source self-driving AI agent which is controlling a car in an open-source world simulator. Saurabh Jha, Shengkun Cui, Timothy Tsai 0002, Siva Kumar Sastry Hari, Michael B. Sullivan 0001, Zbigniew T. Kalbarczyk, Stephen W. Keckler, Ravishankar K. Iyer |
DSN | 8 |
| 2022 | A Mean-Field Game Approach to Cloud Resource Management with Function ApproximationabstractReinforcement learning (RL) has gained increasing popularity for resource management in cloud services such as serverless computing. As self-interested users compete for shared resources in a cluster, the multi-tenancy nature of serverless platforms necessitates multi-agent reinforcement learning (MARL) solutions, which often suffer from severe scalability issues. In this paper, we propose a mean-field game (MFG) approach to cloud resource management that is scalable to a large number of users and applications and incorporates function approximation to deal with the large state-action spaces in real-world serverless platforms. Specifically, we present an online natural actor-critic algorithm for learning in MFGs compatible with various forms of function approximation. We theoretically establish its finite-time convergence to the regularized Nash equilibrium under linear function approximation and softmax parameterization. We further implement our algorithm using both linear and neural-network function approximations, and evaluate our solution on an open-source serverless platform, OpenWhisk, with real-world workloads from production traces. Experimental results demonstrate that our approach is scalable to a large number of users and significantly outperforms various baselines in terms of function latency and resource utilization efficiency. Weichao Mao, Haoran Qiu, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Tamer Basar |
NeurIPS | 6 |
| 2022 | Telogator: a method for reporting chromosome-specific telomere lengths from long readsabstractMOTIVATION: Telomeres are the repetitive sequences found at the ends of eukaryotic chromosomes and are often thought of as a 'biological clock,' with their average length shortening during division in most cells. In addition to their association with senescence, abnormal telomere lengths are well known to be associated with multiple cancers, short telomere syndromes and as risk factors for a broad range of diseases. While a majority of methods for measuring telomere length will report average lengths across all chromosomes, it is known that aberrations in specific chromosome arms are biomarkers for certain diseases. Due to their repetitive nature, characterizing telomeres at this resolution is prohibitive for short read sequencing approaches, and is challenging still even with longer reads. RESULTS: We present Telogator: a method for reporting chromosome-specific telomere length from long read sequencing data. We demonstrate Telogator's sensitivity in detecting chromosome-specific telomere length in simulated data across a range of read lengths and error rates. Telogator is then applied to 10 germline samples, yielding a high correlation with short read methods in reporting average telomere length. In addition, we investigate common subtelomere rearrangements and identify the minimum read length required to anchor telomere/subtelomere boundaries in samples with these haplotypes. AVAILABILITY AND IMPLEMENTATION: Telogator is written in Python3 and is available at github.com/zstephens/telogator. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zachary Stephens, Alejandro Ferrer, Lisa Boardman, Ravishankar K. Iyer, Jean-Pierre A. Kocher |
Bioinform. | 4 |
| 2022 | Data-Driven Application-Oriented Reliability Model of a High-Performance Computing SystemabstractReliability analysis and performance evaluation are complementary methods to quantify nonfunctional aspects of a system. However, a range of factors such as concurrency and heterogeneity quickly exacerbate the state-space explosion problem when attempting detailed system-level modeling and simulation of high-performance computing (HPC) systems. To overcome these impediments to modeling and analysis, this article develops a hierarchical model of an application that implements checkpointing running in an HPC environment subject to application, network, and system-wide outages. The modeling approach ensures that the number of states is linear in the number of checkpoints and possesses a low constant factor for the number of recovery states most relevant to the external influences contributing to degraded application performance. We illustrate the types of analysis enabled by the model through a series of examples with parameters determined empirically from data logs of the Blue Waters supercomputer located at the University of Illinois at Urbana–Champaign. A comprehensive comparative analysis of the model parameters indicates that lowering the failure rate of network nodes would most significantly reduce application downtime. We also discuss how the modeling approach can be used to objectively assess both current and hypothetical future systems to identify competitive designs and enhancements. Bentolhoda Jafary, Saurabh Jha, Lance Fiondella, Ravishankar K. Iyer |
IEEE Trans. Reliab. | 4 |
| 2021 | BayesPerf: minimizing performance monitoring errors using Bayesian statisticsabstractHardware performance counters (HPCs) that measure low-level architectural and microarchitectural events provide dynamic contextual information about the state of the system. However, HPC measurements are error-prone due to non determinism (e.g., undercounting due to event multiplexing, or OS interrupt-handling behaviors). In this paper, we present BayesPerf, a system for quantifying uncertainty in HPC measurements by using a domain-driven Bayesian model that captures microarchitectural relationships between HPCs to jointly infer their values as probability distributions. We provide the design and implementation of an accelerator that allows for low-latency and low-power inference of the BayesPerf model for x86 and ppc64 CPUs. BayesPerf reduces the average error in HPC measurements from 40.1% to 7.6% when events are being multiplexed. The value of BayesPerf in real-time decision-making is illustrated with a simple example of scheduling of PCIe transfers. Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ASPLOS | 4 |
| 2021 | Improved GPU Implementations of the Pair-HMM Forward Algorithm for DNA Sequence AlignmentabstractWith the rise of Next-Generation Sequencing (NGS) technology, clinical sequencing services become more accessible but are also facing new challenges. The surging demand motivates developments of more efficient algorithms for computational genomics and their hardware acceleration. In this work, we use GPU to accelerate the DNA variant calling and its related alignment problem. The Pair-Hidden Markov Model (Pair-HMM) is one of the most popular and compute-intensive models used in variant calling. As a critical part of the Pair-HMM, the forward algorithm is not only a computational but data-intensive algorithm. Multiple previous works have been done in efforts to accelerate the computation of the forward algorithm by the massive parallelization of the workload. In this paper, we bring advanced GPU implementations with various optimizations, such as efficient host-device communication, task parallelization, pipelining, and memory management, to tackle this challenging task. Our design has shown a speedup of 783X comparing to the Java baseline on Intel single-core CPU, 31.88X to the C++ baseline on IBM Power8 multicore CPU, and 1.53X - 2.21X to the previous state-of-the-art GPU implementations over various genomics datasets. Enliang Li, Subho S. Banerjee, Sitao Huang, Ravishankar K. Iyer, Deming Chen |
ICCD | 4 |
| 2021 | Delay sensitivity-driven congestion mitigation for HPC systemsabstractModern high-performance computing (HPC) systems concurrently execute multiple distributed applications that contend for the high-speed network leading to congestion. Consequently, application runtime variability and suboptimal system utilization are observed in production systems. To address these problems, we propose Netscope, a congestion mitigation framework based on a novel delay sensitivity metric. Delay sensitivity of an application is used to quantify the impact of congestion on its runtime. Netscope uses delay sensitivity estimates to drive a congestion mitigation mechanism to selectively throttle applications that are less susceptible to congestion. We evaluate Netscope on two Cray Aries systems, including a production supercomputer, on common scientific applications. Our evaluation shows that Netscope has a low training cost and accurately estimates the impact of congestion on application runtime with a correlation between 0.7 and 0.9. Moreover, Netscope reduces application tail runtime increase by up to 16.3x while improving the median system utility by 12%. Archit Patke, Saurabh Jha, Haoran Qiu, Jim M. Brandt, Ann C. Gentile, Joe Greenseid, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ICS | 8 |
| 2021 | Reinforcement Learning based Disease Progression Model for Alzheimer's DiseaseabstractWe model Alzheimer’s disease (AD) progression by combining differential equations (DEs) and reinforcement learning (RL) with domain knowledge. DEs provide relationships between some, but not all, factors relevant to AD. We assume that the missing relationships must satisfy general criteria about the working of the brain, for e.g., maximizing cognition while minimizing the cost of supporting cognition. This allows us to extract the missing relationships by using RL to optimize an objective (reward) function that captures the above criteria. We use our model consisting of DEs (as a simulator) and the trained RL agent to predict individualized 10-year AD progression using baseline (year 0) features on synthetic and real data. The model was comparable or better at predicting 10-year cognition trajectories than state-of-the-art learning-based models. Our interpretable model demonstrated, and provided insights into, "recovery/compensatory" processes that mitigate the effect of AD, even though those processes were not explicitly encoded in the model. Our framework combines DEs with RL for modelling AD progression and has broad applicability for understanding other neurological disorders. Krishnakant V. Saboo, Anirudh Choudhary, Yurui Cao, Gregory A. Worrell, David T. Jones, Ravishankar K. Iyer |
NeurIPS | 6 |
| 2020 | ML-Driven Malware that Targets AV SafetyabstractEnsuring the safety of autonomous vehicles (AVs) is critical for their mass deployment and public adoption. However, security attacks that violate safety constraints and cause accidents are a significant deterrent to achieving public trust in AVs, and that hinders a vendor's ability to deploy AVs. Creating a security hazard that results in a severe safety compromise (for example, an accident) is compelling from an attacker's perspective. In this paper, we introduce an attack model, a method to deploy the attack in the form of smart malware, and an experimental evaluation of its impact on production-grade autonomous driving software. We find that determining the time interval during which to launch the attack is{ critically} important for causing safety hazards (such as collisions) with a high degree of success. For example, the smart malware caused 33X more forced emergency braking than random attacks did, and accidents in 52.6% of the driving simulations. Saurabh Jha, Shengkun Cui, Subho S. Banerjee, James Cyriac, Timothy Tsai 0002, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 7 |
| 2020 | Inductive-bias-driven Reinforcement Learning For Efficient Schedules in Heterogeneous ClustersabstractThe problem of scheduling of workloads onto heterogeneous processors (e.g., CPUs, GPUs, FPGAs) is of fundamental importance in modern data centers. Current system schedulers rely on application/system-specific heuristics that have to be built on a case-by-case basis. Recent work has demonstrated ML techniques for automating the heuristic search by using black-box approaches which require significant training data and time, which make them challenging to use in practice. This paper presents Symphony, a scheduling framework that addresses the challenge in two ways: (i) a domain-driven Bayesian reinforcement learning (RL) model for scheduling, which inherently models the resource dependencies identified from the system architecture; and (ii) a sampling-based technique to compute the gradients of a Bayesian model without performing full probabilistic inference. Together, these techniques reduce both the amount of training data and the time required to produce scheduling policies that significantly outperform black-box approaches by up to 2.2{\texttimes}. Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ICML | 4 |
| 2020 | AV-FUZZER: Finding Safety Violations in Autonomous Driving SystemsabstractThis paper proposes AV-FUZZER, a testing framework, to find the safety violations of an autonomous vehicle (AV) in the presence of an evolving traffic environment. We perturb the driving maneuvers of traffic participants to create situations in which an AV can run into safety violations. To optimally search for the perturbations to be introduced, we leverage domain knowledge of vehicle dynamics and genetic algorithm to minimize the safety potential of an AV over its projected trajectory. The values of the perturbation determined by this process provide parameters that define participants' trajectories. To improve the efficiency of the search, we design a local fuzzer that increases the exploitation of local optima in the areas where highly likely safety-hazardous situations are observed. By repeating the optimization with significantly different starting points in the search space, AV-FUZZER determines several diverse AV safety violations. We demonstrate AV-FUZZER on an industrial-grade AV platform, Baidu Apollo, and find five distinct types of safety violations in a short period of time. In comparison, other existing techniques can find at most two. We analyze the safety violations found in Apollo and discuss their overarching causes. Guanpeng Li, Saurabh Jha, Timothy Tsai 0002, Michael B. Sullivan 0001, Siva Kumar Sastry Hari, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ISSRE | 8 |
| 2020 | Measuring Congestion in High-Performance Datacenter Interconnects
Saurabh Jha, Archit Patke, Jim M. Brandt, Ann C. Gentile, Benjamin Lim, Michael T. Showerman, Gregory H. Bauer, Larry Kaplan, Zbigniew T. Kalbarczyk, William T. Kramer, Ravishankar K. Iyer |
NSDI | 11 |
| 2020 | FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented Microservices
Haoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
OSDI | 5 |
| 2020 | Live forensics for HPC systems: a case study on distributed storage systemsabstractLarge-scale high-performance computing systems frequently experience a wide range of failure modes, such as reliability failures (e.g., hang or crash), and resource overload-related failures (e.g., congestion collapse), impacting systems and applications. Despite the adverse effects of these failures, current systems do not provide methodologies for proactively detecting, localizing, and diagnosing failures. We present Kaleidoscope, a near real-time failure detection and diagnosis framework, consisting of of hierarchical domain-guided machine learning models that identify the failing components, the corresponding failure mode, and point to the most likely cause indicative of the failure in near real-time (within one minute of failure occurrence). Kaleidoscope has been deployed on Blue Waters supercomputer and evaluated with more than two years of production telemetry data. Our evaluation shows that Kaleidoscope successfully localized 99.3% and pinpointed the root causes of 95.8% of 843 real-world production issues, with less than 0.01% runtime overhead. Saurabh Jha, Shengkun Cui, Subho S. Banerjee, Tianyin Xu, Jeremy Enos, Michael T. Showerman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SC | 8 |
| 2019 | AcMC 2 : Accelerating Markov Chain Monte Carlo Algorithms for Probabilistic ModelsabstractProbabilistic models (PMs) are ubiquitously used across a variety of machine learning applications. They have been shown to successfully integrate structural prior information about data and effectively quantify uncertainty to enable the development of more powerful, interpretable, and efficient learning algorithms. This paper presents AcMC2, a compiler that transforms PMs into optimized hardware accelerators (for use in FPGAs or ASICs) that utilize Markov chain Monte Carlo methods to infer and query a distribution of posterior samples from the model. The compiler analyzes statistical dependencies in the PM to drive several optimizations to maximally exploit the parallelism and data locality available in the problem. We demonstrate the use of AcMC2 to implement several learning and inference tasks on a Xilinx Virtex-7 FPGA. AcMC2-generated accelerators provide a 47-100× improvement in runtime performance over a 6-core IBM Power8 CPU and a 8-18× improvement over an NVIDIA K80 GPU. This corresponds to a 753-1600× improvement over the CPU and 248-463× over the GPU in performance-per-watt terms. Subho S. Banerjee, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ASPLOS | 3 |
| 2019 | A Joint Model for Predicting Structural and Functional Brain Health in Elderly IndividualsabstractThis paper presents a machine-learning-based joint model of brain age and cognitive performance, and demonstrates its superior performance relative to isolated models. Previous studies have chosen to study those two measures of brain health separately for two reasons: 1) although cognition can be measured regardless of an individual's health, brain-age ground-truth can be defined only for healthy individuals; and 2) while brain-age models are developed using neuroimaging data alone, modeling of cognitive performance additionally requires measures of cognitive reserve and biomarkers of cognitive disorders. However, those two measures are biologically related to each other, because they both depend on brain structure. Hence, we developed a joint model by 1) explicitly defining the commonalities and differences between them in a graph, and 2) converting that graph into a multitask-learning model to facilitate learning from population-level data. Our model took as inputs structural neuroimaging data and information related to cognitive reserve and disorders, and predicted brain age and cognitive performance in terms of a Mini-Mental State Examination (MMSE) score. We implemented linear and nonlinear joint models and compared them against isolated models. Our results indicate that joint modeling substantially improves the accuracy of the modeling of individual measures, relative to isolated models. Yogatheesan Varatharajah, Krishnakant V. Saboo, Ravishankar K. Iyer, Scott A. Przybelski, Christopher G. Schwarz, Ronald C. Petersen, Clifford R. Jack Jr., Prashanthi Vemuri |
BIBM | 3 |
| 2019 | ML-Based Fault Injection for Autonomous Vehicles: A Case for Bayesian Fault InjectionabstractThe safety and resilience of fully autonomous vehicles (AVs) are of significant concern, as exemplified by several headline-making accidents. While AV development today involves verification, validation, and testing, end-to-end assessment of AV systems under accidental faults in realistic driving scenarios has been largely unexplored. This paper presents DriveFI, a machine learning-based fault injection engine, which can mine situations and faults that maximally impact AV safety, as demonstrated on two industry-grade AV technology stacks (from NVIDIA and Baidu). For example, DriveFI found 561 safety-critical faults in less than 4 hours. In comparison, random injection experiments executed over several weeks could not find any safety-critical faults. Saurabh Jha, Subho S. Banerjee, Timothy Tsai 0002, Siva Kumar Sastry Hari, Michael B. Sullivan 0001, Zbigniew T. Kalbarczyk, Stephen W. Keckler, Ravishankar K. Iyer |
DSN | 8 |
| 2019 | CAUDIT: Continuous Auditing of SSH Servers To Mitigate Brute-Force Attacks
Phuong Cao, Yuming Wu, Subho S. Banerjee, Justin Azoff, Alexander Withers, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
NSDI | 7 |
| 2019 | Smart Malware that Uses Leaked Control Data of Robotic Applications: The Case of Raven-II Surgical Robots
Key-whan Chung, Peicheng Tang, Zeran Zhu, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Thenkurussi Kesavadas |
RAID | 6 |
| 2019 | Recommendations for performance optimizations when using GATK3.8 and GATK4abstractBACKGROUND: Use of the Genome Analysis Toolkit (GATK) continues to be the standard practice in genomic variant calling in both research and the clinic. Recently the toolkit has been rapidly evolving. Significant computational performance improvements have been introduced in GATK3.8 through collaboration with Intel in 2017. The first release of GATK4 in early 2018 revealed rewrites in the code base, as the stepping stone toward a Spark implementation. As the software continues to be a moving target for optimal deployment in highly productive environments, we present a detailed analysis of these improvements, to help the community stay abreast with changes in performance. RESULTS: We re-evaluated multiple options, such as threading, parallel garbage collection, I/O options and data-level parallelization. Additionally, we considered the trade-offs of using GATK3.8 and GATK4. We found optimized parameter values that reduce the time of executing the best practices variant calling procedure by 29.3% for GATK3.8 and 16.9% for GATK4. Further speedups can be accomplished by splitting data for parallel analysis, resulting in run time of only a few hours on whole human genome sequenced to the depth of 20X, for both versions of GATK. Nonetheless, GATK4 is already much more cost-effective than GATK3.8. Thanks to significant rewrites of the algorithms, the same analysis can be run largely in a single-threaded fashion, allowing users to process multiple samples on the same CPU. CONCLUSIONS: In time-sensitive situations, when a patient has a critical or rapidly developing condition, it is useful to minimize the time to process a single sample. In such cases we recommend using GATK3.8 by splitting the sample into chunks and computing across multiple nodes. The resultant walltime will be nnn.4 hours at the cost of $41.60 on 4 c5.18xlarge instances of Amazon Cloud. For cost-effectiveness of routine analyses or for large population studies, it is useful to maximize the number of samples processed per unit time. Thus we recommend GATK4, running multiple samples on one node. The total walltime will be ∼34.1 hours on 40 samples, with 1.18 samples processed per hour at the cost of $2.60 per sample on c5.18xlarge instance of Amazon Cloud. Jacob R Heldenbrand, Saurabh Baheti, Matthew A. Bockol, Travis M. Drucker, Steven N. Hart, Matthew E. Hudson, Ravishankar K. Iyer, Michael Kalmbach, Katherine Irene Kendig, Eric W. Klee, Nathan R Mattson, Eric D. Wieben, Mathieu Wiepert, Derek E. Wildman, Liudmila S. Mainzer |
BMC Bioinform. | 7 |
| 2019 | Correction to: Recommendations for performance optimizations when using GATK3.8 and GATK4abstractFollowing publication of the original article [1], the author explained that Table 2 is displayed incorrectly. The correct Table 2 is given below. The original article has been corrected. Jacob R Heldenbrand, Saurabh Baheti, Matthew A. Bockol, Travis M. Drucker, Steven N. Hart, Matthew E. Hudson, Ravishankar K. Iyer, Michael Kalmbach, Katherine Irene Kendig, Eric W. Klee, Nathan R Mattson, Eric D. Wieben, Mathieu Wiepert, Derek E. Wildman, Liudmila S. Mainzer |
BMC Bioinform. | 7 |
| 2019 | ASAP: Accelerated Short-Read Alignment on Programmable HardwareabstractThe proliferation of high-throughput sequencing machines ensures rapid generation of up to billions of short nucleotide fragments in a short period of time. This massive amount of sequence data can quickly overwhelm today's storage and compute infrastructure. This paper explores the use of hardware acceleration to significantly improve the runtime of short-read alignment, a crucial step in preprocessing sequenced genomes. We focus on the Levenshtein distance (edit-distance) computation kernel and propose the ASAP accelerator, which utilizes the intrinsic delay of circuits for edit-distance computation elements as a proxy for computation. Our design is implemented on an Xilinx Virtex 7 FPGA in an IBM POWER8 system that uses the CAPI interface for cache coherence across the CPU and FPGA. Our design is$200\times$faster than an equivalent Smith-Waterman-C implementation of the kernel running on the host processor,$40-60\times$faster than an equivalent Landau-Vishkin-C++ implementation of the kernel running on the IBM Power8 host processor, and$2\times$faster for an end-to-end alignment tool for 120–150 base-pair short-read sequences. Further the design represents a$3760\times$improvement over the CPU in performance/Watt terms. Subho S. Banerjee, Mohamed El-Hadedy 0001, Jong Bin Lim, Zbigniew T. Kalbarczyk, Deming Chen, Steven S. Lumetta, Ravishankar K. Iyer |
IEEE Trans. Computers | 7 |
| 2018 | A ML-based Runtime System for Executing Dataflow Graphs on Heterogeneous ProcessorsabstractNo abstract available. Subho S. Banerjee, Arjun P. Athreya, Zbigniew T. Kalbarczyk, Steven S. Lumetta, Ravishankar K. Iyer |
SoCC | 5 |
| 2018 | Characterizing Supercomputer Traffic Networks Through Link-Level AnalysisabstractWe present techniques for characterizing bandwidth and congestion characteristics of supercomputer High-Speed Networks (HSN). By utilizing a link-level perspective, we gain generality over analyses which are tied to specific topologies. We illustrate these techniques using five months of a Blue Waters production dataset consisting of network utilization and congestion counters. We find that: i) execution time of the communication-heavy applications is highly correlated to network stalls observed in the network topology and increase in application runtime can be as high as 1.7x with nominal increase in stalls, ii) heterogeneity in the available link bandwidth in the network can lead to backpressure and congestion even when the network is not underprovisioned, and (iii) links connected to I/O nodes are no more likely to observe congestion during operational hours than any other link in the system. Saurabh Jha, Jim M. Brandt, Ann C. Gentile, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
CLUSTER | 5 |
| 2018 | Hands Off the Wheel in Autonomous Vehicles?: A Systems Perspective on over a Million Miles of Field DataabstractAutonomous vehicle (AV) technology is rapidly becoming a reality on U.S. roads, offering the promise of improvements in traffic management, safety, and the comfort and efficiency of vehicular travel. The California Department of Motor Vehicles (DMV) reports that between 2014 and 2017, manufacturers tested 144 AVs, driving a cumulative 1,116,605 autonomous miles, and reported 5,328 disengagements and 42 accidents involving AVs on public roads. This paper investigates the causes, dynamics, and impacts of such AV failures by analyzing disengagement and accident reports obtained from public DMV databases. We draw several conclusions. For example, we find that autonomous vehicles are 15 - 4000× worse than human drivers for accidents per cumulative mile driven; that drivers of AVs need to be as alert as drivers of non-AVs; and that the AVs' machine-learning-based systems for perception and decision-and-control are the primary cause of 64% of all disengagements. Subho S. Banerjee, Saurabh Jha, James Cyriac, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 5 |
| 2018 | Detection and visualization of complex structural variants from long readsabstractBACKGROUND: With applications in cancer, drug metabolism, and disease etiology, understanding structural variation in the human genome is critical in advancing the thrusts of individualized medicine. However, structural variants (SVs) remain challenging to detect with high sensitivity using short read sequencing technologies. This problem is exacerbated when considering complex SVs comprised of multiple overlapping or nested rearrangements. Longer reads, such as those from Pacific Biosciences platforms, often span multiple breakpoints of such events, and thus provide a way to unravel small-scale complexities in SVs with higher confidence. RESULTS: We present CORGi (COmplex Rearrangement detection with Graph-search), a method for the detection and visualization of complex local genomic rearrangements. This method leverages the ability of long reads to span multiple breakpoints to untangle SVs that appear very complicated with respect to a reference genome. We validated our approach against both simulated long reads, and real data from two long read sequencing technologies. We demonstrate the ability of our method to identify breakpoints inserted in synthetic data with high accuracy, and the ability to detect and plot SVs from NA12878 germline, achieving 88.4% concordance between the two sets of sequence data. The patterns of complexity we find in many NA12878 SVs match known mechanisms associated with DNA replication and structural variant formation, and highlight the ability of our method to automatically label complex SVs with an intuitive combination of adjacent or overlapping reference transformations. CONCLUSIONS: CORGi is a method for interrogating genomic regions suspected to contain local rearrangements using long reads. Using pairwise alignments and graph search CORGi produces labels and visualizations for local SVs of arbitrary complexity. Zachary Stephens, Chen Wang 0001, Ravishankar K. Iyer, Jean-Pierre A. Kocher |
BMC Bioinform. | 3 |
| 2018 | Resiliency of HPC Interconnects: A Case Study of Interconnect Failures and Recovery in Blue WatersabstractAvailability of the interconnection network in high-performance computing (HPC) systems is fundamental to sustaining the continuous execution of applications at scale. When failures occur, interconnect recovery mechanisms orchestrate complex operations to recover network connectivity between the nodes. As the scale and design complexity of HPC systems increase, so does the system's susceptibility to failures during execution of interconnect-recovery procedures. This study characterizes the recovery procedures of the Gemini interconnect network, the largest Gemini network built by Cray, on Blue Waters, a 13.3 petaflop supercomputer at the National Center for Supercomputing Applications (NCSA). We propose a propagation model that captures interconnect failures and recovery procedures to help understand types of failures and their propagation in both the system and applications during recovery. The measurements show that recovery procedures occur very frequently and that the unsuccessful execution of recovery procedures, when additional failures occur during recovery, causes system-wide outages (SWOs, 28 out of 101) and application failures (3.4 percent of all running applications). Saurabh Jha, Valerio Formicola, Catello Di Martino, Mark Dalton, William T. Kramer, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2017 | Unraveling complex local genomic rearrangements from long-read dataabstractIn this paper, we present a graph search approach for identifying arbitrarily complex structural genomic variation. Our method leverages the ability of long reads (e.g. from Pacific Biosciences platforms) to span multiple breakpoints of complicated local rearrangements, allowing us to resolve small-scale complexities that may be overlooked by other tools. We applied our method to a subset of NA12878 germline events using two long read datasets and demonstrate, with a concordance rate of 88.4% between the two sets, an increased ability to denote complex events over baseline calls from short read data. In a majority of the regions analyzed we detected small complexities that flank the breakpoints of larger events, including small insertions, inversions, and duplicated sequences. These patterns of complexity match known mechanisms associated with DNA replication and structural variant formation, and showcase the ability of our approach to efficiently unravel such events. Our method automatically classifies complex structural variant calls as a combination of nested or adjacent reference transformations, allowing users to identify specific structure types of interest. Additionally, an output report is generated for each event with interactive visual representations of the rearrangement. Zachary Stephens, Ravishankar K. Iyer, Chen Wang 0001, Jean-Pierre A. Kocher |
BIBM | 2 |
| 2017 | Data-driven longitudinal modeling and prediction of symptom dynamics in major depressive disorder: Integrating factor graphs and learning methodsabstractThis paper proposes a data-driven longitudinal model that brings together factor graphs and learning methods to demonstrate a significant improvement in predictability in clinical outcomes of patients with major depressive disorder treated with antidepressants. Using data from the Mayo Clinic PGRN-AMPS trial and the STAR*D trial for validation, this work makes two significant contributions in the context of predictability in psychiatric therapeutic outcomes. First, we establish symptom dynamics in response to antidepressants by using the forward algorithm on a factor graph. Symptom dynamics are the changes in the symptom severity that are most likely to occur because of the antidepressants taken during the trial, and the associated clinical outcomes at 4 weeks and 8 weeks into the trial. The structure of the factor graph is inferred by using unsupervised learning to stratify patients by the similarity of their overall symptom severity. Second, by using metabolomics data as an accurate biological measure in addition to symptom survey data and other patient history information, the prediction of clinical outcomes such as response and remission significantly improved from 30% to 68% in men, and from 35% to 72% in women. This work demonstrates a significant difference in how men and women respond to antidepressants in terms of their symptom dynamics, and also shows that top predictors of clinical outcomes for men and women are significantly different and known to play a role in behavioral sciences. Arjun P. Athreya, Subho S. Banerjee, Drew Neavin, Rima Kaddurah-Daouk, A. John Rush, Mark A. Frye, Liewei Wang, Richard M. Weinshilboum, William V. Bobo, Ravishankar K. Iyer |
CIBCB | 10 |
| 2017 | Holistic Measurement-Driven System AssessmentabstractIn high-performance computing systems, application performance and throughput are dependent on a complex interplay of hardware and software subsystems and variable workloads with competing resource demands. Data-driven insights into the potentially widespread scope and propagationof impact of events, such as faults and contention for shared resources, can be used to drive more effective use of resources, for improved root cause diagnosis, and for predicting performance impacts. We present work developing integrated capabilities for holistic monitoring and analysis to understand and characterize propagation of performance-degrading events. These characterizations can be used to determine and invoke mitigating responses by system administrators, applications, and system software. Saurabh Jha, Jim M. Brandt, Ann C. Gentile, Zbigniew T. Kalbarczyk, Gregory H. Bauer, Jeremy Enos, Michael T. Showerman, Larry Kaplan, Brett M. Bode, Annette Greiner, Amanda Bonnie, Mike Mason, Ravishankar K. Iyer, William T. Kramer |
CLUSTER | 13 |
| 2017 | ASAP: Accelerated Short Read Alignment on Programmable Hardware (Abstract Only)
Subho S. Banerjee, Mohamed El-Hadedy 0001, Jong Bin Lim, Daniel Chen 0001, Zbigniew T. Kalbarczyk, Deming Chen, Ravishankar K. Iyer |
FPGA | 7 |
| 2017 | On accelerating pair-HMM computations in programmable hardwareabstractThis paper explores hardware acceleration to significantly improve the runtime of computing the forward algorithm on Pair-HMM models, a crucial step in analyzing mutations in sequenced genomes. We describe 1) the design and evaluation of a novel accelerator architecture that can efficiently process real sequence data without performing wasteful work; and 2) aggressive memoization techniques that can significantly reduce the number of invocations of, and the amount of data transferred to the accelerator. We describe our demonstration of the design on a Xilinx Virtex 7 FPGA in an IBM Power8 system. Our design achieves a 14.85× higher throughput than an 8-core CPU baseline (that uses SIMD and multi-threading) and a 147.49 × improvement in throughput per unit of energy expended on the NA12878 sample. Subho S. Banerjee, Mohamed El-Hadedy 0001, Ching Y. Tan, Zbigniew T. Kalbarczyk, Steven S. Lumetta, Ravishankar K. Iyer |
FPL | 6 |
| 2017 | Trustworthy Services Built on Event-Based Probing for Layered DefenseabstractNumerous event-based probing methods exist for cloud computing environments allowing a hypervisor to gain insight into guest activities. Such event-based probing has been shown to be useful for detecting attacks, system hangs through watchdogs, and for inserting exploit detectors before a system can be patched, among others. Here, we illustrate how to use such probing for trustworthy logging and highlight some of the challenges that existing event-based probing mechanisms do not address. Challenges include ensuring a probe inserted at given address is trustworthy despite the lack of attestation available for probes that have been inserted dynamically. We show how probes can be inserted to ensure proper logging of every invocation of a probed instruction. When combined with attested boot of the hypervisor and guest machines, we can ensure the output stream of monitored events is trustworthy. Using these techniques we build a trustworthy log of certain guest-system-call events. The log powers a cloud-tuned Intrusion Detection System (IDS). New event types are identified that must be added to existing probing systems to ensure attempts to circumvent probes within the guest appear in the log. We highlight the overhead penalties paid by guests to increase guarantees of log completeness when faced with attacks on the guest kernel. Promising results (less that 10% for guests) are shown when a guest relaxes the trade-off between log completeness and overhead. Our demonstrative IDS detects common attack scenarios with simple policies built using our guest behavior recording system. Read Sprabery, Zachary Estrada, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Rakesh Bobba, Roy H. Campbell |
IC2E | 4 |
| 2017 | EEG-GRAPH: A Factor-Graph-Based Model for Capturing Spatial, Temporal, and Observational Relationships in ElectroencephalogramsabstractThis paper presents a probabilistic-graphical model that can be used to infer characteristics of instantaneous brain activity by jointly analyzing spatial and temporal dependencies observed in electroencephalograms (EEG). Specifically, we describe a factor-graph-based model with customized factor-functions defined based on domain knowledge, to infer pathologic brain activity with the goal of identifying seizure-generating brain regions in epilepsy patients. We utilize an inference technique based on the graph-cut algorithm to exactly solve graph inference in polynomial time. We validate the model by using clinically collected intracranial EEG data from 29 epilepsy patients to show that the model correctly identifies seizure-generating brain regions. Our results indicate that our model outperforms two conventional approaches used for seizure-onset localization (5-7% better AUC: 0.72, 0.67, 0.65) and that the proposed inference technique provides 3-10% gain in AUC (0.72, 0.62, 0.69) compared to sampling-based alternatives. Yogatheesan Varatharajah, Min Jin Chong, Krishnakant V. Saboo, Brent M. Berry, Benjamin H. Brinkmann, Gregory A. Worrell, Ravishankar K. Iyer |
NIPS | 7 |
| 2017 | Attack Induced Common-Mode Failures on PLC-Based Safety System in a Nuclear Power Plant: Practical Experience ReportabstractThis paper demonstrates attack induced common-mode failures on an industrial-grade (Tricon) Triple-Modular-Redundant PLC (programmable logic controller) and its impact in a Nuclear Power Plant settings. The attack exploits the fact that during the configuration phase the same control logic is downloaded to all three redundant modules. We describe how an attacker can exploit this vulnerability to embed malicious control logic and how to trigger the attack. The feasibility and the attack impact are evaluated on a testbed, which includes the Tricon PLC as part of a safety protection system in a simulated nuclear power plant. Bernard Lim, Daniel Chen 0001, Yongkyu An, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
PRDC | 5 |
| 2017 | SVAuth - A Single-Sign-On Integration Solution with Runtime Verification
Shuo Chen 0001, Matt McCutchen, Phuong Cao, Shaz Qadeer, Ravishankar K. Iyer |
RV | 5 |
| 2017 | Using OS Design Patterns to Provide Reliability and Security as-a-Service for VM-based CloudsabstractThis paper extends the concepts behind cloud services to offer hypervisor-based reliability and security monitors for cloud virtual machines. Cloud VMs can be heterogeneous and as such guest OS parameters needed for monitoring can vary across different VMs and must be obtained in some way. Past work involves running code inside the VM, which is unacceptable for a cloud environment. We solve this problem by recognizing that there are common OS design patterns that can be used to infer monitoring parameters from the guest OS. We extract information about the cloud user's guest OS with the user's existing VM image and knowledge of OS design patterns as the only inputs to analysis. To demonstrate the range of monitoring functionality possible with this technique, we implemented four sample monitors: a guest OS process tracer, an OS hang detector, a return-to-user attack detector, and a process-based keylogger detector. Zachary Estrada, Read Sprabery, Lok K. Yan, Zhongzhi Yu, Roy H. Campbell, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
VEE | 7 |
| 2017 | Seizure Forecasting and the Preictal State in Canine EpilepsyabstractThe ability to predict seizures may enable patients with epilepsy to better manage their medications and activities, potentially reducing side effects and improving quality of life. Forecasting epileptic seizures remains a challenging problem, but machine learning methods using intracranial electroencephalographic (iEEG) measures have shown promise. A machine-learning-based pipeline was developed to process iEEG recordings and generate seizure warnings. Results support the ability to forecast seizures at rates greater than a Poisson random predictor for all feature sets and machine learning algorithms tested. In addition, subject-specific neurophysiological changes in multiple features are reported preceding lead seizures, providing evidence supporting the existence of a distinct and identifiable preictal state. Yogatheesan Varatharajah, Ravishankar K. Iyer, Brent M. Berry, Gregory A. Worrell, Benjamin H. Brinkmann |
Int. J. Neural Syst. | 2 |
| 2017 | Modeling and Mitigating Impact of False Data Injection Attacks on Automatic Generation ControlabstractThis paper studies the impact of false data injection (FDI) attacks on automatic generation control (AGC), a fundamental control system used in all power grids to maintain the grid frequency at a nominal value. Attacks on the sensor measurements for AGC can cause frequency excursion that triggers remedial actions, such as disconnecting customer loads or generators, leading to blackouts, and potentially costly equipment damage. We derive an attack impact model and analyze an optimal attack, consisting of a series of FDIs that minimizes the remaining time until the onset of disruptive remedial actions, leaving the shortest time for the grid to counteract. We show that, based on eavesdropped sensor data and a few feasible-to-obtain system constants, the attacker can learn the attack impact model and achieve the optimal attack in practice. This paper provides essential understanding on the limits of physical impact of the FDIs on power grids, and provides an analysis framework to guide the protection of sensor data links. For countermeasures, we develop efficient algorithms to detect the attack, estimate which sensor data links are under attack, and mitigate attack impact. Our analysis and algorithms are validated by experiments on a physical 16-bus power system test bed and extensive simulations based on a 37-bus power system model. Rui Tan 0001, Hoang Hai Nguyen, Yi Shyh Eddy Foo, David K. Y. Yau, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Hoay Beng Gooi |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2017 | Failure Diagnosis for Distributed Systems Using Targeted Fault InjectionabstractThis paper introduces a novel approach to automating failure diagnostics in distributed systems by combining fault injection and data analytics. We use fault injection to populate the database of failures for a target distributed system. When a failure is reported from production environment, the database is queried to find “matched” failures generated by fault injections. Relying on the assumption that similar faults generate similar failures, we use information from the matched failures as hints to locate the actual root cause of the reported failures. In order to implement this approach, we introduce techniques for (i) reconstructing end-to-end execution flows of distributed software components, (ii) computing the similarity of the reconstructed flows, and (iii) performing precise fault injection at pre-specified executing points in distributed systems. We have evaluated our approach using an OpenStack cloud platform, a popular cloud infrastructure management system. Our experimental results showed that this approach is effective in determining the root causes, e.g., fault types and affected components, for 71-100 percent of tested failures. Furthermore, it can provide fault locations close to actual ones and can easily be used to find and fix actual root causes. We have also validated this technique by localizing real bugs that occurred in OpenStack. Cuong Pham 0003, Long Wang 0003, Byung-Chul Tak, Salman Baset, Chunqiang Tang, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2017 | Analysis and Diagnosis of SLA Violations in a Production SaaS CloudabstractA software-as-a-service (SaaS) needs to provide its intended service as per its stated service-level agreements (SLAs). While SLA violations in a SaaS platform have been reported, not much work has been done to empirically characterize failures of SaaS. In this paper, we study SLA violations of a production SaaS platform, diagnose the causes, unearth several critical failure modes, and then, suggest various solution approaches to increase the availability of the platform as perceived by the end user. Our approach combines field failure data analysis (FFDA) and fault injection. Our study is based on 283 days of operational logs of the platform. During this time, the platform received business workload from 42 customers spread over 22 countries. We have first developed a set of home-grown FFDA tools to analyze the log, and second implemented a fault injector to automatically inject several runtime errors in the application code written in .NET/C#, and then, collate the injection results. We summarize our finding as: first, system failures have caused 93% of all SLA violations; second, our fault injector has been able to recreate a few cases of bursts of SLA violations that could not be diagnosed from the logs; and third, the fault injection mechanism could recreate several error propagation paths leading to data corruptions that the failure data analysis could not reveal. Finally, the paper presents some system-level implication of this study and how the joint use of fault injection and log analysis may help in improving the reliability of the measured platform. Catello Di Martino, Santonu Sarkar, Rajeshwari Ganesan, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
IEEE Trans. Reliab. | 5 |
| 2016 | Towards longitudinal analysis of a population's electronic health records using factor graphsabstractIn this feasibility study, we demonstrate the use of a factor-graph-based probabilistic graphical model approach to process longitudinal data derived from a population's electronic health records (EHR). Processing of EHR allows for fore-casting patient-specific health complications and inference of population-level statistics on several epidemiological factors. As a case-study, we provide preliminary results and demonstrate feasibility of our approach by processing the EHR of a diabetic cohort in Singapore. Our model passes the feasibility test as we are able to forecast a series of health complications of a new patient based on the factor functions inferred from EHR of 100 diabetic patients spanning 10-years. This forecast gives both the caregivers and the patient a better view of the patient's health in the coming years and increases patient's motivation to stay healthy and conform to medication plan. Furthermore, our approach informs commonly occurring health complications in the population that warrant hospital readmissions, which helps a physician/clinician in decide when to intervene to avoid complications in order to improve the patient's quality of life and minimize the cost of care. Arjun P. Athreya, Kee Yuan Ngiam, Zhaojing Luo, E. Shyong Tai, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
BDCAT | 6 |
| 2016 | Unsupervised single-cell analysis in triple-negative breast cancer: A case studyabstractThis paper demonstrates an unsupervised learning approach to identify genes with significant differential expression across single-cell subpopulations induced by therapeutic treatment. Identifying this set of genes makes it possible to use well-established bioinformatics approaches such as pathway analysis to establish their biological relevance. Then, a biologist can use his/her prior knowledge to investigate in the laboratory, a few particular candidates among the subset of genes overlapping with relevant pathways. Due to the large size of the human genome and limitations in cost and skilled resources, biologists benefit from analytical methods combined with pathway analysis to design laboratory experiments focusing on only a few significant genes. As an example, we show how model-based unsupervised methods can identify a small set of genes (1% of the genome) that have significant differential expression in single-cells and are also highly correlated to pathways (p-value < 1E − 7) with anticancer effects driven by the antidiabetic drug metformin. Further analysis of genes on these relevant pathways reveal three candidate genes previously implicated in several anticancer mechanisms in other cancers, not driven by metformin. Identification of these genes can help biologists and clinicians design laboratory experiments to establish the molecular mechanisms of metformin in triple-negative breast cancer. In a domain where there is no prior knowledge of small biologically significant data, we demonstrate that careful data-driven methods can infer such significant small data to explain biological mechanisms. Arjun P. Athreya, Alan J. Gaglio, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Junmei Cairns, Krishna Rani Kalari, Richard M. Weinshilboum, Liewei Wang |
BIBM | 4 |
| 2016 | Targeted Attacks on Teleoperated Surgical Robots: Dynamic Model-Based Detection and MitigationabstractThis paper demonstrates targeted cyber-physical attacks on teleoperated surgical robots. These attacks exploit vulnerabilities in the robot's control system to infer a critical time during surgery to drive injection of malicious control commands to the robot. We show that these attacks can evade the safety checks of the robot, lead to catastrophic consequences in the physical system (e.g., sudden jumps of robotic arms or system's transition to an unwanted halt state), and cause patient injury, robot damage, or system unavailability in the middle of a surgery. We present a model-based analysis framework that can estimate the consequences of control commands through real-time computation of robot's dynamics. Our experiments on the RAVEN II robot demonstrate that this framework can detect and mitigate the malicious commands before they manifest in the physical system with an average accuracy of 90%. Homa Alemzadeh, Daniel Chen 0001, Thenkurussi Kesavadas, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 6 |
| 2016 | F-DETA: A Framework for Detecting Electricity Theft Attacks in Smart GridsabstractElectricity theft is a major concern for utilities all over the world, and leads to billions of dollars in losses every year. Although improving the communication capabilities between consumer smart meters and utilities can enable many smart grid features, these communications can be compromised in ways that allow an attacker to steal electricity. Such attacks have recently begun to occur, so there is a real and urgent need for a framework to defend against them. In this paper, we make three major contributions. First, we develop what is, to our knowledge, the most comprehensive classification of electricity theft attacks in the literature. These attacks are classified based on whether they can circumvent security measures currently used in industry, and whether they are possible under different electricity pricing schemes. Second, we propose a theft detector based on Kullback-Leibler (KL) divergence to detect cleverly-crafted electricity theft attacks that circumvent detectors proposed in related work. Finally, we evaluate our detector using false data injections based on real smart meter data. For the different attack classes, we show that our detector dramatically mitigates electricity theft in comparison to detectors in prior work. Varun Badrinath Krishna, Kiryung Lee, Gabriel A. Weaver, Ravishankar K. Iyer, William H. Sanders |
DSN | 4 |
| 2016 | A hardware-in-the-loop simulator for safety training in robotic surgeryabstractThis paper presents a simulation-based safety training simulator for robot assisted surgery. While adverse events occur rarely during training, they could be fatal to the patients if they happen during real surgical procedures and are not handled properly by the surgical team. In this work we propose a hardware-in-the-loop robotic surgery simulator with high fidelity of the robot motion in a simulated environment, which is capable of reproducing adverse events during surgery. The proposed simulator is built upon the Raven-II open source surgical robot, integrated with a simulated surgeon console and a safety hazard injection engine, which automatically injects faults into modules of the robot control software. We simulate representative safety hazards seen in the adverse events, related to da Vinci™ robot, reported to the FDA MAUDE database. A novel haptic feedback strategy is provided to the operator when the underlying dynamics differ from the real robot states. Homa Alemzadeh, Daniel Chen 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Thenkurussi Kesavadas |
IROS | 5 |
| 2016 | A Data-Driven Approach to Soil Moisture Collection and PredictionabstractAgriculture has been one of the most under-investigated areas in technology, and the development of Precision Agriculture (PA) is still in its early stages. This paper proposes a data-driven methodology on building PA solutions for collection and data modeling systems. Soil moisture, a key factor in the crop growth cycle, is selected as an example to demonstrate the effectiveness of our data-driven approach. On the collection side, a reactive wireless sensor node is developed that aims to capture the dynamics of soil moisture using MicaZ mote and VH400 soil moisture sensor. The prototyped device is tested on field soil to demonstrate its functionality and the responsiveness of the sensors. On the data analysis side, a unique, site-specific soil moisture prediction framework is built on top of models generated by the machine learning techniques SVM (support vector machine) and RVM (relevance vector machine). The framework predicts soil moisture n days ahead based on the same soil and environmental attributes that can be collected by our sensor node. Due to the large data size required by the machine learning algorithms, our framework is evaluated under the Illinois historical data, not field collected sensor data. It achieves low error rates (15%) and high correlations (95%) between predicted values and actual values across 9 different sites when forecasting soil moisture about 2 weeks ahead. Also, it is shown that the prediction outputs can remain accurate over a long period of time (one year) when reliable data are fed to the model every 45 days. Zhihao Hong, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SMARTCOMP | 3 |
| 2015 | ARIMA-Based Modeling and Validation of Consumption Readings in Power Grids
Varun Badrinath Krishna, Ravishankar K. Iyer, William H. Sanders |
CRITIS | 2 |
| 2015 | Measuring and Understanding Extreme-Scale Application Resilience: A Field Study of 5, 000, 000 HPC Application RunsabstractThis paper presents an in-depth characterization of the resiliency of more than 5 million HPC application runs completed during the first 518 production days of Blue Waters, a 13.1 petaflop Cray hybrid supercomputer. Unlike past work, we measure the impact of system errors and failures on user applications, i.e., the compiled programs launched by user jobs that can execute across one or more XE (CPU) or XK (CPU+GPU) nodes. The characterization is performed by means of a joint analysis of several data sources, which include workload and error/failure logs. In order to relate system errors and failures to the executed applications, we developed LogDiver, a tool to automate the data pre-processing and metric computation. Some of the lessons learned in this study include: i) while about 1.53% of applications fail due to system problems, the failed applications contribute to about 9% of the production node hours executed in the measured period, i.e., the system consumes computing resources, and system-related issues represent a potentially significant energy cost for the work lost, ii) there is a dramatic increase in the application failure probability when executing full-scale applications: 20x (from 0.008 to 0.162) when scaling XE applications from 10,000 to 22,000 nodes, and 6x (from 0.02 to 0.129) when scaling GPU/hybrid applications from 2000 to 4224 nodes, and iii) the resiliency of hybrid applications is impaired by the lack of adequate error detection capabilities in hybrid nodes. Catello Di Martino, William T. Kramer, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 4 |
| 2015 | The future of fault tolerant computingabstractFault tolerant (or dependable) computing has always been an exciting research area in the intersection of computer science and engineering and electrical and electronics engineering. During the last two decades the applicability of the methods and tools that the fault tolerance research community produces has expanded to virtually all application domains. The type of fault tolerance methods employed in a computing system depend on: (a) the faults expected to affect the system, (b) the importance of errors in the system operation, (c) the design, cost and power budgets that can allocated to fault tolerance and reliable operation. New solutions and tools in fault tolerant computing are emerging to deal with the very broad spectrum of values that all (a), (b) and (c) can take in today's computing landscape. Jacob A. Abraham, Ravishankar K. Iyer, Dimitris Gizopoulos, Dan Alexandrescu, Yervant Zorian |
IOLTS | 2 |
| 2015 | Systems-Theoretic Safety Assessment of Robotic Telesurgical Systems
Homa Alemzadeh, Daniel Chen 0001, Zbigniew T. Kalbarczyk, Jaishankar Raman, Nancy G. Leveson, Ravishankar K. Iyer |
SAFECOMP | 7 |
| 2015 | VM-μCheckpoint: Design, Modeling, and Assessment of Lightweight In-Memory VM CheckpointingabstractCheckpointing and rollback techniques enhance reliability and availability of virtual machines and their hosted IT services. This paper proposes VM-μCheckpoint, a light-weight pure-software mechanism for high-frequency checkpointing and rapid recovery for VMs. Compared with existing techniques of VM checkpointing, VM-μCheckpoint tries to minimize checkpoint overhead and speed up recovery by means of copy-on-write, dirty-page prediction and in-place recovery, as well as saving incremental checkpoints in volatile memory. Moreover, VM-μCheckpoint deals with the issue that latency in error detection potentially results in corrupted checkpoints, particularly when checkpointing frequency is high. We also constructed Markov models to study the availability improvements provided by VM-μCheckpoint (from 99 to 99.98 percent on reasonably reliable hypervisors). We designed and implemented VM-μCheckpoint in the Xen VMM. The evaluation results demonstrate that VM-μCheckpoint incurs an average of 6.3 percent overhead (in terms of program execution time) for 50 ms checkpoint intervals when executing the SPEC CINT 2006 benchmark. Error injection experiments demonstrate that VM-μCheckpoint, combined with error detection techniques in RMK, provides high coverage of recovery. Long Wang 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Arun Iyengar |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2014 | Automated Classification of Computer-Based Medical Device Recalls: An Application of Natural Language Processing and Statistical LearningabstractThis paper presents MedSafe, a framework for automated classification of computer-based medical device recalls. The data is collected from the U.S. Food and Drug Administration (FDA) recalls database. We combined techniques in natural language processing and statistical learning to automatically identify the computer-related recalls, by interpreting the natural language semantics of recall descriptions. We evaluated MedSafe on over 16K recall records submitted to the FDA between years 2007-2013. Homa Alemzadeh, Raymond Hoagland, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
CBMS | 4 |
| 2014 | A Performance Evaluation of Sequence Alignment Software in Virtualized EnvironmentsabstractThe prospect of simpler infrastructure management and affordability has garnered interest in cloud computing from bioinformaticians. However, the performance cost of adopting such an infrastructure model for bioinformatics is not fully known. In an effort to help quantify this performance cost, we ran synthetic benchmarks and measured the runtimes of two short-read alignment applications on cloud-like virtualization environments. The environments were implemented utilizing the KVM hypervisor, the Xen hypervisor, and Linux Containers. We compare the runtime in each environment against a physical server and offer discussion and insights. Though the applications perform similar operations, we observe that their performance characteristics differ, as do their performance in the different virtualized environments. We attribute the differences to the way that these programs utilize system resources. We find that the more CPU-bound Novo align is much less sensitive to virtualization environments than BWA is, and has near-physical server performance even when virtualized. Additionally, we find that static CPU pinning can improve performance, and we demonstrate that Linux Containers offer performance comparable to that of a physical server. Zachary Estrada, Zachary Stephens, Cuong Manh Pham, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
CCGRID | 5 |
| 2014 | AHEMS: Asynchronous Hardware-Enforced Memory SafetyabstractThis paper presents AHEMS (Asynchronous Hardware-Enforced Memory Safety), an architectural support for enforcing spatial and temporal memory safety to protect against memory corruption attacks. We integrated AHEMS with the Leon3 open-source processor and prototype on an FPGA. In an evaluation of the detection coverage using 677 security test cases (including spatial and temporal memory errors), selected from the Juliet Test Suite, AHEMS detected all but one memory safety violation. The missed test case involves overflow of a sub-object in a data structure whose detection is not supported by the current prototype. Performance assessment using the Olden benchmarks shows an average 10.6% overhead, and negligible impact on the processor-critical path (0.06% overhead) and power consumption (0.5% overhead). Kuan-Yu Tseng, Dao Lu, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSD | 4 |
| 2014 | Lessons Learned from the Analysis of System Failures at Petascale: The Case of Blue WatersabstractThis paper provides an analysis of failures and their impact for Blue Waters, the Cray hybrid (CPU/GPU) supercomputer at the University of Illinois at Urbana-Champaign. The analysis is based on both manual failure reports and automatically generated event logs collected over 261 days. Results include i) a characterization of the root causes of single-node failures, ii) a direct assessment of the effectiveness of system-level fail over as well as memory, processor, network, GPU accelerator, and file system error resiliency, and iii) an analysis of system-wide outages. The major findings of this study are as follows. Hardware is not the main cause of system downtime. This is notwithstanding the fact that hardware-related failures are 42% of all failures. Failures caused by hardware were responsible for only 23% of the total repair time. These results are partially due to the fact that processor and memory protection mechanisms (x8 and x4 Chip kill, ECC, and parity) are able to handle a sustained rate of errors as high as 250 errors/h while providing a coverage of 99.997% out of a set of more than 1.5 million of analyzed errors. Only 28 multiple-bit errors bypassed the employed protection mechanisms. Software, on the other hand, was the largest contributor to the node repair hours (53%), despite being the cause of only 20% of the total number of failures. A total of 29 out of 39 system-wide outages involved the Lustre file system with 42% of them caused by the inadequacy of the automated fail over procedures. Catello Di Martino, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Fabio Baccanico, Joseph Fullop, William T. Kramer |
DSN | 3 |
| 2014 | Reliability and Security Monitoring of Virtual Machines Using Hardware Architectural InvariantsabstractThis paper presents a solution that simultaneously addresses both reliability and security (RnS) in a monitoring framework. We identify the commonalities between reliability and security to guide the design of Hyper Tap, a hyper visor-level framework that efficiently supports both types of monitoring in virtualization environments. In Hyper Tap, the logging of system events and states is common across monitors and constitutes the core of the framework. The audit phase of each monitor is implemented and operated independently. In addition, Hyper Tap relies on hardware invariants to provide a strongly isolated root of trust. Hyper Tap uses active monitoring, which can be adapted to enforce a wide spectrum of RnS policies. We validate Hyper Tap by introducing three example monitors: Guest OS Hang Detection (GOSHD), Hidden Root Kit Detection (HRKD), and Privilege Escalation Detection (PED). Our experiments with fault injection and real root kits/exploits demonstrate that Hyper Tap provides robust monitoring with low performance overhead. Cuong Manh Pham, Zachary Estrada, Phuong Cao, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 5 |
| 2014 | Analysis and Diagnosis of SLA Violations in a Production SaaS CloudabstractThis paper investigates SLA violations of a production SaaS platform by means of joint use of field failure data analysis (FFDA) and fault injection. The objective of this study is to diagnose the causes of SLA violations, pinpoint critical failure modes under realistic error assumptions and identify potential means to increase the user perceived availability of the platform and assurance of SLA requirements. We base our study on 283 days of logs obtained during the production time of the platform, while it was employed to process business data received by 42 customers in 22 countries. In this paper, we develop a set of tools that include i) a FFDA toolset used to analyze the data extracted from the platform and by the operating system event logs and ii) a. NET/C++ injector able to automate the injection of specific runtime errors in the production code and the collection of results. Major findings include i) 93% of all service level agreement (SLA) violations were due to system failures, ii) there were a few cases of bursts of SLA violations that could not be diagnosed from the logs and were revealed from the performed injections, and iii) the error injection revealed several error propagation paths leading to data corruptions that could not be detected from the analysis of failure data. Catello Di Martino, Daniel Chen 0001, Geetika Goel, Rajeshwari Ganesan, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ISSRE | 6 |
| 2014 | Semantic Security Analysis of SCADA Networks to Detect Malicious Control Commands in Power Grids (Poster)abstractIn this poster, we present a semantic analysis framework based on a collaborative network of intrusion detection systems (IDSes) that we proposed in [3] to detect control-related attacks in power systems. The framework combines system knowledge of both cyber and physical infrastructure in power grids to help the IDS to estimate execution consequences of control commands. We demonstrate the implementation based on Bro IDS [11] and the experimental results on the performance overhead of the semantic analysis framework. Hui Lin 0005, Adam J. Slagell, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SIN | 4 |
| 2014 | An evaluation of zookeeper for high availability in system SabstractZooKeeper provides scalable, highly available coordination services for distributed applications. In this paper, we evaluate the use of ZooKeeper in a distributed stream computing system called System S to provide a resilient name service, dynamic configuration management, and system state management. The evaluation shed light on the advantages of using ZooKeeper in these contexts as well as its limitations. We also describe design changes we made to handle named objects in System S to overcome the limitations. We present detailed experimental results, which we believe will be beneficial to the community. Cuong Manh Pham, Victor Dogaru, Rohit Wagle, Chitra Venkatramani, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ICPE | 6 |
| 2013 | Pluggable Watchdog: Transparent Failure Detection for MPI ProgramsabstractThis paper presents a framework and its techniques that can detect various types of runtime errors and failures in MPI programs. The presented framework offloads its detection techniques to an external device (e.g., extension card). By developing intelligence on the normal behavioral and semantic execution patterns of monitored parallel threads, the presented external error detectors can accurately and quickly detect errors and failures. This architecture allows us to use powerful detectors without directly using the computing power of the monitored system. The separation of hardware of the monitored and monitoring systems offers an extra advantage in terms of system reliability. We have prototyped our system on a parallel computer system by using an FPGA-based PCI extension card as a monitoring device. We have conducted a fault injection experiment to evaluate the presented techniques using eight MPI-based parallel programs. The techniques cover ~98.5% of faults, on average. The average performance overhead is 1.8% for techniques that detect crash and hang failures and 6.6% for techniques that detect SDC failures. Keun Soo Yim, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
IPDPS | 3 |
| 2013 | SymPLFIED: Symbolic Program-Level Fault Injection and Error Detection FrameworkabstractThis paper introduces SymPLFIED, a program-level framework that allows specification of arbitrary error detectors and the verification of their efficacy against hardware errors. SymPLFIED comprehensively enumerates all transient hardware errors in registers, memory, and computation (expressed symbolically as value errors) that potentially evade detection and cause program failure. The framework uses symbolic execution to abstract the state of erroneous values in the program and model checking to comprehensively find all errors that evade detection. We demonstrate the use of SymPLFIED on a widely deployed aircraft collision avoidance application, tcas. Our results show that the SymPLFIED framework can be used to uncover hard-to-detect catastrophic cases caused by transient errors in programs that may not be exposed by random fault injection-based validation. Further, the errors exposed by the framework help us formulate a set of error detectors for the application to avoid the catastrophic case and other incorrect outcomes. Karthik Pattabiraman, Nithin Nakka, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
IEEE Trans. Computers | 4 |
| 2012 | Characterization of the error resiliency of power grid substation devicesabstractWith the advent of modern technologies, microprocessor-based devices are used to monitor and control critical infrastructures, e.g., electric power grids, oil and gas distribution. However, the security and reliability of these microprocessor-based systems is a significant issue, since they are more susceptible to transient errors and malicious attacks. An error in one of these systems could have a cascading and catastrophic impact on the whole infrastructure. This paper explores the error resiliency of power grid substation devices. A software-implemented fault injection technique is used to induce errors/faults inside devices used in power grid substations. The goal is to test the ability of these systems to compute through errors/faults. Our results demonstrate that a single error in a substation device may render the operator in the control center unable to control the operation of a relay in the substation. Kuan-Yu Tseng, Daniel Chen 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 4 |
| 2012 | Architectural impact of secure socket layer on Internet servers: A retrospectabstractSecure socket layer (SSL) is the most popular protocol used in the Internet for facilitating secure communications. In this retrospective, we summarize our original paper which analyzed the performance and architectural impact of SSL on the servers and provided insights into the functioning and acceleration of SSL. In addition, we describe advancements in the area that have occurred on this topic and also discuss future research opportunities. Krishna Kant 0001, Ravishankar K. Iyer, Prasant Mohapatra |
ICCD | 2 |
| 2012 | Architectural impact of secure socket layer on Internet serversabstractSecure socket layer (SSL) is the most popular protocol used in the Internet for facilitating secure communications. In this paper, we analyze the performance and architectural impact of SSL on the servers in terms of various parameters such as throughput, utilization, cache sizes, cache miss ratios, number of processors, control dependencies, file access sizes, bus transactions, network load, etc. The major conclusions from this study are as follows: The use of SSL increases computational cost of the transactions by a factor of 5-7. SSL transactions do not benefit much from a larger L2 cache, but a larger LI cache would be helpful. A complex logic for handling control dependencies is not useful for SSL transaction as the frequency of branches is very low. Because SSL workload is highly CPU bound, it may be possible to enhance SSL performance by using a number of other architectural features as well. Krishna Kant 0001, Ravishankar K. Iyer, Prasant Mohapatra |
ICCD | 2 |
| 2011 | Modeling stream processing applications for dependability evaluationabstractThis paper describes a modeling framework for evaluating the impact of faults on the output of streaming applications. Our model is based on three abstractions: stream operators, stream connections, and tuples. By composing these abstractions within a Stochastic Activity Network, we allow the modeling of complete applications. We consider faults that lead to data loss and to silent data corruption (SDC). Our framework captures how faults originating in one operator propagate to other operators down the stream processing graph. We demonstrate the extensibility of our framework by evaluating three different fault tolerance techniques: checkpointing, partial graph replication, and full graph replication. We show that under crashes that lead to data loss, partial graph replication has a great advantage in maintaining the accuracy of the application output when compared to checkpointing. We also show that SDC can break the no data duplication guarantees of a full graph replication-based fault tolerance technique. Gabriela Jacques-Silva, Zbigniew T. Kalbarczyk, Bugra Gedik, Henrique Andrade, Kun-Lung Wu, Ravishankar K. Iyer |
DSN | 6 |
| 2011 | Improving Log-based Field Failure Data Analysis of multi-node computing systemsabstractLog-based Field Failure Data Analysis (FFDA) is a widely-adopted methodology to assess dependability properties of an operational system. A key step in FFDA is filtering out entries that are not useful and redundant error entries from the log. The latter is challenging: a fault, once triggered, can generate multiple errors that propagate within the system. Grouping the error entries related to the same fault manifestation is crucial to obtain realistic measurements. This paper deals with the issues of the tuple heuristic, used to group the error entries in the log, in multi-node computing systems. We demonstrate that the tuple heuristic can group entries incorrectly; thus, an improved heuristic that adopts statistical indicators is proposed. We assess the impact of inaccurate grouping on dependability measurements by comparing the results obtained with both the heuristics. The analysis encompasses the log of the Mercury cluster at the National Center for Supercomputing Applications. Antonio Pecchia, Domenico Cotroneo, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 4 |
| 2011 | CloudVal: A framework for validation of virtualization environment in cloud infrastructureabstractWe present CloudVal, a framework to validate the reliability of virtualization environment in Cloud Computing infrastructure. A case study, based on injecting faults in the KVM hypervisor and Xen hypervisor, was conducted to show the viability of the framework. The study shows that due to the architectural differences between KVM and Xen, a direct comparison of the two virtualization systems is not feasible. In order to confidently weigh error resiliency of virtualization systems, more comprehensive studies are required. We believe, however, that the fault injection approach and the fault models proposed in this paper are a good starting point towards designing and implementing a benchmark which would enable the assessment of different virtualization infrastructures in a common manner. Cuong Manh Pham, Daniel Chen 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 4 |
| 2011 | Analysis of security data from a large computing organizationabstractThis paper presents an in-depth study of the forensic data on security incidents that have occurred over a period of 5 years at the National Center for Supercomputing Applications at the University of Illinois. The proposed methodology combines automated analysis of data from security monitors and system logs with human expertise to extract and process relevant data in order to: (i) determine the progression of an attack, (ii) establish incident categories and characterize their severity, (iii) associate alerts with incidents, and (iv) identify incidents missed by the monitoring tools and examine the reasons for the escapes. The analysis conducted provides the basis for incident modeling and design of new techniques for security monitoring. Aashish Sharma, Zbigniew T. Kalbarczyk, James Barlow, Ravishankar K. Iyer |
DSN | 4 |
| 2011 | HTAF: Hybrid Testing Automation Framework to Leverage Local and Global Computing Resources
Keun Soo Yim, David Hreczany, Ravishankar K. Iyer |
ICCSA (3) | 3 |
| 2011 | Hauberk: Lightweight Silent Data Corruption Error Detector for GPGPUabstractHigh performance and relatively low cost of GPU-based platforms provide an attractive alternative for general purpose high performance computing (HPC). However, the emerging HPC applications have usually stricter output correctness requirements than typical GPU applications (i.e., 3D graphics). This paper first analyzes the error resiliency of GPGPU platforms using a fault injection tool we have developed for commodity GPU devices. On average, 16-33% of injected faults cause silent data corruption (SDC) errors in the HPC programs executing on GPU. This SDC ratio is significantly higher than that measured in CPU programs (<;2.3%). In order to tolerate SDC errors, customized error detectors are strategically placed in the source code of target GPU programs so as to minimize performance impact and error propagation and maximize recoverability. The presented Hauberk technique is deployed in seven HPC benchmark programs and evaluated using a fault injection. The results show a high average error detection coverage (~87%) with a small performance overhead (~15%). Keun Soo Yim, Cuong Manh Pham, Mushfiq Saleheen, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
IPDPS | 5 |
| 2011 | Identifying Compromised Users in Shared Computing Infrastructures: A Data-Driven Bayesian Network ApproachabstractThe growing demand for processing and storage capabilities has led to the deployment of high-performance computing infrastructures. Users log into the computing infrastructure remotely, by providing their credentials (e.g., username and password), through the public network and using well-established authentication protocols, e.g., SSH. However, user credentials can be stolen and an attacker (using a stolen credential) can masquerade as the legitimate user and penetrate the system as an insider. This paper deals with security incidents initiated by using stolen credentials and occurred during the last three years at the National Center for Supercomputing Applications (NCSA) at the University of Illinois. We analyze the key characteristics of the security data produced by the monitoring tools during the incidents and use a Bayesian network approach to correlate (i) data provided by different security tools (e.g., IDS and Net Flows) and (ii) information related to the users' profiles to identify compromised users, i.e., the users whose credentials have been stolen. The technique is validated with the real incident data. The experimental results demonstrate that the proposed approach is effective in detecting compromised users, while allows eliminating around 80% of false positives (i.e., not compromised user being declared compromised). Antonio Pecchia, Aashish Sharma, Zbigniew T. Kalbarczyk, Domenico Cotroneo, Ravishankar K. Iyer |
SRDS | 5 |
| 2011 | Automated Derivation of Application-Aware Error Detectors Using Static Analysis: The Trusted Illiac ApproachabstractThis paper presents a technique to derive and implement error detectors to protect an application from data errors. The error detectors are derived automatically using compiler-based static analysis from the backward program slice of critical variables in the program. Critical variables are defined as those that are highly sensitive to errors, and deriving error detectors for these variables provides high coverage for errors in any data value used in the program. The error detectors take the form of checking expressions and are optimized for each control-flow path followed at runtime. The derived detectors are implemented using a combination of hardware and software and continuously monitor the application at runtime. If an error is detected at runtime, the application is stopped so as to prevent error propagation and enable a clean recovery. Experiments show that the derived detectors achieve low-overhead error detection while providing high coverage for errors that matter to the application. Karthik Pattabiraman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2011 | Automated Derivation of Application-Specific Error Detectors Using Dynamic AnalysisabstractThis paper proposes a novel technique for preventing a wide range of data errors from corrupting the execution of applications. The proposed technique enables automated derivation of fine-grained, application-specific error detectors based on dynamic traces of application execution. The technique derives a set of error detectors using rule-based templates to maximize the error detection coverage for the application. A probability model is developed to guide the choice of the templates and their parameters for error-detection. The paper also presents an automatic framework for synthesizing the set of detectors in hardware to enable low-overhead, runtime checking of the application. The coverage of the derived detectors is evaluated using fault-injection experiments, while the performance and area overheads of the detectors are evaluated by synthesizing them on reconfigurable hardware. Karthik Pattabiraman, Giacinto Paolo Saggese, Daniel Chen 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2010 | A Soldier Health Monitoring System for Military ApplicationsabstractWith recent advances in technology, various wearable sensors have been developed for the monitoring of human physiological parameters. A Body Sensor Network (BSN) consisting of such physiological and biomedical sensor nodes placed on, near or within a human body can be used for real-time health monitoring. In this paper, we describe an on-going effort to develop a system consisting of interconnected BSNs for real-time health monitoring of soldiers. We discuss the background and an application scenario for this project. We describe the preliminary prototype of the system and present a blast source localization application. Hock-Beng Lim, Di Ma 0001, Bang Wang 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Kenneth L. Watkin |
BSN | 5 |
| 2010 | Fourth workshop on dependable and secure nanocomputingabstractNanocomputing technologies hold the promise for higher performance, lower power consumption as well as increased functionality. However, the dependability of these unprecedentedly small scale devices remains uncertain. The main sources of concern are: • Nanometer devices are expected to be highly sensitive to process variations. The guard-bands used today for avoiding the impact of such variations will not represent a feasible solution in the future. Thus, timing errors may occur more frequently. • New failure modes, specific to new materials, are expected to raise serious challenges to the design and test engineers. • Environment induced errors, like single event upsets (SEU), are likely to occur more frequently than in the case of conventional semiconductor devices. • New hardware redundancy techniques are needed to enable development of energy efficient systems. • The increased complexity of the systems based on nanotechnology will require improved computer aided design (CAD) tools, as well as better validation techniques. • Security of nanocomputing systems may be threatened by malicious attacks targeting new vulnerable areas in the hardware. Jean Arlat, Cristian Constantinescu, Ravishankar K. Iyer, Michael Nicolaidis |
DSN | 3 |
| 2010 | Measurement-based analysis of fault and error sensitivities of dynamic memoryabstractThis paper presents a measurement-based analysis of the fault and error sensitivities of dynamic memory. We extend a software-implemented fault injector to support data-type-aware fault injection into dynamic memory. The results indicate that dynamic memory exhibits about 18 times higher fault sensitivity than static memory, mainly because of the higher activation rate. Furthermore, we show that errors in a large portion of static and dynamic memory space are recoverable by simple software techniques (e.g., reloading data from a disk). The recoverable data include pages filled with identical values (e.g., `0') and pages loaded from files unmodified during the computation. Consequently, the selection of targets for protection should be based on knowledge of recoverability rather than on error sensitivity alone. Keun Soo Yim, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 3 |
| 2010 | Checkpointing virtual machines against transient errorsabstractThis paper proposes VM-μCheckpoint, a lightweight software mechanism for high-frequency checkpointing and rapid recovery of virtual machines. VM-μCheckpoint minimizes checkpoint overhead and speeds up recovery by saving incremental checkpoints in volatile memory and by employing copy-on-write, dirty-page prediction, and in-place recovery. In our approach, knowledge of fault/error latency is used to explicitly address checkpoint corruption, a critical problem, especially when checkpoint frequency is high. We designed and implemented VM-μCheckpoint in the Xen VMM. The evaluation results demonstrate that VM-μCheckpoint incurs an average of 6.3% execution-time overhead for 50ms checkpoint intervals when executing the SPEC CINT 2006 benchmark. Long Wang 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Arun Iyengar |
IOLTS | 3 |
| 2010 | Analysis of Credential Stealing Attacks in an Open Networked EnvironmentabstractThis paper analyses the forensic data on credential stealing incidents over a period of 5 years across 5000 machines monitored at the National Center for Supercomputing Applications at the University of Illinois. The analysis conducted is the first attempt in an open operational environment (i) to evaluate the intricacies of carrying out SSH-based credential stealing attacks, (ii) to highlight and quantify key characteristics of such attacks, and (iii) to provide the system level characterization of such incidents in terms of distribution of alerts and incident consequences. Aashish Sharma, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, James Barlow |
NSS | 3 |
| 2009 | Third workshop on dependable and secure nanocomputingabstractAs already witnessed today, and as foreseen, concerns attached to hardware will play an increasing role in the design and assessment of dependable and secure computing systems. Jean Arlat, Cristian Constantinescu, Ravishankar K. Iyer, Michael Nicolaidis |
DSN | 3 |
| 2009 | An end-to-end approach for the automatic derivation of application-aware error detectorsabstractCritical Variable Recomputation (CVR) based error detection provides high coverage for data critical to an application while reducing the performance overhead associated with detecting benign errors. However, when implemented exclusively in software, the performance penalty associated with CVR based detection is unsuitably high. This paper addresses this limitation by providing a hybrid hardware/software tool chain which allows for the design of efficient error detectors while minimizing additional hardware. Detection mechanisms are automatically derived during compilation and mapped onto hardware where they are executed in parallel with the original task at runtime. When tested using an FPGA platform, results show that our approach incurs an area overhead of 53% while increasing execution time by 27% on average. Galen Lyle, Shelley Cheny, Karthik Pattabiraman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 5 |
| 2009 | Effectiveness of machine checks for error diagnosticsabstractMachine Check Architecture (MCA) is a processor internal architecture subsystem that detects and logs correctable and uncorrectable errors in the data or control paths in each CPU core and the Northbridge. These errors include parity errors associated with caches, TLBs, ECC errors associated with caches and DRAM, and system bus errors. This paper reports on an experimental study on: (i) monitoring a computing cluster for machine checks and using this data to identify patterns that can be employed for error diagnostics and (ii) introducing faults into the machine to understand the resulting machine checks and correlate this data with relevant performance metrics. Nikhil Pandit, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 3 |
| 2009 | Pervasive embedded systems for detection of traumatic brain injuryabstractTransient explosions on the battlefield result in blast injuries that are polytrauma in nature. That is, along with physical wounds additional injury can include cognitive and communication impairments related to traumatic brain injury (TBI). The design of a multisensor system embedded in an advanced combat helmet is presented that is capable of real time tracking of physiological signals (EEG, blast pressure, head acceleration, oxygen saturation and heart rate) and facilitating a reliable and dependable decision making process that provide alerts for potential traumatic brain injury. The cyperphysical system focuses on the use of heterogeneous sensors within the helmet pads and a processing element for real time algorithmic processing. Ajay M. Cheriyan, Albert O. Jarvi, Zbigniew T. Kalbarczyk, Tanya M. Gallagher, Ravishankar K. Iyer, Kenneth L. Watkin |
ICME | 5 |
| 2009 | Quantitative Analysis of Long-Latency Failures in System SoftwareabstractThis paper presents a study on long latency failures using accelerated fault injection. The data collected from the experiments are used to analyze the significance, causes, and characteristics of long latency failures caused by soft errors in the processor and the memory. The results indicate that a non-negligible portion of soft errors in the code and data memory lead to long latency failures. The long latency failures are caused by errors with long fault activation times and errors causing failures only under certain runtime conditions. On the other hand, less than 0.5% of soft errors in the processor registers used in kernel mode lead to a failure with latency longer than a thousand seconds. This is due to a strong temporal locality of the register values. The study shows also that the obtained insight can be used to guide design and placement (in the application code and/or system) of application-specific error detectors. Keun Soo Yim, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
PRDC | 3 |
| 2009 | Discovering Application-Level Insider Attacks Using Symbolic Execution
Karthik Pattabiraman, Nithin Nakka, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SEC | 4 |
| 2008 | Second workshop on dependable and secure nanocomputingabstractThere is still a strong demand for solutions to improve dependability and security. Issues at stake include threats ranging from managing complexity and manufacturing defects to accelerated aging and susceptibility to environment disturbances. These solutions extend beyond hardware, microelectronics or physics; major computing- related issues spanning various aspects are also essential: architecture, networks, communication, synchronization, task parallelization, etc. The workshop is aimed at further characterizing these impairments and threats, as well as distinguishing possible alternative design approaches and operation control paradigms that have to be enforced and/or favored in order to keep achieving dependable and secure computing. Finally, it is worth pointing out that besides design techniques aimed at providing dependable and secure computing, assessment techniques (risk evaluation, validation, testing, etc.) should also play a fundamental role! For all these reasons, we believe that DSN is an ideal forum to address all these issues and foster fruitful exchanges on a long term basis. Jean Arlat, Cristian Constantinescu, Ravishankar K. Iyer, Michael Nicolaidis |
DSN | 3 |
| 2008 | SymPLFIED: Symbolic program-level fault injection and error detection frameworkabstractThis paper introduces SymPLFIED, a program-level framework that allows specification of arbitrary error detectors and the verification of their efficacy against hardware errors. SymPLFIED comprehensively enumerates all transient hardware errors in registers, memory, and computation (expressed as value errors) that potentially evade detection and cause program failure. The framework uses symbolic execution to abstract the state of erroneous values in the program and model checking to comprehensively find all errors that evade detection. We demonstrate the use of SymPLFIED on a widely deployed aircraft collision avoidance application, tcas. Our results show that the SymPLFIED framework can be used to uncover hard-to-detect corner cases caused by transient errors in programs that may not be exposed by random fault-injection based validation. Karthik Pattabiraman, Nithin Nakka, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 4 |
| 2008 | Reliable system design: models, metrics and design techniquesabstractDesign of reliable systems meeting stringent quality, reliability, and availability requirements is becoming increasingly difficult in advanced technologies. The current design paradigm, which assumes that no gate or interconnect will ever operate incorrectly within the lifetime of a product, must change to cope with this situation. Future systems must be designed with built-in mechanisms for failure tolerance, prediction, detection and recovery during normal system operation. This tutorial will focus on models and metrics for designing reliable systems, algorithms and tools for modeling and evaluating such systems, will discuss a broad spectrum of techniques for building such systems with support for concurrent error detection, failure prediction, error correction, recovery, and self-repair. Complex interplay between power, performance and reliability requirements in future systems, and associated constraints will also be discussed. Subhasish Mitra, Ravishankar K. Iyer, Kishor S. Trivedi, James W. Tschanz |
ICCAD | 2 |
| 2008 | Error Behavior Comparison of Multiple Computing Systems: A Case Study Using Linux on Pentium, Solaris on SPARC, and AIX on POWERabstractThis paper presents an approach to conducting experimental studies for the characterization and comparison of the error behavior in different computing systems. The proposed approach is applied to characterize and compare the error behavior of three commercial systems (Linux 2.6 on Pentium 4, Solaris 10 on UltraSPARC IIIi, and AIX 5.3 on POWER 5) under hardware transient faults. The data is obtained by conducting extensive fault injection into kernel code, kernel stack, and system registers with the NFTAPE framework while running the Apache Web server as a workload. The error behavior comparison shows that the Linux system has the highest average crash latency, the Solaris system has the highest hang rate, and the AIX system has the lowest error sensitivity and the least amount of crashes in the more severe categories. Daniel Chen 0001, Gabriela Jacques-Silva, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Bruce G. Mealey |
PRDC | 4 |
| 2008 | Formalizing System Behavior for Evaluating a System Hang DetectorabstractThis paper presents an approach to formally verify the detection capability of a system hang detector. To achieve this goal, an abstract formal model of a typical Linux system is created to thoroughly exercise all execution scenarios that may lead to hangs. The goal is to expose cases (i.e., hang scenarios) that escape detection. Our system model abstracts the basic hardware (e.g., timer, hardware counter) and software (e.g., processes/threads) components present in the Linux system. The model enables: (i) capturing behavior of these components so as to depict execution scenarios that lead to hangs, and (ii) evaluating hang detection coverage. Explicit-state model checking is applied to reason about system behavior and uncover hang scenarios that escape detection. The results indicate that the proposed framework allows identification of corner cases of hang scenarios that escape detection and provides valuable insight to developers for enhancing detection mechanisms. Long Wang 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SRDS | 3 |
| 2007 | Workshop on Dependable and Secure NanocomputingabstractThe continuous advances and progress made in hardware technology makes it possible to foresee a realm of unprecedented performance levels and new application-driven architectural designs, as evidenced by the recent announcement of a 80-core chip [1]. Nevertheless, the evolution of nanotechnologies raises serious challenges with respect to both dependability and security viewpoints. Issues at stake go far beyond developing protections with respect to accidental disturbances in operation, they also relate to the unreliability and variability that will characterize emerging nanoscale devices. Accounting for malicious threats targeting hardware circuits will constitute another increasing concern. Jean Arlat, Ravishankar K. Iyer, Michael Nicolaidis |
DSN | 2 |
| 2007 | How Do Mobile Phones Fail? A Failure Data Analysis of Symbian OS Smart PhonesabstractWhile the new generation of hand-held devices, e.g., smart phones, support a rich set of applications, growing complexity of the hardware and runtime environment makes the devices susceptible to accidental errors and malicious attacks. Despite these concerns, very few studies have looked into the dependability of mobile phones. This paper presents measurement-based failure characterization of mobile phones. The analysis starts with a high level failure characterization of mobile phones based on data from publicly available web forums, where users post information on their experiences in using hand-held devices. This initial analysis is then used to guide the development of a failure data logger for collecting failure-related information on SymbianOS-based smart phones. Failure data is collected from 25 phones (in Italy and USA) over the period of 14 months. Key findings indicate that: (i) the majority of kernel exceptions are due to memory access violation errors (56%) and heap management problems (18%), and (ii) on average users experience a failure (freeze or self shutdown) every 11 days. While the study provide valuable insight into the failure sensitivity of smart-phones, more data and further analysis are needed before generalizing the results. Marcello Cinque, Domenico Cotroneo, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 4 |
| 2007 | Processor-Level Selective ReplicationabstractWe propose a processor-level technique called selective replication, by which the application can choose where in its application stream and to what degree it requires replication. Recent work on static analysis and fault-injection-based experiments on applications reveals that certain variables in the application are critical to its crash- and hang-free execution. If it can be ensured that only the computation of these variables is error-free, then a high degree of crash/hang coverage can be achieved at a low performance overhead to the application. The selective replication technique provides an ideal platform for validating this claim. The technique is compared against complete duplication as provided in current architecture-level techniques. The results show that with about 59% less overhead than full duplication, selective replication detects 97% of the data errors and 87% of the instruction errors that were covered by full duplication. It also reduces the detection of errors benign to the final outcome of the application by 17.8% as compared to full duplication. Nithin Nakka, Karthik Pattabiraman, Ravishankar K. Iyer |
DSN | 3 |
| 2007 | Automated Derivation of Application-aware Error Detectors using Static AnalysisabstractThis paper presents a technique to derive and implement error detectors to protect an application from data errors. The error detectors are derived automatically using compiler-based static analysis from the backward program slice of critical variables in the program. Critical variables are defined as those that are highly sensitive to errors, and deriving error detectors for these variables provides high coverage for errors in any data value used in the program. The error detectors take the form of checking expressions and are optimized for each control flow path followed at runtime. The derived detectors are implemented using a combination of hardware and software. Experiments show that the derived detectors incur low performance overheads while achieving high detection coverage for errors that impact the application. Karthik Pattabiraman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
IOLTS | 3 |
| 2007 | Editorial: Dependability and Security
Ravishankar K. Iyer |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2007 | Inner-Circle Consistency for Wireless Ad Hoc NetworksabstractThis paper proposes and evaluates strategies to build reliable and secure wireless ad hoc networks. Our contribution is based on the notion of inner-circle consistency, where local node interaction is used to neutralize errors/attacks at the source, both preventing errors/attacks from propagating in the network and improving the fidelity of the propagated information. We achieve this goal by combining statistical (a proposed fault-tolerant duster algorithm) and security (threshold cryptography) techniques with application-aware checks to exploit the data/computation that is partially and naturally replicated in wireless applications. We have prototyped an inner-circle framework and used it to demonstrate the idea of inner-circle consistency in two significant wireless scenarios: 1) the neutralization of black hole attacks in AODV networks and 2) the neutralization of sensor errors in a target detection/ localization application executed over a wireless sensor network Claudio Basile, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
IEEE Trans. Mob. Comput. | 3 |
| 2007 | Reliability MicroKernel: Providing Application-Aware Reliability in the OSabstractThis paper describes the reliability MicroKernel (RMK) framework, a loadable kernel module (or a device driver) for providing application-aware reliability, and dynamically configuring reliability mechanisms. Characteristics of application/system execution are exploited transparently through application-aware reliability techniques to achieve low-latency detection, and low-overhead checkpointing. The RMK prototype is implemented in both Linux, and Windows; and it supports detection of application/OS failures, and transparent application checkpointing. Experiment results show that the system hang detection and application hang detection, which exploit characteristics of application, and system behavior, can achieve high coverage (100% observed in our experiments) with a low false positive rate. Moreover, the performance overhead of RMK, and its detection/checkpointing mechanisms, is small: 0.6% for application hang detection, and 0.1% for transparent application checkpointing in the experiments. Long Wang 0003, Zbigniew T. Kalbarczyk, Weining Gu, Ravishankar K. Iyer |
IEEE Trans. Reliab. | 4 |
| 2006 | Security vulnerabilities: from measurements to designabstractThis paper presents a study that uses extensive analysis of real security vulnerabilities to drive the development of: 1) runtime techniques for detection/masking of security attacks and 2) formal source code analysis methods to enable identification and removal of potential security vulnerabilities. The presentation will describe the hardware architecture of a Reliability and Security Engine (RSE) that embodies the proposed techniques, to provide run-time checking at the processor level. Ravishankar K. Iyer |
AsiaCCS | 1 |
| 2006 | An Approach for Detecting and Distinguishing Errors versus Attacks in Sensor NetworksabstractDistributed sensor networks are highly prone to accidental errors and malicious activities, owing to their limited resources and tight interaction with the environment. Yet only a few studies have analyzed and coped with the effects of corrupted sensor data. This paper contributes with the proposal of an on-the-fly statistical technique that can detect and distinguish faulty data from malicious data in a distributed sensor network. Detecting faults and attacks is essential to ensure the correct semantic of the network, while distinguishing faults from attacks is necessary to initiate a correct recovery action. The approach uses hidden Markov models (HMMs) to capture the error/attack-free dynamics of the environment and the dynamics of error/attack data. It then performs a structural analysis of these HMMs to determine the type of error/attack affecting sensor observations. The methodology is demonstrated with real data traces collected over one month of observation from motes deployed on the Great Duck Island Claudio Basile, Meeta Gupta, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 4 |
| 2006 | Application-Aware Reliability and Security: The Trusted ILLIAC ApproachabstractTrusted ILLIAC is a reliable and secure cluster-computing platform being built at the University of Illinois Coordinated Science Laboratory (CSL) and Information Trust Institute (ITI), involving faculty from Electrical and Computer Engineering and Computer Science Departments. Trusted ILLIAC is intended to be a large, demonstrably trustworthy cluster-computing system to support what is variously referred to as on-demand/utility computing or adaptive enterprise computing. Such systems require that a significant number of applications co-exist and share hardware/software resources using a variety of containment boundaries. Current solutions aim at providing hardware and software solutions that can only be described as one-size-fits-all approaches. Today's environments are complex, expensive to implement, and nearly impossible to validate. The challenge is to provide an application-specific level of reliability and security in a totally transparent manner, while delivering optimal performance. A promising approach lies in developing a new set of application-aware methods that provide customized levels of trust (specified by the application) enforced using an integrated approach involving reprogrammable hardware, enhanced compiler methods to extract security and reliability properties, and the support of configurable operating system and middleware. Our approach is to demonstrate such a set of integrated techniques that span entire system hierarchy: processor hardware, operating system, middleware, and application Ravishankar K. Iyer |
NCA | 1 |
| 2006 | An OS-level Framework for Providing Application-Aware ReliabilityabstractThe paper describes the reliability microkernel framework (RMK), a loadable kernel module for providing application-aware reliability and dynamically configuring reliability mechanisms installed in RMK. The RMK prototype is implemented in Linux and supports detection of application/OS failures and transparent application checkpointing. Experiment results show that the OS hang detection, which exploits characteristics of application and system behavior, can achieve high coverage (100% in our experiments) and low false positive rate. Moreover, the performance overhead is negligible because instruction counting is performed in hardware Long Wang 0003, Zbigniew T. Kalbarczyk, Weining Gu, Ravishankar K. Iyer |
PRDC | 4 |
| 2006 | Security Vulnerabilities: From Analysis to Detection and Masking TechniquesabstractThis paper presents a study that uses extensive analysis of real security vulnerabilities to drive the development of: 1) runtime techniques for detection/masking of security attacks and 2) formal source code analysis methods to enable identification and removal of potential security vulnerabilities. A finite-state machine (FSM) approach is employed to decompose programs into multiple elementary activities, making it possible to extract simple predicates to be ensured for security. The FSM analysis pinpoints common characteristics among a broad range of security vulnerabilities: predictable memory layout, unprotected control data, and pointer taintedness. We propose memory layout randomization and control data randomization to mask the vulnerabilities at runtime. We also propose a static analysis approach to detect potential security vulnerabilities using the notion of pointer taintedness. Shuo Chen 0001, Jun Xu 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
Proc. IEEE | 4 |
| 2006 | Editorial: Dependability and Security--Looking Foward to 2006abstractDURING its fi rst two years, TDSC mapped out the challenges and new directions the journal would pursue. With this issue, the fi rst of its third year, I see the journal as continuing to evolve and beginning to gain a high profi le in the dependability and security arenas. A brief overview of this issue’s offerings shows the range and quality of this up-and-coming journal. “From Set Membership to Group Membership: A Separation of Concerns,” by Andre Schiper and Sam Toueg, presents a new way of looking at group membership, giving a simple and succinct specifi cation of the problem and outlining a simple implementation approach based on the state machine paradigm. In “Achieving Privacy in Trust Negotiations with an Ontology-Based Approach,” Anna Squicciarini et al. introduce the notion of privacy-preserving disclosure, that is, a set that does not include attributes, credentials, or combinations of these that may compromise privacy; they propose two techniques based on the notions of substitution and generalization to obtain privacy-preserving disclosure. “An Active Splitter Architecture for Intrusion Detection and Prevention,” by Kostas Xinidis et al., argues that rather than just passively providing generic load distribution, traffi c splitters should implement active operations on the traffi c stream with the goal of reducing the load on the sensors. Gal Badishi et al.’s “Exposing and Eliminating Vulnerabilities to Denial of Service Attacks in Secure Gossip-Based Multicast” proposes a framework and methodology for quantifying the effect of denial-of-service attacks on a distributed system. “A Key Predistribution Scheme for Sensor Networks Using Deployment Knowledge,” by Wenliang Du et al., shows that the performance of sensor networks can be substantially improved using a scheme that recognizes the importance of node-deployment knowledge. “Install-Time Vaccination of Windows Executables to Defend against Stack Smashing Attacks,” by Danny Nebenzahl et al., demonstrates a scheme to detect and vaccinate against smashing attacks, with promising performance results. In Shyh-Yih Wang and Chi-Sung Laih’s “Merging: An Effi cient Solution for a Time-Bound Hierarchical Key Assignment Scheme,” we see that some problems usually addressed by conventional key assignment schemes can be solved directly via merging, with better performance. The past year saw the fi rst special issue of TDSC. The April-June issue featured a special section containing three selected papers from the 2005 IEEE Symposium on Security and Privacy, which took place in Oakland, California. So that the dissemination of these papers would not be held up by publishing schedules, preprints were made available at the conference, where they were well received. We expect to build on this success by publishing another special issue containing papers from the 2006 conference. We are also working on a special issue based on the 2005 International Conference on Dependable Systems and Networks, for which the conference program chairs, in conjunction with TDSC Associate Editors Drs. Arlat and Verissimo, are currently selecting the best papers. The journal plans to continue publishing special issues based on these two premier conferences and to start looking into emerging areas to address via special issues. A statistic of interest is that during the past year TDSC received a total of 179 manuscripts. Six have been received so far in 2006. Last year, we published 39 papers, with an average time from submission to decision of 110 days, a decrease of 28 days from last year. We continue to strive to keep the decision-making time as short as possible, while attempting to attract high-quality submissions. On this second anniversary, I wish to thank all of our 2005 associate editors, submitters, authors, and reviewers, each of whom make the growing success of TDSC possible. I also wish to thank the wonderful publications staff at the IEEE Computer Society Publications offi ce, in particular Suzanne Werner, Selina Flynn, and Joyce Arnold, whose support and dedication have helped us through our second year. Here, at the University of Illinois, special thanks go to Tammi O’Neill, Gerssimoula Kokkosis, and Heidi Leerkamp for providing exceptional support during this past year. Most importantly, my thanks go to the peer community at large—with your support TDSC goes forth into its third year with confi dence and enthusiasm. Ravishankar K. Iyer |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2006 | Active Replication of Multithreaded ApplicationsabstractSoftware-based active replication is expensive in terms of performance overhead. Multithreading can help improve performance; however, thread scheduling is a source of nondeterminism in replica behavior. To achieve strong replica consistency in multithreaded environments, this paper proposes intercepting mutex lock/unlock operations performed by threads on accessing the shared data and contributes with two algorithmic solutions: 1) a loose synchronization algorithm (LSA), which captures the natural concurrency in a leader replica and projects it on follower replicas through interreplica communication, and 2) a preemptive deterministic scheduler (PDS) algorithm, which removes the need for interreplica communication through the notion of round and by suspending threads when it is unable (yet) to schedule them deterministically. Failure behavior and performance of LSA and PDS implementations are evaluated in a triplicated system and compared with existing solutions. A performance evaluation indicates that LSA and PDS outperform existing solutions, with PDS offering lower throughput than LSA. A fault-injection campaign shows that PDS is more robust to errors due to the absence of interreplica communication. Hence, LSA and PDS represent a trade-off between performance and dependability. Finally, LSA and PDS are demonstrated in replicating the Apache Web server, a substantial real-world application. Claudio Basile, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2005 | Neutralization of Errors and Attacks in Wireless Ad Hoc NetworksabstractThis paper proposes and evaluates strategies to build reliable and secure wireless ad hoc networks. Our contribution is based on the notion of inner-circle consistency, where local node interaction is used to neutralize errors/attacks at the source, both preventing errors/attacks from propagating in the network and improving the fidelity of the propagated information. We achieve this goal by combining statistical (a proposed fault-tolerant cluster algorithm) and security (threshold cryptography) techniques with application-aware checks to exploit the data/computation that is partially and naturally replicated in wireless applications. We have prototyped an inner-circle framework with the ns-2 network simulator, and we use it to demonstrate the idea of inner-circle consistency in two significant wireless scenarios: (1) the neutralization of black hole attacks in AODV networks and (2) the neutralization of sensor errors in a target detection/localization application executed over a wireless sensor network. Claudio Basile, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 3 |
| 2005 | Defeating Memory Corruption Attacks via Pointer Taintedness DetectionabstractMost malicious attacks compromise system security through memory corruption exploits. Recently proposed techniques attempt to defeat these attacks by protecting program control data. We have constructed a new class of attacks that can compromise network applications without tampering with any control data. These non-control data attacks represent a new challenge to system security. In this paper, we propose an architectural technique to defeat both control data and non-control data attacks based on the notion of pointer taintedness. A pointer is said to be tainted if user input can be used as the pointer value. A security attack is detected whenever a tainted value is dereferenced during program execution. The proposed architecture is implemented on the SimpleScalar processor simulator and is evaluated using synthetic programs as well as real-world network applications. Our technique can effectively detect both control data and non-control data attacks, and it offers better security coverage than current methods. The proposed architecture is transparent to existing programs. Shuo Chen 0001, Jun Xu 0003, Nithin Nakka, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 5 |
| 2005 | Microprocessor Sensitivity to Failures: Control vs Execution and Combinational vs Sequential LogicabstractThe goal of this study is to characterize the impact of soft errors on embedded processors. We focus on control versus speculation logic on one hand, and combinational versus sequential logic on the other. The target system is a gate-level implementation of a DLX-like processor. The synthesized design is simulated, and transients are injected to stress the processor while it is executing selected applications. Analysis of the collected data shows that fault sensitivity of the combinational logic (4.2% for a fault duration of one clock cycle) is not negligible, even though it is smaller than the fault sensitivity of flip-flops (10.4%). Detailed study of the error impact, measured at the application level, reveals that errors in speculation and control blocks collectively contribute to about 34% of crashes, 34% of fail-silent violations and 69% of application incomplete executions. These figures indicate the increasing need for processor-level detection techniques over generic methods, such as ECC and parity, to prevent such errors from propagating beyond the processor boundaries. Giacinto Paolo Saggese, Anoop Vetteth, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 4 |
| 2005 | Modeling Coordinated Checkpointing for Large-Scale SupercomputersabstractCurrent supercomputing systems consisting of thousands of nodes cannot meet the demands of emerging high-performance scientific applications. As a result, a new generation of supercomputing systems consisting of hundreds of thousands of nodes is being proposed. However, these systems are likely to experience far more frequent failures than today's systems, and such failures must be tackled effectively. Coordinated checkpointing is a common technique to deal with failures in supercomputers. This paper presents a model of a coordinated checkpointing protocol for large-scale supercomputers, and studies its scalability by considering both the coordination overhead and the effect of failures. Unlike most of the existing checkpointing models, the proposed model takes into account failures during checkpointing and recovery, as well as correlated failures. Stochastic activity networks (SANs) are used to model the system, and the model is simulated to study the scalability, reliability, and performance of the system. Long Wang 0003, Karthik Pattabiraman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Lawrence G. Votta, Christopher A. Vick, Alan Wood |
DSN | 4 |
| 2005 | Assessing the Crash-Failure Assumption of Group Communication ProtocolsabstractDesigning and correctly implementing group communication systems (GCSs) is notoriously difficult. Assuming that processes fail only by crashing provides a powerful means to simplify the theoretical development of these systems. When making this assumption, however, one should not forget that clean crash failures provide only a coarse approximation of the effects that errors can have in distributed systems. Ignoring such a discrepancy can lead to complex GCS-based applications that pay a large price in terms of performance overhead yet fail to deliver the promised level of dependability. This paper provides a thorough study of error effects in real systems by demonstrating an error-injection-driven design methodology, where error injection is integrated in the core steps of the design process of a robust fault-tolerant system. The methodology is demonstrated for the Fortika toolkit, a Java-based GCS. Error injection enables us to uncover subtle reliability bottlenecks both in the design of Fortika and in the implementation of Java. Based on the obtained insights, we enhance Fortika's design to reduce the identified bottlenecks. Finally, a comparison of the results obtained for Fortika with the results obtained for the OCAML-based Ensemble system in a previous work, allows us to investigate the reliability implications that the choice of the development platform (Java versus OCAML) can have Sergio Mena, Claudio Basile, Zbigniew T. Kalbarczyk, André Schiper, Ravishankar K. Iyer |
ISSRE | 5 |
| 2005 | Application-Based Metrics for Strategic Placement of DetectorsabstractThe goal of this paper is to provide low-latency detection and prevent error propagation due to value errors. This paper introduces metrics to guide the strategic placement of detectors and evaluates (using fault injection) the coverage provided by ideal detectors embedded at program locations selected using the computed metrics. The computation is represented in the form of a dynamic dependence graph (DDG), a directed-acyclic graph that captures the dynamic dependencies among the values produced during the course of program execution. The DDG is employed to model error propagation in the program and to derive metrics (e.g., value fanout or lifetime) for detector placement. The coverage of the detectors placed is evaluated using fault injections in real programs, including two large SPEC95 integer benchmarks fgcc and perl). Results show that a small number of detectors, strategically placed, can achieve a high degree of detection coverage. Karthik Pattabiraman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
PRDC | 3 |
| 2005 | Editorial: State of the Journal Address
Ravishankar K. Iyer |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2004 | Design and networking considerations for FACEabstractAs we step into the twenty first century, we are surrounded by technology wearing its way into our day-to-day life. making every aspect of it more convenient and organized. Yet, in every household there is one front, i.e. the process of cooking, which has not been sufficiently automated. Even in a commercial environment (restaurant kitchens), the process of cooking remains largely manual. We investigate the potential for a fully automated cooking environment (FACE). We start by describing the architecture of FACE and present an overview of its components and their functionality. Then, we explore communication considerations between various FACE components and discuss existing solutions for addressing the basic networking needs. We also provide a detailed example to describe the workflow in FACE and help illustrate the communication requirements between FACE components. Ultimately, our aim is to provide the vision for a flexible, highly integrated and fully coordinated cooking environment. Anu Tewari, Ravishankar K. Iyer, Vijay Tewari |
CCNC | 2 |
| 2004 | Hierarchical application aware error detection and recoveryabstractProposed is a four-tired approach to develop and integrate detection and recovery support at different levels of the system hierarchy. The proposed mechanisms exploit support provided by (i) embedded hardware, (ii) operating system, (iii) compiler, and (iv) application. Ravishankar K. Iyer |
DAC | 1 |
| 2004 | Error Sensitivity of the Linux Kernel Executing on PowerPC G4 and Pentium 4 ProcessorsabstractThe goals of this study are: (i) to compare Linux kernel (2.4.22) behavior under a broad range of errors on two target processors - the Intel Pentium 4 (P4) running RedHat Linux 9.0 and the Motorola PowerPC (G4) running YellowDog Linux 3.0 - and (ii) to understand how architectural characteristics of the target processors impact the error sensitivity of the operating system. Extensive error injection experiments involving over 115,000 faults/errors are conducted targeting the kernel code, data, stack, and CPU system registers. Analysis of the obtained data indicates significant differences between the two platforms in how errors manifest and how they are detected in the hardware and the operating system. In addition to quantifying the observed differences and similarities, the paper provides several examples to support the insights gained from this research. Weining Gu, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 3 |
| 2004 | An Architectural Framework for Providing Reliability and Security SupportabstractThis paper explores hardware-implemented error-detection and security mechanisms embedded as modules in a hardware-level framework called the reliability and security engine (RSE), which is implemented as an integral part of a modern microprocessor. The RSE interacts with the processor through an input/output interface. The CHECK instruction, a special extension of the instruction set architecture of the processor, is the interface of the application with the RSE. The detection mechanisms described here in detail are: (I) the memory layout randomization (MLR) module, which randomizes the memory layout of a process in order to foil attackers who assume a fixed system layout, (2) the data dependency tracking (DDT) module, which tracks the dependencies among threads of a process and maintains checkpoints of shared memory pages in order to rollback the threads when an offending (potentially malicious) thread is terminated, and (3) the instruction checker module (ICM), which checks an instruction for its validity or the control-flow of the program just as the instruction enters the pipeline for execution. Performance simulations for the studied modules indicate low overhead of the proposed solutions. Nithin Nakka, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Jun Xu 0003 |
DSN | 3 |
| 2004 | Checkpointing of Control Structures in Main Memory Database SystemsabstractThis paper proposes an application-transparent, low-overhead checkpointing strategy for maintaining consistency of control structures in a commercial main memory database (MMDB) system, based on the ARMOR (adaptive reconfigurable mobile object of reliability) infrastructure. Performance measurements and availability estimates show that the proposed checkpointing scheme significantly enhances database availability (an extra nine in improvement compared with major-recovery-based solutions) while incurring only a small performance overhead (less than 2% in a typical workload of real applications). Long Wang 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, H. Vora, T. Chahande |
DSN | 3 |
| 2004 | Formal Reasoning of Various Categories of Widely Exploited Security Vulnerabilities by Pointer Taintedness SemanticsabstractThis paper is motivated by a low level analysis of various categories of severe security vulnerabilities, which indicates that a common characteristic of many classes of vulnerabilities is pointer taintedness. A pointer is said to be tainted if a user input can directly or indirectly be used as a pointer value. In order to reason about pointer taintedness, a memory model is needed. The main contribution of this paper is the formal definition of a memory model using equational logic, which is used to reason about pointer taintedness. The reasoning is applied to several library functions to extract security preconditions, which must be satisfied to eliminate the possibility of pointer taintedness. The results show that pointer taintedness analysis can expose different classes of security vulnerabilities, such as format string, heap corruption and buffer overflow vulnerabilities, leading us to believe that pointer taintedness provides a unifying perspective for reasoning about security vulnerabilities. Shuo Chen 0001, Karthik Pattabiraman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SEC | 4 |
| 2004 | Hardware Support for High Performance, Intrusion- and Fault-Tolerant SystemsabstractThe paper proposes a combined hardware/software approach for realizing high performance, intrusion- and fault-tolerant services. The approach is demonstrated for (yet not limited to) an attribute authority server, which provides a compelling application due to its stringent performance and security requirements. The key element of the proposed architecture is an FPGA-based, parallel crypto-engine providing (1) optimally dimensioned RSA Processors for efficient execution of computationally intensive RSA signatures and (2) a KeyStore facility used as tamper-resistant storage for preserving secret keys. To achieve linear speed-up (with the number of RSA Processors) and deadlock-free execution in spite of resource-sharing and scheduling/synchronization issues, we have resorted to a number of performance enhancing techniques (e.g., use of different clock domains, optimal balance between internal and external parallelism) and have formally modeled and mechanically proved our crypto-engine with the Spin model checker. At the software level, the architecture combines active replication and threshold cryptography, but in contrast with previous work, the code of our replicas is multithreaded so it can efficiently use an attached parallel crypto-engine to compute an attribute authority partial signature (as required by threshold cryptography). Resulting replicated systems that exhibit nondeterministic behavior, which cannot be handled with conventional replication approaches. Our architecture is based on a preemptive deterministic scheduling algorithm to govern scheduling of replica threads and guarantee strong replica consistency. Giacinto Paolo Saggese, Claudio Basile, Luigi Romano, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SRDS | 5 |
| 2004 | Modeling and evaluating the security threats of transient errors in firewall software
Shuo Chen 0001, Jun Xu 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Keith Whisnant |
Perform. Evaluation | 4 |
| 2004 | The Effects of an ARMOR-Based SIFT Environment on the Performance and Dependability of User ApplicationsabstractFew, distributed software-implemented fault tolerance (SIFT) environments have been experimentally evaluated using substantial applications to show that they protect both themselves and the applications from errors. We present an experimental evaluation of a SIFT environment used to oversee spaceborne applications as part of the Remote Exploration and Experimentation (REE) program at the Jet Propulsion Laboratory. The SIFT environment is built around a set of self-checking ARMOR processes running on different machines that provide error detection and recovery services to themselves and to the REE applications. An evaluation methodology is presented in which over 28,000 errors were injected into both the SIFT processes and two representative REE applications. The experiments were split into three groups of error injections, with each group successively stressing the SIFT error detection and recovery more than the previous group. The results show that the SIFT environment added negligible overhead to the application's execution time during failure-free runs. Correlated failures affecting a SIFT process and application process are possible, but the division of detection and recovery responsibilities in the SIFT environment allows it to recover from these multiple failure scenarios. Only 28 cases were observed in which either the application failed to start or the SIFT environment failed to recognize that the application had completed. Further investigations showed that assertions within the SIFT processes-coupled with object-based incremental checkpointing-were effective in preventing system failures by protecting dynamic data within the SIFT processes. Keith Whisnant, Ravishankar K. Iyer, Zbigniew T. Kalbarczyk, Phillip H. Jones, David A. Rennels, Raphael R. Some |
IEEE Trans. Software Eng. | 2 |
| 2003 | A Preemptive Deterministic Scheduling Algorithm for Multithreaded ReplicasabstractSoftware-based active replication is expensive in terms of performance overhead. Multithreading can help improve performance; however, thread scheduling is a source of nondeterminism in replica behavior. This paper presents a Preemptive Deterministic Scheduling (PDS) algorithm for ensuring deterministic replica behavior while preserving concurrency. Threads are synchronized only on updates to the shared state. A replica execution is broken into a sequence of rounds and in a round each thread can acquire up to two mutexes. If a thread cannot acquire a mutex it requests, then it checks if all other threads are suspended. If so, the thread fires a new round; otherwise, the thread is suspended. When a new round fires, all threads' mutex requests are known; thus, it is possible to form a deterministic scheduling of mutex acquisitions in the round. No inter-replica communication is required. The algorithm is formally specified, and the proposed formalism is used to prove its correctness. Failure Claudio Basile, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 3 |
| 2003 | A Data-Driven Finite State Machine Model for Analyzing Security VulnerabilitiesabstractThis paper combines an analysis of data on security vulnerabilities (published in Bugtraq database) and a focused source-code examination to develop a finite state machine (FSM) model to depict and reason about security vulnerabilities. An in-depth analysis of the vulnerability reports and the corresponding source code of the applications leads to three observations: (i) exploits must pass through multiple elementary activities, (ii) multiple vulnerable operations on several objects are involved in exploiting a vulnerability, and (iii) the vulnerability data and corresponding code inspections allow us to derive a predicate for each elementary activity. Each predicate is represented as a primitive FSM (pFSM). Multiple pFSMs are then combined to create an FSM model of vulnerable operations and possible exploits. The proposed FSM methodology is exemplified by analyzing several types of vulnerabilities reported in the data: stack buffer overflow, integer overflow, heap overflow, input validation vulnerabilities, and format string vulnerabilities. For the studied vulnerabilities, we identify three types of pFSMs, which can be used to analyze operations involved in exploiting vulnerabilities and to identify the security checks to be performed at the elementary activity level. A demonstration of the practical usefulness of the FSM modeling approach was the discovery of a new heap overflow vulnerability now published in Bugtraq. Key words: security vulnerabilities, data analysis, finite state machine modeling. 1. Shuo Chen 0001, Zbigniew T. Kalbarczyk, Jun Xu 0003, Ravishankar K. Iyer |
DSN | 4 |
| 2003 | Characterization of Linux Kernel Behavior under ErrorsabstractThis paper describes an experimental study of Linux kernel behavior in the presence of errors that impact the instruction stream of the kernel code. Extensive error injection experiments including over 35,000 errors are conducted targeting the most fre- quently used functions in the selected kernel subsystems. Three types of faults/errors injection campaigns are conducted: (1) ran- dom non-branch instruction, (2) random conditional branch, and (3) valid but incorrect branch. The analysis of the obtained data shows: (i) 95% of the crashes are due to four major causes, namely, unable to handle kernel NULL pointer, unable to handle kernel paging request, invalid opcode, and general protection fault, (ii) less than 10% of the crashes are associated with fault propagation and nearly 40% of crash latencies are within 10 cycles, (iii) errors in the kernel can result in crashes that require reformatting the file system to restore system operation; the process of bringing up the system can take nearly an hour. Subsequently, over 35,000 faults/errors are injected into the kernel functions within four subsystems: architecture- dependent code (arch), virtual file system interface (fs), cen- tral section of the kernel (kernel), and memory management (mm). Three types of fault/error injection campaigns are con- ducted: random non-branch, random conditional branch, and valid but incorrect conditional branch. The data is analyzed to quantify the response of the OS as a whole based on the sub- system and to determine which functions are responsible for error sensitivity. The analysis provides a detailed insight into the OS behavior under faults/errors. The major findings in- clude: • Most crashes (95%) are due to four major causes: unable to handle kernel NULL pointer, unable to handle kernel paging request, invalid opcode, and general protection fault. • Nine errors in the kernel result in crashes (most severe crash category), which require reformatting the file system. The process of bringing up the system can take nearly an hour. • Less than 10% of the crashes are associated with fault propagation, and nearly 40% of crash latencies are within 10 cycles. The closer analysis of the propagation patterns indicates that it is feasible to identify strategic locations for embedding additional assertions in the source code of a given subsystem to detect errors and, hence, to prevent er- ror propagation. Weining Gu, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Zhen-Yu Yang |
DSN | 3 |
| 2003 | Design and Performance of Compressed Interconnects for High Performance ServersabstractAs microprocessors scale rapidly in frequency, the design of fast and efficient interconnects becomes extremely important for low latency data access and high performance. We evaluate a technique for reducing the interconnect width by exploiting the spatial and temporal locality in communication transfers (addresses & data). The width reduction implies a number of other advantages including higher operating frequency, reduced pin-count, lower chip & board cost, etc. We evaluate the effectiveness of the proposed scheme by performing trace-driven simulations for two well-known commercial server workloads (SPECWeb99 and TPC-C). We also study the sensitivity of the compression hit ratio with respect to the number of bits compressed, size of the encoding/decoding table used and the replacement policy. The results indicate that the proposed technique has a potential to reduce address bus width in most cases and data bus widths in some cases while maintaining equal or better performance than in the uncompressed case. Krishna Kant 0001, Ravishankar K. Iyer |
ICCD | 2 |
| 2003 | Error-Injection-Based Failure Characterization of the IEEE 1394 BusabstractThis paper investigates the behavior of the IEEE 1394 bus in the presence of transient errors in the hardware layers of the protocol. Software-implemented error injection is used to introduce errors into the internals of the 1394 bus hardware chipset. Results from this study indicate that the IEEE 1394 bus protocol provides robust network communication in the presence of single-bit errors in the chipset. D. J. Beauregard, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Savio N. Chau, Leon Alkalai |
IOLTS | 3 |
| 2003 | Performance implications of chipset caches in web serversabstractAs Internet usage continues to expand rapidly, careful attention needs to be paid to the design of Internet servers for achieving high performance and end-user satisfaction. In this paper, with the aim of improving memory system performance of Internet servers, we propose and evaluate various design alternatives for "chipset caches", a shared cache layer embedded within a server chipset. Using our trace-based cache simulation framework (CASPER) and SPECweb99 as a representative workload for web servers, we present the performance implications of chipset caches in a front-end dual-processor web server. We start by analyzing the improvement gained by caching the data from processor-initiated requests alone. We study the sensitivity to basic cache parameters (such as cache size and associativity) and also study the impact of prefetching into the chipset cache. We then present the performance implications of routing memory requests initiated by I/O devices through the chipset cache. Finally, we also study the implications of making the chipset cache inclusive. Based on detailed simulation data and its implications on system level performance, this paper shows that chipset caches have significant potential for future Internet servers. Ravishankar K. Iyer |
ISPASS | 1 |
| 2003 | Group Communication Protocols under ErrorsabstractGroup communication protocols constitute a basic building block for highly dependable distributed applications. Designing and correctly implementing a group communication system (GCS) is a difficult task. While many theoretical algorithms have been formalized and proved for correctness, only few research projects have experimentally assessed the dependability of GCS implementations under complex error scenarios. This paper describes a thorough error-injection experimental campaign conducted on Ensemble, a popular GCS. By employing synthetic benchmark applications, we stress selected components of the GCS $the group membership service, the FIFO-ordered reliable multicast - under various error models, including errors in the memory (text and heap segments) and in the network messages. The data show that about 5-6% of the failures are due to an error escaping Ensemble's error-containment mechanism and manifesting as a fail silence violation. This constitutes an impediment to achieving high dependability, the natural objective of GCSs. Our results are derived for a particular system (Ensemble), and more investigation involving other GCSs is required to generalize the conclusions. Nevertheless, through an accurate analysis of the failure causes and the error propagation patterns, this paper offers insights into the design and the implementation of robust GCSs. Claudio Basile, Long Wang 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SRDS | 4 |
| 2003 | Transparent Runtime Randomization for SecurityabstractA large class of security attacks exploit software implementation vulnerabilities such as unchecked buffers. This paper proposes transparent runtime randomization (TRR), a generalized approach for protecting against a wide range of security attacks. TRR dynamically and randomly relocates a program's stack, heap, shared libraries, and parts of its runtime control data structures inside the application memory address space. Making a program's memory layout different each time it runs foils the attacker's assumptions about the memory layout of the vulnerable program and makes the determination of critical address values difficult if not impossible. TRR is implemented by changing the Linux dynamic program loader, hence it is transparent to applications. We demonstrate that TRR is effective in defeating real security attacks, including malloc-based heap overflow, integer overflow, and double-free attacks, for which effective prevention mechanisms are yet to emerge. Furthermore, TRR incurs less than 9% program startup overhead and no runtime overhead. Jun Xu 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SRDS | 3 |
| 2002 | A Framework for Classifying Peer-to-Peer TechnologiesabstractPopularized by Napster and Gnutella file sharing solutions, peer-to-peer (P2P) computing has suddenly emerged at the forefront of Internet computing. The basic notion of cooperative computing and resource sharing has been around for quite some time, although these new applications have opened up possibilities of very flexible web-based information sharing. This article provides a frame-work for classifying current and future P2P technologies. The main motivation for the classification is to identify basic characteristics of P2P applications so that the infrastructure to support P2P computing can concentrate on these basic characteristics. Krishna Kant 0001, Ravishankar K. Iyer, Vijay Tewari |
CCGRID | 2 |
| 2002 | Evaluating the Security Threat of Firewall Data Corruption Caused by Transient ErrorsabstractThis paper experimentally evaluates and models the error-caused security vulnerabilities and the resulting security violations of two Linux kernel firewalls: IPChains and Netfilter. There are two major aspects to this work: to conduct extensive error injection experiments on the Linux kernel and to quantify the possibility of error-caused security violations using a SAN (Stochastic Activity Network) model. The error injection experiments show that about 2% of errors injected into the firewall code segment cause security vulnerabilities. Two types of error-caused security vulnerabilities are distinguished: temporary, which disappear when the error disappears, and permanent, which persist even after the error is removed, as long as the system is not rebooted. Results from simulating the SAN model indicate that under an error rate of 0.1 error/day during a 1-year period in a networked system protected by 20 firewalls, 2 machines (on the average) will experience security violations. This indicates that error-caused security vulnerabilities can be a non-negligible source of a security threats to a highly secure system. Shuo Chen 0001, Jun Xu 0003, Ravishankar K. Iyer, Keith Whisnant |
DSN | 3 |
| 2002 | An Adaptive Architecture for Monitoring and Failure Analysis of High-Speed NetworksabstractDescribes the design of a reconfigurable device using an FPGA (field programmable gate array) whose primary function is high-speed (several Gb/s) network data monitoring and run-time adaptive fault injection and statistics gathering for failure analysis. The device is designed for two types of media: Myrinet SAN and Fibre Channel, and failure analysis can be performed simultaneously over both of these networks. Although the device intercepts and retransmits signals on the network, no impact on the data transfer rate is observed and the latency caused by inserting the device in the network is negligible. The fault injection capabilities are demonstrated on a Myrinet LAN. Fault injection experiments are conducted on data transmitted across the network, including control packets previously inaccessible to software-based techniques. Benjamin Floering, B. Brothers, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 4 |
| 2002 | Joint Panel - IPDS and Workshop on Dependability Benchmarking
Ravishankar K. Iyer, Zbigniew T. Kalbarczyk, Philip Koopman, Henrique Madeira, Gunter Heiner, Karama Kanoun, Haim Levendel, Brendan Murphy, Lawrence G. Votta, Don Wilson |
DSN | 1 |
| 2002 | NFTAPE: Networked Fault Tolerance and Performance EvaluatorabstractThe NFTAPE is a software implemented, highly flexible fault injection environment for conducting automated fault/error injection-based dependability characterization. NFTAPE: (1) enables a user: (i) to specify a fault/error injection plan, (ii) to carry out injection experiments, and (iii) to collect the experimental results for analysis; (2) targets assessment of a broad set of dependability metrics, e.g., availability, reliability, coverage; (3) operates in a distributed environment; (4) can be configured to implement a variety of fault/error injection strategies and thus to serve multiple users and target systems; (5) imposes minimal disturbance of target systems. David T. Stott, Phillip H. Jones, M. Hamman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 5 |
| 2002 | An Experimental Evaluation of the REE SIFT Environment for Spaceborne ApplicationsabstractPresents an experimental evaluation of a software-implemented fault tolerance (SIFT) environment built around a set of self-checking processes called ARMORs running on different machines that provide error detection and recovery services to themselves and to spaceborne scientific applications. The experiments are split into three groups of error injections, with each group successively stressing the SIFT error detection and recovery more than the previous group. The results show that the SIFT environment adds negligible overhead to the application during failure-free runs. Only 11 cases were observed in which either the application failed to start or the SIFT environment failed to recognize that the application had completed. Further investigations showed that assertions within the SIFT processes-coupled with object-based incremental checkpointing-were effective in preventing system failures by protecting dynamic data within the SIFT processes. Keith Whisnant, Ravishankar K. Iyer, Raphael R. Some, David A. Rennels |
DSN | 2 |
| 2002 | Loose Synchronization of Multithreaded ReplicasabstractAlthough multithreading can improve performance, it is a source of nondeterminism in application behavior. Existing approaches to replicating multithreaded applications either synchronize replicas at the interrupt level, at the expense of performance, or use a nonpreemptive deterministic scheduler at the expense of concurrency. This paper presents a loose synchronization algorithm for ensuring deterministic replica behavior while preserving concurrency. The algorithm synchronizes replica threads only on state updates by enforcing an equivalent order of mutex acquisitions across replicas. Claudio Basile, Keith Whisnant, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SRDS | 4 |
| 2001 | A Framework for Database Audit and Control Flow Checking for a Wireless Telephone Network ControllerabstractThe paper presents the design and implementation of a dependability framework for a call-processing environment in a digital mobile telephone network controller. The framework contains a data audit subsystem to maintain the structural and semantic integrity of the database and a preemptive control flow checking technique, PECOS, to protect call-processing clients. Evaluation of the dependability-enhanced system is performed (using NFTAPE, a software-implemented error injection environment). The evaluation shows that for control flow errors in the client, the combination of PECOS and data audit eliminates fail-silence violations, reduces the incidence of client crashes, and eliminates client hangs. For database injections, data audit detects 85% of the errors and reduces the incidence of escaped errors. Evaluation of combined use of data and control checking (with error injection targeting the database and the client) shows coverage increase from 35% to 80% and indicates data flow errors as a key reason for error escapes. Saurabh Bagchi, Keith Whisnant, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Ytzhak H. Levendel, Lawrence G. Votta |
DSN | 5 |
| 2001 | An Experimental Study of Security Vulnerabilities Caused by ErrorsabstractThe paper presents an experimental study which shows that, for the Intel x86 architecture, single-bit control flow errors in the authentication sections of targeted applications can result in significant security vulnerabilities. The experiment targets two well-known Internet server applications: FTP and SSH (secure shell), injecting single-bit control flow errors into user authentication sections of the applications. The injected sections constitute approximately 2-8% of the text segment of the target applications. The results show that out of all activated errors: (a) 1-2% comprised system security (create a permanent window of vulnerability); (b) 43-62% resulted in crash failures (about 8.5% of these errors create a transient window of vulnerability); and (c) 7-12% resulted in fail silence violations. A key reason for the measured security vulnerabilities is that, in the x86 architecture, conditional branch instructions are a minimum of one Hamming distance apart. The design and evaluation of a new encoding scheme that reduces or eliminates this problem is presented. Jun Xu 0003, Shuo Chen 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 4 |
| 2001 | Improving Cache Performance of Network Intensive WorkloadsabstractThe performance of servers for network-intensive workloads such as web services and online transaction processing applications depends on the effective utilization of the processor caches. A detailed analysis of the cache space utilization of web workloads shows us that several memory addresses are referenced only once during their lifetime in the cache. These references frequently reside in the cache for a long time contributing to the pollution of cache. The most commonly adopted least-recently-used (LRU) replacement scheme does not exploit this characteristic. In this paper, we propose an alternative block replacement policy called Single-Touch Aware Replacement (STAR) algorithm. This algorithm predicts blocks that will potentially be referenced only once and replaces them early enough to improve cache efficiency. The STAR scheme was implemented in a trace-driven cache simulator and the performance with several commercial workloads was analyzed. The use of the STAR algorithm results in up to 20% improvement in cache performance for web workloads (SPECweb96, SPECweb99) and up to 5% improvement in online transaction processing (TPC-C) workloads. Udaykiran Vallamsetty, Prasant Mohapatra, Ravishankar K. Iyer, Krishna Kant 0001 |
ICPP | 3 |
| 2001 | Comparing Fail-Sailence Provided by Process Duplication versus Internal Error Detection for DHCP ServerabstractThis paper uses fault injection to compare the ability of two fault-tolerant software architectures to protect an application from faults. These two architectures are Voltan, which uses process duplication, and Chameleon ARMORs, which use self-checking. The target application is a Dynamic Host Configuration Protocol (DHCP) server, a widely used application for managing IP addresses. NFTAPE, a software-based fault injection environment, is used to inject three classes of faults, namely random memory bit-flip, control-flow and high-level target specific faults, into each software architecture and into baseline Solaris and Linux versions. David T. Stott, Neil A. Speirs, Zbigniew T. Kalbarczyk, Saurabh Bagchi, Jun Xu 0003, Ravishankar K. Iyer |
IPDPS | 6 |
| 2001 | Geist: a generator for e-commerce & internet server trafficabstractThis paper describes Geist, a traffic generator for stress testing of web servers. The generator provides a large number of dialable parameters that allow traffic characteristics to range from simple static web-page browsing to the transactional traffic seen by e-commerce front end servers. Unlike other traffic generators, our generator concentrates on the characteristics of the aggregate traffic arriving at the server, which allows for better control of the scaling properties of the traffic and a more scalable generation. In this paper, we describe the traffic characterization and generation process that Geist is based on. We also present the performance of our current implementation, instances of its usage and directions for future work. Krishna Kant 0001, Vijay Tewari, Ravishankar K. Iyer |
ISPASS | 3 |
| 2000 | Architectural Impact of Secure Socket Layer on Internet ServersabstractSecure socket layer (SSL) is the most popular protocol used in the Internet for facilitating secure communications. In this paper, we analyze the performance and architectural impact of SSL on the servers in terms of various parameters such as throughput, utilization, cache sizes, cache miss ratios, number of processors, control dependencies, file access sizes, bus transactions, network load, etc. The major conclusions from this study are as follows: The use of SSL increases computational cost of the transactions by a factor of 5-7. SSL transactions do not benefit much from a larger L2 cache, but a larger L1 cache would be helpful. A complex logic for handling control dependencies is not useful for SSL transaction as the frequency of branches is very low. Because SSL workload is highly CPU bound, it may be possible to enhance SSL performance by using a number of other architectural features as well. Krishna Kant 0001, Ravishankar K. Iyer, Prasant Mohapatra |
ICCD | 2 |
| 2000 | Hierarchical Error Detection in a Software Implemented Fault Tolerance (SIFT) EnvironmentabstractProposes a hierarchical error detection framework for a software-implemented fault tolerance (SIFT) layer of a distributed system. A four-level error detection hierarchy is proposed in the context of Chameleon, a software environment for providing adaptive fault tolerance in an environment of commercial off-the-shelf (COTS) system components and software. The design and implementation of a software-based distributed signature monitoring scheme, which is central to the proposed four-level hierarchy, is described. Both intra-level and inter-level optimizations that minimize the overhead of detection and are capable of adapting to runtime requirements are proposed. The paper presents results from a prototype implementation of two levels of the error detection hierarchy and results of a detailed simulation of the overall environment. The results indicate a substantial increase in availability due to the detection framework and help in understanding the tradeoffs between overhead and coverage for different combinations of techniques. Saurabh Bagchi, Balaji Srinivasan, Keith Whisnant, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2000 | Diagnosing Rediscovered Software Problems Using SymptomsabstractThis paper presents an approach to automatically diagnosing rediscovered software failures using symptoms, in environments in which many users run the same procedural software system. The approach is based on the observation that the great majority of field software failures are rediscoveries of previously reported problems and that failures caused by the same defect often share common symptoms. Based on actual data, the paper develops a small software failure fingerprint, which consists of the procedure call trace, problem detection location, and the identification of the executing software. The paper demonstrates that over 60 percent of rediscoveries can be automatically diagnosed based on fingerprints; less than 10 percent of defects are misdiagnosed. The paper also discusses a pilot that implements the approach. Using the approach not only saves service resources by eliminating repeated data collection for and diagnosis of reoccurring problems, but it can also improve service response time for rediscoveries. Inhwan Lee, Ravishankar K. Iyer |
IEEE Trans. Software Eng. | 2 |
| 1999 | Networked Windows NT System Field Failure Data AnalysisabstractThis paper presents a measurement-based dependability study of a Networked Windows NT system based on field data collected from NT System Logs from 503 servers running in a production environment over a four-month period. The event logs at hand contains only system reboot information. We study individual server failures and domain behavior in order to characterize failure behavior and explore error propagation between servers. The key observations from this study are: (1) system software and hardware failures are the two major contributors to the total system downtime (22% and 10%), (2) recovery from application software failures are usually quick, (3) in many cases, more than one reboots are required to recover from a failure, (4) the average availability of an individual server is over 99%, (5) there is a strong indication of error dependency or error propagation across the network, (6) most (58%) reboots are unclassified indicating the need for better logging techniques, (7) maintenance and configuration contribute to 24% of system downtime. Jun Xu 0003, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
PRDC | 3 |
| 1999 | Failure Data Analysis of a LAN of Windows NT based ComputersabstractThis paper presents results of a failure data analysis of a LAN of Windows NT machines. Data for the study was obtained from event logs collected over a six-month period from the mail routing network of a commercial organization. The study focuses on characterizing causes of machine reboots. The key observations from this study are: 1) most of the problems that lead to reboots are software related; 2) rebooting the machine does not always solve the problem; 3) there are indications of propagated or correlated failures; and 4) though the average availability evaluates to over 99%, the machine downtime lasts (on average) two hours. Since the machines are dedicated mail servers, bringing down one or more of them can potentially disrupt storage, forwarding, reception and delivery of mail. This suggests that the average availability is not a good measure to characterize this type of network service. M. Kalyanakrishnam, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SRDS | 3 |
| 1999 | A Software Multilevel Fault Injection Mechanism: Case Study Evaluating the Virtual Interface ArchitectureabstractThe characteristics of failures occurring in networked computing systems are still poorly understood. As a consequence, this is a rich area for exploration, especially with the arrival of new network interface standards, such as the Virtual Interface Architecture (VIA) adopted by Microsoft, Intel and Compaq. The goal of VIA is to improve the performance of distributed applications by reducing the latency associated with the exchange of critical message between processes in Windows NT-based systems. In this paper, we propose the SMiFI (Software Multilevel Fault Injection) mechanism to evaluate the failure characteristics of networked systems, specifically VIA. The mechanism covers all software protocol layers of the host interface and corrupts both the messages and the computation engines that manipulate the messages. Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SRDS | 3 |
| 1999 | Reliability of Internet Hosts: A Case Study from the End User's Perspective
Mahesh Kalyanakrishnan, Ravishankar K. Iyer, Jaqdish U. Patel |
Comput. Networks | 2 |
| 1999 | Stress-Based and Path-Based Fault InjectionabstractThe objective of fault injection is to mimic the existence of faults and to force the exercise of the fault tolerance mechanisms of the target system. To maximize the efficacy of each injection, the locations, timing, and conditions for faults being injected must be carefully chosen. Faults should be injected with a high probability of being accessed. This paper presents two fault injection methodologies-stress-based injection and path-based injection; both are based on resource activity analysis to ensure that injections cause fault tolerance activity and, thus, the resulting exercise of fault tolerance mechanisms. The difference between these two methods is that stress-based injection validates the system dependability by monitoring the run-time workload activity at the system level to select faults that coincide with the locations and times of greatest workload activity, while path-based injection validates the system from the application perspective by using an analysis of the program flow and resource usage at the application program level to select faults during the program execution. These two injection methodologies focus separately on the system and process viewpoints to facilitate the testing of system dependability. Details of these two injection methodologies are discussed in this paper, along with their implementations, experimental results, and advantages and disadvantages. Timothy K. Tsai, Mei-Chen Hsueh, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
IEEE Trans. Computers | 5 |
| 1999 | Chameleon: A Software Infrastructure for Adaptive Fault ToleranceabstractThis paper presents Chameleon, an adaptive infrastructure, which allows different levels of availability requirements to be simultaneously supported in a networked environment. Chameleon provides dependability through the use of special ARMORs-Adaptive. Reconfigurable, and Mobile Objects for Reliability-that control all operations in the Chameleon environment. Three broad classes of ARMORs are defined: 1) Managers oversee other ARMORs and recover from failures in their subordinates. 2) Daemons provide communication gateways to the ARMORs at the host node. They also make available a host's resources to the Chameleon environment. 3) Common ARMORs implement specific techniques for providing application-required dependability. Employing ARMORs, Chameleon makes available different fault-tolerant configurations and maintains run-time adaptation to changes in the availability requirements of an application. Flexible ARMOR architecture allows their composition to be reconfigured at run-time, i.e., the ARMORs may dynamically adapt to changing application requirements. In this paper, we describe ARMOR architecture, including ARMOR class hierarchy, basic building blocks, ARMOR composition, and use of ARMOR factories. We present how ARMORs can be reconfigured and reengineered and demonstrate how the architecture serves our objective of providing an adaptive software infrastructure. To our knowledge, Chameleon is one of the few real implementations which enables multiple fault tolerance strategies to exist in the same environment and supports fault-tolerant execution of substantially off-the-shelf applications via a software infrastructure only. Chameleon provides fault tolerance from the application's point of view as well as from the software infrastructure's point of view. To demonstrate the Chameleon capabilities, we have implemented a prototype infrastructure which provides set of ARMORs to initialize the environment and to support the dual and TMR application execution modes. Through this testbed environment, we measure the execution overhead and recovery times from failures in the user application, the Chameleon ARMORs, the hardware, and the operating system. Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Saurabh Bagchi, Keith Whisnant |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1999 | Hierarchical Simulation Approach to Accurate Fault Modeling for System Dependability EvaluationabstractThis paper presents a hierarchical simulation methodology that enables accurate system evaluation under realistic faults and conditions. In this methodology, effects of low-level (i.e., transistor or circuit level) faults are propagated to higher levels (i.e., system level) using fault dictionaries. The primary fault models are obtained via simulation of the transistor-level effect of a radiation particle penetrating a device. The resulting current bursts constitute the first-level fault dictionary and are used in the circuit-level simulation to determine the impact on circuit latches and flip-flops. The latched outputs constitute the next level fault dictionary in the hierarchy and are applied in conducting fault injection simulation at the chip-level under selected workloads or application programs. Faults injected at the chip-level result in memory corruptions, which are used to form the next level fault dictionary for the system-level simulation of an application running on simulated hardware. When an application terminates, either normally or abnormally, the overall fault impact on the software behavior is quantified and analyzed. The system in this sense can be a single workstation or a network. The simulation method is demonstrated and validated in the case study of Myrinet (a commercial, high-speed network) based network system. Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Gregory L. Ries, Jaqdish U. Patel, Myeong S. Lee, Yuxiao Xiao |
IEEE Trans. Software Eng. | 2 |
| 1998 | Measurement-based modeling and analysis methodology for characterizing parallel I/O performanceabstractA parallel I/O characterization methodology that consists of a hierarchical modeling and measurement analysis environment for investigating I/O performance is presented. The methodology is illustrated via a case study of a video server workload running under the parallel I/O file system (PIOFS) of IBM SP/2. The measurements demonstrate that for video server and read-intensive workloads, spreading parallel files across all eight I/O servers improves a client's bandwidth performance by 36-52%. With eight clients, the per-client bandwidth performance increases by only 15%-23%. PIOFS-based default file striping results in degradation of bandwidth performance by as much as 25%. Ravishankar K. Iyer |
HiPC | 2 |
| 1998 | The Chameleon Infrastructure for Adaptive, Software Implemented Fault ToleranceabstractThis paper presents Chameleon, an adaptive software infrastructure for supporting different levels of availability requirements in a heterogeneous networked environment. Chameleon provides dependability through the use of ARMORs-Adaptive, Reconfigurable, and Mobile Objects for Reliability. Three broad classes of ARMORs are defined: Managers, Daemons, and Common ARMORs. Key concepts that support adaptive fault tolerance include the construction of fault tolerance execution strategies from a comprehensive set of ARMORs, the creation of ARMORs from a library of reusable basic building blocks, the dynamic adaptation to changing fault tolerance requirements, and the ability to detect and recover from errors in applications and in ARMORs. Saurabh Bagchi, Keith Whisnant, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SRDS | 4 |
| 1998 | Dependability Analysis of a Cache-Based RAID System via Fast Distributed SimulationabstractWe propose a new speculation-based, distributed simulation method for dependability analysis of complex systems in which a detailed functional simulation of a system component is essential to obtain an accurate overall result. Our target example is a networked cluster with compute nodes and a single I/O node. Accurate system dependability characterization is achieved via a combination of detailed simulation of the I/O subsystem behavior in the presence of faults and more abstract simulation of the compute nodes and the switching network. Dependability measures like error coverage, error detection latency and performance measures such as delivery time in the presence of faults are obtained. The approach is implemented on a network of workstations, and experimental results show significant improvements over a Time Warp simulator for the same model. Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SRDS | 3 |
| 1998 | Dependability Analysis of a High-Speed Network Using Software-Implemented Fault Injection and Simulated Fault InjectionabstractThis paper presents a dependability study of high-speed, switched Local Area Networks (LANs) using Myrinet as an example testbed (with theoretical speeds of 2.56 Gbps). The study uses results of two fault injection methods, simulated fault injection and software-implemented fault injection (SWIFI), to analyze the application-level impact of transient faults injected into the network interface hardware. These results include a number of errors, such as dropped or corrupt messages, host interface or host resets, and local or remote host interface hangs. The paper presents the study in two parts: First, the results from the SWIFI method in the real system are used as a basis to validate the simulation and identify the major factors leading to differences between the methods. A comparison between the two injection methods shows that they agree for 83 percent of the fault injections. The results, however, vary greatly, depending on the fault type considered. The study also presents an analysis of the effects of varying workload intensity, host platform, and interface function targeted by the injection. An example of this analysis is to show that the function targeted has a significant impact on the fault activation rate. Finally, the study identifies two mechanisms by which faults may propagate from the interface to other parts of the network; in one example, this propagation caused the interface's host computer to reboot, while another caused a remote interface in the network to hang. David T. Stott, Gregory L. Ries, Mei-Chen Hsueh, Ravishankar K. Iyer |
IEEE Trans. Computers | 4 |
| 1997 | Reliability of Internet Hosts - A Case Study from the End User's PerspectiveabstractThis paper presents the results of a 40-day reliability study on a set of 97 popular Web sites done from an end user's perspective. Data for the study was acquired by periodically attempting to fetch an HTML file from each Web site and recording the outcome of such attempts. Analysis of the acquired data revealed: (i) 94% of the HTML file fetch requests succeed on average; (ii) most failures last less than 15 minutes; (iii) the underlying network plays a dominant role in determining host accessibility: (a) network related-outages account for a major part of the failures, (b) some network-related outages rendered more than 70% of the hosts inaccessible, and (c) host-related failures tend to be shorter than failures that might involve the network; (vi) the network connectivity is high on the average with 93% of the sites being accessible at any given time; and (vii) the mean availability of the hosts is high (0.993). Mahesh Kalyanakrishnan, Ravishankar K. Iyer, Jaqdish U. Patel |
ICCCN | 2 |
| 1997 | Chameleon: Adaptive Fault Tolerance Using Reliable, Mobile AgentsabstractIn networked computing systems, a broad range of commercial and scientific applications that need varying degrees of availability must coexist. It is not cost-effective to develop a reliable platform in each case. It is more efficient to build an infrastructure that provides the required level of dependability for each application's needs. It is also essential that the proposed alternatives should leverage off-the-shelf components. There have been exhaustive studies on fault tolerance strategies capable of providing efficient mechanisms to deal with system operational failures. Most of this work has focused on specific application needs and thus provided only piecemeal solutions. Little work has been done in addressing how to build a reliable networked computing system out of unreliable computation nodes. As a result, there is no comprehensive solution for providing a wide range of fault-tolerant services in a single networked environment. The most feasible way of understanding how such a software environment would fit on top of existing layers (the operating system, the network interfaces, etc.) is to implement an infrastructure for providing a range of reliable services. Fundamental components of the envisioned infrastructure (Chameleon) have been designed so that none of them is a single point of failure. Each of the components is active for a certain period, e.g. during the setting up the system configuration. If a component fails during its active phase, there is a provision for recovery, either by switching to a backup or by regenerating the component. Ravishankar K. Iyer, Zbigniew T. Kalbarczyk, Saurabh Bagchi |
SRDS | 1 |
| 1997 | DEPEND: A Simulation-Based Environment for System Level Dependability AnalysisabstractThe paper presents the rationale for a functional simulation tool, called DEPEND, which provides an integrated design and fault injection environment for system level dependability analysis. The paper discusses the issues and problems of developing such a tool, and describes how DEPEND tackles them. Techniques developed to simulate realistic fault scenarios, reduce simulation time explosion, and handle the large fault model and component domain associated with system level analysis are presented. Examples are used to motivate and illustrate the benefits of this tool. To further illustrate its capabilities, DEPEND is used to simulate the Unix-based Tandem triple-modular-redundancy (TMR) based prototype fault-tolerant system and to evaluate how well it handles near-coincident errors caused by correlated and latent faults. Issues such as memory scrubbing, re-integration policies, and workload dependent repair times, which affect how the system handles near-coincident errors, are also evaluated. Unlike any other simulation-based dependability studies, the accuracy of the simulation model is validated by comparing the results of the simulations with measurements obtained from fault injection experiments conducted on a production Tandem machine. Kumar K. Goswami, Ravishankar K. Iyer, Luke T. Young |
IEEE Trans. Computers | 2 |
| 1996 | Analyze-NOW-an environment for collection and analysis of failures in a network of workstationsabstractThis paper describes Analyze-NOW an environment for collection and analysis of failures/errors in a network of workstations. Descriptions cover the data collection methodology and the tool implemented to facilitate this process. Software tools used for analysis are described, with emphasis on the details of the implementation of the Analyzer, the primary analysis tool. Application of the tools is demonstrated by using them to collect and analyze failure data (for 32 week period) from a network of 69 SunOS-based workstations. Classification based on the source and the effect of faults is used to identify problem areas. Different types of failures encountered on the machines and the network are highlighted to develop a proper understanding of failures in a network environment. Lastly, a case is made for using the results from the analysis tool to pinpoint the problem areas in the network. Anshuman Thakur, Ravishankar K. Iyer |
ISSRE | 2 |
| 1996 | A Gate-Level Simulation Environment for Alpha-Particle-Induced Transient FaultsabstractMixed analog and digital mode simulators have been available for accurate /spl alpha/-particle-induced transient fault simulation. However, they are not fast enough to simulate a large number of transient faults on a relatively large circuit in a reasonable amount of time. In this paper, we describe a gate-level transient fault simulation environment which has been developed based on realistic fault models. Although the environment was developed for /spl alpha/-particle-induced transient faults, the methodology can be used for any transient fault which can be modeled as a transient pulse of some width. The simulation environment uses a gate level timing fault simulator as well as a zero-delay parallel fault simulator. The timing fault simulator uses logic level models of the actual transient fault phenomenon and latch operation to accurately propagate the fault effects to the latch outputs, after which point the zero-delay parallel fault simulator is used to speed up the simulation without any loss in accuracy. The environment is demonstrated on a set of ISCAS-89 sequential benchmark circuits. Hungse Cha, Elizabeth M. Rudnick, Janak H. Patel, Ravishankar K. Iyer, Gwan S. Choi |
IEEE Trans. Computers | 4 |
| 1996 | Analyze-NOW-an environment for collection and analysis of failures in a network of workstationsabstractThis paper describes Analyze-NOW, an environment for the collection and analysis of failures/errors in a network of workstations. Descriptions cover the data collection methodology and the tool implemented to facilitate this process. Software tools used for analysis are described, with emphasis on the details of the implementation of the Analyzer, the primary analysis tool. Application of the tools is demonstrated by using them to collect and analyze failure data (for 32-week period) from a network of 69 SunOS-based workstations. Classification based on the source and effect of faults is used to identify problem areas. Different types of failures encountered on the machines and network are highlighted to develop a proper understanding of failures in a network environment. The results from the analysis tool should be used to pinpoint the problem areas in the network. The results obtained from using Analyze-NOW on failure data from the monitored network reveal some interesting behavior of the network. Nearly 70% of the failures were network-related, whereas disk errors were few. Network-related failures were 75% of all hard-failures (failures that make a workstation unusable). Half of the network-related failures were due to servers not responding to clients, and half were performance-related and others. Problem areas in the network were found using this tool. The authors' approach was compared to the method of using the network architecture to locate problem areas. This comparison showed that locating problem areas using network architecture over-estimates the number of problem areas. Anshuman Thakur, Ravishankar K. Iyer |
IEEE Trans. Reliab. | 2 |
| 1995 | Analysis of failures in the Tandem NonStop-UX Operating SystemabstractThe paper presents results from an investigation of failures in several releases of Tandem's NonStop-UX Operating System, which is based on Unix System V. The analysis covers software failures from the field and failures reported by Tandem's test center. Fault classification is based on the status of the reported failures, the detection point of the errors in the operating system code, the panic message generated by the systems, the module that was found to be faulty, and the type of programming mistake. This classification reveals which modules in the operating system generate the most faults and the modules in which most errors are detected. We also present distributions of the failure and repair times including inter arrival time of unique failures and time between duplicate failures. These distributions, unlike generic time distributions, such as time between failures, help characterize the software quality. Distribution of the repair times emphasizes the repair process and the factors influencing repair. Distribution of up time of the systems before the panic reveals the factors triggering the panic. Anshuman Thakur, Ravishankar K. Iyer, Luke T. Young, Inhwan Lee |
ISSRE | 2 |
| 1995 | A Measurement-Based Model to Predict the Performance Impact of System Modifications: A Case StudyabstractThe paper presents a performance case study of parallel jobs executing in real multi user workloads. The study is based on a measurement based model capable of predicting the completion time distribution of the jobs executing under real workloads. The model constructed is also capable of predicting the effects of system design changes on application performance. The model is a finite state, discrete time Markov model with rewards and costs associated with each state. The Markov states are defined from real measurements and represent system/workload states in which the machine has operated. The paper places special emphasis on choosing the correct number of states to represent the workload measured. Specifically, the performance of computationally bound, parallel applications executing in real workloads on an Alliant FX/80 is evaluated. The constructed model is used to evaluate scheduling policies, the performance effects of multiprogramming overhead, and the scalability of the Alliant FX/8O in real workloads. The model identifies a number of available scheduling policies which would improve the response time of parallel jobs. In addition, the model predicts that doubling the number of processors in the current configuration would only improve response time for a typical parallel application by 25%. The model recommends a different processor configuration to more fully utilize extra processors. The paper also presents empirical results which validate the model created.> Robert T. Dimpsey, Ravishankar K. Iyer |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1995 | Software Dependability in the Tandem GUARDIAN SystemabstractBased on extensive field failure data for Tandem's GUARDIAN operating system, the paper discusses evaluation of the dependability of operational software. Software faults considered are major defects that result in processor failures and invoke backup processes to take over. The paper categorizes the underlying causes of software failures and evaluates the effectiveness of the process pair technique in tolerating software faults. A model to describe the impact of software faults on the reliability of an overall system is proposed. The model is used to evaluate the significance of key factors that determine software dependability and to identify areas for improvement. An analysis of the data shows that about 77% of processor failures that are initially considered due to software are confirmed as software problems. The analysis shows that the use of process pairs to provide checkpointing and restart (originally intended for tolerating hardware faults) allows the system to tolerate about 75% of reported software faults that result in processor failures. The loose coupling between processors, which results in the backup execution (the processor state and the sequence of events) being different from the original execution, is a major reason for the measured software fault tolerance. Over two-thirds (72%) of measured software failures are recurrences of previously reported faults. Modeling, based on the data, shows that, in addition to reducing the number of software faults, software dependability can be enhanced by reducing the recurrence rate.> Inhwan Lee, Ravishankar K. Iyer |
IEEE Trans. Software Eng. | 2 |
| 1994 | Impact of Loop Granularity and Self-Preemtion on the Performance of Loop Parallel Applications on a Multiprogrammed Shared-Memory MultiprocessorabstractThis study uses real system measurements to investigate the relationships between loop granularity, parallel loop distribution and barrier wait times, and their impact on the multiprogramming performance of loop parallel applications on the CEDAR shared-memory multiprocessor The overhead due to multiprogramming varies from 5% for applications with large loop granularity to 140% for applications with very fine-gram loops. This is because applications with fine-gram loops have unequal parallel work distribution among the clusters m multiprogrammed environments, while the parallel work in applications with large loop granularity is equally distributed. Moreover, increased barrier wait times of the mam task and wait-for-work times of the helper tasks also contribute to the multi-programming performance degradation of the fine-grain loop parallel applications. We propose and implement a self-preemption technique to address the problem of met eased barrier wait times and wait-for-work times. Using this technique, the overhead due to multiprogramming is reduced by as much as 100%, and speedups of 1.1 to 1.7 are obtained. Chitra Natarajan, Ravishankar K. Iyer |
ICPP (2) | 3 |
| 1994 | Measurement-Based Characterization of Global Memory and Network Contention, Operating System and Parallelization Overheads: A Case Study on Shared-Memory MultiprocessorabstractPresents a characterization of (1) the global memory and interconnection network contention overhead, (2) the operating system overheads, and (3) the runtime system parallelization overheads for the Cedar shared-memory multiprocessor. The measurements were obtained using five representative compute-intensive, scientific, loop parallel applications from the Perfect Benchmark Suite. The overheads were measured for a range of Cedar configurations from 1 processor to the full 4-cluster/32-processor configuration, thus characterizing the effect of this scaling on the overheads. For the full 4-cluster Cedar, the operating system overhead was found to constitute 5-21%: of the total completion time of an application. The parallelization overhead accounts for 10-25% of the application, completion time and the overhead due to global memory and network contention contributes 8-21% of the application completion time.> Chitra Natarajan, Ravishankar K. Iyer |
ISCA | 3 |
| 1993 | Fault behavior dictionary for simulation of device-level transientsabstractThe paper presents a methodology for the simulation of massive number of device-level transient faults. Fault injection locations and the gate around those locations are extracted and evaluated with SPICE. The extracted sub-circuits are exercised exhaustively while fault-injections are performed. Faulty behavior at the outputs of each sub-circuit is recorded in a dictionary, along with the associated input vector, fault-injection time, and location. The recorded logical errors are injected concurrent transient simulator is developed to allow simultaneous evaluation of a massive number of fault-injections, in a single simulation pass. The methodology is illustrated by a case study of MC68000 microprocessor. Gwan S. Choi, Ravishankar K. Iyer, Daniel G. Saab |
ICCAD | 2 |
| 1993 | Panel: Field Failures And Reliability In Operation
Ram Chillarege, Ravishankar K. Iyer, Jean-Claude Laprie, John D. Musa |
ISSRE | 2 |
| 1993 | MEASURE+ - A Measurement-Based Dependability Analysis PackageabstractMost existing dependability modeling and evaluation tools are designed for building and solving commonly used models with emphasis on solution techniques, not for identifying realistic models from measurements. In this paper, a measurement-based dependability analysis package, MEASURE+, is introduced. Given measured data from real systems in a specified format MEASURE+ can generate appropriate dependability models and measures including Markov and semi-Markov models, k-out-of-n availability models, failure distribution and hazard functions, and correlation parameters. These models and measures obtained from data are valuable for understanding actual error/failure characteristics, identifying system bottlenecks, evaluating dependability for real systems, and verifying assumptions made in analytical models. The paper illustrates MEASURE+ by applying it to the data from a VAXcluster multicomputer system. Models of field failure behavior identified by MEASURE+ indicate that both traditional models assuming failure independence and those few taking correlation into account are not representative of the actual occurrence process of correlated failures. Ravishankar K. Iyer |
SIGMETRICS | 2 |
| 1993 | Dependability Measurement and Modeling of a Multicomputer SystemabstractA measurement-based analysis of error data collected from a DEC VAXcluster multicomputer system is presented. Basic system dependability characteristics such as error/failure distributions and hazard rate are obtained for both the individual machine and the entire VAXcluster. Markov reward models are developed to analyze error/failure behavior and to evaluate performance loss due to errors/failures. Correlation analysis is then performed to quantify relationships of error/failures across machines and across time. It is found that shared resources constitute a major reliability bottleneck. It is shown that for measured system, the homogeneous Markov model, which assumes constant failure rates, overestimates the transient reward rate for the short-term operation, and underestimates it for the long-term operation. Correlation analysis shows that errors are highly correlated across machines and across time. The failure correlation coefficient is low. However, its effect on system unavailability is significant.> Ravishankar K. Iyer |
IEEE Trans. Computers | 2 |
| 1993 | Prediction-Based Dynamic Load-Sharing HeuristicsabstractPresents dynamic load-sharing heuristics that use predicted resource requirements of processes to manage workloads in a distributed system. A previously developed statistical pattern-recognition method is employed for resource prediction. While nonprediction-based heuristics depend on a rapidly changing system status, the new heuristics depend on slowly changing program resource usage patterns. Furthermore, prediction-based heuristics can be more effective since they use future requirements rather than just the current system state. Four prediction-based heuristics, two centralized and two distributed, are presented. Using trace driven simulations, they are compared against random scheduling and two effective nonprediction based heuristics. Results show that the prediction-based centralized heuristics achieve up to 30% better response times than the nonprediction centralized heuristic, and that the prediction-based distributed heuristics achieve up to 50% improvements relative to their nonpredictive counterpart.> Kumar K. Goswami, Murthy V. Devarakonda, Ravishankar K. Iyer |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 1993 | FINE: A Fault Injection and Monitoring Environment for Tracing the UNIX System Behavior under FaultsabstractThe authors present a fault injection and monitoring environment (FINE) as a tool to study fault propagation in the UNIX kernel. FINE injects hardware-induced software errors and software faults into the UNIX kernel and traces the execution flow and key variables of the kernel. FINE consists of a fault injector, a software monitor, a workload generator, a controller, and several analysis utilities. Experiments on SunOS 4.1.2 are conducted by applying FINE to investigate fault propagation and to evaluate the impact of various types of faults. Fault propagation models are built for both hardware and software faults. Transient Markov reward analysis is performed to evaluate the loss of performance due to an injected fault. Experimental results show that memory and software faults usually have a very long latency, while bus and CPU faults tend to crash the system immediately. About half of the detected errors are data faults, which are detected when the system is tries to access an unauthorized memory location. Only about 8% of faults propagate to other UNIX subsystems. Markov reward analysis shows that the performance loss incurred by bus faults and CPU faults is much higher than that incurred by software and memory faults. Among software faults, the impact of pointer faults is higher than that of nonpointer faults.> Wei-lun Kao, Ravishankar K. Iyer |
IEEE Trans. Software Eng. | 2 |
| 1992 | A User-Oriented Synthetic Workload GeneratorabstractA user-oriented synthetic workload generator that simulates user file access behavior on the basis of real workload characterizations is described. The workload generator is designed for experiments and simulations related to file system analysis. The model for this workload generator is user-oriented and job-specific, represents file I/O operations at the system call level, allows general distributions for usage measures, and assumes independence in the file I/O operation stream. The workload generator consists of three parts that handle specification of distributions, creation of an initial file system, and selection and execution file I/O operations. Results from experiments on a network file system (NFS) verify that the workload generator can produce realistic workloads and demonstrate the application of the generator.> Wei-lun Kao, Ravishankar K. Iyer |
ICDCS | 2 |
| 1992 | Analysis of large system black-box test dataabstractStudies black box testing and verification of large systems. Testing data is collected from several test teams. A flat, integrated database of test, fault, repair, and source file information is built. A new analysis methodology based on the black box test design and white box analysis is proposed. The methodology is intended to support the reduction of testing costs and enhancement of software quality by improving test selection, eliminating test redundancy, and identifying error prone source files. Using example data from AT&T systems, the improved analysis methodology is demonstrated.> Kent C. Clapp, Ravishankar K. Iyer, Ytzhak H. Levendel |
ISSRE | 2 |
| 1992 | Analysis of software halts in the tandem GUARDIAN operating systemabstractA systematic methodology is given to investigate the dependability of operational software. The methodology combines several techniques. Time series analysis is used to characterize the occurrence of software failures. Markov reward modeling is used to determine the loss in service due to failures of software components, and to identify major bottlenecks. The effectiveness of built-in fault tolerance is also evaluated. The methodology is illustrated using the software halt data from the Tandem GUARDIAN operating system. The results show that the occurrences of software halts are not correlated with each other in time. Interrupt a handling and memory management are found to be the major bottlenecks in the measured system. The fault tolerance in the measured system was shown to reduce the service loss by nearly 90%.> Inhwan Lee, Ravishankar K. Iyer |
ISSRE | 2 |
| 1992 | Analysis of the VAX/VMS error logs in multicomputer environments-a case study of software dependabilityabstractAn analysis is given of the software error logs produced by the VAX/VMS operating system from two VAXcluster multicomputer environments. Basic error characteristics are identified by statistical analysis. Correlations between software and hardware errors, and among software errors on different machines are investigated. Finally, reward analysis and reliability growth analysis are performed to evaluate software dependability. Results show that major software problems in the measured systems are from program flow control and I/O management. The network-related software is suspected to be a reliability bottleneck. It is shown that a multicomputer software 'time between error' distribution can be modeled by a 2-phase hyperexponential random variable: a lower error rate pattern which characterizes regular errors, and a higher error rate pattern which characterizes error bursts and concurrent errors on multiple machines.> Ravishankar K. Iyer |
ISSRE | 2 |
| 1992 | FOCUS: An Experimental Environment for Fault Sensitivity AnalysisabstractFOCUS, a simulation environment for conducting fault-sensitivity analysis of chip-level designs, is described. The environment can be used to evaluate alternative design tactics at an early design stage. A range of user specified faults is automatically injected at runtime, and their propagation to the chip I/O pins is measured through the gate and higher levels. A number of techniques for fault-sensitivity analysis are proposed and implemented in the FOCUS environment. These include transient impact assessment on latch, pin and functional errors, external pin error distribution due to in-chip transients, charge-level sensitivity analysis, and error propagation models to depict the dynamic behavior of latch errors. A case study of the impact of transient faults on microprocessor-based jet-engine controller is used to identify the critical fault propagation paths, the module most sensitive to fault propagation, and the module with the highest potential for causing external errors.> Gwan S. Choi, Ravishankar K. Iyer |
IEEE Trans. Computers | 2 |
| 1992 | Analysis and Modeling of Correlated Failures in Multicomputer SystemsabstractBased on the measurements from two DEC VAX-cluster multicomputer systems, the issue of correlated failures is addressed. In particular, the characteristics of correlated failures, their impact and their modelling on dependability, are discussed. It is found from the data that most correlated failures are related to errors in shared resources and propagate from one machine to another. Comparisons between measurement-based models and analytical models that assume failure independence show that the impact of correlated failures on dependability is significant. Two validated models. the c-dependent model and the p-dependent model, are developed to evaluate the dependability of systems with correlated failures.> Ravishankar K. Iyer |
IEEE Trans. Computers | 2 |
| 1992 | Guest Editors' Introduction
Ravishankar K. Iyer, Kishor S. Trivedi |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1991 | Performance Prediction and Tuning on a MultiprocessorabstractArticle Free Access Share on Performance prediction and tuning on a multiprocessor Authors: R. T. Dimpsey Center for Reliable and High-Performance Computing, Coordinated Science Laboratory, University of Illinois at Urbana-Champaign, 1101 W. Springfield Ave., Urbana, IL Center for Reliable and High-Performance Computing, Coordinated Science Laboratory, University of Illinois at Urbana-Champaign, 1101 W. Springfield Ave., Urbana, ILView Profile , R. K. Iyer Center for Reliable and High-Performance Computing, Coordinated Science Laboratory, University of Illinois at Urbana-Champaign, 1101 W. Springfield Ave., Urbana, IL Center for Reliable and High-Performance Computing, Coordinated Science Laboratory, University of Illinois at Urbana-Champaign, 1101 W. Springfield Ave., Urbana, ILView Profile Authors Info & Claims ISCA '91: Proceedings of the 18th annual international symposium on Computer architectureApril 1991 Pages 190–199https://doi.org/10.1145/115952.115972Published:01 April 1991Publication History 9citation315DownloadsMetricsTotal Citations9Total Downloads315Last 12 Months10Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Robert T. Dimpsey, Ravishankar K. Iyer |
ISCA | 2 |
| 1991 | Modeling and Measuring Multiprogramming and System Overheads on a Shared-Memory Multiprocessor: Case Study
Robert T. Dimpsey, Ravishankar K. Iyer |
J. Parallel Distributed Comput. | 2 |
| 1990 | Performance degradation due to multiprogramming and system overheads in real workloads: case study on a shared memory multiprocessorabstractIn this paper, performance degradation specifically due to the multiprogramming (MP) overhead in a parallel execution environment is quantified. In addition, total system overhead is also measured. A methodology, which estimates the MP overhead present in real workloads, is illustrated with real measurements taken on an Alliant FX/80 running Xylem (Cedar's operating system). It is found that MP overhead usually consumes between 10% and 23% of the processing power available to parallel programs. Total system overhead usually consumes between 12% and 30% of the parallel environment processing power, but is found to be as high as 82.1%. The mean MP overhead is determined to be 16% which is well over half the total system overhead executed on the system (the mean system overhead is determined to be 24% of the processing power). It is found that MP overhead, total system overhead, and application completion time are all moderately correlated. Relationships between the characteristics of a workload and the overhead measurements indicate that processor utilization and to a lesser degree paging are moderately correlated with the overhead present in the workload. It is also found that MP overhead is statistically independent of the number of parallel jobs in the system, while total system overhead is not. Robert T. Dimpsey, Ravishankar K. Iyer |
ICS | 2 |
| 1990 | Automatic Recognition of Intermittent Failures: An Experimental Study of Field DataabstractA methodology is proposed for recognizing the symptoms of persistent problems in large systems. The system error rate is used to identify the error states among which relationships may exist. Statistical techniques are used to validate and quantify the strength of the relationship among these error states. As input, the approach takes the raw error logs containing a single entry for each error that is detected as an isolated event. As output, it produces a list of symptoms that characterize persistent errors. Thus, given a failure, it is determined whether the failure is an intermittent manifestation of a common fault or whether it is an isolated (transient) incident. The technique is shown to work on two CYBER systems and on IBM 3081 multiprocessor system. Comparisons to real failure/repair information obtained from field engineers show that, in about 85% of the cases, the error symptoms recognized by this approach correspond to real problems. The remaining 15% of the cases, although not directly supported by field data, are confirmed as being valid problems.> Ravishankar K. Iyer, Luke T. Young, P. V. Krishna Iyer |
IEEE Trans. Computers | 1 |
| 1990 | Guest Editor's Introduction Experimental Computer Science
Ravishankar K. Iyer |
IEEE Trans. Software Eng. | 1 |
| 1989 | FOCUS: an experimental environment for validation of fault-tolerant systems - case study of a jet-engine controllerabstractA simulation environment that allows the run-time injection of transient and permanent faults and the assessment of their impact in complex systems is described. The error data from the simulation are automatically fed into the analysis software in order to quantify the fault-tolerance of the system under test. The features of the environment are illustrated with case study of a fault-tolerant, dual-configuration real-time jet engine controller. The entire controller, described at the logic and functional levels, is simulated, and transient fault injections are performed. In the controller, fault detection and reconfiguration are performed by transactions over the communication links. The simulation consists of the instructions specifically designed to exercise this cross-channel communication. The level of effectiveness of the dual configuration of the system to single and multiple transient errors is measured. The results are used to identify critical design aspects from a fault-tolerance viewpoint.> Gwan S. Choi, Ravishankar K. Iyer, Victor Carreno |
ICCD | 2 |
| 1989 | Multiprogramming Performance Degradation: Case Study on a Shared Memory Multiprocesor
Robert T. Dimpsey, Ravishankar K. Iyer |
ICPP (2) | 2 |
| 1989 | An Experimental Study of Memory Fault LatencyabstractThe difficulty with the measurement of fault latency is due to the lack of observability of the fault occurrence and error generation instants in a production environment. The authors describe an experiment, using data from a VAX 11/780 under real workload, to study fault latency in the memory subsystem accurately. Fault latency distributions are generated for stuck-at-zero (s-a-0) and stuck-at-one (s-a-1) permanent fault models. The results show that the mean fault latency of an s-a-0 fault is nearly five times that of the s-a-1 fault. An analysis of variance is performed to quantify the relative influence of different workload measures on the evaluated latency.> Ram Chillarege, Ravishankar K. Iyer |
IEEE Trans. Computers | 2 |
| 1989 | Predictability of Process Resource Usage: A Measurement-Based Study on UNIXabstractA statistical approach is developed for predicting the CPU time, the file I/O, and the memory requirements of a program at the beginning of its life, given the identity of the program. Initially, statistical clustering is used to identify high-density regions of process resource usage. The identified regions form the states for building a state-transition model to characterize the resource usage of each program in its past executions. The prediction scheme uses the knowledge of the program's resource usage in its last execution together with its state-transition model to predict the resource usage in its next execution. The prediction scheme is shown to work using process resource-usage data collected from a VAX 11/780 running 4.3 BSD Unix. The results show that the predicted values correlate strongly with the actual; the coefficient of correlation between the predicted and actual values for CPU time is 0.84. The errors in prediction are mostly small and are heavily skewed toward small values.> Murthy V. Devarakonda, Ravishankar K. Iyer |
IEEE Trans. Software Eng. | 2 |
| 1988 | Transient fault behavior in a microprocessor - A case studyabstractThe authors describe an experimental analysis to study the susceptibility of a microprocessor-based jet engine controller (an HS 1602) to upsets caused by current and voltage transients. A design automation environment which allows the run-time injection of transients and their tracing from the device to the pin-level is described. The resulting error data are categorized by the charge levels of the injected transients by location and by their potential to cause logic upsets, latched errors, and pin errors. The results show a 3-pC threshold below which the transients have little impact. An ALU (arithmetic logic unit) transient is most likely to result in logic upsets and pin errors (i.e. impact the external environment). The transients in the countdown unit are serious since they can result in latched errors thus causing latent faults.> P. Duba, Ravishankar K. Iyer |
ICCD | 2 |
| 1988 | Performance Analysis of a Shared Memory Multiprocessor: Case Study
Robert T. Dimpsey, Ravishankar K. Iyer |
ICPP (1) | 2 |
| 1988 | Measurement-Based Analysis of Multiple Latent Errors and Near-coincident Fault Discovery in a Shared Memory Multiprocessor
Subhasish G. Mitra, Ravishankar K. Iyer |
ICPP (1) | 2 |
| 1988 | Performability Modeling Based on Real Data: A Case StudyabstractA measurement-based performability model is described that is based on error and resource-usage data collected on a multiprocessor system. A method for identifying the model structure is introduced, and the resulting model is validated against real data. Model development from the collection of raw data to the estimation of the expected reward is described. Both normal behavior and error behavior of the system are characterized. The measured data show that the holding times in key operational and error states are not simple exponentials and that a semi-Markov process is necessary to model the system behavior. A reward function, which is based on the service rate and the error rate in each state, is defined in order to estimate the performability of the system and to depict the cost of different types of errors.> Mei-Chen Hsueh, Ravishankar K. Iyer, Kishor S. Trivedi |
IEEE Trans. Computers | 2 |
| 1988 | Accurate Low-Cost Methods for Performance Evaluation of Cache Memory SystemsabstractTrace-driven simulation is a simple way of evaluating cache memory systems with varying hardware parameters. But to evaluate realistic workloads, simulating even a few million addresses is not adequate and such large scale simulation is impractical from the consideration of space and time requirements. New methods of simulation based on statistical techniques are proposed for decreasing the need for large trace measurements and for predicting true program behavior. In the method, sampling techniques are applied while collecting the address trace from a workload. This drastically reduces the space and time needed to collect the trace. New simulation techniques are developed to use the sample data not only to predict the mean miss rate of the cache, but also to provide an empirical estimate of its actual distribution. Finally, a new concept of primed cache is introduced to simulate large caches by the sampling-based method.> Subhasis Laha, Janak H. Patel, Ravishankar K. Iyer |
IEEE Trans. Computers | 3 |
| 1987 | A Measurement-Based Study of Concurrency in a Multiprocessor
P. G. McGuire, Ravishankar K. Iyer |
ICPP | 2 |
| 1987 | Measurement-Based Analysis of Error LatencyabstractThis paper demonstrates a practical methodology for the study of error latency under a real workload. The method is illustrated with sampled data on the physical memory activity, gathered by hardware instrumentation on a VAX 11/780 during the normal workload cycle of the installation. These data are used to simulate fault occurrence and to reconstruct the error discovery process in the system. The technique provides a means to study the system under different workloads and for multiple days. An approach to determine the percentage of undiscovered errors is also developed and a verification of the entire methodology is performed. This study finds that the mean error latency, in the memory containing the operating system, varies by a factor of 10 to 1 (in hours) between the low and high workloads. It is found that of all errors occurring within a day, 70 percent are detected in the same day, 82 percent within the following day, and 91 percent within the third day. The increase in failure rate due to latency is not so much a function of remaining errors but is dependent on whether or not there is a latent error. Ram Chillarege, Ravishankar K. Iyer |
IEEE Trans. Computers | 2 |
| 1986 | Error Propagation in a Digital Avionic Processor: A Simulation-Based Study
D. Lomelino, Ravishankar K. Iyer |
RTSS | 2 |
| 1986 | A Measurement-Based Model for Workload Dependence of CPU ErrorsabstractThis paper proposes and validates a methodology to measure explicitly the increase in the risk of a processor error with increasing workload. By relating the occurrence of a CPU related error to the system activity just prior to the occurrence of an error, the approach measures the dynamic CPU workload/failure relationship. The measurements show that the probability of a CPU related error (the load hazard) increases nonlinearly with increasing workload; i.e., the CPU rapidly deteriorates as end points are reached. The load hazard is observed to be most sensitive to system CPU utilization, the I/O rate, and the interrupt rates. The results are significant because they indicate that it may not be useful to push a system close to its performance limits (the previously accepted operating goal) since what we gain in slightly improved performance is more than offset by the degradation in reliability. Importantly, they also indicate that conventional reliability models need to be reevaluated so as to take system work-load explicity into account. Ravishankar K. Iyer, David J. Rossetti |
IEEE Trans. Computers | 1 |
| 1986 | Measurement and Modeling of Computer Reliability as Affected by System ActivityabstractThis paper demonstrates a practical approach to the study of the failure behavior of computer systems. Particular attention is devoted to the analysis of permanent failures. A number of important techniques, which may have general applicability in both failure and workload analysis, are brought together in this presentation. These include: smeared averaging of the workload data, clustering of like failures, and joint analysis of workload and failures. Approximately 17 percent of all failures affecting the CPU were estimated to be permanent. The manifestation of a permanent failure was found to be strongly correlated with the level and type of workload prior to the failure. Although, in strict terms, the results only relate to the manifestation of permanent failures and not to their occurrence, there are strong indications that permanent failures are both caused and discovered by increased activity. More measurements and experiments are necessary to determine their respective contributions to the measured workload/failure relationship. Ravishankar K. Iyer, David J. Rossetti, Mei-Chen Hsueh |
ACM Trans. Comput. Syst. | 1 |
| 1985 | The Effect of System Workload on Error Latency: An Experimental StudyabstractIn this paper, a methodology for determining and characterizing error latency is developed. The method is based on real workload data, gathered by an experiment instrumented on a VAX 11/780 during the normal workload cycle of the installation. This is the first attempt at jointly studying error latency and workload variations in a full production system. Distributions of error latency were generated by simulating the occurrence of faults under varying workload conditions. A family of error latency distributions so generated illustrate that error latency is not so much a function of when in time a fault occurred but rather a function of the workload that followed the failure. The study finds that the mean error latency varies by a 1 to 8 (hours) ratio between high and low workloads. The method is general and can be applied to any system. Ram Chillarege, Ravishankar K. Iyer |
SIGMETRICS | 2 |
| 1985 | Effect of System Workload on Operating System Reliability: A Study on IBM 3081abstractThis paper presents an analysis of operating system failures on an IBM 3081 running VM/SP. We find three broad categories of software failures: error handling (ERH), program control or logic (CTL), and hardware related (HS); it is found that more than 25 percent of software failures occur in the hardware/software interface. Measurements show that results on software reliability cannot be considered representative unless the system workload is taken into account. For example, it is shown that the risk of a software failure increases in a nonlinear fashion with the amount of interactive processing, as measured by parameters such as the paging rate and the amount of overhead (operating system CPU time). The overall CPU execution rate, although measured to be close to 100 percent most of the time, is not found to correlate strongly with the occurrence of failures. The paper discusses possible reasons for the observed workload failure dependency based on detailed investigations of the failure data. Ravishankar K. Iyer, David J. Rossetti |
IEEE Trans. Software Eng. | 1 |
| 1985 | Hardware-Related Software Errors: Measurement and AnalysisabstractThis paper describes an analysis of hardware-related software (HW/SW) errors on an MVS/SP operating system at Stanford University. The analysis procedure demonstrates a methodology for evaluating the interaction between hardware and software as it relates to system reliability. The paper examines the operating system's handling of HW/SW errors and also the effectiveness of recovery management. Nearly 35 percent of all observed software failures were found to be hareware-related. The analysis shows that the operating system is seldom able to diagnose that a software error may be hardware-related. The impact of HW/SW errors on the system is evaluated by measuring the effectiveness of system recovery in containing the propagation of HW/SW errors. The system failure probability for HW/SW errors is close to three times that for software errors in general. The observed HW/SW errors are seen to have a specific pattern, suggesting the possibility of the use of such error patterns for intelligent error prediction and recovery. Ravishankar K. Iyer, Paola Velardi |
IEEE Trans. Software Eng. | 1 |
| 1984 | Reliability Evaluation of Fault-Tolerant Systems - Effect of Variability in Failure RatesabstractIn this correspondence models for the variation in system reliability due to uncertainty in failure rate estimation are developed. Two techniques are proposed. The first is exact and is based on the complete distribution of the failure rate. The second is an approximation and employs only the first and second moments. The application of these models in reliability analysis is then discussed and illustrated with numerical examples. Ravishankar K. Iyer |
IEEE Trans. Computers | 1 |
| 1984 | A Study of Software Failures and Recovery in the MVS Operating SystemabstractThis paper describes an analysis of system detected software errors on the MVS operating system at the Center for Information Technology (CIT), Stanford University. The analysis procedure demonstrates a methodology by which systems with automatic recovery features can be evaluated. Most common error categories are determined and related to the program in execution at the time of the error. The severity of the error is measured by evaluating the criticality of the program for continued system operation. The system recovery and error correction features are then analyzed and an estimate of the system fault tolerance to errors of different levels of severity is made. Paola Velardi, Ravishankar K. Iyer |
IEEE Trans. Computers | 2 |
| 1982 | A Statistical Failure/Load Relationship: Results of a Multicomputer StudyabstractIn this correspondence we present a statistical model which relates mean computer failure rates to level of system activity. Our analysis reveals a strong statistical dependency of both hardware and software component failure rates on several common measures of utilization (specifically CPU utilization, I/O initiation, paging, and job-step initiation rates). We establish that this effect is not dominated by a specific component type, but exists across the board in the two systems studied. Our data covers three years of normal operation (including significant upgrades and reconfigurations) for two large Stanford University computer complexes. The complexes, which are composed of IBM mainframe equipment of differing models and vintage, run similar operating systems and provide the same interface and capability to their users. The empirical data comes from identically structured and maintained failure logs at the two sites along with IBM OS/VS2 operating system performance/load records. Ravishankar K. Iyer, Steven E. Butner, Edward J. McCluskey |
IEEE Trans. Computers | 1 |