VLDB 2026 Research / reviewers in the wild / expert
Saurabh Jha
dblp:130/2298
· DBLP profile ↗
33ranked-venue papers
12as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 8 first-author · 12 since 2021Security and privacy · 7 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Praxis: Integrating Program Analysis with Observability for Root-Cause AnalysisabstractUnresolved production cloud incidents cost an average of over $2M per hour. This paper introduces PRAXIS, an orchestrator that manages and deploys an agentic workflow for diagnosing code- and configuration-caused cloud incidents. PRAXIS employs an LLM-driven structured traversal over two types of graph: (1) a service dependency graph (SDG) that captures microservice-level dependencies; and (2) a hammock-block program dependence graph (PDG) that captures code-level dependencies for each microservice. Compared to state-of-the-art ReAct baselines, PRAXIS improves RCA accuracy by up to 6.3x while reducing token consumption by 5.3x. PRAXIS is demonstrated on a set of 30 comprehensive real-world incidents that is being compiled into an RCA benchmark. Shengkun Cui, Rahul Krishna, Saurabh Jha, Ravishankar K. Iyer |
DSN | 3 |
| 2025 | CPU-Limits kill Performance: Time to rethink Resource Control
Chirag C. Shetty, Sarthak Chakraborty, Hubertus Franke, Larisa Shwartz, Chandrasekhar Narayanaswami 0001, Indranil Gupta, Saurabh Jha |
SoCC | 7 |
| 2025 | ITBench: Evaluating AI Agents across Diverse Real-World IT Automation TasksabstractRealizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a framework that offers a systematic methodology for benchmarking AI agents to address real-world IT automation tasks. Our initial release targets three key areas: Site Reliability Engineering (SRE), Compliance and Security Operations (CISO), and Financial Operations (FinOps). The design enables AI researchers to understand the challenges and opportunities of AI agents for IT automation with push-button workflows and interpretable metrics. IT-Bench includes an initial set of 102 real-world scenarios, which can be easily extended by community contributions. Our results show that agents powered by state-of-the-art models resolve only 11.4% of SRE scenarios, 25.2% of CISO scenarios, and 25.8% of FinOps scenarios (excluding anomaly detection). For FinOps-specific anomaly detection (AD) scenarios, AI agents achieve an F1 score of 0.35. We expect ITBench to be a key enabler of AI-driven IT automation that is correct, safe, and fast. IT-Bench, along with a leaderboard and sample agent implementations, is available at https://github.com/ibm/itbench. Saurabh Jha, Rohan R. Arora, Yuji Watanabe, Takumi Yanagawa, Yinfang Chen, Jackson Clark, Bhavya, Mudit Verma, Hirokuni Kitahara, Noah Zheutlin, Saki Takano, Divya Pathak, Felix George, Xinbo Wu, Bekir O. Turkkan, Gerard Vanloo, Michael Nidd, Oishik Chatterjee, Pranjal Gupta, Suranjana Samanta, Pooja Aggarwal, Rong Lee, Jae-wook Ahn, Debanjana Kar, Amit M. Paradkar, Yu Deng 0004, Pratibha Moogi, Prateeti Mohapatra, Naoki Abe, Chandrasekhar Narayanaswami 0001, Tianyin Xu, Lav R. Varshney, Ruchi Mahindru, Anca Sailer, Larisa Shwartz, Daby M. Sow, Nicholas C. Fuller, Ruchir Puri |
ICML | 1 |
| 2025 | Page Migration for Hardware Memory Disaggregation Across a NetworkabstractHardware memory disaggregation (HMD) is an emerging technology that enables access to remote memory, thereby creating expansive memory pools and reducing memory underutilization in datacenters.However, a significant challenge arises when accessing remote memory over a network: increased contention that can lead to severe application performance degradation.To reduce the performance penalty of using remote memory, the operating system uses page migration to promote frequently accessed pages closer to the processor.However, previously proposed page migration mechanisms do not achieve the best performance in HMD systems because of obliviousness to variable page transfer costs that occur due to network contention.To address these limitations, we present INDIGO: a network-aware page migration framework that uses novel page telemetry and a learning-based approach for network adaptation.We implemented INDIGO in the Linux kernel and evaluated it with common cloud and HPC applications on a real disaggregated memory system prototype.Our evaluation shows that INDIGO offers up to 50-70% improvement in application performance compared to other state-of-the-art page migration policies and reduces network traffic up to 2×. Archit Patke, Christian Pinto, Saurabh Jha, Haoran Qiu, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ICS | 3 |
| 2025 | STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern CloudsabstractIn cloud-scale systems, failures are the norm. A distributed computing cluster exhibits hundreds of machine failures and thousands of disk failures; software bugs and misconfigurations are reported to be more frequent. The demand for autonomous, AI-driven reliability engineering continues to grow, as existing human-in-the-loop practices can hardly keep up with the scale of modern clouds. This paper presents STRATUS, an LLM-based multi-agent system for realizing autonomous Site Reliability Engineering (SRE) of cloud services. STRATUS consists of multiple specialized agents (e.g., for failure detection, diagnosis, mitigation), organized in a state machine to assist system-level safety reasoning and enforcement. We formalize a key safety specification of agentic SRE systems like STRATUS, termed Transactional No-Regression (TNR), which enables safe exploration and iteration. We show that TNR can effectively improve autonomous failure mitigation. STRATUS significantly outperforms state-of-the-art SRE agents in terms of success rate of failure mitigation problems in AIOpsLab and ITBench (two SRE benchmark suites), by at least 1.5 times across various models. STRATUS shows a promising path toward practical deployment of agentic systems for cloud reliability. Yinfang Chen, Jackson Clark, Yiming Su, Noah Zheutlin, Bhavya, Rohan R. Arora, Yu Deng 0004, Saurabh Jha, Tianyin Xu |
NeurIPS | 9 |
| 2025 | Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUsabstractThis study characterizes GPU resilience in Delta, a large-scale AI system that consists of 1,056 A100 and H100 GPUs, with over 1,300 petaflops of peak throughput. We used 2.5 years of operational data (11.7 million GPU hours) on GPU errors. Our major findings include: (i) H100 GPU memory resilience is worse than A100 GPU memory, with 3.2x lower per-GPU MTBE for memory errors, (ii) The GPU memory error-recovery mechanisms on H100 GPUs are insufficient to handle the increased memory capacity, (iii) H100 GPUs demonstrate significantly improved GPU hardware resilience over A100 GPUs with respect to critical hardware components, (iv) GPU errors on both A100 and H100 GPUs frequently result in job failures due to the lack of robust recovery mechanisms at the application level, and (v) We project the impact of GPU node availability on larger-scales and find that significant overprovisioning of 5% is necessary to handle GPU failures. Shengkun Cui, Archit Patke, Aditya Ranjan, Ziheng Chen 0006, Phuong Cao, Gregory H. Bauer, Brett M. Bode, Catello Di Martino, Saurabh Jha, Chandrasekhar Narayanaswami 0001, Daby M. Sow, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SC | 10 |
| 2024 | SAM: Subseries Augmentation-Based Meta-Learning for Generalizing AIOps Models in Multi-Cloud MigrationabstractIn the context of cloud computing, enterprises are increasingly adopting multi-cloud strategies to enhance performance, ensure cost efficiency, and avoid vendor lock-in. This trend presents a significant challenge for the migration of AI for IT operations (AIOps) models across different cloud providers due to variations in architecture, performance, and data distribution. Traditional methods of re-training AIOps models for new cloud environments are labor-intensive and delay deployment. To address this issue, we introduce a novel framework called SAM (Subseries Augmentation-based Meta-learning), which facilitates seamless model migration between clouds without the need for re-training from scratch. SAM leverages data augmentation and meta-learning to efficiently adapt AIOps models to new cloud environments. It has proven effective in adapting anomaly detectors across various config-urations over both public and simulated datasets. We believe that SAM can also be adapted to other AI models used for automating IT tasks such as alerting and resource scaling. Paulito Palmes, Saurabh Jha, Bekir O. Turkkan, Gerard Vanloo, Frank Bagehorn, Chandrasekhar Narayanaswami 0001, Larisa Shwartz, Naoki Abe, Yu Deng 0004, Daby M. Sow |
CLOUD | 3 |
| 2024 | Optimizing IT FinOps and Sustainability through Unsupervised Workload CharacterizationabstractThe widespread adoption of public and hybrid clouds, along with elastic resources and various automation tools for dynamic deployment, has accelerated the rapid provisioning of compute resources as needed. Despite these advancements, numerous resources persist unnecessarily due to factors such as poor digital hygiene, risk aversion, or the absence of effective tools, resulting in substantial costs and energy consumption. Existing threshold-based techniques prove inadequate in effectively addressing this challenge. To address this issue, we propose an unsupervised machine learning framework to automatically identify resources that can be de-provisioned completely or summoned on a schedule. Application of this approach to enterprise data has yielded promising initial results, facilitating the segregation of productive workloads with recurring demands from non-productive ones. Rohan R. Arora, Saurabh Jha, Chandrasekhar Narayanaswami 0001, Cheuk Lam, Jerrold Leichter, Yu Deng 0004, Daby M. Sow |
AAAI | 3 |
| 2024 | Queue Management for SLO-Oriented Large Language Model ServingabstractLarge language model (LLM) serving is becoming an increasingly critical workload for cloud providers. Existing LLM serving systems focus on interactive requests, such as chatbots and coding assistants, with tight latency SLO requirements. However, when such systems execute batch requests that have relaxed SLOs along with interactive requests, it leads to poor multiplexing and inefficient resource utilization. To address these challenges, we propose QLM, a queue management system for LLM serving. QLM maintains batch and interactive requests across different models and SLOs in a request queue. Optimal ordering of the request queue is critical to maintain SLOs while ensuring high resource utilization. To generate this optimal ordering, QLM uses a Request Waiting Time (RWT) Estimator that estimates the waiting times for requests in the request queue. These estimates are used by a global scheduler to orchestrate LLM Serving Operations (LSOs) such as request pulling, request eviction, load balancing, and model swapping. Evaluation on heterogeneous GPU devices and models with real-world LLM serving dataset shows that QLM improves SLO attainment by 40-90% and throughput by 20-400% while maintaining or improving device utilization compared to other state-of-the-art LLM serving systems. QLM's evaluation is based on the production requirements of a cloud provider. QLM is publicly available at https://www.github.com/QLM-project/QLM. Archit Patke, Dhemath Reddy, Saurabh Jha, Haoran Qiu, Christian Pinto, Chandrasekhar Narayanaswami 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SoCC | 3 |
| 2024 | iPrism: Characterize and Mitigate Risk by Quantifying Change in Escape Routes
Shengkun Cui, Saurabh Jha, Ziheng Chen 0006, Zbigniew T. Kalbarczvk, Ravishankar K. Iyer |
DSN | 2 |
| 2024 | Power-aware Deep Learning Model Serving with μ-Serve
Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Tamer Basar, Ravishankar K. Iyer |
USENIX ATC | 5 |
| 2024 | Blue Waters system and component reliabilityabstractSummary The Blue Waters system, installed in 2012 at NCSA, has the largest component count of any system Cray has built. Blue Waters includes a mix of dual‐socket CPU (XE) and single‐socket CPU, single GPU (XK) nodes. The primary storage is provided by Cray's Sonexion/ClusterStor Luster storage system delivering 35 PB (raw) storage at 1 TB/s. The statistical failure rates over time for each component including CPU, DIMM, GPU, disk drive, power supply, blower, etc and their impact on higher level failure rates for individual nodes and the systems as a whole are presented in detail, with a particular emphasis on identifying any increases in rate that might indicate the right‐side of the expected bathtub curve has been reached. Strategies employed by NCSA and Cray for minimizing the impact of component failure, such as the preemptive removal of suspect disk drives, are also presented. Brett M. Bode, David King, Celso L. Mendes, William T. Kramer, Saurabh Jha, Roger Ford, Justin Davis, Steven Dramstad |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | Fault Injection Based Interventional Causal Learning for Distributed ApplicationsabstractWe apply the machinery of interventional causal learning with programmable interventions to the domain of applications management. Modern applications are modularized into interdependent components or services (e.g. microservices) for ease of development and management. The communication graph among such components is a function of application code and is not always known to the platform provider. In our solution we learn this unknown communication graph solely using application logs observed during the execution of the application by using fault injections in a staging environment. Specifically, we have developed an active (or interventional) causal learning algorithm that uses the observations obtained during fault injections to learn a model of error propagation in the communication among the components. The “power of intervention” additionally allows us to address the presence of confounders in unobserved user interactions. We demonstrate the effectiveness of our solution in learning the communication graph of well-known microservice application benchmarks. We also show the efficacy of the solution on a downstream task of fault localization in which the learned graph indeed helps to localize faults at runtime in a production environment (in which the location of the fault is unknown). Additionally, we briefly discuss the implementation and deployment status of a fault injection framework which incorporates the developed technology. Qing Wang 0016, Jesus Rios, Saurabh Jha, Karthikeyan Shanmugam 0001, Frank Bagehorn, Robert Filepp, Naoki Abe, Larisa Shwartz |
AAAI | 3 |
| 2022 | Localizing and Explaining Faults in Microservices Using Distributed TracingabstractFinding the exact location of a fault in a large distributed microservices application running in containerized cloud environments can be very difficult and time-consuming. We present a novel approach that uses distributed tracing to automatically detect, localize and aid in explaining application-level faults. We demonstrate the effectiveness of our proposed approach by injecting faults into a well-known microservice-based benchmark application. Our experiments demonstrated that the proposed fault localization algorithm correctly detects and localize the microservice with the injected fault. We also compare our approach with other fault localization methods. In particular, we empirically show that our method outperforms methods in which a graph model of error propagation is used for inferring fault locations using error logs. Our work illustrates the value added by distributed tracing for localizing and explaining faults in microservices. Jesus Rios, Saurabh Jha, Larisa Shwartz |
CLOUD | 2 |
| 2022 | Exploiting Temporal Data Diversity for Detecting Safety-critical Faults in AV Compute SystemsabstractSilent data corruption caused by random hardware faults in autonomous vehicle (AV) computational elements is a significant threat to vehicle safety. Previous research has explored design diversity, data diversity, and duplication techniques to detect such faults in other safety-critical domains. However, these are challenging to use for AVs in practice due to significant resource overhead and design complexity. We propose, DiverseAV, a low-cost data-diversity-based redundancy technique for detecting safety-critical random hardware faults in computational elements. DiverseAV introduces data-diversity between the redundant agents by exploiting the temporal semantic consistency available in the AV sensor data. DiverseAV is a black-box technique that offers a plug-and-play solution as it requires no knowledge of the internals of the AI agent responsible for executing driving decisions, requiring little to no modification to the agent itself for achieving high coverage of transient and permanent hardware faults. It is commercially viable because it avoids software modifications to agents that are costly in terms of development and testing time. Specifically, DiverseAV distributes the sensor data between the two software agents in a round-robin manner. As a result, the sensor data for two consecutive time steps are semantically similar in terms of their worldview but significantly different at the bit level, thus ensuring the state and data diversity between the two agents necessary for detecting faults. We demonstrate DiverseAV using an open-source self-driving AI agent which is controlling a car in an open-source world simulator. Saurabh Jha, Shengkun Cui, Timothy Tsai 0002, Siva Kumar Sastry Hari, Michael B. Sullivan 0001, Zbigniew T. Kalbarczyk, Stephen W. Keckler, Ravishankar K. Iyer |
DSN | 1 |
| 2022 | A fault injection platform for learning AIOps modelsabstractIn today’s IT environment with a growing number of costly outages, increasing complexity of the systems, and availability of massive operational data, there is a strengthening demand to effectively leverage Artificial Intelligence and Machine Learning (AI/ML) towards enhanced resiliency. In this paper, we present an automatic fault injection platform to enable and optimize the generation of data needed for building AI/ML models to support modern IT operations. The merits of our platform include the ease of use, the possibility to orchestrate complex fault scenarios and to optimize the data generation for the modeling task at hand. Specifically, we designed a fault injection service that (i) combines fault injection with data collection in a unified framework, (ii) supports hybrid and multi-cloud environments, and (iii) does not require programming skills for its use. Our current implementation covers the most common fault types both at the application and infrastructure levels. The platform also includes some AI capabilities. In particular, we demonstrate the interventional causal learning capability currently available in our platform. We show how our system is able to learn a model of error propagation in a micro-service application in a cloud environment (when the communication graph among micro-services is unknown and only logs are available) for use in subsequent applications such as fault localization. Frank Bagehorn, Jesus Rios, Saurabh Jha, Robert Filepp, Larisa Shwartz, Naoki Abe |
ASE | 3 |
| 2022 | An evolutionary algorithm based feature selection and fuzzy rule reduction technique for the prediction of skin cancerabstractSummary In current years, the death rate from skin cancers (SCs) tends to develop pretty. Various research verified that SC rank third as a deadliest disease, after breast and lung cancer. It will become vital to diagnose this malignancy at an early stage. The objective of this research is to mix machine learning and soft computing techniques to gain higher accuracy within the prediction of SC. To play out the exploration work, we utilized two data sets, one from “Save Life Hospital,” India, and the other is the UCI repository skin cancer data set. In this article, three meta‐heuristic algorithms, the FS_GA, the FS_PSO, and the FS_ACO, were used to select the best features from the data set provided to it. The AFRG_algorithm generates a set of fuzzy rules automatically and the RR_algorithm reduces certain fuzzy rules from the fuzzy system. For the SCC_dataset, the end accuracy obtained was 97.67%, 98.45%, and 99.22%. For the UCI_dataset, the end accuracy obtained was 98.81%, 99.72%, and 99.67%. Experimental results on the used datasets show that the proposed method strikingly improves the forecast exactitude of skin malignancy. Saurabh Jha, Ashok Kumar Mehta |
Concurr. Comput. Pract. Exp. | 1 |
| 2022 | Data-Driven Application-Oriented Reliability Model of a High-Performance Computing SystemabstractReliability analysis and performance evaluation are complementary methods to quantify nonfunctional aspects of a system. However, a range of factors such as concurrency and heterogeneity quickly exacerbate the state-space explosion problem when attempting detailed system-level modeling and simulation of high-performance computing (HPC) systems. To overcome these impediments to modeling and analysis, this article develops a hierarchical model of an application that implements checkpointing running in an HPC environment subject to application, network, and system-wide outages. The modeling approach ensures that the number of states is linear in the number of checkpoints and possesses a low constant factor for the number of recovery states most relevant to the external influences contributing to degraded application performance. We illustrate the types of analysis enabled by the model through a series of examples with parameters determined empirically from data logs of the Blue Waters supercomputer located at the University of Illinois at Urbana–Champaign. A comprehensive comparative analysis of the model parameters indicates that lowering the failure rate of network nodes would most significantly reduce application downtime. We also discuss how the modeling approach can be used to objectively assess both current and hypothetical future systems to identify competitive designs and enhancements. Bentolhoda Jafary, Saurabh Jha, Lance Fiondella, Ravishankar K. Iyer |
IEEE Trans. Reliab. | 2 |
| 2021 | BayesPerf: minimizing performance monitoring errors using Bayesian statisticsabstractHardware performance counters (HPCs) that measure low-level architectural and microarchitectural events provide dynamic contextual information about the state of the system. However, HPC measurements are error-prone due to non determinism (e.g., undercounting due to event multiplexing, or OS interrupt-handling behaviors). In this paper, we present BayesPerf, a system for quantifying uncertainty in HPC measurements by using a domain-driven Bayesian model that captures microarchitectural relationships between HPCs to jointly infer their values as probability distributions. We provide the design and implementation of an accelerator that allows for low-latency and low-power inference of the BayesPerf model for x86 and ppc64 CPUs. BayesPerf reduces the average error in HPC measurements from 40.1% to 7.6% when events are being multiplexed. The value of BayesPerf in real-time decision-making is illustrated with a simple example of scheduling of PCIe transfers. Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ASPLOS | 2 |
| 2021 | Delay sensitivity-driven congestion mitigation for HPC systemsabstractModern high-performance computing (HPC) systems concurrently execute multiple distributed applications that contend for the high-speed network leading to congestion. Consequently, application runtime variability and suboptimal system utilization are observed in production systems. To address these problems, we propose Netscope, a congestion mitigation framework based on a novel delay sensitivity metric. Delay sensitivity of an application is used to quantify the impact of congestion on its runtime. Netscope uses delay sensitivity estimates to drive a congestion mitigation mechanism to selectively throttle applications that are less susceptible to congestion. We evaluate Netscope on two Cray Aries systems, including a production supercomputer, on common scientific applications. Our evaluation shows that Netscope has a low training cost and accurately estimates the impact of congestion on application runtime with a correlation between 0.7 and 0.9. Moreover, Netscope reduces application tail runtime increase by up to 16.3x while improving the median system utility by 12%. Archit Patke, Saurabh Jha, Haoran Qiu, Jim M. Brandt, Ann C. Gentile, Joe Greenseid, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ICS | 2 |
| 2020 | ML-Driven Malware that Targets AV SafetyabstractEnsuring the safety of autonomous vehicles (AVs) is critical for their mass deployment and public adoption. However, security attacks that violate safety constraints and cause accidents are a significant deterrent to achieving public trust in AVs, and that hinders a vendor's ability to deploy AVs. Creating a security hazard that results in a severe safety compromise (for example, an accident) is compelling from an attacker's perspective. In this paper, we introduce an attack model, a method to deploy the attack in the form of smart malware, and an experimental evaluation of its impact on production-grade autonomous driving software. We find that determining the time interval during which to launch the attack is{ critically} important for causing safety hazards (such as collisions) with a high degree of success. For example, the smart malware caused 33X more forced emergency braking than random attacks did, and accidents in 52.6% of the driving simulations. Saurabh Jha, Shengkun Cui, Subho S. Banerjee, James Cyriac, Timothy Tsai 0002, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 1 |
| 2020 | Inductive-bias-driven Reinforcement Learning For Efficient Schedules in Heterogeneous ClustersabstractThe problem of scheduling of workloads onto heterogeneous processors (e.g., CPUs, GPUs, FPGAs) is of fundamental importance in modern data centers. Current system schedulers rely on application/system-specific heuristics that have to be built on a case-by-case basis. Recent work has demonstrated ML techniques for automating the heuristic search by using black-box approaches which require significant training data and time, which make them challenging to use in practice. This paper presents Symphony, a scheduling framework that addresses the challenge in two ways: (i) a domain-driven Bayesian reinforcement learning (RL) model for scheduling, which inherently models the resource dependencies identified from the system architecture; and (ii) a sampling-based technique to compute the gradients of a Bayesian model without performing full probabilistic inference. Together, these techniques reduce both the amount of training data and the time required to produce scheduling policies that significantly outperform black-box approaches by up to 2.2{\texttimes}. Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ICML | 2 |
| 2020 | AV-FUZZER: Finding Safety Violations in Autonomous Driving SystemsabstractThis paper proposes AV-FUZZER, a testing framework, to find the safety violations of an autonomous vehicle (AV) in the presence of an evolving traffic environment. We perturb the driving maneuvers of traffic participants to create situations in which an AV can run into safety violations. To optimally search for the perturbations to be introduced, we leverage domain knowledge of vehicle dynamics and genetic algorithm to minimize the safety potential of an AV over its projected trajectory. The values of the perturbation determined by this process provide parameters that define participants' trajectories. To improve the efficiency of the search, we design a local fuzzer that increases the exploitation of local optima in the areas where highly likely safety-hazardous situations are observed. By repeating the optimization with significantly different starting points in the search space, AV-FUZZER determines several diverse AV safety violations. We demonstrate AV-FUZZER on an industrial-grade AV platform, Baidu Apollo, and find five distinct types of safety violations in a short period of time. In comparison, other existing techniques can find at most two. We analyze the safety violations found in Apollo and discuss their overarching causes. Guanpeng Li, Saurabh Jha, Timothy Tsai 0002, Michael B. Sullivan 0001, Siva Kumar Sastry Hari, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
ISSRE | 3 |
| 2020 | Measuring Congestion in High-Performance Datacenter Interconnects
Saurabh Jha, Archit Patke, Jim M. Brandt, Ann C. Gentile, Benjamin Lim, Michael T. Showerman, Gregory H. Bauer, Larry Kaplan, Zbigniew T. Kalbarczyk, William T. Kramer, Ravishankar K. Iyer |
NSDI | 1 |
| 2020 | FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented Microservices
Haoran Qiu, Subho S. Banerjee, Saurabh Jha, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
OSDI | 3 |
| 2020 | Live forensics for HPC systems: a case study on distributed storage systemsabstractLarge-scale high-performance computing systems frequently experience a wide range of failure modes, such as reliability failures (e.g., hang or crash), and resource overload-related failures (e.g., congestion collapse), impacting systems and applications. Despite the adverse effects of these failures, current systems do not provide methodologies for proactively detecting, localizing, and diagnosing failures. We present Kaleidoscope, a near real-time failure detection and diagnosis framework, consisting of of hierarchical domain-guided machine learning models that identify the failing components, the corresponding failure mode, and point to the most likely cause indicative of the failure in near real-time (within one minute of failure occurrence). Kaleidoscope has been deployed on Blue Waters supercomputer and evaluated with more than two years of production telemetry data. Our evaluation shows that Kaleidoscope successfully localized 99.3% and pinpointed the root causes of 95.8% of 843 real-world production issues, with less than 0.01% runtime overhead. Saurabh Jha, Shengkun Cui, Subho S. Banerjee, Tianyin Xu, Jeremy Enos, Michael T. Showerman, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
SC | 1 |
| 2019 | ML-Based Fault Injection for Autonomous Vehicles: A Case for Bayesian Fault InjectionabstractThe safety and resilience of fully autonomous vehicles (AVs) are of significant concern, as exemplified by several headline-making accidents. While AV development today involves verification, validation, and testing, end-to-end assessment of AV systems under accidental faults in realistic driving scenarios has been largely unexplored. This paper presents DriveFI, a machine learning-based fault injection engine, which can mine situations and faults that maximally impact AV safety, as demonstrated on two industry-grade AV technology stacks (from NVIDIA and Baidu). For example, DriveFI found 561 safety-critical faults in less than 4 hours. In comparison, random injection experiments executed over several weeks could not find any safety-critical faults. Saurabh Jha, Subho S. Banerjee, Timothy Tsai 0002, Siva Kumar Sastry Hari, Michael B. Sullivan 0001, Zbigniew T. Kalbarczyk, Stephen W. Keckler, Ravishankar K. Iyer |
DSN | 1 |
| 2018 | Characterizing Supercomputer Traffic Networks Through Link-Level AnalysisabstractWe present techniques for characterizing bandwidth and congestion characteristics of supercomputer High-Speed Networks (HSN). By utilizing a link-level perspective, we gain generality over analyses which are tied to specific topologies. We illustrate these techniques using five months of a Blue Waters production dataset consisting of network utilization and congestion counters. We find that: i) execution time of the communication-heavy applications is highly correlated to network stalls observed in the network topology and increase in application runtime can be as high as 1.7x with nominal increase in stalls, ii) heterogeneity in the available link bandwidth in the network can lead to backpressure and congestion even when the network is not underprovisioned, and (iii) links connected to I/O nodes are no more likely to observe congestion during operational hours than any other link in the system. Saurabh Jha, Jim M. Brandt, Ann C. Gentile, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
CLUSTER | 1 |
| 2018 | Hands Off the Wheel in Autonomous Vehicles?: A Systems Perspective on over a Million Miles of Field DataabstractAutonomous vehicle (AV) technology is rapidly becoming a reality on U.S. roads, offering the promise of improvements in traffic management, safety, and the comfort and efficiency of vehicular travel. The California Department of Motor Vehicles (DMV) reports that between 2014 and 2017, manufacturers tested 144 AVs, driving a cumulative 1,116,605 autonomous miles, and reported 5,328 disengagements and 42 accidents involving AVs on public roads. This paper investigates the causes, dynamics, and impacts of such AV failures by analyzing disengagement and accident reports obtained from public DMV databases. We draw several conclusions. For example, we find that autonomous vehicles are 15 - 4000× worse than human drivers for accidents per cumulative mile driven; that drivers of AVs need to be as alert as drivers of non-AVs; and that the AVs' machine-learning-based systems for perception and decision-and-control are the primary cause of 64% of all disengagements. Subho S. Banerjee, Saurabh Jha, James Cyriac, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
DSN | 2 |
| 2018 | Resiliency of HPC Interconnects: A Case Study of Interconnect Failures and Recovery in Blue WatersabstractAvailability of the interconnection network in high-performance computing (HPC) systems is fundamental to sustaining the continuous execution of applications at scale. When failures occur, interconnect recovery mechanisms orchestrate complex operations to recover network connectivity between the nodes. As the scale and design complexity of HPC systems increase, so does the system's susceptibility to failures during execution of interconnect-recovery procedures. This study characterizes the recovery procedures of the Gemini interconnect network, the largest Gemini network built by Cray, on Blue Waters, a 13.3 petaflop supercomputer at the National Center for Supercomputing Applications (NCSA). We propose a propagation model that captures interconnect failures and recovery procedures to help understand types of failures and their propagation in both the system and applications during recovery. The measurements show that recovery procedures occur very frequently and that the unsuccessful execution of recovery procedures, when additional failures occur during recovery, causes system-wide outages (SWOs, 28 out of 101) and application failures (3.4 percent of all running applications). Saurabh Jha, Valerio Formicola, Catello Di Martino, Mark Dalton, William T. Kramer, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2017 | Holistic Measurement-Driven System AssessmentabstractIn high-performance computing systems, application performance and throughput are dependent on a complex interplay of hardware and software subsystems and variable workloads with competing resource demands. Data-driven insights into the potentially widespread scope and propagationof impact of events, such as faults and contention for shared resources, can be used to drive more effective use of resources, for improved root cause diagnosis, and for predicting performance impacts. We present work developing integrated capabilities for holistic monitoring and analysis to understand and characterize propagation of performance-degrading events. These characterizations can be used to determine and invoke mitigating responses by system administrators, applications, and system software. Saurabh Jha, Jim M. Brandt, Ann C. Gentile, Zbigniew T. Kalbarczyk, Gregory H. Bauer, Jeremy Enos, Michael T. Showerman, Larry Kaplan, Brett M. Bode, Annette Greiner, Amanda Bonnie, Mike Mason, Ravishankar K. Iyer, William T. Kramer |
CLUSTER | 1 |
| 2015 | Improving Main Memory Hash Joins on Intel Xeon Phi Processors: An Experimental ApproachabstractModern processor technologies have driven new designs and implementations in main-memory hash joins. Recently, Intel Many Integrated Core (MIC) co-processors (commonly known as Xeon Phi) embrace emerging x86 single-chip many-core techniques. Compared with contemporary multi-core CPUs, Xeon Phi has quite different architectural features: wider SIMD instructions, many cores and hardware contexts, as well as lower-frequency in-order cores. In this paper, we experimentally revisit the state-of-the-art hash join algorithms on Xeon Phi co-processors. In particular, we study two camps of hash join algorithms: hardware-conscious ones that advocate careful tailoring of the join algorithms to underlying hardware architectures and hardware-oblivious ones that omit such careful tailoring. For each camp, we study the impact of architectural features and software optimizations on Xeon Phi in comparison with results on multi-core CPUs. Our experiments show two major findings on Xeon Phi, which are quantitatively different from those on multi-core CPUs. First, the impact of architectural features and software optimizations has quite different behavior on Xeon Phi in comparison with those on the CPU, which calls for new optimization and tuning on Xeon Phi. Second, hardware oblivious algorithms can outperform hardware conscious algorithms on a wide parameter window. These two findings further shed light on the design and implementation of query processing on new-generation single-chip many-core technologies. Saurabh Jha, Bingsheng He, Mian Lu, Xuntao Cheng, Huynh Phung Huynh |
Proc. VLDB Endow. | 1 |
| 2013 | Exploiting data parallelism in the yConvex hypergraph algorithm for image representation using GPGPUsabstractTo define and identify a region-of-interest (ROI) in a digital image, the shape descriptor of the ROI has to be described in terms of its boundary characteristics. To address the generic issues of contour tracking, the yConvex Hypergraph (yCHG) model was proposed by Kanna et al [1]. This yCHG model represents any connected region as a finite set of disjoint yConvex hyperedges (yCHE), which helps to perform the contour tracking precisely without retracing the same contour. We observe that the serial implementation of the yCHG is quite costly in terms of memory and computation for high resolution images. These issues motivated us to exploit the high level data parallelism available on Graphic Processing Units (GPUs). In this work, we propose a parallel approach to implement yCHG model by exploiting massively parallel cores of NVIDIA Compute Unified Device Architecture (CUDA). We perform our experiments on the MODIS satellite image database by NASA, and based on our analysis we observe that the performance of the serial implementation is better on smaller images, but once the threshold is achieved in terms of image resolution, the parallel implementation outperforms its sequential counterpart by 2 to 10 times (2x-10x). We also conclude that an increase in the number of hyperedges in ROI of given size does not impact the performance of the overall algorithm. Saurabh Jha, Tejaswi Agarwal, B. Rajesh Kanna |
ICS | 1 |