EDBT 2026 Demo / reviewers in the wild / expert
Saurabh Hukerikar
dblp:117/8021
· DBLP profile ↗
9ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-2612-2001ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorSecurity and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Invited Paper: Hardware-Software Co-Design for Highly Optimized, Customized, and Reliable AI SystemsabstractOver the past decade, AI has been rapidly integrated into our daily life, coming in every shape and size and working across systems from big clouds to IoT. As a result, AI systems are increasingly requiring enhancements in model efficiency, hardware acceleration, and memory systems to satisfy stringent constraints on efficiency, reliability, and security. However, advancing across these fronts is challenging as compute demand outpaces Moore’s-law efficiency, hardening into an AI compute wall and an AI energy wall. Breaking through requires a unified AI co-design loop that co-optimizes algorithms and hardware, including efficient AI-to-hardware mapping, so that ongoing goals (accuracy, sparsity, latency) align with concrete hardware choices (precision modes, interconnects, memory hierarchies) and AI-specific execution and memory-reuse patterns. This paper details the principal co-design challenges, presents complementary strategies, and outlines a practical roadmap toward highly optimized, efficient, reliable, and secure AI systems. Jörg Henkel, Mehdi Baradaran Tahoori, Heba Khdr, Hassan Nassar, Vincent Meyers, Deming Chen, Selin Yildirim, Yingbing Huang, Nirmal Saxena, Saurabh Hukerikar, Srivi Dhruvanarayan |
ICCAD | 10 |
| 2022 | Runtime Fault Diagnostics for GPU Tensor CoresabstractTensor cores in NVIDIA GPUs are important computational engines that accelerate diverse AI deep neural networks and algorithms for perception, mapping, localization and path planning in autonomous drive systems. The occurrence of random hardware faults, particularly permanent faults in these computational units have potentially catastrophic consequences. This paper describes software-based runtime diagnostics for tensor cores that leverage universal test patterns (UTP) to achieve high diagnostic fault coverage with very low latency enabling GPU-based systems to meet the ASIL B functional safety targets outlined in the ISO 26262 standard. Saurabh Hukerikar, Nirmal Saxena |
ITC | 1 |
| 2021 | Characterizing and Mitigating Soft Errors in GPU DRAMabstractGPUs are used in high-reliability systems, including high-performance computers and autonomous vehicles. Because GPUs employ a high-bandwidth, wide-interface to DRAM and fetch each memory access from a single DRAM device, implementing full-device correction through ECC is expensive and impractical. This challenge is compounded by worsening relative rates of multi-bit DRAM errors and increasing GPU memory capacities. This paper first presents high-energy neutron beam testing results for the HBM2 memory on a compute-class GPU. These results uncovered unexpected intermittent errors that we determine to be caused by cell damage from the high-intensity beam. As these errors are an artifact of the testing apparatus, we provide best-practice guidance on how to identify and filter them from the results of beam testing campaigns. Second, we use the soft error beam testing results to inform the design and evaluation of system-level error protection mechanisms by reporting the relative error rates and error patterns from soft errors in GPU DRAM. We observe locality in the multi-bit errors, which we attribute to the underlying structure of the HBM2 memory. Based on these error patterns, we propose several novel ECC schemes to decrease the silent data corruption risk by up to five orders of magnitude relative to SEC-DED ECC, while also reducing the number of uncorrectable errors by up to 7.87 ×. We compare novel binary and symbol-based ECC organizations that differ in their design complexity, hardware overheads, and permanent error correction abilities, ultimately recommending two promising organizations. These schemes replace SEC-DED ECC with no additional redundancy, likely no performance impacts, and modest area and complexity costs. Michael B. Sullivan 0001, Nirmal Saxena, Mike O'Connor, Donghyuk Lee, Paul Racunas, Saurabh Hukerikar, Timothy Tsai 0002, Siva Kumar Sastry Hari, Stephen W. Keckler |
MICRO | 6 |
| 2020 | PLEXUS: A Pattern-Oriented Runtime System Architecture for Resilient Extreme-Scale High-Performance Computing SystemsabstractFor high-performance computing (HPC) system designers and users, meeting the myriad challenges of next-generation exascale supercomputing systems requires rethinking their approach to application and system software design. Among these challenges, providing resiliency and stability to the scientific applications in the presence of high fault rates requires new approaches to software architecture and design. As HPC systems become increasingly complex, they require intricate solutions for detection and mitigation for various modes of faults and errors that occur in these large-scale systems, as well as solutions for failure recovery. These resiliency solutions often interact with and affect other system properties, including application scalability, power and energy efficiency. Therefore, resilience solutions for HPC systems must be thoughtfully engineered and deployed. In previous work, we developed the concept of resilience design patterns, which consist of templated solutions based on well-established techniques for detection, mitigation and recovery. In this paper, we use these patterns as the foundation to propose new approaches to designing runtime systems for HPC systems. The instantiation of these patterns within a runtime system enables flexible and adaptable end-to-end resiliency solutions for HPC environments. The paper describes the architecture of the runtime system, named Plexus, and the strategies for dynamically composing and adapting pattern instances under runtime control. This runtime-based approach enables actively balancing the cost-benefit trade-off between performance overhead and protection coverage of the resilience solutions. Based on a prototype implementation of PLEXUS, we demonstrate the resiliency and performance gains achieved by the pattern-based runtime system for a parallel linear solver application. Saurabh Hukerikar, Christian Engelmann |
PRDC | 1 |
| 2019 | Resiliency of automotive object detection networks on GPU architecturesabstractSafety is the most important aspect of an autonomous driving platform. Deep neural networks (DNNs) play an increasingly critical role in localization, perception, and control in these systems. The object detection and classification inference are of particular importance to construct a precise picture of a vehicle's surrounding objects. Graphics Processing Units (GPU) are well-suited to accelerate such DNN-based inference applications since they leverage data and thread-level parallelism in GPU architectures. Understanding the vulnerability of such DNNs to random hardware faults (including transient and permanent faults) in GPU-based systems is essential to meet the safety requirements of auto safety standards such as the ISO 26262, as well as to influence the design of hardware and software-based safety features in current and future generations of GPU architectures and GPU-based automotive platforms. In this paper, we assess the vulnerability of object detection and classification DNNs to permanent and transient faults using fault injection experiments and accelerated neutron beam testing respectively. We also evaluate the effectiveness of chip-level safety mechanisms in GPU architectures, such as ECC and parity, in detecting these random hardware faults. Our studies demonstrate that such object detection networks tend to be vulnerable to random hardware faults, which cause incorrect or mispredicted object detection outcomes. The neutron beam experiments show that existing chip-level protections successfully mitigate all silent data corruption events caused by transient faults. For permanent faults, while ECC and parity are effective in some cases, our results suggest the need for exploring other complementary detection methods, such as periodic online and offline diagnostic testing. Atieh Lotfi, Saurabh Hukerikar, Keshav Balasubramanian, Paul Racunas, Nirmal Saxena, Richard Bramley, Yanxiang Huang |
ITC | 2 |
| 2018 | Shrink or Substitute: Handling Process Failures in HPC Systems Using In-Situ RecoveryabstractEfficient utilization of today's high-performance computing (HPC) systems with complex software and hardware components requires that the HPC applications are designed to tolerate process failures at runtime. With low mean-time-to-failure (MTTF) of current and future HPC systems, long running simulations on these systems requires capabilities for gracefully handling process failures by the applications themselves. In this paper, we explore the use of fault tolerance extensions to Message Passing Interface (MPI) called user-level failure mitigation (ULFM) for handling process failures without the need to discard the progress made by the application. We explore two alternative recovery strategies, which use ULFM along with application-driven in-memory checkpointing. In the first case, the application is recovered with only the surviving processes, and in the second case, spares are used to replace the failed processes, such that the original configuration of the application is restored. Our experimental results demonstrate that graceful degradation is a viable alternative for recovery in environments where spares may not be available. Rizwan A. Ashraf, Saurabh Hukerikar, Christian Engelmann |
PDP | 2 |
| 2018 | Pattern-based Modeling of Multiresilience Solutions for High-Performance ComputingabstractResiliency is the ability of large-scale high-performance computing (HPC) applications to gracefully handle errors, and recover from failures. In this paper, we propose a pattern-based approach to constructing resilience solutions that handle multiple error modes. Using resilience patterns, we evaluate the performance and reliability characteristics of detection, containment and mitigation techniques for transient errors that cause silent data corruptions and techniques for fail-stop errors that result in process failures. We demonstrate the design and implementation of the multiresilience solution based on patterns instantiated across multiple layers of the system stack. The patterns are integrated to work together to achieve resiliency to different error types in a performance-efficient manner. Rizwan A. Ashraf, Saurabh Hukerikar, Christian Engelmann |
ICPE | 2 |
| 2017 | Big Data Meets HPC Log Analytics: Scalable Approach to Understanding Systems at Extreme ScaleabstractToday's high-performance computing (HPC) systems are heavily instrumented, generating logs containing information about abnormal events, such as critical conditions, faults, errors and failures, system resource utilization, and about the resource usage of user applications. These logs, once fully analyzed and correlated, can produce detailed information about the system health, root causes of failures, and analyze an application's interactions with the system, providing valuable insights to domain scientists and system administrators. However, processing HPC logs requires a deep understanding of hardware and software components at multiple layers of the system stack. Moreover, most log data is unstructured and voluminous, making it more difficult for system users and administrators to manually inspect the data. With rapid increases in the scale and complexity of HPC systems, log data processing is becoming a big data challenge. This paper introduces a HPC log data analytics framework that is based on a distributed NoSQL database technology, which provides scalability and high availability, and the Apache Spark framework for rapid in-memory processing of the log data. The analytics framework enables the extraction of a range of information about the system so that system administrators and end users alike can obtain necessary insights for their specific needs. We describe our experience with using this framework to glean insights from the log data about system behavior from the Titan supercomputer at the Oak Ridge National Laboratory. Byung H. Park, Saurabh Hukerikar, Ryan Adamson, Christian Engelmann |
CLUSTER | 2 |
| 2016 | Rolex: resilience-oriented language extensions for extreme-scale systems
Saurabh Hukerikar, Robert F. Lucas |
J. Supercomput. | 1 |