EDBT 2026 Demo / reviewers in the wild / expert
Gautham Vunnam
dblp:223/7030
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0009-1875-9043ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PinDrop: Breaking the Silence on SDCs in a Large-Scale FleetabstractSilent Data Corruptions (SDCs) pose a significant and often hidden threat to the reliability of large-scale computing infrastructure, as they can silently compromise data integrity without immediate detection. Detecting such behaviors at hyperscale is often challenging due to their intermittent nature and the vast diversity of hardware and workloads in a large-scale fleet. This work addresses these challenges by introducing PinDrop, a characterization methodology that leverages continuous, high-frequency testing infrastructure across millions of servers to gather information about SDCs at scale. By leveraging extensive test-suites tailored to mimic real-world applications and exercise a wide range of CPU features, we provide the most comprehensive characterization of SDC failures to date, analyzing over 500 million test executions across millions of devices. Our findings reveal that 0.035% of tested machines suffer from at least one SDC failure during their lifetime. Examining years of data (rather than just a testing snapshot in time), we observe SDCs emerging long after initial deployment and persisting over time. Detailed analysis shows that an average of 0.0024% of tested machines begin failing in each quarter they are tested beyond an initial burn-in period, confirming a fundamental need for continuous testing. Our findings also provide further insights into SDC behaviors at-scale, including failure breakdowns across architectures, test families, specific core IDs, and output-level behaviors. Peter W. Deutsch, Harish Dattatraya Dixit, Gautham Vunnam, Carl Moran, Eleanor Ozer, Sriram Sankar |
HPCA | 3 |
| 2025 | Hardware Sentinel: Protecting Software Applications from Hardware Silent Data CorruptionsabstractSilent Data Corruptions (SDCs) pose a significant challenge in large-scale infrastructures, affecting data center applications unpredictably and reducing service reliability. Primarily caused by silicon defects, traditional hardware testing methods are insufficient to prevent SDC propagation. SDCs are influenced by various factors, including data randomization, workload characteristics, environmental conditions, and aging, necessitating top-down approaches from the application layer. In this paper, we introduce Hardware Sentinel, a novel framework that detects SDCs through typical software failure indicators such as segmentation faults, core dumps, application crashes, and logs. We have validated our framework in a large-scale data center fleet, across diverse application, kernel, and hardware configurations, achieving a high success rate of SDC detection. Hardware Sentinel has uncovered novel instances of SDCs, surpassing the detection capabilities of published testing techniques. Our analysis of over 6 years' worth of application and system failure data within a large-scale infrastructure has successfully identified hundreds of defective CPUs that triggered SDCs. Notably, the Hardware Sentinel flow increases effective coverage over existing hardware-testing methods like Fleetscanner (out-of-production testing) by 1.74x and Ripple (in-production testing) by 1.92x. We share the top kernel exceptions with the highest correlation to silent data corruption failures. We present results spanning 7 CPU generations from multiple semiconductor manufacturers, 13 large-scale workloads, and 27 data center regions, providing insights into the trade-offs involved in detection and fleet deployment. Rhea Dutta, Harish Dattatraya Dixit, Rik van Riel, Gautham Vunnam, Sriram Sankar |
ASPLOS (2) | 4 |
| 2025 | CP-Bench: A PyTorch Test Suite to Detect AI Hardware Failure, Performance Degradation, and Silent Data CorruptionabstractThe growing complexity in manufacturing and operating the hardware in AI clusters leads to significant challenges in reliability. Hyperscalars have reported various AI hardware failures during high-stake jobs such as GenAI model training, where one GPU failure could bring down the entire training job. To tackle this issue, we present CP-Bench, an open-source, Configurable and Parameterizable, PyTorch-level test suite designed to test AI hardware failure, performance degradation, and silent data corruption (SDC). Built upon open-source projects, CP-Bench contains 30+ AI workloads (e.g., Llama), and implements various checks (e.g., SDC check) within these workloads. We have deployed CP-Bench throughout Meta’s AI hardware lifecycle, spanning manufacturing, in-production diagnostics, and device RMA; CP-Bench identified various hardware issues, some of which were not caught by vendor’s tooling. Notably, vendor has acknowledged to establish CP-Bench as a valid RMA criteria and plan to integrate CP-Bench into its tooling. CP-Bench is open-sourced at https://github.com/facebookincubator/CP-Bench. Sunny Yang, Suman Gumudavelli, Shreya Varshini, Abhinav Pandey, Abhinav Jauhri, Francesco Caggioni, Gautham Vunnam, Harish Dattatraya Dixit, Jason Liang, Philip Henzler, Sameeksha Gupta, Tyler Graf, Venkat Ramesh, Fan Fred Lin |
ITC | 8 |