VLDB 2026 Research / reviewers in the wild / expert
Harish Dattatraya Dixit
dblp:223/7108
· DBLP profile ↗
15ranked-venue papers
2as first author
14since 2021 · last 2026
0009-0001-1163-5568ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 1 first-author · 13 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SEVI: Silent Data Corruption of Vector Instructions in Hyper-Scale DatacentersabstractSilent Data Corruption (SDC) poses a reliability threat in modern datacenters. These insidious errors evade detections and propagate incorrect results throughout the system. Companies including Google, Meta, and Alibaba have reported SDC incidents affecting their production. In this paper, we present the first comprehensive instruction- and application-level analysis of vector instruction SDCs in hyper-scale datacenters using a two-stage approach. We perform over 78 trillion test rounds in more than 14 billion CPU seconds. Our observations reveal undocumented SDC patterns that provide insights into possible underlying causes and inspire new mitigation strategies. Based on these findings, we propose a low-overhead SDC detection mechanism leveraging in-application algorithm-based fault tolerance. Our method achieves 88% to 100% SDC machine detection rate with a time overhead of only 1.35% even for modestly sized inputs. Yixuan Mei, Shreya Varshini, Harish Dattatraya Dixit, Sriram Sankar, K. V. Rashmi |
ASPLOS (2) | 3 |
| 2026 | PinDrop: Breaking the Silence on SDCs in a Large-Scale FleetabstractSilent Data Corruptions (SDCs) pose a significant and often hidden threat to the reliability of large-scale computing infrastructure, as they can silently compromise data integrity without immediate detection. Detecting such behaviors at hyperscale is often challenging due to their intermittent nature and the vast diversity of hardware and workloads in a large-scale fleet. This work addresses these challenges by introducing PinDrop, a characterization methodology that leverages continuous, high-frequency testing infrastructure across millions of servers to gather information about SDCs at scale. By leveraging extensive test-suites tailored to mimic real-world applications and exercise a wide range of CPU features, we provide the most comprehensive characterization of SDC failures to date, analyzing over 500 million test executions across millions of devices. Our findings reveal that 0.035% of tested machines suffer from at least one SDC failure during their lifetime. Examining years of data (rather than just a testing snapshot in time), we observe SDCs emerging long after initial deployment and persisting over time. Detailed analysis shows that an average of 0.0024% of tested machines begin failing in each quarter they are tested beyond an initial burn-in period, confirming a fundamental need for continuous testing. Our findings also provide further insights into SDC behaviors at-scale, including failure breakdowns across architectures, test families, specific core IDs, and output-level behaviors. Peter W. Deutsch, Harish Dattatraya Dixit, Gautham Vunnam, Carl Moran, Eleanor Ozer, Sriram Sankar |
HPCA | 2 |
| 2026 | Server Life Extensions at Scale
Rhea Dutta, Harish Dattatraya Dixit, Sriram Sankar |
IOLTS | 2 |
| 2026 | Innovative Practices Session: Recent Approaches in Dealing with Silent Data Corruption
Harish Dattatraya Dixit, Arani Sinha, Nithya Jagannathan, David P. Lerner, Francesco Angione |
VTS | 1 |
| 2025 | Hardware Sentinel: Protecting Software Applications from Hardware Silent Data CorruptionsabstractSilent Data Corruptions (SDCs) pose a significant challenge in large-scale infrastructures, affecting data center applications unpredictably and reducing service reliability. Primarily caused by silicon defects, traditional hardware testing methods are insufficient to prevent SDC propagation. SDCs are influenced by various factors, including data randomization, workload characteristics, environmental conditions, and aging, necessitating top-down approaches from the application layer. In this paper, we introduce Hardware Sentinel, a novel framework that detects SDCs through typical software failure indicators such as segmentation faults, core dumps, application crashes, and logs. We have validated our framework in a large-scale data center fleet, across diverse application, kernel, and hardware configurations, achieving a high success rate of SDC detection. Hardware Sentinel has uncovered novel instances of SDCs, surpassing the detection capabilities of published testing techniques. Our analysis of over 6 years' worth of application and system failure data within a large-scale infrastructure has successfully identified hundreds of defective CPUs that triggered SDCs. Notably, the Hardware Sentinel flow increases effective coverage over existing hardware-testing methods like Fleetscanner (out-of-production testing) by 1.74x and Ripple (in-production testing) by 1.92x. We share the top kernel exceptions with the highest correlation to silent data corruption failures. We present results spanning 7 CPU generations from multiple semiconductor manufacturers, 13 large-scale workloads, and 27 data center regions, providing insights into the trade-offs involved in detection and fleet deployment. Rhea Dutta, Harish Dattatraya Dixit, Rik van Riel, Gautham Vunnam, Sriram Sankar |
ASPLOS (2) | 2 |
| 2025 | From Gates to SDCs: Understanding Fault Propagation Through the Compute StackabstractSilent Data Corruption (SDC) is the most severe effect of a silicon defect in a CPU or other computing chip. The arithmetic units of a CPU are, usually, unprotected and are, thus, the ones that most likely produce SDCs (as well as visible malfunctions of programs such as crashes). In this work, we shed light on the traversal of silicon defects from their point of origin deep inside arithmetic units of complex CPUs towards the program result. We employ microarchitecture-level fault injection enhanced with gate-level designs of the arithmetic units of interest. The hybrid setup combines (i) the accuracy of the hardware and fault modeling and (ii) the speed of program simulation to run long programs to end (thus observing SDC incidents); the analysis that this combination delivers is impossible at other abstraction layers which are either hardware-agnostic (software level) or extremely slow (gate-level). We quantify the effects of faults in two stages and with multiple metrics: (a) how faults propagate to the outputs of the arithmetic units when individual instructions are executed, and (b) how faults eventually affect the outcome of the program generating SDCs, crashes, or being masked. Our fine-grain findings can be utilized for informed fault detection and tolerance strategies at the hardware or the software levels. Odysseas Chatzopoulos, George Papadimitriou 0001, Dimitris Gizopoulos, Harish Dattatraya Dixit, Sriram Sankar |
DATE | 4 |
| 2025 | Veritas - Demystifying Silent Data Corruptions: μArch-Level Modeling and Fleet Data of Modern x86 CPUsabstractHyperscalers have reported unexpectedly high numbers of defective CPU chips, with a defect rate of 1 in a 1000, leading to Silent Data Corruptions (SDCs) in their computing fleets. However, there is no public data on the rate of SDC incidents (corrupted program executions) in large fleets, nor nor any detailed information on which CPU units, microarchitectures, or workloads are more likely to generate SDCs due to silicon defects. While CPU array structures have been studied for fault effects, arithmetic units like integer and floating-point units have not been thoroughly analyzed as potential root causes of SDCs. This paper addresses this critical gap by accurately modeling hardware faults in the arithmetic units of modern x86 CPUs and measuring the probability and rates of SDCs. Using a full-system gem5-based fault injector, the paper examines SDC trends across five recent $x 86$ microarchitectures, various arithmetic units, and instruction classes. By integrating real-world defect rates from large-scale datacenter experiments with early-stage modeling and simulation, the paper provides critical insights into SDC incident rates across different systems. This information is essential for guiding hardware-based or software-based fault protection methods and is the paper’s primary contribution to minimizing the impact of silent data corruptions in computing. Odysseas Chatzopoulos, Nikos Karystinos, George Papadimitriou 0001, Dimitris Gizopoulos, Harish Dattatraya Dixit, Sriram Sankar |
HPCA | 5 |
| 2025 | Meta's Second Generation AI Chip: Model-Chip Co-Design and Productionization ExperiencesabstractThe rapid growth of AI workloads at Meta has motivated our inhouse development of AI chips, aiming to significantly reduce the total cost of ownership and mitigate risks posed by unpredictable GPU supplies.At ISCA'23, we presented Meta's first-generation AI chip, MTIA 1.This paper describes its successor, MTIA 2i, now deployed at scale and serving billions of users.MTIA 2i significantly improves upon MTIA 1, reducing total cost of ownership by 44% compared to GPUs while delivering competitive performance per watt.A key differentiator is its memory hierarchy: instead of costly HBM, it uses large SRAM alongside LPDDR.Although there has been a proliferation of publications on AI chips, they often focus on architectural design and overlook three critical aspects:(1) co-designing and optimizing ML models to work effectively with the AI chip; (2) demonstrating sufficient flexibility to support a wide range of models; and (3) during the productionization process, addressing challenges unanticipated or decisions deferred at design time, such as dealing with memory errors, safe overclocking, reducing provisioned power, and implementing real-time firmware updates to mitigate silicon design defects.A key contribution of this paper is sharing our experience with these aspects, based on our journey of productionizing MTIA 2i at scale. Joel Coburn, Chunqiang Tang, Sameer Abu Asal, Neeraj Agrawal, Raviteja Chinta, Harish Dattatraya Dixit, Brian Dodds, Saritha Dwarakapuram, Amin Firoozshahian, Cao Gao, Kaustubh Gondkar, Tyler Graf, Junhan Hu, Sterling Hughes, Adam Hutchin, Bhasker Jakka, Guoqiang Jerry Chen, Indu Kalyanaraman, Ashwin Kamath, Pankaj Kansal, Erum Kazi, Roman Levenstein, Mahesh Maddury, Alex Mastro, Siji Medaiyese, Pritesh Modi, Jack Montgomery, Nadathur Satish, Amit Nagpal, Ashwin Narasimha, Maxim Naumov, Eleanor Ozer, Jongsoo Park, Poorvaja Ramani, Harikrishna Reddy, David Reiss, Deboleena Roy, Sathish Sekar, Pavan Shetty, Aravind Sukumaran-Rajam, Eran Tal, Mike Tsai, Shreya Varshini, Richard Wareing, Olívia Wu, Xiaolong Xie, Hangchen Yu, Tanmay Zargar, Zitong Zeng, Feixiong Zhang, Ajit Mathews, Jiyuan Zhang 0008, Emmanuel Menage, Truls Edvard Stokke, Mohammed Sourouri |
ISCA | 6 |
| 2025 | CP-Bench: A PyTorch Test Suite to Detect AI Hardware Failure, Performance Degradation, and Silent Data CorruptionabstractThe growing complexity in manufacturing and operating the hardware in AI clusters leads to significant challenges in reliability. Hyperscalars have reported various AI hardware failures during high-stake jobs such as GenAI model training, where one GPU failure could bring down the entire training job. To tackle this issue, we present CP-Bench, an open-source, Configurable and Parameterizable, PyTorch-level test suite designed to test AI hardware failure, performance degradation, and silent data corruption (SDC). Built upon open-source projects, CP-Bench contains 30+ AI workloads (e.g., Llama), and implements various checks (e.g., SDC check) within these workloads. We have deployed CP-Bench throughout Meta’s AI hardware lifecycle, spanning manufacturing, in-production diagnostics, and device RMA; CP-Bench identified various hardware issues, some of which were not caught by vendor’s tooling. Notably, vendor has acknowledged to establish CP-Bench as a valid RMA criteria and plan to integrate CP-Bench into its tooling. CP-Bench is open-sourced at https://github.com/facebookincubator/CP-Bench. Sunny Yang, Suman Gumudavelli, Shreya Varshini, Abhinav Pandey, Abhinav Jauhri, Francesco Caggioni, Gautham Vunnam, Harish Dattatraya Dixit, Jason Liang, Philip Henzler, Sameeksha Gupta, Tyler Graf, Venkat Ramesh, Fan Fred Lin |
ITC | 9 |
| 2025 | Special Session: Trustworthy Hardware-AI at the CloudabstractNowadays, AI applications are becoming extremely popular in our everyday life as well as for the industry. Recent incidents involving hyperscalers have revealed that even cloud-based datacenter hardware can experience failures leading to Silent Data Corruptions (SDCs), also called Silent Data Errors (SDEs). This Special Session delves into the implications of such failures on AI workloads, both during training and inference, and explores methodologies for efficiently detecting SDCs or SDEs through dedicated monitoring phases. Francesco Angione, Paolo Bernardi 0002, Alberto Bosio, Harish Dattatraya Dixit, Salvatore Pappalardo, Annachiara Ruospo, Ernesto Sánchez 0001, Arani Sinha, Vittorio Turco |
VTS | 4 |
| 2024 | Silent Data Corruptions in Computing Systems: Early Predictions and Large-Scale MeasurementsabstractSilent Data Corruptions (SDCs) due to defects in computing chips (CPUs, GPUs, AI accelerators) is a critical threat to the quality of large-scale computing in different application domains: cloud computing, high-performance computing, edge computing. Recent public reports by cloud hyperscalers have emphasized that apart from the usual suspects for SDCs (memory, storage, network), the heart of the computations, the processing elements of all types generate an unexpectedly large rate of SDCs which can cause erroneous calculations and severe information loss. We report, in a consolidated form, recent efforts to correlate early microarchitecture-level simulation-based predictions about the likelihood, rates, severity, and root causes of SDCs and large-scale in-field studies in cloud data centers. Early microarchitecture-level prediction of SDC characteristics (susceptible units, workloads, instructions) can shed light to the cryptic problem of SDCs. The findings of a diligent pre-silicon analysis can assist better understanding of SDCs and can thus drive effective protection decisions either at the hardware or at the software levels at deployment stages. Dimitris Gizopoulos, George Papadimitriou 0001, Odysseas Chatzopoulos, Nikos Karystinos, Harish Dattatraya Dixit, Sriram Sankar |
ETS | 5 |
| 2024 | Silent Data Corruption: Test or Reliability Problem?abstractRecently, companies such as Google, Meta (Facebook), and Microsoft reported in the mainstream press about seemingly random errors which, initially undetected ("silently"), had crept into their large cloud data centers. These reports mentioned that very specific instructions were intermittently incorrectly executed, propagated through the operating system, and would potentially manifest themselves as application-level errors. Are the root causes of these so-called silent data errors test escapes and/or reliability issues? Why are they only noticed now? Is that only the case because such large server farms bring together larger numbers of CPUs than ever seen before? And what counter measures can we take against them? Erik Jan Marinissen, Harish Dattatraya Dixit, R. D. (Shawn) Blanton, Aaron Kuo, Wei Li 0159, Subhasish Mitra, Chris Nigh, Ruben Purdy, Ben Kaczer, Dishant Sangani, Pieter Weckx, Philippe Roussel, Georges Gielen |
ETS | 2 |
| 2023 | Silent Data Corruptions: The Stealthy Saboteurs of Digital IntegrityabstractSilent Data Corruptions (SDCs) pose a significant threat to the integrity of digital systems. These stealthy saboteurs silently corrupt data, remaining undetected by traditional error handling mechanisms. The silent nature of SDCs makes them challenging to trace at the hardware level, as they evade error reporting systems. Instead, their effects manifest at the application level, potentially causing data loss and system-wide issues. Detecting and measuring SDCs present unique challenges. Their low occurrence rates, dependence on hardware structure and software workloads, and correlation to environmental factors make accurate measurement complex. Addressing SDCs requires proactive measures to prevent data corruption and ensure digital integrity. Software redundancy methods provide a means to tolerate SDCs by introducing duplication or triplication of application resources. However, these methods come with their own limitations, including increased code size, altered execution patterns, and potential vulnerability to other types of failures. Understanding the nature of SDCs and developing effective mitigation strategies are crucial for maintaining digital integrity in large-scale infrastructure services. This paper sheds light on the stealthy saboteurs that silently corrupt data, emphasizes the need for comprehensive measurement techniques, and explores the limitations of existing mitigation approaches. By addressing the challenges posed by SDCs, we can fortify digital systems against these hidden threats and ensure the reliability and integrity of our digital infrastructure. George Papadimitriou 0001, Dimitris Gizopoulos, Harish Dattatraya Dixit, Sriram Sankar |
IOLTS | 3 |
| 2022 | Efficient Soft-Error Detection for Low-precision Deep Learning Recommendation ModelsabstractSoft error, namely silent corruption of signal or datum in a computer system, cannot be caverlierly ignored as compute and communication density grow exponentially. Soft error detection has been studied in the context of enterprise computing, high-performance computing and more recently in convolutional neural networks related to autonomous driving.Deep learning recommendation systems (DLRMs) have by now become ubiquitous and serve billions of users per day. Nevertheless, DLRM-specific soft error detection methods are hitherto missing. To fill the gap, this paper presents the first set of soft-error detection methods for low-precision quantized-arithmetic operators in DLRM including general matrix multiplication (GEMM) and EmbeddingBag. A practical method must detect error and do so with low overhead lest reduced inference speed degrades user experience. Exploiting the characteristics of both quantized arithmetic and the operators, we achieved more than 95% detection accuracy for GEMM with an overhead below 20%. For EmbeddingBag, we achieved 99% effectiveness in significant-bit-flips detection with less than 10% of false positives, while keeping overhead below 26%. Sihuan Li, Ping Tak Peter Tang, Daya Shanker Khudia, Jongsoo Park, Harish Dattatraya Dixit, Zizhong Chen |
IEEE Big Data | 6 |
| 2020 | Optimizing Interrupt Handling Performance for Memory Failures in Large Scale Data CentersabstractIntermittent hardware failures are generally non-catastrophic and typical large-scale service infrastructures are designed to tolerate them while still serving user traffic. However, intermittent errors cause performance aberrations if they are not handled appropriately. System error reporting mechanisms send hardware interrupts to the Central Processing Unit (CPU) for handling the hardware errors. This disrupts the CPU's normal operation, which impacts the performance of the server. In this paper, we describe common intermittent hardware errors observed on server systems in a large-scale data center environment. We discuss two methodologies of handling interrupts in server systems - System Management Interrupt (SMI) and Corrected Machine Check Interrupt (CMCI). We characterize the performance of these methods in live environments as compared to prior studies that used error injection to simulate error behavior. Our experience shows that error injection methods are not reflective of production behavior. We also present a hybrid approach for handling error interrupts that achieves better performance, while preserving monitoring granularity, in large scale data center environments. Harish Dattatraya Dixit, Fan Fred Lin, Bill Holland, Matt Beadon, Zhengyu Yang 0003, Sriram Sankar |
ICPE | 1 |