VLDB 2026 Research / reviewers in the wild / expert
Abraham Chan
dblp:201/3037
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-7260-4124ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Security and privacy · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Thinking Inside the Box: Injecting Realistic Radiation Faults in ML Accelerators
Bruno Loureiro Coelho, Mani Sadati, Abraham Chan, Alex Hands, Karthik Pattabiraman, Paolo Rech |
DSN | 3 |
| 2026 | The Statistical Assessment of Bayes-"sub"Optimal Binary Machine Learning Classifier Risk
Abraham Chan, Ilir Gashi, Sathish Gopalakrishnan, Karthik Pattabiraman, Kizito Salako |
SAFECOMP | 1 |
| 2025 | ReMlX: Resilience for ML Ensembles using XAI at Inference against Faulty Training DataabstractSafety-critical domains, such as healthcare and autonomous vehicles, employ machine learning (ML), where mis-predictions can cause severe repercussions. Training datasets may contain faults, thereby compromising ML accuracy. Ensembles, where multiple ML models vote on predictions, are effective at maintaining predictive capability, and thus resilient against faulty training data because individual models focus on diverse input features. Nevertheless, ensemble diversity varies per input. Hence, weighted ensembles can bolster resilience by assigning unique weights to constituent models. While existing weighted ensembles focus on output-space diversity, we propose leveraging their feature-space diversity to better capture model independence and achieve greater resilience. Therefore, we present ReMlX, which applies explainable artificial intelligence to extract the feature-space diversity of ensemble models, and adjusts their weights to maximize resilience. Compared to its most competitive baseline, ReMlX is 12% more resilient but 15% slower than dynamic weighted ensembles based on stacking. Abraham Chan, Arpan Gujarati, Karthik Pattabiraman, Sathish Gopalakrishnan |
DSN | 1 |
| 2025 | DLAFI: Software-Based Fault Injection for Permanent Faults in Deep Learning AcceleratorsabstractDeep learning accelerators (DLAs) are used in safety-critical applications, making their reliability an important goal. Permanent faults arising due to wear and tear and manufacturing defects are a particular concern for the reliability of DLAs. Unfortunately, existing permanent fault injection methods are either slow (hardware simulations) or inaccurate (software-level). We introduce DLAFI, an LLVM-based fault injection framework that accurately simulates the hardware behavior of systolic arrays (SAs)-the core compute components of DLAs, while achieving comparable speed as software-level injection. DLAFI models the SA’s scheduling strategy to dynamically map machine learning (ML) operations to the SA’s processing elements. Compared with hardware simulation-based fault injection, DLAFI enables the analysis of higher complexity ML applications such as object detection and large language models, and is three orders of magnitude faster overall. Using DLAFI, we evaluate the resilience of various ML workloads across SA sizes and scheduling strategies, and find that larger SAs reduce fault impact, balanced schedulers can reduce resilience, faults in final layers exhibit higher vulnerability, and vision models are more resilient than language models. SeyedMani Sadati, Abraham Chan, Udit Kumar Agarwal, Karthik Pattabiraman |
ISSRE | 2 |
| 2023 | Resilience Assessment of Large Language Models under Transient Hardware FaultsabstractLarge Language Models (LLMs) are transforming the field of natural language processing and revolutionizing the way machines interact with humans. LLMs like ChatGPT and Google’s Bard have already made significant strides in conversational AI, enabling machines to understand natural language and respond in a more human-like manner. In addition to typical applications like sentiment analysis and text generation, LLMs are also used in safety-critical applications such as code generation and speech comprehension in autonomous driving vehicles, where reliability is important.In this work, we investigate the resilience of LLMs under transient hardware faults. Specifically, we used IR-level fault injection (FI) to assess the reliability of five popular LLMs, including Bert, GPT2, and T5, under transient hardware faults. Moreover, we also investigate how the resilience of LLMs varies with different pre-training, fine-tuning objectives, and the number of encoder and decoder blocks. We find that LLMs are quite resilient to transient faults overall. We also find that the behavior of the LLM under transient faults varies significantly with the input, LLM’s architecture, and the type of task (e.g., translation vs. fill-in-the-blank). Finally, we find that the Silent Data Corruption (SDC) rate varies with different fine-tuning objectives, and for the fill-mask fine-tuning objective, the SDC rate also increases with the model size. Overall, our findings indicate that the use of LLMs in safety-critical applications needs further investigation. Udit Kumar Agarwal, Abraham Chan, Karthik Pattabiraman |
ISSRE | 2 |
| 2023 | Evaluating the Effect of Common Annotation Faults on Object Detection TechniquesabstractMachine learning (ML) is applied in many safety-critical domains such as autonomous driving and medical diagnosis. Many ML applications in such domains require object detection, which includes both classification and localization, to provide additional context. To ensure high accuracy, state-of-the-art object detection (OD) systems require large quantities of correctly annotated images for training. However, creating such datasets is non-trivial, may involve significant human effort, and is hence inevitably prone to annotation faults. We evaluate the effect of such faults on OD applications. We present ODFI, which can inject five different types of common annotation faults into any COCO-formatted dataset. We then use ODFI to inject these faults into two road traffic and one medical X-ray imaging datasets. Finally, using these faulty datasets, we systematically evaluate and compare the efficacy of existing OD techniques that are designed to be robust against such faults. To do so, we introduce a new metric that evaluates the robustness of OD models in the presence of faults. We find that (1) single-stage detectors trained with faulty annotations perform better in scenes with more objects, (2) redundant bounding boxes have the least impact on robustness, and (3) ensembles have the highest overall robustness among the robust OD techniques considered. Abraham Chan, Arpan Gujarati, Karthik Pattabiraman, Sathish Gopalakrishnan |
ISSRE | 1 |
| 2023 | Mixed precision support in HPC applications: What about reliability?
Alessio Netti, Patrik Omland, Michael Paulitsch, Jorge Parra, Gustavo Espinosa, Udit Kumar Agarwal, Abraham Chan, Karthik Pattabiraman |
J. Parallel Distributed Comput. | 8 |
| 2022 | The Fault in Our Data Stars: Studying Mitigation Techniques against Faulty Training Data in Machine Learning ApplicationsabstractMachine learning (ML) has been adopted in many safety-critical applications like automated driving and medical diagnosis. Incorrect decisions by ML models can lead to catastrophic consequences, such as vehicle crashes and inappropriate medical procedures, thereby endangering our lives. The correct behaviour of a ML model is contingent upon the availability of well-labelled training data. However, obtaining large and high-quality training datasets for safety-critical applications is difficult, often resulting in the use of faulty training data.We compare the efficacy of five different error mitigation techniques, derived from a survey of more than 200 related articles, which are designed to tolerate noisy/faulty training data. We experimentally find that the error mitigation capabilities of these techniques vary across datasets, ML models, and different kinds of faults. We further find that ensemble learning offers the highest resilience among all the techniques across different configurations, followed by label smoothing. Abraham Chan, Arpan Gujarati, Karthik Pattabiraman, Sathish Gopalakrishnan |
DSN | 1 |
| 2022 | LLTFI: Framework Agnostic Fault Injection for Machine Learning Applications (Tools and Artifact Track)abstractAs machine learning (ML) has become more preva-lent across many critical domains, so has the need to understand ML applications' resilience. While prior work like TensorFI [1], MindFI [2], and PyTorchFI [3] has focused on building ML fault injectors for specific ML frameworks, there has been little work on performing fault injection (FI) for ML applications written in multiple frameworks. We present LLTFI, a framework-agnostic fault injection tool for ML applications, allowing users to run FI experiments on ML applications at the LLVM IR level. LLTFI provides users with finer FI granularity at the level of instructions, and a better understanding of how faults manifest and propagate between different ML components. We evaluate LLTFI on six ML programs and compare it with TensorFI. We found significant differences in the Silent Data Corruption (SDC) rates for similar faults between the two tools. Finally, we use LLTFI to evaluate the efficacy of selective instruction duplication - an error mitigation technique - for ML programs. Udit Kumar Agarwal, Abraham Chan, Karthik Pattabiraman |
ISSRE | 2 |
| 2021 | Understanding the Resilience of Neural Network Ensembles against Faulty Training DataabstractMachine learning is becoming more prevalent in safety-critical systems like autonomous vehicles and medical imaging. Faulty training data, where data is either misla-belled, missing, or duplicated, can increase the chance of misclassification, resulting in serious consequences. In this paper, we evaluate the resilience of ML ensembles against faulty training data, in order to understand how to build better ensembles. To support our evaluation, we develop a fault injection framework to systematically mutate training data, and introduce two diversity metrics that capture the distribution and entropy of predicted labels. Our experiments find that ensemble learning is more resilient than any individual model and that high accuracy neural networks are not necessarily more resilient to faulty training data. Further, we find that simple majority voting suffices in most cases for resilience in ML ensembles. Finally, we observe diminishing returns for resilience as we increase the number of models in an ensemble. These findings can help machine learning developers build ensembles that are both more resilient and more efficient. Abraham Chan, Niranjhana Narayanan, Arpan Gujarati, Karthik Pattabiraman, Sathish Gopalakrishnan |
QRS | 1 |
| 2017 | IPA: Error Propagation Analysis of Multi-Threaded Programs Using Likely InvariantsabstractError Propagation Analysis (EPA) is a technique forunderstanding how errors affect a program's execution and resultin program failures. For this purpose, EPA usually compares thetraces of a fault-free (golden) run with those from a faulty run ofthe program. This makes existing EPA approaches brittle for multithreadedprograms, which do not typically have a deterministicgolden run. In this paper, we study the use of likely invariantsgenerated by automated approaches as alternatives for goldenrun based EPA in multithreaded programs. We present InvariantPropagation Analysis (IPA), an approach and a framework forautomatically deriving invariants for multithreaded programs, and using the invariants for EPA. We evaluate the invariantsderived by IPA in terms of their coverage for different faulttypes across six representative programs through fault injectionexperiments. We find that stable invariants can be inferred in allsix programs, although their coverage of faults depends on theapplication and the fault type. Abraham Chan, Stefan Winter 0001, Habib Saissi, Karthik Pattabiraman, Neeraj Suri |
ICST | 1 |