EDBT 2026 Demo / reviewers in the wild / expert
Arpan Gujarati
dblp:135/0577
· DBLP profile ↗
21ranked-venue papers
11as first author
10since 2021 · last 2025
0000-0002-2949-6826ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Security and privacy · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ReMlX: Resilience for ML Ensembles using XAI at Inference against Faulty Training DataabstractSafety-critical domains, such as healthcare and autonomous vehicles, employ machine learning (ML), where mis-predictions can cause severe repercussions. Training datasets may contain faults, thereby compromising ML accuracy. Ensembles, where multiple ML models vote on predictions, are effective at maintaining predictive capability, and thus resilient against faulty training data because individual models focus on diverse input features. Nevertheless, ensemble diversity varies per input. Hence, weighted ensembles can bolster resilience by assigning unique weights to constituent models. While existing weighted ensembles focus on output-space diversity, we propose leveraging their feature-space diversity to better capture model independence and achieve greater resilience. Therefore, we present ReMlX, which applies explainable artificial intelligence to extract the feature-space diversity of ensemble models, and adjusts their weights to maximize resilience. Compared to its most competitive baseline, ReMlX is 12% more resilient but 15% slower than dynamic weighted ensembles based on stacking. Abraham Chan, Arpan Gujarati, Karthik Pattabiraman, Sathish Gopalakrishnan |
DSN | 2 |
| 2025 | EtherTime: Cross-Vendor Evaluation of PTP/NTP on Ethernet-Based COTS Embedded Platforms
Vincent Bode, William Shen, Arpan Gujarati |
RTCSA | 3 |
| 2025 | Faster, Exact, More General Response-Time Analysis for NVIDIA Holoscan ApplicationsabstractWe present a scalable method to compute exact worst-case end-to-end latency in applications built on the NVIDIA Holoscan SDK, a framework increasingly adopted for soft real-time ML workloads in medical devices, surgical instruments, and robotics. Holoscan applications are structured as directed acyclic graphs of non-preemptible task threads (operators) connected by FIFO queues, where execution depends not only on input availability but also on downstream buffer capacity - an atypical backpressure mechanism not captured by standard dataflow or middleware models. Existing analyses either lack convergence guarantees or rely on restrictive assumptions (e.g., fixed execution times, unit-sized buffers), resulting in overly conservative bounds. We show that Holoscan's scheduling semantics can be faithfully reduced to homogeneous synchronous dataflow graphs (HSDFGs), enabling exact end-to-end latency analysis. Building on this insight, we introduce a dynamic algorithm that computes tight upper bounds on response time across infinite input streams under variable task runtimes and arbitrary buffer sizes. We prove its correctness and convergence, and demonstrate that it outperforms HSDFG model checking with UPPAAL in runtime while avoiding the pessimism of prior Holoscan-specific analyses. Experiments on real Holoscan applications from NVIDIA HoloHub and large synthetic graphs confirm its scalability and precision. Philip Schowitz, Shubhaankar Sharma, Siddharth Balodi, Soham Sinha 0001, Bruce Shepherd, Arpan Gujarati |
RTSS | 6 |
| 2024 | RABIT, a Robot Arm Bug Intervention Tool for Self-Driving LabsabstractSelf-driving labs are transforming scientific research and accelerating experimentation using software-controlled lab equipment. These labs are exposed to human errors by inexperienced researchers working in the lab (e.g., setting incorrect target location could cause a robot arm to collide with an expensive piece of equipment). We present RABIT, a Robot Arm Bug Intervention Tool, which (i) allows systematically specifying safety rules across diverse devices and (ii) evaluates and enforces these rules using simulation, a low-fidelity testbed, and a production environment. We report our experience adapting RABIT for the Hein Lab, a state-of-the-art research lab that blends advanced robotics with synthetic organic chemistry. Zainab Saeed Wattoo, Petal Vitis, Ruizhe Zhu, Noah Depner, Ivory Zhang, Jason Hein, Arpan Gujarati, Margo I. Seltzer |
DSN | 7 |
| 2024 | Response-Time Analysis of a Soft Real-time NVIDIA Holoscan ApplicationabstractNVIDIA Holoscan SDK is a novel edge and embedded software development framework designed for NVIDIA System-on-Chips (SoCs), primarily targeting medical device applications. This SDK facilitates complex data processing workflows using Directed Acyclic Graphs (DAGs) composed of functional units termed operators. These operators, running in separate threads, are usually interconnected with intricate execution dependencies influenced by both upstream and downstream conditions on communication data buffers. Current methods to measure the response time of a complex Holoscan application rely on empirical benchmarking, which can be costly, time-consuming, and unreliable – limitations that are particularly critical in sectors where safety and certification concerns are paramount. This paper introduces a novel static analysis methodology to determine worst-case end-to-end response times in NVIDIA Holoscan applications. Our approach overcomes the drawbacks of existing empirical tools by providing a response-time analysis capable of handling complex operator interactions and communication buffering mechanisms inherent in Holoscan’s architecture. Through rigorous theoretical analysis and empirical validation, our method not only ensures predictability in system behavior but also aids developers in identifying performance bottlenecks and optimizing system design. Evaluation using real-world NVIDIA HoloHub applications demonstrates the efficiency and accuracy of our analysis, achieving theoretical response times as close as $0.3 \%$ of empirically measured numbers on NVIDIA hardware using less than 1ms computation time. Philip Schowitz, Soham Sinha 0001, Arpan Gujarati |
RTSS | 3 |
| 2023 | Evaluating the Effect of Common Annotation Faults on Object Detection TechniquesabstractMachine learning (ML) is applied in many safety-critical domains such as autonomous driving and medical diagnosis. Many ML applications in such domains require object detection, which includes both classification and localization, to provide additional context. To ensure high accuracy, state-of-the-art object detection (OD) systems require large quantities of correctly annotated images for training. However, creating such datasets is non-trivial, may involve significant human effort, and is hence inevitably prone to annotation faults. We evaluate the effect of such faults on OD applications. We present ODFI, which can inject five different types of common annotation faults into any COCO-formatted dataset. We then use ODFI to inject these faults into two road traffic and one medical X-ray imaging datasets. Finally, using these faulty datasets, we systematically evaluate and compare the efficacy of existing OD techniques that are designed to be robust against such faults. To do so, we introduce a new metric that evaluates the robustness of OD models in the presence of faults. We find that (1) single-stage detectors trained with faulty annotations perform better in scenes with more objects, (2) redundant bounding boxes have the least impact on robustness, and (3) ensembles have the highest overall robustness among the robust OD techniques considered. Abraham Chan, Arpan Gujarati, Karthik Pattabiraman, Sathish Gopalakrishnan |
ISSRE | 2 |
| 2022 | The Fault in Our Data Stars: Studying Mitigation Techniques against Faulty Training Data in Machine Learning ApplicationsabstractMachine learning (ML) has been adopted in many safety-critical applications like automated driving and medical diagnosis. Incorrect decisions by ML models can lead to catastrophic consequences, such as vehicle crashes and inappropriate medical procedures, thereby endangering our lives. The correct behaviour of a ML model is contingent upon the availability of well-labelled training data. However, obtaining large and high-quality training datasets for safety-critical applications is difficult, often resulting in the use of faulty training data.We compare the efficacy of five different error mitigation techniques, derived from a survey of more than 200 related articles, which are designed to tolerate noisy/faulty training data. We experimentally find that the error mitigation capabilities of these techniques vary across datasets, ML models, and different kinds of faults. We further find that ensemble learning offers the highest resilience among all the techniques across different configurations, followed by label smoothing. Abraham Chan, Arpan Gujarati, Karthik Pattabiraman, Sathish Gopalakrishnan |
DSN | 2 |
| 2022 | Arming IDS Researchers with a Robotic Arm DatasetabstractIndustry 4.0 is rapidly transforming traditional manufacturing practices. Smart manufacturing technologies that automate research and development using a combination of robotic arms and domain-specific cyber-physical systems are at the core of this transformation. Unfortunately, dependence on networked communication increases the risk of security attacks, which must be mitigated using either platforms that are secure by design or intrusion detection and prevention systems. We report on an ongoing project to design and develop intrusion detection systems (IDS) for the Hein Lab, a smart manufacturing research lab in the chemical sciences domain. Designing effective IDS requires large datasets and high-quality, domain-specific benchmarks, which are difficult to obtain. To address this gap, we present the Robotic Arm Dataset (RAD), which we collected at the Hein Lab over a three-month period. We also present our non-intrusive tracing framework RATracer, which can be retrofitted onto any existing Python-based automation pipeline, and two sets of preliminary analyses based on the command and power data in RAD. Arpan Gujarati, Zainab Saeed Wattoo, Maryam Raiyat Aliabadi, Sean Clark 0004, Parisa Shiri, Amee Trivedi, Ruizhe Zhu, Jason Hein, Margo I. Seltzer |
DSN | 1 |
| 2022 | In-ConcReTeS: Interactive Consistency meets Distributed Real-Time Systems, Again!abstractThe problem of replica coordination is fundamental to building Byzantine fault-tolerant (BFT) distributed systems. Seminal BFT architectures for safety-critical real-time systems from the eighties and nineties relied on custom processors and networks, and are hence not readily usable today. Modern-day deployments on cloud platforms do not “scale down” to embedded platforms and are not designed around timeliness. Recent work on real-time BFT protocols focuses on simulations and reliability analyses. In short, there exist no easily programmable BFT libraries that can be conveniently retrofitted onto real-time applications with deadlines and that perform well on embedded platforms. We propose In-ConcReTeS, a BFT key-value store designed for building highly reliable control applications on commodity embedded platforms. At its core, In-ConcReTeS is a real-time friendly redesign and an efficient implementation of a BFT protocol used by seminal fault-tolerant architectures. We evaluated In-ConcReTeS using an inverted pendulum simulation and an automotive benchmark on a cluster of four Raspberry Pis connected over Ethernet. Our results show that, unlike Redis and etcd, In-ConcReTeS can repeatedly synchronize hundreds of key-value pairs, while tolerating faults, every tens of milliseconds. Arpan Gujarati, Ningfeng Yang, Björn B. Brandenburg |
RTSS | 1 |
| 2021 | Understanding the Resilience of Neural Network Ensembles against Faulty Training DataabstractMachine learning is becoming more prevalent in safety-critical systems like autonomous vehicles and medical imaging. Faulty training data, where data is either misla-belled, missing, or duplicated, can increase the chance of misclassification, resulting in serious consequences. In this paper, we evaluate the resilience of ML ensembles against faulty training data, in order to understand how to build better ensembles. To support our evaluation, we develop a fault injection framework to systematically mutate training data, and introduce two diversity metrics that capture the distribution and entropy of predicted labels. Our experiments find that ensemble learning is more resilient than any individual model and that high accuracy neural networks are not necessarily more resilient to faulty training data. Further, we find that simple majority voting suffices in most cases for resilience in ML ensembles. Finally, we observe diminishing returns for resilience as we increase the number of models in an ensemble. These findings can help machine learning developers build ensembles that are both more resilient and more efficient. Abraham Chan, Niranjhana Narayanan, Arpan Gujarati, Karthik Pattabiraman, Sathish Gopalakrishnan |
QRS | 3 |
| 2020 | Serving DNNs like Clockwork: Performance Predictability from the Bottom Up
Arpan Gujarati, Reza Karimi, Safya Alzayat, Antoine Kaufmann, Ymir Vigfusson, Jonathan Mace |
OSDI | 1 |
| 2020 | Real-Time Replica Consistency over Ethernet with Reliability BoundsabstractEthernet is expected to play a key role in the development of the next generation of safety-critical distributed real-time systems. Unfortunately, the use of switched Ethernet in place of traditional field buses such as CAN exposes systems to the risk of Byzantine errors (or inconsistent broadcasts) due to environmentally-induced transient faults.Byzantine fault tolerance (BFT) protocols can mitigate such errors to a large extent. However, no BFT protocol has yet been investigated from the perspective of hard real-time predictability. Classical Byzantine safety guarantees (e.g., 3f+ 1 processes can tolerate up to f Byzantine faults) are oblivious to non-uniform fault rates across different system components that arise due to environmental disturbances. Furthermore, existing analyses abstract from the underlying network topology despite its strong influence on actual failure rates.In this work, we present (i) a hard real-time interactive consistency protocol that allows distributed processes to agree on a common state despite Byzantine errors; and (ii) the first quantitative, real-time-aware reliability analysis of such a protocol deployed over switched Ethernet in the presence of stochastic transient faults. Our analysis is free of reliability anomalies and, as we show in our evaluation, can be used for a reliability-aware design space exploration of different fault tolerance alternatives. Arpan Gujarati, Sergey Bozhko, Björn B. Brandenburg |
RTAS | 1 |
| 2019 | From Iteration to System Failure: Characterizing the FITness of Periodic Weakly-Hard SystemsabstractEstimating metrics such as the Mean Time To Failure (MTTF) or its inverse, the Failures-In-Time (FIT), is a central problem in reliability estimation of safety-critical systems. To this end, prior work in the real-time and embedded systems community has focused on bounding the probability of failures in a single iteration of the control loop, resulting in, for example, the worst-case probability of a message transmission error due to electromagnetic interference, or an upper bound on the probability of a skipped or an incorrect actuation. However, periodic systems, which can be found at the core of most safety-critical real-time systems, are routinely designed to be robust to a single fault or to occasional failures (case in point, control applications are usually robust to a few skipped or misbehaving control loop iterations). Thus, obtaining long-run reliability metrics like MTTF and FIT from single iteration estimates by calculating the time to first fault can be quite pessimistic. Instead, overall system failures for such systems are better characterized using multi-state models such as weakly-hard constraints. In this paper, we describe and empirically evaluate three orthogonal approaches, PMC, Mart, and SAp, for the sound estimation of system’s MTTF, starting from a periodic stochastic model characterizing the failure in a single iteration of a periodic system, and using weakly-hard constraints as a measure of system robustness. PMC and Mart are exact analyses based on Markov chain analysis and martingale theory, respectively, whereas SAp is a sound approximation based on numerical analysis. We evaluate these techniques empirically in terms of their accuracy and numerical precision, their expressiveness for different definitions of weakly-hard constraints, and their space and time complexities, which affect their scalability and applicability in different regions of the space of weakly-hard constraints. Arpan Gujarati, Mitra Nasri, Rupak Majumdar, Björn B. Brandenburg |
ECRTS | 1 |
| 2019 | Correspondence article: a correction of the reduction-based schedulability analysis for APA scheduling
Arpan Gujarati, Felipe Cerqueira, Björn B. Brandenburg, Geoffrey Nelissen |
Real Time Syst. | 1 |
| 2018 | Quantifying the Resiliency of Fail-Operational Real-Time Networked Control SystemsabstractIn time-sensitive, safety-critical systems that must be fail-operational, active replication is commonly used to mitigate transient faults that arise due to electromagnetic interference (EMI). However, designing an effective and well-performing active replication scheme is challenging since replication conflicts with the size, weight, power, and cost constraints of embedded applications. To enable a systematic and rigorous exploration of the resulting tradeoffs, we present an analysis to quantify the resiliency of fail-operational networked control systems against EMI-induced memory corruption, host crashes, and retransmission delays. Since control systems are typically robust to a few failed iterations, e.g., one missed actuation does not crash an inverted pendulum, traditional solutions based on hard real-time assumptions are often too pessimistic. Our analysis reduces this pessimism by modeling a control system's inherent robustness as an (m,k)-firm specification. A case study with an active suspension workload indicates that the analytical bounds closely predict the failure rate estimates obtained through simulation, thereby enabling a meaningful design-space exploration, and also demonstrates the utility of the analysis in identifying non-trivial and non-obvious reliability tradeoffs. Arpan Gujarati, Mitra Nasri, Björn B. Brandenburg |
ECRTS | 1 |
| 2018 | Tableau: a high-throughput and predictable VM scheduler for high-density workloadsabstractIn the increasingly competitive public-cloud marketplace, improving the efficiency of data centers is a major concern. One way to improve efficiency is to consolidate as many VMs onto as few physical cores as possible, provided that performance expectations are not violated. However, as a prerequisite for increased VM densities, the hypervisor's VM scheduler must allocate processor time efficiently and in a timely fashion. As we show in this paper, contemporary VM schedulers leave substantial room for improvements in both regards when facing challenging high-VM-density workloads that frequently trigger the VM scheduler. As root causes, we identify (i) high runtime overheads and (ii) unpredictable scheduling heuristics. To better support high VM densities, we propose Tableau, a VM scheduler that guarantees a minimum processor share and a maximum bound on scheduling delay for every VM in the system. Tableau combines a low-overhead, core-local, table-driven dispatcher with a fast on-demand table-generation procedure (triggered on VM creation/teardown) that employs scheduling techniques typically used in hard real-time systems. In an evaluation of Tableau and three current Xen schedulers on a 16-core Intel Xeon machine, Tableau is shown to improve tail latency (e.g., a 17X reduction in maximum ping latency compared to Credit) and throughput (e.g., 1.6X peak web server throughput compared to RTDS when serving 1 KiB files with a 100 ms SLA). Manohar Vanga, Arpan Gujarati, Björn B. Brandenburg |
EuroSys | 2 |
| 2017 | Swayam: distributed autoscaling to meet SLAs of machine learning inference services with resource efficiencyabstractDevelopers use Machine Learning (ML) platforms to train ML models and then deploy these ML models as web services for inference (prediction). A key challenge for platform providers is to guarantee response-time Service Level Agreements (SLAs) for inference workloads while maximizing resource efficiency. Swayam is a fully distributed autoscaling framework that exploits characteristics of production ML inference workloads to deliver on the dual challenge of resource efficiency and SLA compliance. Our key contributions are (1) model-based autoscaling that takes into account SLAs and ML inference workload characteristics, (2) a distributed protocol that uses partial load information and prediction at frontends to provision new service instances, and (3) a backend self-decommissioning protocol for service instances. We evaluate Swayam on 15 popular services that were hosted on a production ML-as-a-service platform, for the following service-specific SLAs: for each service, at least 99% of requests must complete within the response-time threshold. Compared to a clairvoyant autoscaler that always satisfies the SLAs (i.e., even if there is a burst in the request rates), Swayam decreases resource utilization by up to 27%, while meeting the service-specific SLAs over 96% of the time during a three hour window. Microsoft Azure's Swayam-based framework was deployed in 2016 and has hosted over 100,000 services. Arpan Gujarati, Sameh Elnikety, Yuxiong He, Kathryn S. McKinley, Björn B. Brandenburg |
Middleware | 1 |
| 2015 | When Is CAN the Weakest Link? A Bound on Failures-in-Time in CAN-Based Real-Time SystemsabstractA method to bound the Failures In Time (FIT) rate of a CAN-based real-time system, i.e., the expected number of failures in one billion operating hours, is proposed. The method leverages an analysis, derived in the paper, of the probability of a correct and timely message transmission despite host and network failures due to electromagnetic interference (EMI). For a given workload, the derived FIT rate can be used to find an optimal replication factor, which is demonstrated with a case study based on a message set taken from a simple mobile robot. Arpan Gujarati, Björn B. Brandenburg |
RTSS | 1 |
| 2015 | Multiprocessor real-time scheduling with arbitrary processor affinities: from practice to theory
Arpan Gujarati, Felipe Cerqueira, Björn B. Brandenburg |
Real Time Syst. | 1 |
| 2014 | Linux's Processor Affinity API, Refined: Shifting Real-Time Tasks Towards Higher SchedulabilityabstractVirtually all major real-time operating systems such as QNX, VxWorks, LynxOS, and most real-time variants of Linux expose processor affinity APIs to restrict task migrations. Initially motivated by throughput and isolation reasons, the ability to flexibly control migrations on a per-task basis has also proved to be useful from a schedulability perspective. However, as the motivation to use processor affinities is highly application-specific, the two interests can conflict, i.e., The fixed, user-specified processor affinities chosen for non-schedulability reasons can actually limit any possible gains in schedulability. This paper specifically addresses the scenario where processor affinities are given as input, and investigates the following question: while maintaining API compatibility (i.e., Without changing the interface exposed to the programmer), is it possible to improve schedulability beyond what Linux and Linux-like systems currently offer, without violating the original affinity restrictions? To answer this question, we explore the similarities between priority-based scheduling with processor affinities and the assignment problem with seniority and job priority constraints, studied previously by Caron et al. In an operations-research context, to derive a more generic model of migrations. Based on vertex-weighted bipartite matchings, the proposed model exploits the idea of shifting high-priority tasks among processors in their affinity set, in order to accommodate lower-priority tasks that have more constrained processor affinities. The proposed approach is analyzed with a novel shifting-aware schedulability analysis based on linear programming. An empirical evaluation in terms of schedulability shows shifting to be effective, although performance naturally degrades if migration overheads are high. Felipe Cerqueira, Arpan Gujarati, Björn B. Brandenburg |
RTSS | 2 |
| 2013 | Outstanding Paper Award: Schedulability Analysis of the Linux Push and Pull Scheduler with Arbitrary Processor AffinitiesabstractContemporary multiprocessor real-time operating systems, such as VxWorks, LynxOS, QNX, and real-time variants of Linux, allow a process to have an arbitrary processor affinity, that is, a process may be pinned to an arbitrary subset of the processors in the system. Placing such a hard constraint on process migrations can help to improve cache performance of specific multi-threaded applications, achieve isolation among components, and aid in load-balancing. However, to date, the lack of schedulability analysis for such systems prevents the use of arbitrary processor affinities in predictable hard real-time applications. In this paper, it is shown that job-level fixed-priority scheduling with arbitrary processor affinities is strictly more general than global, clustered, and partitioned job-level fixed-priority scheduling. The Linux push and pull scheduler is studied as a reference implementation and techniques for the schedulability analysis of hard real-time tasks with arbitrary processor affinity masks are presented. The proposed tests work by reducing the scheduling problem to ``global-like'' sub-problems to which existing global schedulability tests can be applied. Schedulability experiments show the proposed techniques to be effective. Arpan Gujarati, Felipe Cerqueira, Björn B. Brandenburg |
ECRTS | 1 |