VLDB 2026 Research / reviewers in the wild / expert
Chandramouli N. Amarnath
dblp:271/9975
· DBLP profile ↗
24ranked-venue papers
7as first author
22since 2021 · last 2026
0000-0001-9938-2157ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 7 first-author · 22 since 2021Software engineering, systems software and programming languages · 8 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Error Resilient Transformers: A Novel Soft Error Vulnerability Guided Approach to Error Checking and SuppressionabstractTransformer networks have achieved remarkable success in Natural Language Processing (NLP) and Computer Vision applications. However, the underlying large volumes of Transformer computations demand high reliability and resilience to soft errors in processor hardware. The objective of this research is to develop efficient techniques for design of error resilient Transformer architectures. To enable this, we first perform a soft error vulnerability analysis of every fully connected layers in Transformer computations. Based on this study, error detection and suppression modules are selectively introduced into datapaths to restore Transformer performance under anticipated error rate conditions. Memory access errors and neuron output errors are detected using checksums of linear Transformer computations. Correction consists of determining output neurons with out-of-range values and suppressing the same to zero. For a Transformer with nominal BLEU score of 52.7, such vulnerability guided selective error suppression can recover language translation performance from a BLEU score of 0 to 50.774 with as much as 0.001 probability of activation error, incurring negligible memory and computation overheads. Kwondo Ma, Chandramouli N. Amarnath, Jackson Isenberg, Abhijit Chatterjee |
J. Electron. Test. | 2 |
| 2025 | EGIS: Entropy Guided Image Synthesis for Dataset-Agnostic Testing of RRAM-Based DNNsabstractWhile resistive random access memory (RRAM) based deep neural networks (DNN) are important for low-power inference in IoT and edge applications, they are vulnerable to the effects of manufacturing process variations that degrade their performance (classification accuracy). However, to test the same post-manufacture, the (image) dataset used to train the associated machine learning applications may not be available to the RRAM crossbar manufacturer for privacy reasons. As such, the performance of DNNs needs to be assessed with carefully crafted dataset-agnostic synthetic test images that expose anomalies in the crossbar manufacturing process to the maximum extent possible. In this work, we propose a dataset-agnostic post-manufacture testing framework for RRAM-based DNNs using Entropy Guided Image Synthesis (EGIS). We first create a synthetic image dataset such that the DNN outputs corresponding to the synthetic images minimize an entropy-based loss metric. Next, a small subset (consisting of 10–20 images) of the synthetic image dataset, called the compact image dataset, is created to expedite testing. The response of the device under test (DUT) to the compact image dataset is passed to a machine learning based outlier detector for pass/fail labeling of the DUT. It is seen that the test accuracy using such synthetic test images is very close to that of contemporary test methods. Anurup Saha, Chandramouli N. Amarnath, Kwondo Ma, Abhijit Chatterjee |
DATE | 2 |
| 2025 | European Test Symposium Teams: an Anniversary SnapshotabstractThe IEEE European Test Symposium (ETS) has been facilitating progress in electronic systems testing since its launch in 1996. On the occasion of its 30th anniversary, this collaborative paper gathers sections by 21 ETS teams to outline their influential ideas and milestones. Each team’s section highlights historical perspective, current research, frameworks and projects as well as forward-looking research agendas in the area of electronic-based circuits and systems testing, reliability, safety, security and validation. This anniversary summary documents how research of various ETS teams, exemplifying the test community, has been evolving and transitioning from concepts to practical standards and Electronic Design Automation (EDA) tools and flows. This legacy is a strong base to drive the next generation of advances in electronic systems testing. Maksim Jenihhin, Jaan Raik, Artur Jutman, Natalia Cherezova, Raimund Ubar, Liviu Miclea, Szilárd Enyedi, Iulia Stefan, Ovidiu Stan, Cosmina Corches, Zebo Peng, Petru Eles, Rolf Drechsler, S. Eggersglüß, Görschwin Fey, Andreas Glowatz, Daniel Tille, Georges Gielen, Anthony Coyette, Wim Dobbelaere, Ronny Vanhooren, Po-Yao Chuang, Erik Jan Marinissen, Giorgio Di Natale, M. Barragan, Paolo Maistri, S. Mir, Vatajelu I. Vatajelu, Paolo Bernardi 0002, Stefano Di Carlo, Paolo Prinetto, Matteo Sonza Reorda, Massimo Violante, Haralampos-G. D. Stratigopoulos, M. K. Michael, Stelios Neophytou, Stavros Hadjitheophanous, Kyriakos Christou, M. Skitsas, Alberto Bosio, Bastien Deveautour, Patrick Girard 0001, Marcello Traiola, Arnaud Virazel, Fernando Santos 0001, Angeliki Kritikakou, Gioele Casagranda, Marzio Vallero, Flavio Vella, Paolo Rech, Letícia Maria Veiras Bolzani, Milos Krstic, Marko S. Andjelkovic, Fabian Vargas 0001, Grigor Tshagharyan, Gurgen Harutunyan, Valery A. Vardanian, Samvel K. Shoukourian, Yervant Zorian, Jennifer Dworak, Kundan Nepal, Theodore W. Manikas, Mottaqiallah Taouil, Moritz Fieback, Anteneh Gebregiorgis, Rajendra Bishnoi, Said Hamdioui, Abhijit Chatterjee, Anurup Saha, Suhasini Komarraju, K. Ma, Chandramouli N. Amarnath, Mehdi Baradaran Tahoori, Mahta Mayahinia, Maryam Rajabalipanah, Katayoon Basharkhah, N. Nosrati, Zahra Jahanpeima, Zainalabedin Navabi, Hans-Joachim Wunderlich, Sybille Hellebrand |
ETS | 72 |
| 2025 | Confidence Driven Compact Testing of Compute-in-Memory Based Language Models
Anurup Saha, Chandramouli N. Amarnath, Kwondo Ma, Abhijit Chatterjee |
ETS | 2 |
| 2025 | Adaptive Testing of Compute-in-Memory Based CNNs Using Probabilistic Test Acceptance LimitsabstractCompute-in-memory (CiM) based convolutional neural network (CNN) accelerators achieve low-power inference, utilizing memristive crossbar arrays for matrix multiplications. However, inherent conductance variations within the crossbar introduce computational errors. These errors propagate to the CNN output and cause image misclassification, leading to substantial accuracy degradation. This paper addresses the critical challenge of efficient and reliable post-manufacture testing for CiM-based CNN accelerators. We propose a novel test image sampling methodology, which iteratively applies sampled images from the CNN's testing dataset using progressive random sampling (PRS) to a device under test (DUT) and estimates a confidence interval for the DUT accuracy. Based on the confidence interval and the acceptable accuracy threshold, the test labels a DUT as “pass” or “fail”. Furthermore, if we have access to an initial set of DUTs, we apply the images from the CNN's testing dataset to these DUTs and leverage the DUT outputs to rank-order test images. We develop a sequential estimation test (SET) framework, where the images from the CNN's testing dataset are sequentially applied according to a predetermined rank and the test terminates when a DUT can be confidently labeled as “pass” or “fail” based on the applied images. In each case, the number of applied test images adapts to the quality of the DUT. Experiments show that PRS and SET achieve$2.2\times$and$4.6\times$speedup compared to state-of-the-art test methodologies. Anurup Saha, Kwondo Ma, Chandramouli N. Amarnath, Moinuddin K. Qureshi, Abhijit Chatterjee |
IOLTS | 3 |
| 2025 | Error Resilient Online Reinforcement Learning Using Adaptive Statistical ChecksabstractOnline deep reinforcement learning (deep RL)-based systems are being increasingly deployed in a variety of safety-critical applications. Due to the dynamic nature of the environments they work in, onboard reinforcement learning (RL) hardware is vulnerable to soft errors from radiation, thermal effects and electrical noise that corrupts the results of computations. Existing approaches to on-line error resilience in machine learning systems have relied on the availability of large training datasets to configure resilience parameters. This is not always feasible for online RL systems. Similarly, other approaches involving specialized hardware or modifications to training algorithms are difficult to implement for onboard RL applications. In contrast, we present a novel error resilience approach for online RL that leverages running statistics of neuron output values collected across the (real-time) RL training process to configure error detection thresholds (called checks) for the deep RL forward pass. Similarly, we formulate checks on the deep RL backward pass using running statistical thresholds on reduced-dimension checksums of online learning weight updates to rapidly detect and correct errors in online deep RL training. In this methodology, statistical concentration bounds leveraging running statistics are used to diagnose neuron outputs or weights as erroneous. The use of running statistics allows the checks to adapt to changes caused by continual on-line RL training. Erroneous neurons are set to zero (suppressed) in the forward pass. Erroneous weight updates are frozen, allowing nonerroneous weight updates to proceed and allowing online learning without rerunning training episodes. Our approach is compared against the state of the art and validated on several RL algorithms as well as a hardware validation platform. Chandramouli N. Amarnath, Jackson Isenberg, Abhijit Chatterjee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | OATT: Outlier-Oriented Alternative Testing and Post-Manufacture Tuning of Analog/Mixed-Signal CircuitsabstractModern analog mixed-signal (AMS) devices manufactured in advanced CMOS processes pose significant testing and post-manufacture tuning challenges. Measurement of the specifications of AMS components is generally difficult as this requires the use of a range of dedicated tests while defect-based testing on the other hand, requires extensive defect simulations that are compute-intensive. To overcome these limitations, this research proposes OATT; a testing and post-manufacture tuning approach for AMS circuits that is designed to stress the performance of the device under test (DUT), formalize a statistical (multidimensional Gaussian) distribution of the expected response of known “good” devices (inliers), and use test limits grounded in theoretical statistics to classify all out-of-distribution devices (outliers) as “bad.” It is an alternative test approach in that it does not explicitly target simulation of defect mechanisms. Tuning is performed to transform individual outlier DUT responses to those resembling inlier devices by modulating hardware tuning knobs, such as bias voltages and currents, using a reinforcement learning algorithm. Circuit simulations and hardware results demonstrate the viability and efficiency of the proposed approach. Suhasini Komarraju, Akhil Tammana, Chandramouli N. Amarnath, Abhijit Chatterjee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | Signature Driven Post-Manufacture Testing and Tuning of RRAM Spiking Neural Networks for Yield RecoveryabstractResistive random access Memory (RRAM) based spiking neural networks (SNN) are becoming increasingly attractive for pervasive energy-efficient classification tasks. However, such networks suffer from degradation of performance (as determined by classification accuracy) due to the effects of process variations on fabricated RRAM devices resulting in loss of manufacturing yield. To address such yield loss, a two-step approach is developed. First, an alternative test framework is used to predict the performance of fabricated RRAM based SNNs using the SNN response to a small subset of images from the test image dataset, called the SNN response signature (to minimize test cost). This diagnoses those SNNs that need to be performance-tuned for yield recovery. Next, SNN tuning is performed by modulating the spiking thresholds of the SNN neurons on a layer-by-layer basis using a trained regressor that maps the SNN response signature to the optimal spiking threshold values during tuning. The optimal spiking threshold values are determined by an off-line optimization algorithm. Experiments show that the proposed framework can reduce the number of out-of-spec SNN devices by up to 54% and improve yield by as much as 8.6%. Anurup Saha, Chandramouli N. Amarnath, Kwondo Ma, Abhijit Chatterjee |
ASPDAC | 2 |
| 2024 | Learning Assisted Post-Manufacture Testing and Tuning of RRAM-Based DNNs for Yield RecoveryabstractVariability-induced accuracy degradation of RRAM-based DNNs is of great concern due to their significant potential for use in future energy-efficient machine learning architectures. To address this, we propose a two-step process. First, an enhanced testing procedure is used to predict DNN accuracy from a set of compact test stimuli (images). This test response (signature) is simply the concatenated vectors of output neurons of intermediate and final DNN layers over the compact test images applied. DNNs with a predicted accuracy below a threshold are then tuned based on this signature vector. Using a clustering based approach, the signature is mapped to the optimal tuning parameter values of the DNN (determined using off-line training of the DNN via back-propagation) in a single step, eliminating any post-manufacture training of the DNN weights (expensive). The tuning parameters themselves consist of the gains and offsets of the ReLU activation of neurons of the DNN on a per-layer basis and can be tuned digitally. Tuning is achieved in less than a second of tuning time, with yield improvements of over 45% with a modest accuracy reduction of 4% compared to digital DNNs. Kwondo Ma, Anurup Saha, Chandramouli N. Amarnath, Abhijit Chatterjee |
DATE | 3 |
| 2024 | AMS Test Stimulus Generation and Response Analysis Using Hyperdimensional Clustering: Minimizing Misclassification RateabstractPrevalent test strategies for analog/mixed-signal systems rely on either (a) prediction of device-under-test (DUT) design specifications from observed test responses to carefully crafted alternate test stimulus, or (b) detecting outliers from known optimized test response statistics of devices subjected to expected manufacturing process variations. In both of these test paradigms, misclassification of DUTs (false positives and false negatives) is not explicitly considered during test generation itself due to computational complexity, but rather based on post-test determination of test acceptance thresholds. In this paper, we propose a novel test generation approach based on hyperdimensional clustering, that explicitly targets DUT misclassification rate during test stimulus generation itself. The use of hyperdimensional vectors for clustering good and bad devices along with a set of simple vector operations for training and inference allows fast determination of misclassification rate within the test generation procedure itself. Experimental results show that the test generation times are reduced by 15X with significant improvements in DUT misclassification rate. Suhasini Komarraju, Akhil Tammana, Gowsika Dharmaraj, Chandramouli N. Amarnath, Abhijit Chatterjee |
ETS | 5 |
| 2024 | Post-Manufacture Criticality-Aware Gain Tuning of Timing Encoded Spiking Neural Networks for Yield RecoveryabstractTime-to-first-spike (TTFS) encoded spiking neural networks (SNNs), implemented using memristive crossbar arrays (MCA), achieve higher inference speed and energy efficiency compared to artificial neural networks (ANNs) and rate encoded SNNs. However, memristive crossbar arrays are vulnerable to conductance variations in the embedded memristor cells. These degrade the performance of TTFS encoded SNNs, namely their classification accuracy, with adverse impact on the yield of manufactured chips. To combat this yield loss, we propose a post-manufacture testing and tuning framework for these SNNs. In the testing phase, a timing encoded signature of the SNN, which is statistically correlated to the SNN performace, is extracted. In the tuning phase, this signature is mapped to optimal values of the tuning knobs (gain parameters), one parameter per layer, using a trained regressor, allowing very fast tuning (about 150ms). To further reduce the tuning overhead, we rank order hidden layer neurons based on their criticality and show that adding gain programmability only to 50% of the neurons is sufficient for performance recovery. Experiments show that the proposed framework can improve yield by up to 34% and average accuracy of memristive SNNs by up to 9%. Anurup Saha, Kwondo Ma, Chandramouli N. Amarnath, Abhijit Chatterjee |
ETS | 3 |
| 2024 | Efficient Optimized Testing of Resistive RAM Based Convolutional Neural NetworksabstractResistive random access memory (RRAM) based memristive crossbar arrays enable low power and low latency inference for convolutional neural networks (CNNs), making them suitable for deployment in IoT and edge devices. However RRAM cells within a crossbar suffer from conductance variations, making RRAM-based CNNs vulnerable to degradation of their classification accuracy. To address this, the classification accuracy of RRAM based CNN chips can be estimated using predictive tests, where a trained regressor predicts the accuracy of a CNN chip from the CNN’s response to a compact test dataset. In this research, we present a framework for co-optimizing the pixels of the compact test dataset and the regressor. The novelty of the proposed approach lies in the ability to co-optimize individual image pixels, overcoming barriers posed by the computational complexity of optimizing the large numbers of pixels in an image using state-of-the-art techniques. The co-optimization problem is solved using a three step process: a greedy image down selection followed by backpropagation driven image optimization and regressor fine-tuning. Experiments show that the proposed test approach reduces the CNN classification accuracy prediction error by $31 \%$ compared to the state of the art. It is seen that a compact test dataset with only 2-4 images is needed for testing, making the scheme suitable for built-in test applications. Anurup Saha, Kwondo Ma, Chandramouli N. Amarnath, Abhijit Chatterjee |
IOLTS | 3 |
| 2024 | Error Resilient Hyperdimensional Computing Using Hypervector Encoding and Cross-ClusteringabstractEmerging brain-inspired hyperdimensional computing (HDC) algorithms are vulnerable to timing and soft errors in associative memory used to store high-dimensional data representations. Such errors can significantly degrade HDC performance. A key challenge is error correction after an error in computation is detected. This work presents two novel error resilience frameworks for hyperdimensional computing systems. The first, called the checksum hypervector encoding (CHE) framework, relies on creation of a single additional hypervector that is a checksum of all the class hypervectors of the HDC system. For error resilience, elementwise validation of the checksum property is performed and those elements across all class vectors for which the property fails are removed from consideration. For an HDC system with K class hypervectors of dimension D, the second cross-hypervector clustering (CHC) framework clusters $D , K$- dimensional vectors consisting of the i-th element of each of the K HDC class hypervectors, $1 \le \quad i \quad \le \quad K$. Statistical properties of these vector clusters are checked prior to each hypervector query and all the elements of all K-dimensional vectors corresponding to statistical outlier vectors are removed as before. The choice of which framework to use is dictated by the complexity of the dataset to classify. Up to three orders of magnitude better resilience to errors than the state-of-the-art across multiple HDC high-dimensional encoding (representation) systems is demonstrated.11Our codes and data are available at https://github.com/mmejri3/er-hdc Chandramouli N. Amarnath, Abhijit Chatterjee |
VTS | 2 |
| 2024 | Error Resilience in Deep Neural Networks Using Neuron Gradient StatisticsabstractModern deep neural networks (DNNs) are deployed across a wide range of applications, from medical robotics to autonomous driving, where safety and reliability are key concerns. The complexity, speed, and low-power operation of the underlying hardware makes them vulnerable to soft errors that corrupt the results of computations and memory accesses. Existing approaches to error resilience are either expensive in terms of overhead, require DNN retraining or applicable to only specific hardware domains. In contrast, we present a novel error resilience approach that does not require DNN retraining and scales across computation as well as weight parameter errors. In the proposed methodology, the statistics of gradients of neuron output values relative to adjacent neurons in an ordering of neurons allow tight theoretically grounded thresholding of neuron outputs to diagnose erroneous neuron outputs. These are then set to zero (suppressed) for error resilience. A low-overhead error diagnosis module is used for this purpose and is designed using gradient statistics collected across the training dataset of the DNN. Our approach is compared against state of the art error resilience techniques and validated on multiple datasets, networks and error scenarios as well a hardware test case. Chandramouli N. Amarnath, Kwondo Ma, Abhijit Chatterjee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Error Resilient Transformers: A Novel Soft Error Vulnerability Guided Approach to Error Checking and Suppression
Kwondo Ma, Chandramouli N. Amarnath, Abhijit Chatterjee |
ETS | 2 |
| 2023 | A Resilience Framework for Synapse Weight Errors and Firing Threshold Perturbations in RRAM Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) can be implemented with power-efficient digital as well as analog circuitry. However, in Resistive RAM (RRAM) based SNN accelerators, synapse weights programmed into the crossbar can differ from their ideal values due to defects and programming errors, degrading inference accuracy. In addition, circuit nonidealities within analog spiking neurons that alter the neuron spiking rate (modeled by variations in neuron firing threshold) can degrade SNN inference accuracy when the value of inference time steps (ITSteps) of SNN is set to a critical minimum that maximizes network throughput. We first develop a recursive linearized check to detect synapse weight errors with high sensitivity. This triggers a correction methodology which sets out-of-range synapse values to zero. For correcting the effects of firing threshold variations, we develop a test methodology that calibrates the extent of such variations. This is then used to proportionally increase inference time steps during inference for chips with higher variation. Experiments on a variety of SNNs prove the viability of the proposed resilience methods. Anurup Saha, Chandramouli N. Amarnath, Abhijit Chatterjee |
ETS | 2 |
| 2023 | A Novel Approach to Error Resilience in Online Reinforcement LearningabstractOnline reinforcement learning (RL) based systems are being increasingly deployed in a variety of safety-critical applications ranging from drone control to medical robotics. These systems typically use RL onboard rather than relying on remote operation from high-performance datacenters. Due to the dynamic nature of the environments they work in, onboard RL hardware is vulnerable to soft errors from radiation, thermal effects and electrical noise that corrupt the results of computations. Existing approaches to on-line error resilience in machine learning systems have relied on availability of the large training datasets to configure resilience parameters, which is not necessarily feasible for online RL systems. Similarly, other approaches involving specialized hardware or modifications to training algorithms are difficult to implement for onboard RL applications. In contrast, we present a novel error resilience approach for online RL that makes use of running statistics collected across the (real-time) RL training process to configure error detection thresholds without the need to access a reference training dataset. In this methodology, statistical concentration bounds leveraging running statistics are used to diagnose neuron outputs as erroneous. These erroneous neurons are then set to zero (suppressed). Our approach is compared against the state of the art and validated on several RL algorithms involving the use of multiple concentration bounds on CPU as well as GPU hardware. Chandramouli N. Amarnath, Abhijit Chatterjee |
IOLTS | 1 |
| 2023 | OATT: Outlier Oriented Alternative Testing and Post-Manufacture Tuning of Mixed-Signal/RF Circuits and SystemsabstractPrevalent specification-based AMS testing techniques require the use of complex test circuits or regressors that are difficult to implement on-chip as well as suffer from coverage loss when devices are under-specified. Complementary defect based testing techniques require the simulation of explosively large defect sets under assumed failure mechanisms. We overcome these limitations in our proposed approach OATT; an Outlier-oriented Alternative Testing and Tuning methodology. OATT maximizes the number and magnitude of the statistical principal components (PCA) of the time-domain DUT test response vectors across diverse manufacturing process corners. This allows construction of a multi-dimensional Gaussian probability density model that characterizes the distribution of DUT responses in the principal components domain. Outliers of this probability density model are classified as defective devices using calibrated confidence ellipses, implicitly detecting devices with parametric as well as hard defects. The embedded DUT response is acquired using coherent undersampling and does not require explicit signal reconstruction. Post-manufacture tuning is performed by minimizing the statistical distance of the DUT response in the PCA domain from the nominal Gaussian model using multi-arm bandit reinforcement learning. Simulation results demonstrate the viability and promise of the proposed approach. Suhasini Komarraju, Akhil Tammana, Chandramouli N. Amarnath, Abhijit Chatterjee |
ITC | 3 |
| 2022 | Soft Error Resilient Deep Learning Systems Using Neuron Gradient StatisticsabstractDeep learning techniques have been widely adopted in daily life with applications ranging from face recognition to recommender systems. The substantial overhead of conventional error tolerance techniques precludes their widespread use, while approaches involving median filtering and invariant generation rely on alterations to DNN training that may be difficult to achieve for larger networks on larger datasets. To address this issue, this paper presents a novel approach taking advantage of the statistics of neuron output gradients to identify and suppress erroneous neuron values. By using the statistics of neurons’ gradients with respect to their neighbors, tighter statistical thresholds are obtained compared to the use of neuron output values alone. This approach is modular and is combined with accurate, low-overhead error detection methods to ensure it is used only when needed, further reducing its cost. Deep learning models can be trained using standard methods and our error correction module is fit to a trained DNN, achieving comparable or superior performance compared to baseline error correction methods while incurring comparable hardware overhead without needing to modify DNN training or utilize specialized hardware architectures. Chandramouli N. Amarnath, Kwondo Ma, Abhijit Chatterjee |
IOLTS | 1 |
| 2022 | Efficient Low Cost Alternative Testing of Analog Crossbar Arrays for Deep Neural NetworksabstractAnalog crossbar arrays have recently attracted significant attention due to their usefulness for deep neural net (DNN) computations with ultra-low power consumption. However, recent studies have shown that DNNs implemented with such crossbar arrays suffer from as high as 30% degradation in performance due to the effects of manufacturing process variability effects resulting in degradation of their functional safety. One way to test these DNNs is to apply an exhaustive set of test images to each device to ascertain its performance. This is expensive and time-consuming. We propose an alternative test scheme in which a small subset of test images is applied to each DNN and the classification accuracy of the DNN is predicted directly from observation of the final layer outputs of the network. This saves test cost while allowing binning of DNNs for performance. Experimental results for a variety of test cases are presented and show test efficiency improvements of 10.3X over testing with the exhaustive test image set. Kwondo Ma, Anurup Saha, Chandramouli N. Amarnath, Abhijit Chatterjee |
ITC | 3 |
| 2021 | Addressing Soft Error and Security Threats in DNNs Using Learning Driven Algorithmic ChecksabstractThe reliability of Deep Neural Networks (DNNs) is of great concern due to their widespread use in safety-critical applications. Prior research has focused on adaptation of algorithm based fault tolerance schemes for error detection in the dot product (linear) computations of DNNs. In this research, we show that compact machine-learned algorithmic checks inspired by prior work on linear checksums but adapted to the overall nonlinear nature of DNN computations can be used to detect both soft errors and image-triggered trojans in real-time. Experiments indicate that the method incurs low computation overhead (4%-20%), while achieving high coverage (up to 90% for soft errors and 99% for image-triggered attacks). Chandramouli N. Amarnath, Md Imran Momtaz, Abhijit Chatterjee |
IOLTS | 1 |
| 2021 | Hierarchical Failure Modeling and Machine Learning Assisted Correction of Electro-Mechanical Subsystem Failures in Autonomous VehiclesabstractAutonomous systems that rely on multiple interacting subsystems require a high degree of reliability and resilience to a wide range of failures in those subsystems. In this work the effects of electro-mechanical failures in the steer-by-wire, brake-by-wire and vehicle controller subsystems of autonomous vehicles on subsystem and vehicle level performance are studied. A machine learning assisted correction approach using Gaussian Processes to learn fault dynamics on-line is developed and its efficacy is demonstrated under a variety of vehicle maneuvers and failure conditions at the subsystem and vehicle levels. Chandramouli N. Amarnath, Md Imran Momtaz, Abhijit Chatterjee |
ITC | 1 |
| 2020 | Encoded Check Driven Concurrent Error Detection in Particle Filters for Nonlinear State EstimationabstractIn this paper we propose a framework for concurrent detection of soft computation errors in particle filters which are finding increasing use in robotics applications. The particle filter works by sampling the multi-variate probability distribution of the states of a system (samples called particles, each particle representing a vector of states) and projecting these into the future using appropriate nonlinear mappings. We propose the addition of a `check' state to the system as a linear combination of the system states for error detection. The check state produces an error signal corresponding to each particle, whose statistics are tracked across a sliding time window. Shifts in the error statistics across all particles are used to detect soft computation errors as well as anomalous sensor measurements. Simulation studies indicate that errors in particle filter computations can be detected with high coverage and low latency. Chandramouli N. Amarnath, Md Imran Momtaz, Abhijit Chatterjee |
IOLTS | 1 |
| 2020 | Concurrent Error Detection in Embedded Digital Control of Nonlinear Autonomous Systems Using Adaptive State Space ChecksabstractThe advent of pervasive autonomous systems such as self-driving cars and drones has raised questions about their safety and trustworthiness. This is particularly relevant in the event of on-board subsystem errors or failures. In this research, we show how encoded Extended Kalman Filter can be used to detect anomalous behaviors of critical components of nonlinear autonomous systems: sensors, actuators, state estimation algorithms and control software. As opposed to prior work that is limited to linear systems or requires the use of cumbersome machine learned checks with fixed detection thresholds, the proposed approach necessitates the use of time-varying checks with dynamically adaptive thresholds. The method is lightweight in comparison to existing methods (does not rely on machine learning paradigms) and achieves high coverage as well as low detection latency of errors. A quadcopter and an automotive steer-by-wire system are used as test vehicles for the research and simulation and hardware results indicate the overhead, coverage and error detection latency benefits of the proposed approach. Md Imran Momtaz, Chandramouli N. Amarnath, Abhijit Chatterjee |
ITC | 2 |