Suvadeep Banerjee

dblp:60/9443 · DBLP profile ↗
← Back
34ranked-venue papers
11as first author
15since 2021 · last 2025
0000-0001-5188-1651ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 32 · 9 first-author · 14 since 2021Software engineering, systems software and programming languages · 9 · 3 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 Machine Learning-Driven STL Generation for Enhancing Functional Safety of E/E Systems
abstract
The increasing complexity of safety-critical hardware systems demands advanced methods for ensuring functional safety (FuSa). Traditional techniques like ATPG and BIST are intrusive, requiring additional hardware and disrupting operations, making them unsuitable for in-field testing. To address this, for the first time, we propose a machine learning (ML)-driven automated Self-Test Library (STL) generation for seamless in-field testing during idle periods, ensuring uninterrupted fault detection and high system performance. Utilizing reinforcement learning, the STL generates design-specific test patterns, achieving up to $57.57 \%$ improvement in fault coverage and up to $85 \%$ efficiency compared to existing pattern-based testing, enhancing FuSa in mission-critical applications.
Sanjay Das, Swastik Bhattacharya, Anand Menon, Shamik Kundu, Pooja Madhusoodhanan, Prasanth Viswanathan Pillai, Rubin A. Parekhji, Arnab Raha, Suvadeep Banerjee, Suriyaprakash Natarajan, Kanad Basu
DAC9
2025 Enhancing AMS Circuit Reliability: An Anomaly Dataset for Functional Safety Research in Automotive SoCs
Sanjay Das, Anand Menon, Omar Abiola Abioye, Afreen Fatimah Khazi-Syed, Jonathan Edward Lee, Ayush Arunachalam, Shamik Kundu, Pooja Madhusoodhanan, Prasanth Viswanathan Pillai, Rubin A. Parekhji, Arnab Raha, Suvadeep Banerjee, Suriyaprakash Natarajan, Kanad Basu
ACM Great Lakes Symposium on VLSI12
2025 OpenAssert: Towards Secure Assertion Generation using Large Language Models
abstract
Assertions are critical components used in hardware verification, ensuring robust functionality, fortifying design security, and providing essential verification features. Traditional hardware assertion methods are not automated, complicate security audits, and require effort, causing prolonged development cycles. Recent studies have highlighted the potential of commercial Large Language Models (LLMs) to generate security-focused assertions by leveraging textual data from design specifications. However, reliance on proprietary models like GPT-4 severely jeopardizes IP privacy and data confidentiality, undermining transparency and accountability in data handling practices. In this paper, we address secure hardware assertion generation by proposing a practical approach to significantly enhance the feasibility of open-source LLMs. Our proposed method, OpenAssert, involves fine-tuning existing models to be utilized locally at the user’s end without compromising confidentiality. Additionally, we employ Retrieval Augmentation Generation to refine these models, mitigating hallucinations and security-related errors. OpenAssert demonstrates improvements, achieving up to a 44% increase in rouge-1 score, a 49% improvement in cosine similarity, and a 43.4% reduction in word error rate for security-critical designs compared to open-source models.
Anand Menon, Samit Shahnawaz Miftah, Amisha Srivastava, Shamik Kundu, Shovik Kundu, Arnab Raha, Suvadeep Banerjee, Deepak Mathaikutty, Kanad Basu
VTS7
2024 Graph Learning-based Fault Criticality Analysis for Enhancing Functional Safety of E/E Systems
abstract
The increasing complexity of Electrical and Electronic (E/E) systems underscores the need for protective measures to ensure functional safety (FuSa) in high-assurance environments. This entails the identification and fortification of vulnerable nodes to enhance system reliability during mission-critical scenarios. Traditionally, the assessment of E/E system reliability has relied on fault injection (FI) techniques and simulations. However, FI faces challenges in coping with escalating design complexity, including resource demands and timing overheads. Furthermore, it falls short in identifying critical components that may lead to functional failures. To address these challenges, we propose a Machine Learning (ML)-based framework for predicting critical nodes in hardware designs. The process begins with constructing a graph from the design netlist, forming the foundation for training a Graph Convolutional Network (GCN). The GCN model utilizes graph node attributes, node labels, and edge connections to learn and predict critical nodes in the circuit. The model furnishes up to 93.7% accuracy in identifying vulnerable circuit nodes during evaluation on diverse designs such as Synchronous Dynamic Random Access Memory (SDRAM) controller, OpenRISC 1200 (OR1200) modules. Furthermore, we incorporate an explainability analysis to interpret individual node predictions. This analysis discerns the critical design factors influencing fault criticality in the design. Moreover, to the best of our knowledge, we, for the first time, perform a regression analysis to generate node criticality scores, quantifying the degrees of criticality, that can enable prioritizing resources towards critical nodes.
Sanjay Das, Shamik Kundu, Pooja Madhusoodhanan, Prasanth Viswanathan Pillai, Rubin A. Parekhji, Arnab Raha, Suvadeep Banerjee, Suriyaprakash Natarajan, Kanad Basu
DAC7
2024 DiagNNose: Toward Error Localization in Deep Learning Hardware-Based on VTA-TVM Stack
abstract
Low-level hardware faults manifested in a Deep learning (DL) accelerator usher in graceless degradation of high-level classification accuracy, which can eventuate to catastrophic circumstances. This violates the crucial Functional Safety (FuSa) of the DL accelerator, maintaining which is imperative in high-assurance applications. Conventional techniques for error localization incur high-test efforts, without regards to the unique challenges posed by DL systems. In this direction, we propose DiagNNose, a two-tier machine learning-based error localization framework for on-line fault management in DL accelerators. We develop a novel diagnostic pattern selection algorithm to obtain a minimal subset of functional test patterns, that are executed in the accelerator in mission mode. By extracting and analyzing dataflow-based features from the intermediate computations of the general matrix multiply (GEMM) core, a lightweight multilayer perceptron accomplishes bit-level error localization in 8-bit, 16-bit, and 32-bit datapath units with high fidelity. We have limited ourselves to a single accelerator design, i.e., the versatile tensor accelerator (VTA) architecture to evaluate our proposed DiagNNose framework. On executing state-of-the-art deep neural networks trained on ImageNet; error localization using only 30 diagnostic functional test patterns demonstrate up to 98.4% diagnosability, thereby demonstrating an improvement of 54.63% over a random test pattern set, with as low as 4.95% overhead in the DL accelerator in mission mode.
Shamik Kundu, Suvadeep Banerjee, Arnab Raha, Suriyaprakash Natarajan, Kanad Basu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Reusing GEMM Hardware for Efficient Execution of Depthwise Separable Convolution on ASIC-Based DNN Accelerators
abstract
Deep learning (DL) accelerators are optimized for standard convolution. However, lightweight convolutional neural networks (CNNs) use depthwise convolution (DwC) in key layers, and the structural difference between DwC and standard convolution leads to significant performance bottleneck in executing lightweight CNNs on such platforms. This work reuses the fast general matrix-vector multiplication (GEMM) core of DL accelerators by mapping DwC to channel-wise parallel matrix-vector multiplications. An analytical framework is developed to guide pre-RTL hardware choices, and new hardware modules and software support are developed for end-to-end evaluation of the solution. This GEMM-based DwC execution strategy offers substantial performance gains for lightweight CNNs: 7× speedup and 1.8× lower off-chip communication for MobileNet-v1 over a conventional DL accelerator, and 74× speedup over a CPU, and even 1.4× speedup over a power-hungry GPU.
Susmita Dey Manasi, Suvadeep Banerjee, Abhijit Davare, Anton A. Sorokin, Steven M. Burns, Desmond Kirkpatrick, Sachin S. Sapatnekar
ASP-DAC2
2023 Enhanced ML-Based Approach for Functional Safety Improvement in Automotive AMS Circuits
abstract
The extensive adoption of safety-critical applications in high-assurance environments, such as the automotive domain, has laid emphasis on safeguarding the reliability and Functional Safety (FuSa) of the Electrical and/or Electronic (E/E) components constituting such systems. Most modern automotive Systems-on-Chips (SoCs) comprise Analog and Mixed Signal (AMS) circuits, which are more susceptible to faults than their digital equivalents. However, their attributes of operating in the continuous signal region can be leveraged to perform early anomaly detection, which could facilitate the subversion of the eventual hardware failure state, thereby improving the FuSa of the system. To this end, we had proposed a novel unsupervised learning-based early anomaly detection framework catered to automotive AMS circuits (in ITC 2022). However, existing approaches to AMS FuSa violation detection are limited by pre-specified feature inputs, and lack rationale for identifying signals to be monitored to perform anomaly detection. To address these issues as well as further augment our original solution, in this paper, we propose a novel anomaly detection strategy that involves: (1) a genetic algorithm-based feature selection approach, (2) a novel signal selection algorithm that ascertains the best intermediate circuit signal, for furnishing enhanced anomaly detection accuracy, while reducing the associated detection latency, and (3) an explainable AI (XAI)-based framework that boosts user interpretability and transparency of the anomaly detection framework. This XAI approach, in turn, can be provided as feedback to the designer during circuit design and validation. The proposed approach is evaluated using case studies of two representative AMS circuits, which are prevalent in automotive SoCs. Our experimental analyses demonstrate that the proposed approach furnishes up to 100% detection accuracy and 2.3× reduction in detection time compared to our existing framework, in addition to providing insights by improving transparency of the anomaly detection framework, thereby exhibiting the efficacy of our solution.
Ayush Arunachalam, Sanjay Das, Monikka Rajan, Xiankun Jin, Suvadeep Banerjee, Arnab Raha, Suriyaprakash Natarajan, Kanad Basu
ITC6
2023 Trouble-Shooting at GAN Point: Improving Functional Safety in Deep Learning Accelerators
abstract
The proliferation of Deep Neural Networks (DNNs) in real-time mission critical applications has promoted the implementation of custom-built DNN inference accelerators. These accelerators require a considerable amount of on-chip memory to store millions of trained DNN parameters for executing inference at the edge. Drastic technology scaling in recent years have made these memory circuits highly vulnerable to faults due to various reasons like aging, latent defects, single event upsets, etc. Such faults are highly detrimental to the classification accuracy of the DNN accelerator, leading to the crucial Functional Safety (FuSa) violation. This can eventuate to catastrophic circumstances, when used in mission-critical applications. In order to detect such violations in mission mode, we propose to generate a set of functional test patterns by leveraging the concept of Generative Adversarial Networks (GANs), that are independent of the DNN model and the accelerator characteristics. Our experimental results demonstrate that, the generated test patterns significantly improve FuSa violation detection coverage by up to 130.28%, compared to existing techniques. To the best of our knowledge, this is the first work that generates GAN-based test patterns in order to perform FuSa violation detection in mission-critical DNN accelerators.
Shamik Kundu, Suvadeep Banerjee, Arnab Raha, Suriyaprakash Natarajan, Kanad Basu
IEEE Trans. Computers2
2023 A Novel Low-Power Compression Scheme for Systolic Array-Based Deep Learning Accelerators
abstract
The proliferation of deep learning algorithms has catalyzed their utilization to solve a multitude of real-world problems. Algorithms such as deep neural networks (DNNs) are compute- and power-intensive, thereby accentuating the development of hardware platforms like DNN inference accelerators. However, inference execution of large DNNs in resource-constrained environments induces energy bottlenecks in these accelerators. Since large DNNs consist of hundreds of millions of trained parameters, accessing them from the accelerator memory incurs substantial energy. To address this challenge, we propose HardCompress, which, to the best of our knowledge, is the first low-power solution that uses traditional compression strategies pertaining to commercial DNN accelerators in resource-constrained IoT edge devices. The three-step approach involves hardware-based post-quantization trimming of weights, followed by their dictionary-based compression and subsequent decompression by a low-power hardware engine during inference in the accelerator. We evaluate the proposed solution on lightweight networks trained on the MNIST dataset, the compact model trained on the CIFAR-10 dataset, and large DNNs trained on the ImageNet dataset. Performance of HardCompress at different quantization levels has been analyzed. Furthermore, to quantify the effectiveness of the proposed solution, an energy framework that contrasts the DRAM energies of the original and HardCompressed models has been developed. Finally, a fault injection framework which compares the fault resilience of the original model with its HardCompressed counterpart is also proposed. Our results exhibit that HardCompress, without any performance degradation in large DNNs, furnishes a maximum compression of 99.27%, equivalent to$137\times $reduction in memory footprint and 0.07 J for 8-bit quantization in the systolic array-based DNN accelerator. Furthermore, our proposed low-power decompression engine incurs an area overhead of only 0.02%; thus, enabling HardCompress’ utilization in resource-constrained environments.
Ayush Arunachalam, Shamik Kundu, Arnab Raha, Suvadeep Banerjee, Suriyaprakash Natarajan, Kanad Basu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 Unsupervised Learning-based Early Anomaly Detection in AMS Circuits of Automotive SoCs
abstract
With the proliferation of safety-critical applications in the automotive domain, it is imperative to guarantee the functional safety of circuits and components constituting automotive systems, e.g., the electrical and/or electronic subsystems in automotive vehicles. Analog and Mixed-Signal (AMS) circuits, prevalent in such systems, are more susceptible to faults than their digital counterparts, due to advanced manufacturing nodes, parametric perturbations, environmental stress, etc. However, their continuous signal characteristics provide an opportunity for early anomaly detection, which in turn, facilitates the deployment of safety mechanisms to prevent eventual system failure. Towards this end, we propose a novel unsupervised machine learning-based framework to perform early anomaly detection in AMS circuits. Our approach involves anomaly injection in various circuit locations and individual components to develop a training dataset encompassing a wide range of possible anomalous scenarios, feature extraction from observation signals, and clustering algorithms to facilitate anomaly detection. To this end, we propose a novel centroid selection technique for the unsupervised learning algorithms, which is tailored for detecting anomalies in AMS circuits. This approach furnishes high fidelity anomaly detection by identifying the ideal cluster centers corresponding to anomalous and non-anomalous signals. Furthermore, time series-based analysis is proposed to improve and expedite the anomaly detection performance. We evaluated our solution using a case study of two AMS circuits commonly present in automotive systems-on-chips. Our experimental results exhibit that the proposed approach furnishes up to 100% accuracy. Additionally, the time series-based technique reduces the anomaly detection latency by 5×, thereby demonstrating the efficacy of our solution.
Ayush Arunachalam, Athulya Kizhakkayil, Shamik Kundu, Arnab Raha, Suvadeep Banerjee, Robert Jin, Kanad Basu
ITC5
2022 DEFCON: Defect Acceleration through Content Optimization
abstract
During manufacturing of integrated circuits, it is imperative for cost and quality that defects that occur on die are screened early in the test process, preferably before packaging. As part of screening, stress steps are performed to accelerate latent defects so that they become observable and are detected by subsequent test steps. Traditionally, the levers for applying stress have been increased supply voltage and temperature while concurrently running a sliver of content that had been created to “test” defects. In this paper, we describe a methodology to generate content specifically targeting stress at latent defects by maximizing electrical activity. Our goal is to efficiently accelerate all classes of latent defects. This paper will give details of the technology and content generation methodology for scanned digital logic. Silicon results are provided on a recent client product. Ongoing work that extends this to address latent defects in other areas of the die are outlined.
Suriyaprakash Natarajan, Abhijit Sathaye, Chaitali Oak, Nipun Chaplot, Suvadeep Banerjee
ITC5
2022 Special Session: Effective In-field Testing of Deep Neural Network Hardware Accelerators
abstract
Ongoing research to obtain high performance Deep Neural Network (DNN) executions have led to the development of customized purpose-built deep learning inference accelerators. DNN accelerators are susceptible to faults, due to high-energy particles, process variations, temperature and structural deformities manifesting as latent defects. These faults can introduce misclassification, thereby jeopardizing the Functional Safety (FuSa) of the accelerator in mission mode, which can eventuate to disastrous consequences, including loss of human lives. In this paper, we explore the impact of such faults on the FuSa of a DNN accelerator by varying the network parameters, position and characteristics of the injected fault across multiple exhaustive datasets. Furthermore, we analyze the efficiency of a software-based self test scheme to detect FuSa violations in the accelerator in mission mode, that employs functional test patterns, akin to instances in the application dataset. The test patterns, selected from the dataset of the DNN, furnish up to 100% coverage with cardinality as low as 0.1% of the entire test dataset.
Shamik Kundu, Suvadeep Banerjee, Arnab Raha, Kanad Basu
VTS2
2021 Online Fast Detection and Diagnosis of Power Grid Security Attacks Using State Checksums
abstract
State information transmitted across communication links of distributed power grids can be compromised by security attacks on these links causing power fluctuations and outage. Encryption of transmitted data is a first line of defense against such attacks but the entire network can be compromised if the keys are broken such as through side-channel attacks on encryption hardware. We propose the use of low cost state checksums as a critical second line of defense against such attacks. It is seen that stealth and replay attacks that can defeat prior residual filter based attack detection methods, can be detected using state checksums. Further, the proposed attack detection approach is highly resilient to adversarial counterattacks. A key benefit is low latency diagnosis of compromised communication links (and thereby sub-networks) not possible through use of encryption techniques alone. Diagnosis is driven by a dynamic binary search driven partitioning of system states into compromised vs. nominal subsets. We demonstrate through simulation of IEEE benchmark systems and two high-voltage power grids that our proposed scheme significantly reduces the latency of attack mitigation with lower complexity in comparison to prior research.
Suvadeep Banerjee, Abhijit Chatterjee
IOLTS1
2021 Real-Time Error Detection in Nonlinear Control Systems Using Machine Learning Assisted State-Space Encoding
abstract
Successful deployment of autonomous systems in a wide range of societal applications depends on error-free operation of the underlying signal processing and control functions. Real-time error detection in nonlinear systems has mostly relied on redundancy at the component or algorithmic level causing expensive area and power overheads. This paper describes a real-time error detection methodology for nonlinear control systems for detecting sensor and actuator degradations as well as malfunctions due to soft errors in the execution of the control algorithm on a digital processor. Our approach is based on creation of a redundant check state in such a way that its value can be computed from the current states of the system as well as from a history of prior observable state values and inputs (via machine learning algorithms). By checking for consistency between the two, errors are detected with low latency. The method is demonstrated on two test case simulations - an inverted pendulum balancing problem and a sliding mode controller driven brake-by-wire (BBW) system. In addition, hardware results from error injection experiments in an ARM core representation on an FPGA and artificial sensor degradations on a self-balancing robot prove the practical feasibility of implementation.
Suvadeep Banerjee, Balavinayagam Samynathan, Jacob A. Abraham, Abhijit Chatterjee
IEEE Trans. Dependable Secur. Comput.1
2021 Toward Functional Safety of Systolic Array-Based Deep Learning Hardware Accelerators
abstract
High accuracy and ever-increasing computing power have made deep neural networks (DNNs) the algorithm of choice for various machine learning, computer vision, and image processing applications across the computing spectrum. To this end, Google developed the tensor processing unit (TPU) to accelerate the computationally intensive matrix multiplication operation of a DNN on its systolic array architecture. Faults manifested in the datapath of such a systolic array due to latent manufacturing defects or single-event effects may lead to functional safety (FuSa) violation. Although DNNs are known to resist minor perturbations with their inherent fault-tolerant characteristics, we show that the classification accuracy of the model plummets from 97.4% to 7.75% with a minimal fault rate of 0.0003% in the accelerator, implying catastrophic circumstances when deployed across mission-critical systems. Hence, to ensure FuSa of such accelerators, this article provides an extensive FuSa assessment of the accelerator exposed to faults in the datapath, by varying the network parameters, position, and characteristics of the induced error across multiple exhaustive data sets. Furthermore, we propose two novel strategies to obtain a diminutive set of functional test patterns to detect FuSa violation in a DNN accelerator. Our experimental results demonstrate that the obtained test sets can achieve an average of 92.63% (in some cases, up to 100%) fault coverage with cardinality as low as 0.1% of the entire test data set.
Shamik Kundu, Suvadeep Banerjee, Arnab Raha, Suriyaprakash Natarajan, Kanad Basu
IEEE Trans. Very Large Scale Integr. Syst.2
2019 Mixed Signal Design Validation Using Reinforcement Learning Guided Stimulus Generation for Behavior Discovery
abstract
High operating speeds and use of aggressive fabrication technologies necessitate validation of mixed-signal electronic systems at every stage of top-down design: behavioral to netlist to physical design to silicon. At each step, design validation establishes the equivalence of lower level design descriptions against their higher level specifications. Prior research has leveraged state reachability analysis, nonconvex optimization, or performance specifications in order to generate tests. In contrast, we reformulate the systems under validation as a Markov decision process and examine the use of reinforcement-learning to provide a globally convergent solution, a means of “storing” the valuable information created during stimulus generation, and low-cost iterated generation. The integration of the proposed design validation methodology with deep-Q learning software and the suite of Cadence simulation tools is presented, validation results for selected design bugs in representative designs are analyzed, and the quality and efficiency of the proposed design validation methodology is discussed.
Barry John Muldrey, Suvadeep Banerjee, Abhijit Chatterjee
VTS2
2019 ALERA: Accelerated Reinforcement Learning Driven Adaptation to Electro-Mechanical Degradation in Nonlinear Control Systems Using Encoded State Space Error Signatures
abstract
The successful deployment of autonomous real-time systems is contingent on their ability to recover from performance degradation of sensors, actuators, and other electro-mechanical subsystems with low latency. In this article, we introduce ALERA, a novel framework for real-time control law adaptation in nonlinear control systems assisted by system state encodings that generate an error signal when the code properties are violated in the presence of failures. The fundamental contributions of this methodology are twofold—first, we show that the time-domain error signal contains perturbed system parameters’ diagnostic information that can be used for quick control law adaptation to failure conditions and second, this quick adaptation is performed via reinforcement learning algorithms that relearn the control law of the perturbed system from a starting condition dictated by the diagnostic information, thus achieving significantly faster recovery. The fast (up to 80X faster than traditional reinforcement learning paradigms) performance recovery enabled by ALERA is demonstrated on an inverted pendulum balancing problem, a brake-by-wire system, and a self-balancing robot.
Suvadeep Banerjee, Abhijit Chatterjee
ACM Trans. Intell. Syst. Technol.1
2018 ReiNN: Efficient error resilience in artificial neural networks using encoded consistency checks
abstract
In this research, a low cost error detection and correction approach is developed for multilayer perceptron networks, where checker neurons are used to encode hidden layer functions using independent training experiments. Error detection and correction is predicated on validating consistency properties of the encoded checks and shows that high coverage of injected errors can be achieved with extremely low computational overhead.
Sujay Pandey, Suvadeep Banerjee, Abhijit Chatterjee
ETS2
2018 Cross-Layer Control Adaptation for Autonomous System Resilience
abstract
The last decade has seen tremendous advances in the transformation of ubiquitous control, computing and communication platforms that are anytime, anywhere. These platforms allow humans to interact with machines through sensing, control and actuation functions in ways not imaginable a few decades ago. While robust control techniques aim to maintain autonomous system performance in the presence of bounded modeling errors, they are not designed to manage large multi- parameter variations and internal component failures that are inevitable during lengthy periods of field deployment. To address the trustworthiness of autonomous systems in the field, we propose a cross-layer error resilience approach in which errors are detected and corrected at appropriate levels of the design (hardware-through software) with the objective of minimizing the latency of error recovery while maintaining high failure coverage. At the control processor level, soft errors in the digital control processor are considered. At the system level, sensor and actuator failures are analyzed. These impairments define the health of the system. A methodology for adapting the control procedure of the autonomous system to compensate for degraded system health is proposed. It is shown how this methodology can be applied to simple linear and nonlinear control systems to maintain system performance in the presence of internal component failures. Experimental results demonstrate the feasibility of the proposed methodology.
Md Imran Momtaz, Suvadeep Banerjee, Sujay Pandey, Jacob A. Abraham, Abhijit Chatterjee
IOLTS2
2018 Error Resilient Neuromorphic Networks Using Checker Neurons
abstract
The last decade has seen tremendous advances in the application of artificial neural networks to solving problems that mimic human intelligence. Many of these systems are implemented using traditional digital compute engines where errors can occur during memory accesses or during numerical computation. While such networks are inherently error resilient, specific errors can result in incorrect decisions. This work develops a low overhead error detection and correction approach for multilayer artificial neural networks, here the hidden layer functions are approximated using checker neurons. Experimental results show that a high coverage of injected errors can be achieved with extremely low computational overhead using consistency properties of the encoded checks. A key side benefit is that the checks can flag errors when the network is presented outlier data that do not correspond to data with which the network is trained to operate.
Sujay Pandey, Suvadeep Banerjee, Abhijit Chatterjee
IOLTS2
2017 Real-time self-learning for control law adaptation in nonlinear systems using encoded check states
abstract
With the wide proliferation of autonomous sense-and-control real-time systems (such as robots and self-driven cars), a key research objective is rapid recovery from the effects of anomalies and impairments arising from performance degradation of sensors and actuators and electro-mechanical subsystems due to field wear and tear. This must be achieved with minimal impact on system performance while maintaining low implementation overhead and high coverage of multi-parameter failure mechanisms. In this work, we propose a reinforcement learning framework for on-line control law adaptation in autonomous nonlinear systems assisted by system state encodings. These encodings are exploited to generate time-varying error signals whose (transient) waveforms in relation to the input stimulus, contain root-cause diagnostic information. This establishes a statistical correlation between the transient waveforms and the parameters of the optimal nonlinear controller under arbitrary multi-parameter perturbations of sensor/actuator and subsystem performances. Consequently this correlation is tapped, using pre-deployment supervised learning algorithms, to predict near-optimal controller parameter values whenever sufficiently large parameter deviations are detected (due to non-zero error signals). From these near-optimal starting conditions, an actor-critic reinforcement learning controller for nonlinear systems quickly converges to the optimal control law for the parameter-perturbed system (up to 10× faster than for systems not assisted by the diagnostic information provided by the state encoding driven error signal above). We implement the proposed methodology on two nonlinear systems demonstrating fast performance recovery in real time.
Suvadeep Banerjee, Abhijit Chatterjee
ETS1
2017 Design of efficient error resilience in signal processing and control systems: From algorithms to circuits
abstract
The proliferation of cyber physical systems in society, from the smart grid to sensor networks and robots has raised the importance of error resilience in signal processing and control systems to unprecedented levels. Resilience to errors in sensing and control algorithm execution in processors all the way down to circuits for sensing and actuation is of critical importance in safety-critical applications where undetected errors can have disastrous consequences. In this presentation, we describe how ideas in the domain of algorithm-based fault tolerance developed in the mid-80s for signal processing and matrix computations can be applied to a vast domain of circuits and systems in electrical engineering; from digital and analog filters to complex nonlinear autonomous control systems. The key insight is that electrical systems can be fundamentally represented by linear and nonlinear differential equations with equivalent matrix representations. These representations can be encoded with extra check states that bear a known relationship with all the observable states of the system independent of the system driving inputs. By checking for the validity of this relationship, errors can be detected and mitigated in real-time with near-zero latency with minimal hardware overhead. The broad vision of the proposed methodology is illustrated with examples from different electrical engineering domains.
Jacob A. Abraham, Suvadeep Banerjee, Abhijit Chatterjee
IOLTS2
2017 Probabilistic error detection and correction in switched capacitor circuits using checksum codes
abstract
In the past, techniques for error detection in linear digital and analog circuits using checksum codes have been developed and shown to be highly efficient. While error detection is a solved problem, error correction has proved to be difficult due to the time and area overheads involved in diagnosing failed system states and correcting them in real-time. To solve the correction problem, real-time probabilistic correction mechanisms have been proposed for digital circuits that correct for state errors in a probabilistic manner, circumventing the process of accurate error diagnosis. Such a technique is difficult to apply to continuous-time analog circuits without altering the analog transfer function, due to the nature of error feedback mechanisms involved. However, switched-capacitor circuits offer intrinsic advantages; they replicate analog continuous-time behavior while retaining the benefits of a digital clock. In this work, we show how errors in switched-capacitor circuits can be detected and corrected, using probabilistic correction algorithms, by taking advantage of the separation in time afforded by the use of a digital clock between error-free and error-affected clock cycles of the circuit. By probabilistically correcting errors in real-time before the onset of future clock cycles, the advantages offered by digital clocks are exploited to deliver high-fidelity analog performance in switched-capacitor filters resulting in significant SNR benefits at low cost.
Md Imran Momtaz, Suvadeep Banerjee, Abhijit Chatterjee
IOLTS2
2017 On-line diagnosis and compensation for parametric failures in linear state variable circuits and systems using time-domain checksum observers
abstract
A large class of real-time circuits and systems can be expressed in linear state variable form. Of particular interest are systems with sensors and actuators that can degrade over time or circuits that can suffer from parametric deviations due to electrical degradation. In the past, multiple checksum codes have been used for error detection and correction in linear systems. In this research, it is shown for the first time that under multi-parameter failures (variations), the transient checksum response (single checksum) to a system contains multi-parameter diagnostic information about the parametric failure. In other words, when a single checksum transient response is time-sampled, the resulting sampled values can be mapped to the critical parameters of the system that affect its performance under multi-parameter variations. The resulting information can be used to rapidly actuate adaption of relevant linear control mechanisms for the system (e.g. PID control) in order to restore system performance with very low correction latency. We make a case for this new “checksum observer” paradigm using motors, generators and a simple biquadratic filter. Preliminary results are presented and demonstrate the viability of the proposed ideas.
Md Imran Momtaz, Suvadeep Banerjee, Abhijit Chatterjee
VTS2
2016 Concurrent error detection and tolerance in Kalman filters using encoded state and statistical covariance checks
abstract
The Kalman filter is a versatile tool used in control and signal processing systems to predict statistically significant data from noisy measurements. In many practical control systems, not all the system states are directly controllable and observable. From noisy measurements of a limited subset of the observable system states, the Kalman filter predicts the mean values and covariances of the complete set of continuously evolving system states using specialized matrix arithmetic. Our goal is to detect errors in any underlying arithmetic computation (e.g. addition/multiplication) involved in the operation of the Kalman filter. While prior linear state checksum methods can be used to detect errors in a subset of the matrix operations of the Kalman filter, they do not suffice for detecting errors in the majority of calculations involved in determining the state covariances. To solve this problem, we develop the notion of statistical state covariance checks. Two applications of a Kalman filter, a trajectory tracking system and a linearized control system for an inverted pendulum are used to demonstrate the proposed approach. A simple state restoration approach is used to compensate for detected errors allowing the complete system to tolerate errors as and when they affect system operation.
Sujay Pandey, Suvadeep Banerjee, Abhijit Chatterjee
IOLTS2
2016 Efficient cross-layer concurrent error detection in nonlinear control systems using mapped predictive check states
abstract
The rapid proliferation of sensor networks and robots in a wide range of societal applications has focused renewed attention on error-free operation of their underlying signal processing and control functions for reasons of safety and reliability. While real-time error detection in linear systems has been investigated in the past, error detection in nonlinear control functions has largely relied on implementing redundancy in components, units, or subsystems resulting in excessive area/performance overheads. In this paper, we introduce a realtime error detection methodology for nonlinear control state space systems that uses mapped predictive check states for detecting sensor and actuator malfunctions and transient errors in the execution of the control algorithm on the underlying processor. In our approach, the check state at time t bears a known relationship with the corresponding states of the nonlinear system. This check state can also be predicted from knowledge of the prior system states and inputs using nonlinear mappings. Consistency between the prior known relationship and its predicted value above, is used to check for errors in system function. We demonstrate the proposed approach on two test cases - a classical nonlinear inverted pendulum balancing problem using a moving cart and a nonlinear sliding mode controller driven electromagnetic brake-by-wire (BBW) system. Simulation results show the effectiveness of the proposed approach for detecting degradation of the sensor and actuator functions and soft errors in the execution of the control algorithms.
Suvadeep Banerjee, Abhijit Chatterjee, Jacob A. Abraham
ITC1
2016 Infant mortality tests for analog and mixed-signal circuits
abstract
A new methodology for extracting compact test stimuli from functional tests for infant mortality testing of analog circuits is proposed. The test stimuli are extracted such that they produce extremal electrical activity in the circuit to push latent defects over the edge to become hard defects. Results on analog modules from the receiver sub-system of a high speed serial interface show that tests that are about 1/5th the size of the original functional test can be extracted while achieving consistently at least 80% of the electrical activity of the functional test.
Suvadeep Banerjee, Suriyaprakash Natarajan
VTS1
2016 Real-time DC motor error detection and control compensation using linear checksums
abstract
The correct operation of transducers such as electric motors, is becoming increasingly important in autonomous systems that depend on the reliability of the underlying electronics to deliver Quality of Service to the end customer. In this paper, a methodology for detecting errors in DC motor operation and adapting its control parameters to compensate for the same using continuous linear checksums, is developed. The approach is different from prior application of checksum codes to analog circuits and servomotor systems, in that none of the states of the DC motor are directly controllable. Accordingly, from measurements of the observable states of the system, the control law for the motor is modified almost instantaneously, to recover motor performance in the most optimal manner possible in the presence of parametric deviations due to wear and tear. We focus on wear out susceptible electromechanical anomalies such as loss of torque due to increased ballbearing friction, etc. Simulation results supported by a hardware prototype support the viability of the proposed error detection and compensation methodology.
Md Imran Momtaz, Suvadeep Banerjee, Abhijit Chatterjee
VTS2
2015 Concurrent error detection in nonlinear digital filters using checksum linearization and residue prediction
abstract
Soft errors due to alpha particles, neutrons and environmental noise are of increasing concern due to aggressive technology scaling. While prior work has focused mostly on error resilience of linear signal processing algorithms, there is increasing need to address the same for nonlinear systems used in emerging applications for sensing and control. In this paper, a new approach for detecting errors in nonlinear digital filters is developed that does not require full duplication of all the nonlinear operations in the filter. First, a checksum of the linear least squares fit to the nonlinear function of the filter is derived that is ideally zero when the filter nonlinearities are not excited. Next, in residue prediction, linear predictive codes are used to predict the nonzero checksum error values that result exclusively from filter nonlinearity excitation. This allows fine granularity soft error detection at low hardware cost. Simulation experiments on a nonlinear Volterra filter prove the viability of the proposed concurrent error detection methodology.
Suvadeep Banerjee, Md Imran Momtaz, Abhijit Chatterjee
IOLTS1
2014 Error Resilient Real-Time State Variable Systems for Signal Processing and Control
abstract
The advent of sensor networks, robots, autonomous vehicles and the smart grid have made the dependability of circuits and systems that control them critical to society and national defense. While significant advances in the design of linear and nonlinear control systems have been made to allow modes of operation not possible in the past, the problem of resilience to errors induced by hostile operating environments remains largely unexplored even though the probability of such errors occurring during real-time operation has increased. In this talk we propose mechanisms for detecting transient errors in control systems and circuitry as well as diagnosing and correcting for their effects on overall system operation. It is shown how real-number checksum encodings of circuit function can be used to detect and correct errors in the plant and feedback subsystems of linear control systems. Applications to signal processing and control algorithms are described. It is shown how errors in motor control electronics can be detected and corrected using the proposed methodology. Finally, extensions to nonlinear control systems are presented.
Suvadeep Banerjee, Álvaro Gómez-Pau, Abhijit Chatterjee, Jacob A. Abraham
ATS1
2014 Design of low cost fault tolerant analog circuits using real-time learned error compensation
abstract
Analog checksum based fault tolerance for linear circuits has been proposed in the past but remains a theoretical artifact due to the high cost and complexity of error compensation while other redundancy based methods have prohibitive overheads. To resolve this, new low cost error compensation methods for widely used linear analog circuits are developed in this research. Trial and error based compensation learning methods combined with the use of less than minimum distance codes are used for failure tolerance. This results in significant hardware savings over prior correction schemes with minimal increase in error correction latency. It is shown how dual failures in analog circuits, not possible with existing techniques, can be compensated using the proposed fault-learning approach.
Suvadeep Banerjee, Álvaro Gómez-Pau, Abhijit Chatterjee
ETS1
2014 Real-time transient error and induced noise cancellation in linear analog filters using learning-assisted adaptive analog checksums
abstract
Analog circuits are sensitive to signal aggressions and power supply noise, crosstalk coupling and alpha particle strikes can cause significant degradation of circuit's SNR. This research proposes a novel approach to real-time transient error and induced noise cancellation in linear analog circuits using analog checksums. It is based on the use of state space representations of analog filters and is a significant advancement over prior research that addressed only hard parametric deviations. A key innovation is the use of less than minimum distance checksum codes for error detection and correction using real-time learning of the likely source of transient errors and noise within the analog circuit. By running a simple hardware-directed search algorithm, the circuit “learns” how best to compensate for the injected signal disturbances with low overhead under the assumption that the source of the injected errors/noise and the error/noise statistics are stationary over time. Successful simulations and preliminary experimental results demonstrate almost complete compensation of injected noise, therefore validating the proposal.
Álvaro Gómez-Pau, Suvadeep Banerjee, Abhijit Chatterjee
IOLTS2
2013 Enhanced Resolution Time-Domain Reflectometry for High Speed Channels: Characterizing Spatial Discontinuities with Non-ideal Stimulus
abstract
In the recent past, there has been steady growth in the data transfer rate of modern digital serial communication systems. Consequently, accurate characterization of high-speed signal transmission lines is necessary for ensuring high signal integrity. Time domain reflectometry (TDR) has been widely used in prior research to characterize high speed interconnect. The accuracy of the characterization depends on the sampling rate and the slew rate of the TDR input excitation signal. At high speeds it is not always possible to deliver "perfect" (impulse/step) TDR stimulus. In this paper, an algorithm is presented to compensate for the inherent distortion in nonideal TDR stimulus to improve the accuracy of interconnect characterization. The algorithm is applied to the problem of micro strip transmission line characterization for identifying discontinuities in signal interconnect. Hardware measurements validate the effectiveness of the proposed technique.
Suvadeep Banerjee, Hyun Woo Choi, David C. Keezer, Abhijit Chatterjee
Asian Test Symposium1
2013 Real-time checking of linear control systems using analog checksums
abstract
In the recent past, there has been a proliferation of complex control problems in sensor network design, multi-agent systems such as autonomous vehicles and robotics, to name a few. While prior research has focused on the design of optimal controllers for real-time systems, in the future it will become increasingly difficult to perform periodic maintenance of such systems due to their mobile and autonomous nature. Moreover, in safety-critical real-time applications it will become increasingly necessary to perform real-time monitoring of the plant as well as its controller functions for reasons of reliability and safety. In this paper, we develop, for the first time, a theory for implementing low-overhead and high coverage detection of transient errors and permanent faults in linear control systems consisting of the plant and its controller using analog checksums. The approach is demonstrated on a servo-motor control problem. It is shown that small parametric perturbations as well as transient errors are detected in real-time using the proposed checking methodology.
Suvadeep Banerjee, Aritra Banerjee, Abhijit Chatterjee, Jacob A. Abraham
IOLTS1