VLDB 2026 Research / reviewers in the wild / expert
Amit Ranjan Trivedi
dblp:116/7561
· DBLP profile ↗
44ranked-venue papers
7as first author
31since 2021 · last 2026
0000-0001-5436-7922ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 7 first-author · 22 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EigenShield: Inference-Time, Model-Agnostic Jailbreaking Defense via Causal Subspace FilteringabstractLarge Language Models (LLMs) and Vision-Language Models (VLMs) remain highly vulnerable to adversarial attacks despite widespread adoption. Existing defenses typically require retraining, rely on heuristics, or fail under adaptive and out-of-distribution (OOD) conditions. We introduce EigenShield, a principled, inference-time, architecture-agnostic defense that leverages Random Matrix Theory (RMT) to suppress adversarial noise in high-dimensional embeddings. EigenShield uses spiked covariance modeling and a Robustness-based Nonconformity Score (RbNS) with quantile thresholding to isolate and preserve causal eigenvectors, filtering out adversarial components without model access or adversarial training. We develop a theoretical framework establishing conditions for asymptotic noise suppression and demonstrate effectiveness in both unimodal and multimodal settings. Empirically, EigenShield consistently improves robustness across threat models, reducing attack success rates (ASR) by up to 48% over state-of-the-art defenses, including adversarial training, UNIGUARD, CIDER, and input transformations. On jailbreak attacks, EigenShield lowers LLM ASR by up to 92.9% relative to undefended models. Under multimodal adversarial attacks, it reduces VLM ASR by up to 76.5%. Against adaptive attacks on LLMs, it achieves ASR reductions of up to 77.7%. In OOD settings, EigenShield maintains strong performance, reducing ASR by up to 88.4% for LLMs and 80.4% for VLMs. Nastaran Darabi, Devashri Naik, Sina Tayebati, Dinithi Jayasuriya, Ranganath Krishnan, Amit Ranjan Trivedi |
AAAI | 6 |
| 2026 | Resilience in Ambient Multi-Agent LLMs via Decentralized Bio-Autonomic Control and Immune-Inspired Anomaly DetectionabstractLarge Language Model (LLM) agents are now widely deployed in Ambient Intelligence (AmI) environments, where autonomous agents must sense, act, and coordinate at scale. As agent capabilities and interdependence increase, traditional reliability strategies such as isolated adaptive control, anomaly detection, or trust modeling have proven inadequate due to their fragmented and scenario-specific nature. Comprehensive architectures that enable integrated self-management, collective anomaly response, robust information dissemination, and privacy-preserving adaptation remain scarce. We propose a bio-autonomic framework for decentralized resilience in multi-agent LLM systems where a unified architecture systematically applies principles from biological autonomic systems to LLM-based multi-agent environments. Specifically, each agent implements an autonomic control loop, formally structured as Monitor-Analyze-Plan-Execute over a shared Knowledge base (MAPE-K), for self-regulation. At the system level, the framework integrates immune-inspired anomaly detection using the Dendritic Cell Algorithm, probabilistic computational trust, decentralized gossip for robust information sharing, and federated learning with homomorphic encryption for collaborative, privacy-preserving adaptation. This holistic approach enables LLM agent ecosystems to self-organize, detect and isolate faults, and collectively adapt as system complexity increases. Empirical evaluations show that our framework achieves substantially improved resilience and recovery compared to state-of-the-art multi-agent baselines. Nastaran Darabi, Devashri Naik, Sina Tayebati, Dinithi Jayasuriya, Amit Ranjan Trivedi |
AAAI | 5 |
| 2026 | CaDRO: Causal-Guided Dimensionality Reduction for Scalable Multi-Objective Pareto OptimizationabstractMulti-objective optimization of analog circuits is hindered by high-dimensional parameter spaces, strong feedback couplings, and expensive transistor-level simulations. Evolutionary algorithms such as Non-dominated Sorting Genetic Algorithm II (NSGA-II) are widely used but treat all parameters equally, thereby wasting effort on variables with little impact on performance, which limits their scalability. We introduce CaDRO, a causal-guided dimensionality reduction framework that embeds causal discovery into the optimization pipeline. CaDRO builds a quantitative causal map through a hybrid observational-interventional process, ranking parameters by their causal effect on the objectives. Low-impact parameters are fixed to values from high-quality solutions, while critical drivers remain active in the search. The reduced design space enables focused evolutionary optimization without modifying the underlying algorithm. Across amplifiers, regulators, and RF circuits, CaDRO converges up to 10× faster than NSGA-II while preserving or improving Pareto quality. For instance, on the Folded-Cascode Amplifier, hypervolume improves from 0.56 to 0.94, and on the LDO regulator from 0.65 to 0.81, with large gains in non-dominated solutions. Dinithi Jayasuriya, Divake Kumar, Sureshkumar Senthilkumar, Devashri Naik, Nastaran Darabi, Amit Ranjan Trivedi |
DATE | 6 |
| 2025 | Intelligent Sensing-to-Action for Robust Autonomy at the Edge: Opportunities and ChallengesabstractAutonomous edge computing in robotics, smart cities, and autonomous vehicles relies on the seamless integration of sensing, processing, and actuation for real-time decision-making in dynamic environments. At its core is the sensing-to-action loop, which iteratively aligns sensor inputs with computational models to drive adaptive control strategies. These loops can adapt to hyper-local conditions, enhancing resource efficiency and responsiveness, but also face challenges such as resource constraints, synchronization delays in multimodal data fusion, and the risk of cascading errors in feedback loops. This article explores how proactive, context-aware sensing-to-action and action-to-sensing adaptations can enhance efficiency by dynamically adjusting sensing and computation based on task demands, such as sensing a very limited part of the environment and predicting the rest. By guiding sensing through control actions, action-to-sensing pathways can improve task relevance and resource use, but they also require robust monitoring to prevent cascading errors and maintain reliability. Multi-agent sensing-action loops further extend these capabilities through coordinated sensing and actions across distributed agents, optimizing resource use via collaboration. Additionally, neuromorphic computing, inspired by biological systems, provides an efficient framework for spike-based, event-driven processing that conserves energy, reduces latency, and supports hierarchical control-making it ideal for multi-agent optimization. This article highlights the importance of end-to-end co-design strategies that align algorithmic models with hardware and environmental dynamics, improve cross-layer inter-dependencies to improve throughput, precision, and adaptability for energy-efficient edge autonomy in complex environments. Amit Ranjan Trivedi, Sina Tayebati, Hemant Kumawat, Nastaran Darabi, Divake Kumar, Adarsh Kosta, Yeshwanth Venkatesha, Dinithi Jayasuriya, Nethmi Jayasinghe, Priyadarshini Panda, Saibal Mukhopadhyay, Kaushik Roy 0001 |
DATE | 1 |
| 2025 | Generative Sensing: Pre-training LiDAR with Masked Autoencoders for Ultra-Frugal PerceptionabstractWe propose a disruptively frugal generative sensing approach for LiDAR that generates, rather than senses, parts of the environment that are either predictable based on extensive training or have limited impact on overall prediction accuracy. Our generative pre-training strategy for this purpose, radially masked autoencoding (R-MAE), also allows focusing on radial segments of the data, which captures spatial relationships and distances between objects more effectively than conventional procedures. As a result, the proposed methodology not only reduces sensing energy but also improves prediction accuracy. Our evaluations on Waymo and nuScenes show that our approach achieves over a 5% average precision improvement in detection tasks across datasets and over a 4% accuracy improvement when transferring domains from Waymo and nuScenes to KITTI. Our method achieves up to 3.17% and 2.31% improvements in mean average precision (mAP) and NDS, respectively, in nuScenes. Even with 90% radial masking, it surpasses baseline models by up to 5.59% in mAP and mean average precision with heading (mAPH) across all object classes in the Waymo dataset. Resultantly, R-MAE reduces the sensing energy by ∼9.1× than the conventional LiDAR processing despite computational overheads. Codes are available: https://github.com/sinatayebati/Radial_MAE. Sina Tayebati, Theja Tulabandhula, Amit Ranjan Trivedi |
ICASSP | 3 |
| 2025 | Hercules: Hardware accElerator foR stoChastic schedULing in hEterogeneous SystemsabstractEfficient workload scheduling is a critical challenge in modern heterogeneous computing environments, particularly in high-performance computing (HPC) systems. Traditional software-based schedulers struggle to efficiently balance workload distribution due to high scheduling overhead, lack of adaptability to dynamic workloads, and suboptimal resource utilization. These pitfalls are compounded in heterogeneous systems, where differing computational elements can have vastly different performance profiles. To resolve these hindrances, we present a novel FPGA-based accelerator for stochastic online scheduling (SOS). We modify a greedy cost selection assignment policy by adapting existing cost equations to engage with discretized time before implementing them into a hardware accelerator design. Our design leverages hardware parallelism, precalculation, and precision quantization to reduce job scheduling latency. By introducing a hardware-accelerated approach to real-time scheduling, this paper establishes a new paradigm for adaptive scheduling mechanisms in heterogeneous computing systems. The proposed design achieves high throughput, low latency, and energy-efficient operation, offering a scalable alternative to traditional software scheduling methods. Experimental results demonstrate consistent workload distribution, fair machine utilization, and up to 1060× speedup over single-threaded software scheduling policy implementations. This makes the SOS accelerator a strong candidate for deployment in high-performance computing system, deep-learning pipelines, and other performance-critical applications. Vairavan Palaniappan, Adam H. Ross, Amit Ranjan Trivedi, Debjit Pal |
ICCAD | 3 |
| 2025 | Enhancing 3D Robotic Vision Robustness by Minimizing Adversarial Mutual Information through Curriculum TrainingabstractAdversarial attacks exploit vulnerabilities in a model's decision boundaries through small, carefully crafted perturbations that lead to significant mispredictions. In 3D vision, the high dimensionality and sparsity of data greatly expand the attack surface, making 3D vision particularly vulnerable for safety-critical robotics. To enhance 3D vision's adversarial robustness, we propose a training objective that simultaneously minimizes prediction loss and mutual information (MI) under adversarial perturbations to contain the upper bound of misprediction errors. This approach simplifies handling adversarial examples compared to conventional methods, which require explicit searching and training on adversarial samples. However, minimizing prediction loss conflicts with minimizing MI, leading to reduced robustness and catastrophic forgetting. To address this, we integrate curriculum advisors in the training setup that gradually introduce adversarial objectives to balance training and prevent models from being overwhelmed by difficult cases early in the process. The advisors also enhance robustness by encouraging training on diverse MI examples through entropy regularizers. We evaluated our method on ModelNet40 and KITTI using PointNet, DGCNN, SECOND, and PointTransformers, achieving 2-5% accuracy gains on ModelNet40 and a 5-10% mAP improvement in object detection. Our code is publicly available at https://github.com/nstrndrbi/Mine-N-Learn. Nastaran Darabi, Dinithi Jayasuriya, Devashri Naik, Theja Tulabandhula, Amit Ranjan Trivedi |
ICRA | 5 |
| 2025 | Uncertainty-Aware Deep Reinforcement Learning with Calibrated Quantile Regression and Evidential LearningabstractWe present a novel statistical approach to incorporate uncertainty awareness in model-free distributional deep reinforcement learning for mission and safety-critical robotics. Deep learning predictions are influenced by uncertainties in the data, termed as aleatoric uncertainties, as well as uncertainties in the learning process and model structure, known as epistemic uncertainties. The proposed algorithm, called as Calibrated Evidential Quantile Regression in Deep-Q Networks (CEQR-DQN), addresses key challenges associated with separately estimating aleatoric and epistemic uncertainty in stochastic robotic environments. It combines deep evidential learning with quantile calibration based on the principles of conformal inference to provide explicit, sample-free computations of global uncertainty as opposed to local estimates based on simple variance. Thereby, the proposed approach overcomes limitations of traditional methods in computational and statistical efficiency and handling of out-of-distribution (OOD) observations. Tested on a suite of representative miniaturized Atari games (i.e., MinAtar), CEQR-DQN is shown to surpass similar existing frameworks in scores and learning speed. Its ability to rigorously evaluate uncertainties improves exploration strategies and can serve as a blueprint for other uncertainty-aware robotic algorithms. Alex C. Stutts, Danilo Erricolo, Theja Tulabandhula, Mohit Mittal, Amit Ranjan Trivedi |
ICRA | 5 |
| 2025 | SPARC: Subspace-Aware Prompt Adaptation for Robust Continual Learning in LLMsabstractWe propose SPARC, a lightweight continual learning framework for large language models (LLMs) that enables efficient task adaptation through prompt tuning in a lower-dimensional space. By leveraging principal component analysis (PCA), we identify a compact subspace of the training data. Optimizing prompts in this lower-dimensional space enhances training efficiency, as it focuses updates on the most relevant features while reducing computational overhead. Furthermore, since the model’s internal structure remains unaltered, the extensive knowledge gained from pretraining is fully preserved, ensuring that previously learned information is not compromised during adaptation. Our method achieves high knowledge retention in both task-incremental and domain-incremental continual learning setups while fine-tuning only 0.04% of the model’s parameters. Additionally, by integrating LoRA, we enhance adaptability to computational constraints, allowing for a tradeoff between accuracy and training cost. Experiments on the SuperGLUE benchmark demonstrate that our PCA-based prompt tuning combined with LoRA maintains full knowledge retention while improving accuracy, utilizing only 1% of the model’s parameters. These results establish our approach as a scalable and resource-efficient solution for continual learning in LLMs. Dinithi Jayasuriya, Sina Tayebati, Davide Ettori, Ranganath Krishnan, Amit Ranjan Trivedi |
IJCNN | 5 |
| 2025 | Neural Precision Polarization: Simplifying Neural Network Inference with Dual-Level PrecisionabstractWe introduce a precision polarization scheme for DNN inference that utilizes only very low and very high precision levels, assigning low precision to the majority of network weights and activations while reserving high precision paths for targeted error compensation. This separation allows for distinct optimization of each precision level, thereby reducing memory and computation demands without compromising model accuracy. In the discussed approach, a floating-point model can be trained in the cloud and then downloaded to an edge device, where network weights and activations are directly quantized to meet the edge devices’ desired level (such as NF4/INT8). To address accuracy loss from quantization, surrogate paths are introduced, leveraging low-rank approximations on a layer-by-layer basis. These paths are trained with a sensitivity-based metric on minimal training data to recover accuracy loss under quantization as well as due to process variability such as when the main prediction path is implemented using analog acceleration. Our simulation results show that neural precision polarization enables ~464 TOPS/W MAC efficiency and reliability by integrating rank-8 error recovery paths with highly efficient, though potentially unreliable, bitplane-wise compute-in-memory processing. Dinithi Jayasuriya, Nastaran Darabi, Maeesha Binte Hashem, Amit Ranjan Trivedi |
ISCAS | 4 |
| 2025 | Tutorial: Autonomy with Neuromorphic SystemabstractThis tutorial will discuss how to incorporate neuromorphic circuits, computing, and sensing into an autonomous system at the edge, and present an end-to-end system analysis showing their integration. Amit Ranjan Trivedi, Priyadarshini Panda, Kaushik Roy 0001, Saibal Mukhopadhyay |
ISLPED | 1 |
| 2025 | From Signals to Features to Insights: Multi-Level Novelty Detection for Fast Scientific DiscoveryabstractMost scientific discoveries depend on identifying novel signals hidden in massive, noisy datasets generated by modern experiments. Traditional novelty detection methods are often insufficient in speed, robustness, and adaptability to resource-constrained environments. We discuss a perspective on a hierarchical framework for multi-level novelty detection spanning sensor signals, feature representations, and model outputs. At the signal level, we discuss analog circuits that extract statistical densities and moments in real-time, enabling interpretable and energy-efficient filtering. At the feature level, we introduce Likelihood Regret, an unsupervised measure that detects anomalies by retraining generative models on shared latent representations, with optimizations for embedded deployment. At the output level, we leverage predictive uncertainty, applying both compute-efficient Monte Carlo reuse and Monte Carlo-free techniques like evidential learning and conformal inference. Our framework demonstrates how integrating novelty detection across the sensing-to-inference can accelerate insights in domains such as high-energy physics. Devashri Naik, Nastaran Darabi, Sina Tayebati, Dinithi Jayasuriya, Shamma Nasrin, Danush Shekar, Corrinne Mills, Benjamin Parpillon, Farah Fahim, Mark S. Neubauer, Amit Ranjan Trivedi |
VTS | 11 |
| 2025 | MOSAIC: Collaborative Compute-in-Memory µArrays for Flexible and Scalable Deep LearningabstractCompute-in-Memory (CiM) architectures, particularly those leveraging SRAM-based arrays, present significant opportunities for accelerating deep learning by mitigating data movement bottlenecks in traditional von Neumann systems. SRAM-based CiM offers notable advantages, including speed, seamless CMOS integration, and compatibility with existing System-on-Chip (SoC) designs. However, challenges persist, primarily stemming from analog-domain computations that necessitate analog-to-digital (A/D) and digital-to-analog (D/A) conversions, leading to reduced accuracy, increased power overhead, and rigid operational constraints. To overcome these limitations, we propose MOSAIC, a novel CiM architecture designed around three foundational principles: (1) co-designing deep learning operators for CiM such as employing multiplication-free computations and frequency-domain processing to minimize/eliminate DAC/ADC overheads; (2) leveraging a memory-immersed digitization approach that utilizes parasitic bit-lines as capacitive DACs within CiM arrays, thereby significantly reducing peripheral complexity and enhancing scalability; and (3) orchestrating inference over a network of compact CiM µArrays by dynamically interconnecting them to provide flexibility, efficiency, and minimized computational overhead for varied inference workload characteristics. Collectively with these fundamental innovations, MOSAIC addresses critical bottlenecks in accuracy, scalability, and flexibility, unlocking CiM’s full potential for efficient deep learning in embedded systems. Amit Ranjan Trivedi, Shamma Nasrin, Priyesh Shukla, Nastaran Darabi, Divake Kumar, Dinithi Jayasuriya, Nethmi Jayasinghe |
VTS | 1 |
| 2024 | Invited: Conformal Inference meets Evidential Learning: Distribution-Free Uncertainty Quantification with Epistemic and Aleatoric SeparabilityabstractThis paper introduces a lightweight framework for quantifying uncertainty in deep learning models deployed at the edge, addressing the challenge of making reliable predictions under computational constraints. By integrating conformal inference and evidential learning into a novel approach called Conformalized Evidential Quantile Regression (CEQR), it offers a practical solution for models to assess and communicate their confidence in predictions. The method efficiently distinguishes between aleatoric and epistemic uncertainties, ensuring statistical robustness and real-time applicability on resource-limited devices. This work paves the way for safer, more reliable AI applications in critical areas by enabling models to recognize when they don't know. Alex C. Stutts, Divake Kumar, Theja Tulabandhula, Amit Ranjan Trivedi |
DAC | 4 |
| 2024 | Navigating the Unknown: Uncertainty-Aware Compute-in-Memory Autonomy of Edge RoboticsabstractThis paper addresses the challenging problem of energy-efficient and uncertainty-aware pose estimation in insect-scale drones, which is crucial for tasks such as surveillance in constricted spaces and for enabling non-intrusive spatial intelligence in smart homes. Since tiny drones operate in highly dynamic environments, where factors like lighting and human movement impact their predictive accuracy, it is crucial to deploy uncertainty-aware prediction algorithms that can account for environmental variations and express not only the prediction but also confidence in the prediction. We address both of these challenges with Compute-in-Memory (CIM) which has become a pivotal technology for deep learning acceleration at the edge. While traditional CIM techniques are promising for energy-efficient deep learning, to bring in the robustness of uncertainty-aware predictions at the edge, we introduce a suite of novel techniques: First, we discuss CIM-based acceleration of Bayesian filtering methods uniquely by leveraging the Gaussian-like switching current of CMOS inverters along with co-design of kernel functions to operate with extreme parallelism and with extreme energy efficiency. Secondly, we discuss the CIM-based acceleration of variational inference of deep learning models through probabilistic processing while unfolding iterative computations of the method with a compute reuse strategy to significantly minimize the workload. Overall, our co-design methodologies demonstrate the potential of CIM to improve the processing efficiency of uncertainty-aware algorithms by orders of magnitude, thereby enabling edge robotics to access the robustness of sophisticated prediction frameworks within their extremely stringent area/power resources. Nastaran Darabi, Priyesh Shukla, Dinithi Jayasuriya, Divake Kumar, Alex C. Stutts, Amit Ranjan Trivedi |
DATE | 6 |
| 2024 | Conformalized Multimodal Uncertainty Regression and ReasoningabstractThis paper introduces a lightweight uncertainty estimator capable of predicting multimodal (disjoint) uncertainty bounds by integrating conformal prediction with a deep-learning regressor. We specifically discuss its application for visual odometry (VO), where environmental features such as flying domain symmetries and sensor measurements under ambiguities and occlusion can result in multimodal uncertainties. Our simulation results show that uncertainty estimates in our framework adapt sample-wise against challenging operating conditions such as pronounced noise, limited training data, and limited parametric size of the prediction model. We also develop a reasoning framework that leverages these robust uncertainty estimates and incorporates optical flow-based reasoning to improve prediction prediction accuracy. Thus, by appropriately accounting for predictive uncertainties of data-driven learning and closing their estimation loop via rule-based reasoning, our methodology consistently surpasses conventional deep learning approaches on all these challenging scenarios–pronounced noise, limited training data, and limited model size–reducing the prediction error by 2–3×. Mimmo Parente, Nastaran Darabi, Alex C. Stutts, Theja Tulabandhula, Amit Ranjan Trivedi |
ICASSP | 5 |
| 2024 | Mutual Information-calibrated Conformal Feature Fusion for Uncertainty-Aware Multimodal 3D Object Detection at the EdgeabstractIn the expanding landscape of AI-enabled robotics, robust quantification of predictive uncertainties is of great importance. Three-dimensional (3D) object detection, a critical robotics operation, has seen significant advancements; however, the majority of current works focus only on accuracy and ignore uncertainty quantification. Addressing this gap, our novel study integrates the principles of conformal inference (CI) with information theoretic measures to perform lightweight, Monte Carlo-free uncertainty estimation within a multimodal framework. Through a multivariate Gaussian product of the latent variables in a Variational Autoencoder (VAE), features from RGB camera and LiDAR sensor data are fused to improve the prediction accuracy. Normalized mutual information (NMI) is leveraged as a modulator for calibrating uncertainty bounds derived from CI based on a weighted loss function. Our simulation results show an inverse correlation between inherent predictive uncertainty and NMI throughout the model’s training. The framework demonstrates comparable or better performance in KITTI 3D object detection benchmarks to similar methods that are not uncertainty-aware, making it suitable for real-time edge robotics. Alex C. Stutts, Danilo Erricolo, Sathya Ravi, Theja Tulabandhula, Amit Ranjan Trivedi |
ICRA | 5 |
| 2024 | STARNet: Sensor Trustworthiness and Anomaly Recognition via Lightweight Likelihood Regret for Robust Edge AutonomyabstractComplex sensors such as LiDAR, RADAR, and event cameras have proliferated in autonomous robotics to enhance perception and understanding of the environment. Meanwhile, these sensors are also vulnerable to diverse failure mechanisms that can intricately interact with their operational environment. In parallel, the limited availability of training data on complex sensors affects the reliability of their deep learning-based prediction flow, when their prediction models fail to generalize to environments not adequately captured in the training set. To address these reliability concerns, this paper introduces STARNet, a Sensor Trustworthiness and Anomaly Recognition Network designed to detect untrustworthy sensor streams that may arise from sensor malfunctions and/or challenging environments. STARNet employs the concept of likelihood regret (LR) for continuous evaluation of the trustworthiness of sensor streams. We tailor the framework to resource-constrained edge devices with two settings: a gradient-free framework suit- able for low-complexity hardware with fixed-point precision capabilities, and a low-rank tunability of underlying models for LR extraction, reducing the extraction workload. Through extensive simulations, we demonstrate the efficacy of STARNet in detecting untrustworthy sensor streams in unimodal and multimodal settings. In particular, the network shows superior performance in addressing internal sensor failures, such as cross-sensor interference and crosstalk. In diverse test scenarios involving adverse weather and sensor malfunctions, we show that STARNet enhances prediction accuracy by approximately 15% by filtering out untrustworthy sensor streams. STARNet is publicly available at https://github.com/nstrndrbi/STARNet. Nastaran Darabi, Sina Tayebati, Sureshkumar Senthilkumar, Dinithi Jayasuriya, Sathya Ravi, Theja Tulabandhula, Amit Ranjan Trivedi |
IJCNN | 7 |
| 2024 | ADC/DAC-Free Analog Acceleration of Deep Neural Networks With Frequency TransformationabstractThe edge processing of deep neural networks (DNNs) is becoming increasingly important due to its ability to extract valuable information directly at the data source to minimize latency and energy consumption. Although pruning techniques are commonly used to reduce model size for edge computing, they have certain limitations. Frequency-domain model compression, such as with the Walsh–Hadamard transform (WHT), has been identified as an efficient alternative. However, the benefits of frequency-domain processing are often offset by the increased multiply-accumulate (MAC) operations required. This article proposes a novel approach to an energy-efficient acceleration of frequency-domain neural networks by utilizing analog-domain frequency-based tensor transformations. Our approach offers unique opportunities to enhance computational efficiency, resulting in several high-level advantages, including array microarchitecture with parallelism, analog-to-digital converter (ADC)/digital-to-analog converter (DAC)-free analog computations, and increased output sparsity. Our approach achieves more compact cells by eliminating the need for trainable parameters in the transformation matrix. Moreover, our novel array microarchitecture enablesadaptive stitchingof cells column-wise and row-wise, thereby facilitating perfect parallelism in computations. Additionally, our scheme enables ADC/DAC-free computations by training against highly quantized matrix-vector products, leveraging the parameter-free nature of matrix multiplications. Another crucial aspect of our design is its ability to handle signed-bit processing for frequency-based transformations. This leads to increased output sparsity and reduced digitization workload. On a$16 \ttimes 16$crossbars, for 8-bit input processing, the proposed approach achieves the energy efficiency of 801 tera operations per second per Watt (TOPS/W) without early termination strategy and 2655 TOPS/W with early termination strategy at VDD$=$0.85 V for 16-nm predictive technology models (PTM). Nastaran Darabi, Maeesha Binte Hashem, Hongyi Pan, A. Enis Çetin, Wilfred Gomes, Amit Ranjan Trivedi |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2023 | Robust Monocular Localization of Drones by Adapting Domain Maps to Depth Prediction InaccuraciesabstractWe present a novel monocular localization framework by jointly training deep learning-based depth prediction and Bayesian filtering-based pose reasoning. The proposed cross-modal framework significantly outperforms deep learning-only predictions with respect to model scalability and tolerance to environmental variations. Specifically, we show little-to-no degradation of pose accuracy even with extremely poor depth estimates from a lightweight depth predictor. Our framework also maintains high pose accuracy in extreme lighting variations compared to standard deep learning, even without explicit domain adaptation. By openly representing the map and intermediate feature maps (such as depth estimates), our framework also allows for faster updates and reusing intermediate predictions for other tasks, such as obstacle avoidance, resulting in much higher resource efficiency. Priyesh Shukla, Sureshkumar S., Alex C. Stutts, Sathya Ravi, Theja Tulabandhula, Amit Ranjan Trivedi |
ICASSP | 6 |
| 2023 | Energy Efficient Memory-based Inference of LSTM by Exploiting FPGA OverlayabstractThe fourth industrial revolution (a.k.a. Industry 4.0) relies on intelligent machines that are fully autonomous and can diagnose and resolve operational issues without human intervention. Therefore, embedded computing platforms enabling the necessary computations for intelligent machines are critical for the ongoing industrial revolution. Especially field programmable gate arrays (FPGAs) are highly suited for such embedded computing due to their high performance and easy reconfigurability. Many Industry 4.0 applications, such as predictive maintenance, critically depend on real-time and reliable processing of time-series data using recurrent neural network models, especially long short-term memory (LSTM). Therefore, the FPGA-based acceleration of LSTM is imperative for many Industry 4.0 applications. Existing LSTM models for FPGAs incur significant resources and power and are not energy efficient. Moreover, prior works focusing on reducing latency and power mainly adhere to model pruning, which compromises the accuracy. Comparatively, we propose a memory-based energy-efficient inference of LSTM by exploiting overlay in FPGA. In our methodology, we pre-compute predominant operations and store them in the available embedded memory blocks (EMBs) of an FPGA. On-demand, these pre-computed results are accessed to minimize the necessary workload. Via this methodology, we obtained lower latency, lower power, and better energy efficiency than state-of-the-art LSTM models without any loss of accuracy. Specifically, when implemented on the ZynQ XCU104 evaluation board, a 3 x reduction in latency and 5 x reduction in power is obtained then the reference 16-bit LSTM model. Krishnendu Guha, Amit Ranjan Trivedi, Swarup Bhunia |
IJCNN | 2 |
| 2023 | Lightweight, Uncertainty-Aware Conformalized Visual OdometryabstractData-driven visual odometry (VO) is a critical subroutine for autonomous edge robotics, and recent progress in the field has produced highly accurate point predictions in complex environments. However, emerging autonomous edge robotics devices like insect-scale drones and surgical robots lack a computationally efficient framework to estimate VO's predictive uncertainties. Meanwhile, as edge robotics continue to proliferate into mission-critical application spaces, awareness of the model's predictive uncertainties has become crucial for risk-aware decision-making. This paper addresses this challenge by presenting a novel, lightweight, and statistically robust framework that leverages conformal inference (CI) to extract VO's uncertainty bands. Our approach represents the uncertainties using flexible, adaptable, and adjustable prediction intervals that, on average, guarantee the inclusion of the ground truth across all degrees of freedom (DOF) of pose estimation. We discuss the architectures of generative deep neural networks for estimating multivariate uncertainty bands along with point (mean) prediction. We also present techniques to improve the uncertainty estimation accuracy, such as leveraging Monte Carlo dropout (MC-dropout) for data augmentation. Finally, we propose a novel training loss function that combines interval scoring and calibration loss with traditional training metrics-mean-squared error and KL-divergence-to improve uncertainty-aware learning. Our simulation results demonstrate that the presented framework consistently captures true uncertainty in pose estimations across different datasets, estimation models, and applied noise types, indicating its wide applicability. Alex C. Stutts, Danilo Erricolo, Theja Tulabandhula, Amit Ranjan Trivedi |
IROS | 4 |
| 2023 | Readout IC with 40 MSPS in-pixel ADC for future vertex detector upgrades of Large Hadron ColliderabstractWe present a prototype smart pixel concept test chip, designed in CMOS 28 nm bulk process, showing a proof-of-concept readout integrated circuit (ROIC) designed for a future Phase III high luminosity upgrade of the large Hadron collider. The presented design employs a synchronous analog-to-digital converter (ADC) for the frontend design, with signal processing and data conversion within a single bunch crossing of 25 ns. It is, therefore, capable of accurately detecting hits occurring in consecutive bunch crossings without off-time registration of events and pileup insensitivity, making it particularly suitable for the innermost layers of the vertex detector. The ROIC consists of a matrix of$32\times 16$pixels, each$25\times 25\ \mu \mathrm{m}^{2}$in size. Each pixel contains a charge-sensitive preamplifier with leakage current compensation and three auto-zero comparators for a 2-bit flash-type ADC. The total power consumption is$\sim 4\ \mu\mathrm{W}$per pixel. The measured noise at the output of all the hit comparators across the ROIC is$< 45\ \mathrm{e}_{\text{RMS}}^{-}$, which allows an in-time threshold setting of ≈ 475 e−. Benjamin Parpillon, Amit Ranjan Trivedi, Farah Fahim |
ISCAS | 2 |
| 2023 | MC-CIM: Compute-in-Memory With Monte-Carlo Dropouts for Bayesian Edge IntelligenceabstractWe propose MC-CIM, a compute-in-memory (CIM) framework for robust, yet low power, Bayesian edge intelligence. Deep neural networks (DNN) with deterministic weights cannot express their prediction uncertainties, thereby pose critical risks for applications where the consequences of mispredictions are fatal such as surgical robotics. To address this limitation, Bayesian inference of a DNN has gained attention. Using Bayesian inference, not only the prediction itself, but the prediction confidence can also be extracted for planning risk-aware actions. However, Bayesian inference of a DNN is computationally expensive, ill-suited for real-time and/or edge deployment. An approximation to Bayesian DNN using Monte Carlo Dropout (MC-Dropout) has shown high robustness along with low computational complexity. Enhancing the computational efficiency of the method, we discuss a novel CIM module that can perform in-memory probabilistic dropout in addition to in-memory weight-input scalar product to support the method. We also propose a compute-reuse reformulation of MC-Dropout where each successive instance can utilize the product-sum computations from the previous iteration. Even more, we discuss how the random instances can be optimally ordered to minimize the overall MC-Dropout workload by exploiting combinatorial optimization methods. Application of the proposed CIM-based MC-Dropout execution is discussed for MNIST character recognition and visual odometry (VO) of autonomous drones. The framework reliably gives prediction confidence amidst non-idealities imposed by MC-CIM to a good extent. Proposed MC-CIM with$16\times 31$SRAM array, 0.85 V supply, 16nm low-standby power (LSTP) technology consumes 32 pJ for 30 MC-Dropout instances of probabilistic inference in its most optimal computing and peripheral configuration, saving$\sim 34$% energy compared to typical execution. Priyesh Shukla, Shamma Nasrin, Nastaran Darabi, Wilfred Gomes, Amit Ranjan Trivedi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | Non van-Neumann Anomaly Detection in Multi-Channel Time-Series using Charge Trap Transistor CrossbarsabstractMany of the existing techniques to detect anomalies in multi-channel time-series lack flexibility and incur significant processing-overheads; therefore, real-time, flexible anomaly detection in resource-constrained edge devices is still an open problem. Addressing the above challenges, we present an ultra-low-power non von-Neumann framework for statistical modeling based anomaly detection in multi-channel time-series. To disruptively minimize the power dissipation for anomaly detection and enhance scalability to complex time-series statistics, we pursue a co-designed approach where the anomaly statistics are modelled using an unconventional Harmonic-Mean of Gaussian-like (HMG) functions. We show that multivariate HMG functions can be implemented simply by exploiting the short-circuit current of multi-input inverters. A non-von Neumann crossbar of charge trap inverters implements a mixture of HMG function and stores model weights. On Yahoo real time-series dataset, our anomaly detection approach achieves an f1-score higher than 0.85 even in the presence of significant process variation and consumes 181fJ/sensor sample for a three-channel time-series. Compared to baseline digital and analog approaches for anomaly detection, our framework is $\sim 40 \times$ and $\sim 6 \times$ more energy efficient, respectively, while being more scalable to time-series dimension. Ahish Shylendra, Priyesh Shukla, Amit Ranjan Trivedi |
ISCAS | 3 |
| 2022 | RIHANN: Remote IoT Hardware Authentication With Intrinsic IdentifiersabstractThe heterogeneous array of edge devices in an Internet of Things (IoT) infrastructure is increasingly vulnerable to physical in-field tampering attacks. These devices can significantly benefit from a difficult-to-clone and tamper-immune intrinsic identifier that can verify the authenticity or integrity of the physical components. In this article, we develop an intrinsic device identifier,RIHANN, that captures the state of the electronic hardware in an IoT device. This state can adequately reflect any physical tampering of the hardware components by transforming the intrinsic delay variations in the electronic components of an edge device into unique and robust signatures. Our proposed authentication approach utilizes the boundary scan architecture (BSA) in printed circuit boards (PCBs). BSA is a prevalent design for test (DFT) structure used in most PCBs in IoT edge devices. This technique supports an extensive array of heterogeneous devices and can seamlessly operate during the device’s runtime. We measure the boundary scan path delays using the parallel scan delay-measurement (PSDM) technique for commercially available ICs. We perform practical experiments on 20 devices, generate signatures, and evaluate their uniqueness, robustness, randomness, and resistance to aging. We also introduce a security protocol for the cloud server, owner/verifier, or other IoT devices connected to a network to verify their identity remotely. The policy prevents attacks from extracting the device’s secret keys using an efficient moving target defense mechanism that periodically updates and evolves the challenge–response database. Shubhra Deb Paul, Fengchao Zhang, Patanjali SLPSK, Amit Ranjan Trivedi, Swarup Bhunia |
IEEE Internet Things J. | 4 |
| 2022 | Ferroelectric FET-Based Implementation of FitzHugh-Nagumo Neuron ModelabstractFerroelectric field-effect transistor (FeFET)-based circuit implementation mimicking FitzHugh-Nagumo neuron is proposed in this work. The proposed circuit is shown to mimic biological neuron properties, such as excitation block and anodal break excitation which are not mimicked by an integrate and fire neuron model. We also show a winner-take-all circuit that can be used with this proposed neuron implementation. The neuron implementation requires just one FeFET, three baseline field-effect transistors, and one capacitor, making it area and energy-efficient. The neuron circuit, with minimum sized transistors, consumes approximately 10 pJ per spike. The neuron’s energy consumption per spike can be reduced to as low as 100 fJ by designing some of the transistors with aspect ratio less than one. Dinesh Rajasekharan, Amol D. Gaidhane, Amit Ranjan Trivedi, Yogesh Singh Chauhan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Ultralow-Power Localization of Insect-Scale Drones: Interplay of Probabilistic Filtering and Compute-in-MemoryabstractWe propose a novel compute-in-memory (CIM)-based ultralow-power framework for probabilistic localization of insect-scale drones. Localization is a critical subroutine for path planning and rotor control in drones, where a drone is required to continuously estimate its pose (position and orientation) in flying space. The conventional probabilistic localization approaches rely on the 3-D Gaussian mixture model (GMM)-based representation of a 3-D map. A GMM model with hundreds of mixture functions is typically needed to adequately learn and represent the intricacies of the map. Meanwhile, localization using complex GMM map models is computationally intensive. Since insect-scale drones operate under extremely limited area/power budget, continuous localization using GMM models entails much higher operating energy, thereby limiting flying duration and/or size of the drone due to a larger battery. Addressing the computational challenges of localization in an insect-scale drone using a CIM approach, we propose a novel framework of 3-D map representation using a harmonic mean of the “Gaussian-like” mixture (HMGM) model. We show thatshort-circuit currentof a multiinput floating-gate CMOS-based inverter follows the harmonic mean of a Gaussian-like function. Therefore, the likelihood function useful for drone localization can be efficiently implemented by connecting many multiinput inverters in parallel, each programmed with the parameters of the 3-D map model represented as HMGM. When the depth measurements are projected to the input of the implementation, the summed current of the inverters emulates the likelihood of the measurement. We have characterized our approach on an RGB-D scenes dataset. The proposed localization framework is$\sim 25\times $energy-efficient than the traditional, 8-bit digital GMM-based processor paving the way for tiny autonomous drones. Priyesh Shukla, Ankith Muralidhar, Nick Iliev, Theja Tulabandhula, Sawyer B. Fuller, Amit Ranjan Trivedi |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2021 | Compute-in-Memory Upside Down: A Learning Operator Co-Design Perspective for ScalabilityabstractThis paper discusses the potential of model-hardware co-design to simplify the implementation complexity of compute-in-SRAM deep learning considerably. Although compute-in-SRAM has emerged as a promising approach to improve the energy efficiency of DNN processing, current implementations suffer due to complex and excessive mixed-signal peripherals, such as the need for parallel digital-to-analog converters (DACs) at each input port. Comparatively, our approach inherently obviates complex peripherals by co-designing learning operators to SRAM's operational constraints. For example, our co-designed implementation is DAC-free even for multibit precision DNN processing. Additionally, we also discuss the interaction of our compute-in-SRAM operator with Bayesian inference of DNNs. We show a synergistic interaction of Bayesian inference with our framework, where Bayesian methods allow achieving similar accuracy with much smaller network size. Although each iteration of sample-based Bayesian inference is computationally expensive, the cost is minimized by our compute-in-SRAM approach. Meanwhile, by reducing the network size, Bayesian methods reduce the footprint cost of compute-in-SRAM implementation, which is a crucial concern for the method. We characterize this interaction for deep learning-based pose (position and orientation) estimation for a drone. Shamma Nasrin, Priyesh Shukla, Shruthi Jaisimha, Amit Ranjan Trivedi |
DATE | 4 |
| 2021 | MF-Net: Compute-In-Memory SRAM for Multibit Precision Inference Using Memory-Immersed Data Conversion and Multiplication-Free OperatorsabstractWe propose a co-design approach for compute-in-memory inference for deep neural networks (DNN). We use multiplication-free function approximators based on l1norm along with a co-adapted processing array and compute flow. Using the approach, we overcame many deficiencies in the current art of in-SRAM DNN processing such as the need for digital-to-analog converters (DACs) at each operating SRAM row/column, the need for high precision analog-to-digital converters (ADCs), limited support for multi-bit precision weights, and limited vector-scale parallelism. Our co-adapted implementation seamlessly extends to multi-bit precision weights, it doesn't require DACs, and it easily extends to higher vector-scale parallelism. We also propose an SRAM-immersed successive approximation ADC (SA-ADC), where we exploit the parasitic capacitance of bit lines of SRAM array as a capacitive DAC. Since the dominant area overhead in SA-ADC comes due to its capacitive DAC, by exploiting the intrinsic parasitic of SRAM array, our approach allows low area implementation of within-SRAM SA-ADC. Our 8×62 SRAM macro, which requires a 5-bit ADC, achieves ~105 tera operations per second per Watt (TOPS/W) with 8-bit input/weight processing at 45 nm CMOS. Our 8×30 SRAM macro, which requires a 4-bit ADC, achieves ~84 TOPS/W. SRAM macros that require lower ADC precision are more tolerant of process variability, however, have lower TOPS/W as well. We evaluated the accuracy and performance of our proposed network for MNIST, CIFAR10, and CIFAR100 datasets. We chose a network configuration which adaptively mixes multiplication-free and regular operators. The network configurations utilize the multiplication-free operator for more than 85% operations from the total. The selected configurations are 98.6% accurate for MNIST, 90.2% for CIFAR10, and 66.9% for CIFAR100. Since most of the operations in the considered configurations are based on proposed SRAM macros, our compute-in-memory's efficiency benefits broadly translate to the system-level. Shamma Nasrin, Diaa Badawi, A. Enis Çetin, Wilfred Gomes, Amit Ranjan Trivedi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2021 | On Database-Free Authentication of Microelectronic ComponentsabstractCounterfeit integrated circuits (ICs) have become a significant security concern in the semiconductor industry as a result of the increasingly complex and distributed nature of the supply chain. These counterfeit chips may result in performance degradation, profit reduction, and reputation risk for the manufacturer. Therefore, developing effective countermeasures against such malpractices is becoming severely crucial. Physical unclonable function (PUF)-based authentication methods have the potential to mitigate these challenges. However, PUF-based solutions are restrained by several factors, such as additional design efforts and significant area/power overhead, struggle to maintain and update challenge-response pairs (CRPs) database, and the vulnerability to machine learning (ML) attacks. In this article, we address these challenges by developing a novel database-free and enrolment-free hardware authentication approaches, i.e., a digital watermark metric for ICs. To enable efficient database-free hardware integrity verification without enrolment, first, we transform the intrinsic variations in circuit parameters, e.g., boundary scan chain (BSC) path delays in the joint test action group (JTAG) chain into robust digital signatures. Then, we perform statistical analysis on a small pilot unit of authentic chips to create a robust watermark for a complete batch of chips, which jointly captures the characteristics of the physical layout, the manufacturing process, and the foundry. The increasing complexity in the current state-of-the-art designs makes it extremely hard for an adversary to perfectly clone such statistical characterization of circuit parameters using counterfeit or compromised hardware. Besides, the proposed approach requires no additional design or hardware overhead in IC design since it utilizes an embedded structure, which inherently exists within the chips. It also obviates the design house from characterizing each manufactured chip instance, reducing overall testing cost. A path-delay measurement method at a high resolution based on clock phase sweep is introduced to measure the delay values effectively. The proposed intrinsic identifier-based authentication approach is validated by performing emulation on FPGAs and also by conducting physical measurements on custom-made printed circuit boards (PCBs). The reliability of the generated watermarks is evaluated with environmental temperature fluctuations and the aging effect. Fengchao Zhang, Shubhra Deb Paul, Patanjali SLPSK, Amit Ranjan Trivedi, Swarup Bhunia |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2020 | Supported-BinaryNet: Bitcell Array-Based Weight Supports for Dynamic Accuracy-Energy Trade-Offs in SRAM-Based Binarized Neural NetworkabstractIn this work, we introduce bitcell array-based support parameters to improve the prediction accuracy of SRAM-based binarized neural network (SRAM-BNN). Our approach enhances the training weight space of SRAM-BNN while requiring minimal overheads to a typical design. More flexibility of the weight space leads to higher prediction accuracy in our design. We adapt row digital-to-analog (DAC) converter, and computing flow in SRAM-BNN for bitcell array-based weight supports. Using the discussed interventions, our scheme also allows a dynamic trade-off of accuracy against energy to address dynamic energy constraints in typical real-time applications. Our approach reduces classification error in MNIST from 1.4% to 0.91%. To reduce the power overheads, we propose a dynamic drop out of support parameters, which also reduces the processing energy of the in-SRAM weight-input product Our architecture can dropout 52% of the bitcell array-based support parameters with only minimal accuracy degradation. We also characterize our design under varying degrees of process variability in the transistors. Shamma Nasrin, Srikanth Ramakrishna, Theja Tulabandhula, Amit Ranjan Trivedi |
ISCAS | 4 |
| 2020 | MC2RAM: Markov Chain Monte Carlo Sampling in SRAM for Fast Bayesian InferenceabstractThis work discusses the implementation of Marko Chain Monte Carlo (MCMC) sampling from an arbitrary Gaussian mixture model (GMM) within SRAM. We show a novel architecture of SRAM by embedding it with random number generators (RNGs), digital-to-analog converters (DACs), and analog-to-digital converters (ADCs) so that SRAM arrays can be used for high performance Metropolis-Hastings (MH) algorithm-based MCMC sampling. Most of the expensive computations are performed within the SRAM and can be parallelized for high speed sampling. Our iterative compute flow minimizes data movement during sampling. We characterize power-performance trade-off of our design by simulating on 45 nm CMOS technology. For a two-dimensional, two mixture GMM, the implementation consumes ~ 91μW power per sampling iteration and produces 500 samples in 2000 clock cycles on an average at 1 GHz clock frequency. Our study highlights interesting insights on how low-level hardware non-idealities can affect high-level sampling characteristics, and recommends ways to optimally operate SRAM within area/power constraints for high performance sampling. Priyesh Shukla, Ahish Shylendra, Theja Tulabandhula, Amit Ranjan Trivedi |
ISCAS | 4 |
| 2020 | Low Power Unsupervised Anomaly Detection by Nonparametric Modeling of Sensor StatisticsabstractThis article presents anomaly detection by examining sensor stream statistics (AEGIS), a novel mixed-signal framework for real-time AEGIS. AEGIS utilizes kernel density estimation (KDE)-based nonparametric density estimation to generate a real-time statistical model of the sensor data stream. The likelihood estimate of the sensor data point can be obtained based on the generated statistical model to detect outliers. We present CMOS Gilbert Gaussian cell-based design to realize Gaussian kernels for KDE. For outlier detection, the decision boundary is defined in terms of kernel standard deviation (σKernel) and likelihood threshold (PThres). We adopt a sliding window to update the detection model in real time. We use time-series data set provided from Yahoo to benchmark the performance of AEGIS. A f1-score higher than 0.87 is achieved by optimizing parameters such as length of the sliding window and decision thresholds which are programmable in AEGIS. Discussed architecture is designed using 45-nm technology node and our approach on average consumes ~75-μW power at a sampling rate of 2 MHz while using ten recent inlier samples for density estimation. Ahish Shylendra, Priyesh Shukla, Saibal Mukhopadhyay, Swarup Bhunia, Amit Ranjan Trivedi |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2019 | An Intrinsic and Database-Free Authentication by Exploiting Process Variation in Back-End CapacitorsabstractDetection of counterfeit chips has emerged as a crucial concern. Physically unclonable function (PUF)-based techniques are widely used for authentication; however, these require dedicated hardware and large signature database. In this paper, we show intrinsic and database-free authentication using back-end capacitors. The discussed technique simplifies authentication setup and reduces the test cost. We show that an analog-to-digital converter (ADC) can be modified for back-end capacitor-based authentication in addition to its regular functionality; hence, a dedicated authentication module is not necessary. Moreover, since back-end capacitors are quite insensitive to temperature and aging-induced variations than transistors, the discussed technique results in a more reliable authentication than transistor PUF-based authentication. Discussed authentication scheme manifests significant resilience against power supply instability. The modifications to conventional ADC incur 3.2% power overhead and 75% active-area overhead; however, arguably, the advantages of the discussed intrinsic and database-free authentication outweigh the overheads. Ahish Shylendra, Swarup Bhunia, Amit Ranjan Trivedi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Ultralow power acoustic feature-scoring using gaussian I-V transistorsabstractThis paper discusses energy-efficient acoustic feature-scoring using transistors with Gaussian-shaped Ids-Vgs. Acoustic feature-scoring is a critical step in speech recognition tasks such as speaker recognition. Suited to the transistor, we discuss a novel acoustic model based on the harmonic mean of Gaussian functions, yielding a much-simplified physical implementation. Compared to digital CMOS, the discussed implementation reduces energy dissipation of a 13-dimensional mixture function by 45×. For a test-case recognizing among seven speakers, the design reduces power dissipation by 7.8× than digital CMOS. Exploiting the simplified mixture function, we present algorithmic overdesigning for resiliency against hardware imperfections. Amit Ranjan Trivedi, Ahish Shylendra |
DAC | 1 |
| 2018 | Intrinsic and Database-free Watermarking in ICs by Exploiting Process and Design Dependent Variability in Metal-Oxide-Metal CapacitancesabstractAuthentication of integrated circuits (IC) to verify their integrity has emerged as a critical need to address increasing concerns associated with counterfeit ICs in the supply chain. In this paper, novel SAR-ADC based intrinsic and database-free authentication scheme has been proposed. Proposed technique utilizes mismatch in back end of line (BEOL) capacitors used in charge-redistribution SAR ADC to generate authentication signature. BEOL metal-oxide-metal (MOM) capacitors form a reliable source of process variation information and are less sensitive to aging & temperature induced variations. Line edge roughness is the primary source of mismatch in BEOL capacitors and thus, capacitor mismatch variation has been analyzed in terms of LER and geometric parameters. Resource overhead incurred by the proposed modifications to the ADC architecture to incorporate authentication ability is minimal and existing on-chip calibration circuitry is used to extract signature. Proposed technique does not require sophisticated test setup, thereby, simplifying the authentication procedure. Ahish Shylendra, Swarup Bhunia, Amit Ranjan Trivedi |
ISLPED | 3 |
| 2017 | Low Power Spatial Localization of Mobile Sensors with Recurrent Neural NetworkabstractThis work proposes a novel low power/area digital architecture for a Linear Program (LP) solver based on a recurrent (non-linear feedback) Neural Network (RNN), which is applied for spatial localization of sensors. Training data is not needed in our approach. We solve the primal and dual optimization problems for spatial localization with a single multi-functional data path which does not require matrix inversions. FPGA and ASIC implementations are presented which target a sensor (microrobot) localization problem in 2D using angle-of-arrival (AOA) measurements. The results show that the estimated locations (2D coordinates) are very close to the ground truth values in all tested scenarios. The proposed RNN has input and output layers, and a hidden layer with four neurons. FPGA implementation of the localizer in 180 nm process dissipates 180 mW of power at 1.5V and 31.25 MHz. When scaled to 128 neurons the performance is 13 Mop/sec/W for the same FPGA technology. The design has been also simulated in an ASIC 45 nm (PDK45 1V VDD) standard cell technology. With 128 neurons in the hidden layer the ASIC consumes 196.9 mW power at 516 MHz, which is equivalent to 677.165 Mop/sec/W. Nick Iliev, Amit Ranjan Trivedi |
ICCD | 2 |
| 2016 | Gate/source-overlapped heterojunction Tunnel FET-based LAMSTAR neural network and its Application to EEG Signal ClassificationabstractThis paper explores reduced complexity physical implementation of self-organizing-map (SOM) and LAMSTAR (Large Scale Memory Storage and Retrieval) neural network. Unique Gaussian IDS-VGScharacteristic of emerging gate/source-overlapped heterojunction Tunnel FET (SO-HTFET) is utilized to simplify the complexity of a SOM. For a given pattern, SO-HTFET-based SOM performs associative processing between the applied pattern feature and the stored neuron states. SO-HTFET reduces the SOM computing cell to just a single transistor. This is remarkable considering that a conventional digital SOM cell will require more than 100 transistors. IDS-VGS variance of SO-HTFET is modulated by varying its drain-to-source voltage (VDS). This enables dynamic adaptation of distance measures in SO-HTFET-based SOM. Various SOM-modules are combined in a LAMSTAR network with link weights to facilitate deep learning and integration of various features of the applied pattern in a decision making process. Electroencephalogram (EEG) classification is studied using SO-HTFET-based LAMSTAR. SO-HTFET enables a higher number of hidden neurons in LAMSTAR by reducing the complexity of SOM and thereby, improves classification accuracy than a conventional design. EEG classification accuracy is specifically evaluated for fixed neuron and dynamic neuron approaches. The optimal variance of SO-HTFET IDS-VGSis extracted for these approaches. Susmita Dey Manasi, Amit Ranjan Trivedi |
IJCNN | 2 |
| 2015 | Optimization of FinFET-based circuits using a dual gate pitch techniqueabstractSource/drain stressors in FinFET-based circuits lose their effectiveness at smaller contacted gate pitches. To improve circuit performance, a dual gate pitch technique is proposed in this work, where standard cells with double the gate pitch are selectively used on the gates of the circuit critical paths, at minimal area and power costs. A stress-aware library characterization is performed for FinFET-based standard cells by obtaining stress distributions using finite element simulations on a subset of structures. The stresses are then employed to create look-up tables for mobility multipliers and threshold voltage shifts, for subsequent performance characterization of FinFET-based standard cells. Finally, a circuit delay optimizer is applied using the dual gate pitch approach and is compared with an alternative gate sizing approach. Using a combination of gate sizing and the dual gate pitch approach, it is shown that the average power delay product improves by 12.9% and 15.9% in 14nm and 10nm technologies, respectively. Sravan K. Marella, Amit Ranjan Trivedi, Saibal Mukhopadhyay, Sachin S. Sapatnekar |
ICCAD | 2 |
| 2014 | Ultra-low power electronics with Si/Ge tunnel FETabstractSi/Ge Tunnel FET (TFET) with its subthermal subthreshold swing is attractive for low power analog and digital designs. Greater Ion/Ioffratio of TFET can reduce the dynamic power in digital designs, while higher gm/IDScan lower the bias power of analog amplifier. However, the above benefits of TFET are eclipsed by MOSFET at a higher power/performance point. Ultra low power scalability of the key analog and digital circuits, SRAM and operational transconductance amplifier (OTA), with TFET is demonstrated. Analyzing a TFET based cellular neural network, this work shows the feasibility of ultra-low-power neuromorphic computing with TFET. Amit Ranjan Trivedi, Mohammad Faisal Amir, Saibal Mukhopadhyay |
DATE | 1 |
| 2013 | Exploring tunnel-FET for ultra low power analog applications: a case study on operational transconductance amplifierabstractThis work studies the potentials and challenges of designing ultra-low-power analog circuits exploiting unique characteristics of Tunnel-FET (TFET). TFET can achieve ultra-low quiescent current (~pA). In the subthreshold operation, TFET exhibit subthreshold swing lower than 60mV/decade, and hence higher transconductance per bias current than the MOSFET. TFET also exhibit very weak temperature dependence, and higher output resistance. Among several challenges, TFET demonstrate higher Shot noise at low biasing current. Through design of TFET based Operational Transconductance Amplifier (OTA) these challenges and opportunities are discussed. For implantable bio-medical applications, TFET OTA based neural amplifier design is studied. Amit Ranjan Trivedi, Sergio Carlo, Saibal Mukhopadhyay |
DAC | 1 |
| 2012 | Self-adaptive power gating with test circuit for on-line characterization of energy inflection activityabstractA test circuit is presented for post-silicon and on-line characterization of the energy-inflection activity of power-gated circuits (the activity when overhead energy is equal to leakage savings) under static (process) and dynamic (voltage/temperature/input) variations. The test circuit is applied to design self-adaptive power-gating for energy-efficient SRAM. Amit Ranjan Trivedi, Saibal Mukhopadhyay |
VTS | 1 |
| 2012 | On the parametric failures of SRAM in a 3D-die stack considering tier-to-tier supply cross-talkabstractThis paper analyzes the supply crosstalk between logic cores and SRAMs on separate tiers in a 3D die-stack using a distributed RLC based 3D power grid model. The analysis shows that due to the supply cross-talk power variation in cores modulates the performances and parametric failures in SRAM. Wen Yueh, Subho Chatterjee, Amit Ranjan Trivedi, Saibal Mukhopadhyay |
VTS | 3 |