EDBT 2026 Demo / reviewers in the wild / expert
Saibal Mukhopadhyay
dblp:66/1210
· DBLP profile ↗
199ranked-venue papers
18as first author
54since 2021 · last 2026
0000-0002-8894-3390ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 146 · 18 first-author · 22 since 2021Artificial intelligence and machine learning · 37 · 24 since 2021Software engineering, systems software and programming languages · 30 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 since 2021Computer networks · 3 · 1 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Top-Down Design Methodology for Accuracy-to-Noise Mapping in Analog Compute-in-Memory Systems: Evaluation on Mamba
Isha Chakraborty, Laith A. Shamieh, Wei-Chun Wang 0001, Han Cho, Saibal Mukhopadhyay |
ISLPED | 5 |
| 2026 | BB-CIM: A Back-Bias Tuned Analog Compute In-Memory in 22nm FD-SOI for Improved Power-efficiency
Saideep Cherukuri, Apurba Prasad Padhy, Narasimha Vasishta Kidambi, Saibal Mukhopadhyay |
ISLPED | 4 |
| 2026 | Memory-Augmented Representation for Efficient Event-based Visuomotor Policy Learning with Adaptive Perception and ControlabstractEvent-based cameras are well-suited for fast and agile autonomous navigation due to their ultra-fast, microsecond-level temporal resolution. However, fully leveraging this potential requires highly efficient processing algorithms capable of asynchronous, event-by-event representations and policy updates. Current methods employ synchronous dense representation or process events in a fixed-rate time windows, leading to inefficiencies via redundant computation. We address this by proposing an end-to-end framework for event-to-control policy learning designed for reactive navigation tasks. Our method consists of a memory-augmented perception module that updates the representation asynchronously and adaptively selects the number of events to process. Using the memory representation, a lightweight policy module is jointly optimized with the perception module, and learns to predict control commands at rates that dynamically adjust to scene complexity in an event-based reinforcement learning setting. Evaluations on simulated drone navigation tasks demonstrate higher sample efficiency and robustness compared to dense frame-based methods. Moreover, our approach significantly reduces computational complexity by minimizing processing steps and event counts while maintaining competitive performance against state-of-the-art event-based methods. Uday Kamal, Saibal Mukhopadhyay |
WACV | 2 |
| 2026 | SSMRadNet : A Sample-wise State-Space Framework for Efficient and Ultra-Light Radar Segmentation and Object DetectionabstractWe introduce SSMRadNet, the first multi-scale State Space Model (SSM) based detector for Frequency Modulated Continuous Wave (FMCW) radar that sequentially processes raw ADC samples through two SSMs. One SSM learns a chirp-wise feature by sequentially processing samples from all receiver channels within one chirp, and a second SSM learns a representation of a frame by sequentially processing chirp-wise features. The latent representations of a radar frame are decoded to perform segmentation and detection tasks. Comprehensive evaluations on the RADIal dataset show SSMRadNet has 10-33× fewer parameters and 60-88× less computation (GFLOPs) while being 3.7× faster than state-of-the-art transformer and convolution-based radar detectors at competitive performance for segmentation tasks. Anuvab Sen, Mir Sayeed Mohammad, Saibal Mukhopadhyay |
WACV | 3 |
| 2026 | Towards Streaming LiDAR Object Detection with Point Clouds as Egocentric SequencesabstractAccurate and low-latency 3D object detection is essential for autonomous driving, where safety hinges on both rapid response and reliable perception. While rotating LiDAR sensors are widely adopted for their robustness and fidelity, current detectors face a trade-off: streaming methods process partial polar sectors on the fly for fast updates but suffer from limited visibility, cross-sector dependencies, and distortions from retrofitted Cartesian designs, whereas full-scan methods achieve higher accuracy but are bottlenecked by the inherent latency of a LiDAR revolution. We propose Polar-Fast-Cartesian-Full (PFCF), a hybrid detector that combines fast polar processing for intra-sector feature extraction with accurate Cartesian reasoning for full-scene understanding. Central to PFCF is a custom Mamba SSM-based streaming backbone with dimensionally-decomposed convolutions that avoids distortion-heavy planes, enabling parameter-efficient, translation-invariant, and distortion-robust polar representation learning. Local sector features are extracted via this backbone, then accumulated into a sector feature buffer to enable efficient inter-sector communication through a full-scan backbone. PFCF establishes a new Pareto frontier on the Waymo Open dataset, surpassing prior streaming baselines by 10% mAP and matching full-scan accuracy at twice the update rate. Code is available at https://github.com/meilongzhang/Polar-Hierarchical-Mamba. Mellon M. Zhang, Glen Chou, Saibal Mukhopadhyay |
WACV | 3 |
| 2025 | Low-Latency Digital Feedback for Stochastic Quantum Calibration Using Cryogenic CMOSabstractIn order to develop quantum computing systems towards practically useful applications, their physical quantum bits (qubits) must be able to operate with minimal error. Recent work has demonstrated stochastic gate calibration protocols for quantum systems which are meant to track drifting control parameters and tune gate operations to high fidelity. These protocols critically rely on low-latency feedback between the quantum system and its classical control hardware, which is impossible without on-board classical compute from FPGAs or ASICs. In this work, we analyze the performance of a single-shot stochastic calibration protocol for indefinite outcome quantum circuits under various latency conditions based on timing considerations from experimental quantum systems. We also demonstrate the benefits that can be achieved with ASIC implementation of the protocol by synthesizing the classical control logic in a 28 nm CMOS design node, with simulations extended to 14 nm FinFET and at both room and cryogenic temperatures. We show that these classes of quantum calibration protocols can be easily implemented within contemporary control system architectures for low-latency performance without significant power or resource utilization, allowing for the rapid tuning and drift control of any gate-model quantum system towards fault-tolerant computation. Nathan Eli Miller, Laith A. Shamieh, Saibal Mukhopadhyay |
DATE | 3 |
| 2025 | Intelligent Sensing-to-Action for Robust Autonomy at the Edge: Opportunities and ChallengesabstractAutonomous edge computing in robotics, smart cities, and autonomous vehicles relies on the seamless integration of sensing, processing, and actuation for real-time decision-making in dynamic environments. At its core is the sensing-to-action loop, which iteratively aligns sensor inputs with computational models to drive adaptive control strategies. These loops can adapt to hyper-local conditions, enhancing resource efficiency and responsiveness, but also face challenges such as resource constraints, synchronization delays in multimodal data fusion, and the risk of cascading errors in feedback loops. This article explores how proactive, context-aware sensing-to-action and action-to-sensing adaptations can enhance efficiency by dynamically adjusting sensing and computation based on task demands, such as sensing a very limited part of the environment and predicting the rest. By guiding sensing through control actions, action-to-sensing pathways can improve task relevance and resource use, but they also require robust monitoring to prevent cascading errors and maintain reliability. Multi-agent sensing-action loops further extend these capabilities through coordinated sensing and actions across distributed agents, optimizing resource use via collaboration. Additionally, neuromorphic computing, inspired by biological systems, provides an efficient framework for spike-based, event-driven processing that conserves energy, reduces latency, and supports hierarchical control-making it ideal for multi-agent optimization. This article highlights the importance of end-to-end co-design strategies that align algorithmic models with hardware and environmental dynamics, improve cross-layer inter-dependencies to improve throughput, precision, and adaptability for energy-efficient edge autonomy in complex environments. Amit Ranjan Trivedi, Sina Tayebati, Hemant Kumawat, Nastaran Darabi, Divake Kumar, Adarsh Kosta, Yeshwanth Venkatesha, Dinithi Jayasuriya, Nethmi Jayasinghe, Priyadarshini Panda, Saibal Mukhopadhyay, Kaushik Roy 0001 |
DATE | 11 |
| 2025 | Has the Deep Neural Network learned the Stochastic Process? An Evaluation ViewpointabstractThis paper presents the first systematic study of evaluating Deep Neural Networks (DNNs) designed to forecast the evolution of stochastic complex systems. We show that traditional evaluation methods like threshold-based classification metrics and error-based scoring rules assess a DNN's ability to replicate the observed ground truth but fail to measure the DNN's learning of the underlying stochastic process. To address this gap, we propose a new evaluation criteria called _Fidelity to Stochastic Process (F2SP)_, representing the DNN's ability to predict the system property _Statistic-GT_—the ground truth of the stochastic process—and introduce an evaluation metric that exclusively assesses F2SP. We formalize F2SP within a stochastic framework and establish criteria for validly measuring it. We formally show that Expected Calibration Error (ECE) satisfies the necessary condition for testing F2SP, unlike traditional evaluation methods. Empirical experiments on synthetic datasets, including wildfire, host-pathogen, and stock market models, demonstrate that ECE uniquely captures F2SP. We further extend our study to real-world wildfire data, highlighting the limitations of conventional evaluation and discuss the practical utility of incorporating F2SP into model assessment. This work offers a new perspective on evaluating DNNs modeling complex systems by emphasizing the importance of capturing underlying the stochastic process. Beomseok Kang, Biswadeep Chakraborty, Saibal Mukhopadhyay |
ICLR | 4 |
| 2025 | A Dynamical Systems-Inspired Pruning Strategy for Addressing Oversmoothing in Graph Attention NetworksabstractGraph Neural Networks (GNNs) face a critical limitation known as oversmoothing, where increasing network depth leads to homogenized node representations, severely compromising their expressiveness. We present a novel dynamical systems perspective on this challenge, revealing oversmoothing as an emergent property of GNNs’ convergence to low-dimensional attractor states. Based on this insight, we introduce DYNAMO-GAT, which combines noise-driven covariance analysis with Anti-Hebbian learning to dynamically prune attention weights, effectively preserving distinct attractor states. We provide theoretical guarantees for DYNAMO-GAT’s effectiveness and demonstrate its superior performance on benchmark datasets, consistently outperforming existing methods while requiring fewer computational resources. This work establishes a fundamental connection between dynamical systems theory and GNN behavior, providing both theoretical insights and practical solutions for deep graph learning. Biswadeep Chakraborty, Saibal Mukhopadhyay |
ICML | 3 |
| 2025 | AdaCred: Adaptive Causal Decision Transformers with Feature Crediting
Hemant Kumawat, Saibal Mukhopadhyay |
AAMAS | 2 |
| 2025 | Adaptive Graph Structure Inference for Learning Multivariate Point Processes using Spiking Neural NetworksabstractAccurate modeling and prediction of temporal point processes (TPPs) are crucial across domains such as neuroscience, epidemiology, finance, and social media analysis. We introduce the Spiking Dynamic Graph Network (SDGN), which integrates spiking neural networks (SNNs) with local spike-timing-dependent plasticity (STDP) to learn, online and in an event-driven fashion, the evolving spatio-temporal graph underlying a stream of timestamped events. SDGN relies on adaptive time-stepping, surrogate-gradient smoothing, and priority-queue updates to achieve O(log N) complexity per spike, ensuring both stability and efficiency. On synthetic benchmarks and four large-scale real-world datasets (NYC Taxi, 911 dispatches, Reddit posts, and Stack Overflow events), SDGN attains up to 15% higher held-out log-likelihood and 2× faster inference compared to state-of-the-art baselines. Ablation studies quantify the impact of each core component, and we discuss extensions for handling very dense graphs and heavy-tailed inter-event distributions. Biswadeep Chakraborty, Hemant Kumawat, Beomseok Kang, Saibal Mukhopadhyay |
IJCNN | 4 |
| 2025 | Tutorial: Autonomy with Neuromorphic SystemabstractThis tutorial will discuss how to incorporate neuromorphic circuits, computing, and sensing into an autonomous system at the edge, and present an end-to-end system analysis showing their integration. Amit Ranjan Trivedi, Priyadarshini Panda, Kaushik Roy 0001, Saibal Mukhopadhyay |
ISLPED | 4 |
| 2025 | FLAME: Fast Long-context Adaptive Memory for Event-based VisionabstractWe propose Fast Long-range Adaptive Memory for Event (FLAME), a novel scalable architecture that combines neuro-inspired feature extraction with robust structured sequence modeling
to efficiently process asynchronous and sparse event camera data. As a departure from conventional input encoding methods, FLAME presents Event Attention Layer, a novel feature extractor that leverages neuromorphic dynamics (Leaky Integrate-and-Fire (LIF)) to directly capture multi-timescale features from event streams. The feature extractor is integrates with a structured state-space model with a novel Event-Aware HiPPO (EA-HiPPO) mechanism that dynamically adapts memory retention based on inter-event intervals to understand relationship across varying temporal scales and event sequences. A Normal Plus Low Rank (NPLR) decomposition reduces the computational complexity of state update from $\mathcal{O}(N^2)$ to $\mathcal{O}(Nr)$, where $N$ represents the dimension of the core state vector and $r$ is the rank of a low-rank component (with $r \ll N$). FLAME demonstrates state-of-the-art accuracy for event-by-event processing on complex event camera datasets. Biswadeep Chakraborty, Saibal Mukhopadhyay |
NeurIPS | 2 |
| 2024 | A Hardware Accelerated Autoencoder for RF Communication Using Short-Time-Fourier- Transform Assisted Convolutional Neural NetworkabstractThis paper presents a hardware-accelerated autoencoder (AE) for wireless communication using a Short-Time-Fourier-Transform Assisted Convolutional Neural Network (STFT-CNN-AE). The design aims to reduce the autoencoder's resource requirements and power dissipation while maintaining its performance even in low Signal-to-Noise Ratio (SNR) wireless channels. The STFT-CNN-AE was implemented and tested on a Zynq UltraScale+ FPGA platform. Prototype measurements show that the STFT-CNN-AE achieves 3.5 times higher throughput at 2.6 times faster frequency, consumes 59% less power, and requires 76% fewer hardware resources (LUT and DSP) compared to a prior Multi-Layer Perceptron-based AE (MLP-AE). These improvements were achieved while maintaining comparable performance in low SNR (<7.5dB) channels. Kuchul Jung, Jongseok Woo, Saibal Mukhopadhyay |
DATE | 3 |
| 2024 | Cognitive Sensing for Energy-Efficient Edge IntelligenceabstractEdge platforms in autonomous systems integrate multiple sensors to interpret their environment. The high-resolution and high-bandwidth pixel arrays of these sensors improve sensing quality but also generate a vast, and arguably unnecessary, volume of real-time data. This challenge, often referred to as the analog data deluge, hinders the deployment of high-quality sensors in resource-constrained environments. This paper discusses the concept of cognitive sensing, which learns to extract low-dimensional features directly from high-dimensional analog signals, thereby reducing both digitization power and generated data volume. First, we discuss design methods for analog-to-feature extraction (AFE) using mixed-signal compute-in-memory. We then present examples of cognitive sensing, incorporating signal processing or machine learning, for various sensing modalities including vision, Radar, and Infrared. Subsequently, we discuss the reliability challenges in cognitive sensing, taking into account hardware and algorithmic properties of AFE. The paper concludes with discussions on future research directions in this emerging field of cognitive sensors. Minah Lee, Sudarshan Sharma, Wei-Chun Wang 0001, Hemant Kumawat, Nael Mizanur Rahman, Saibal Mukhopadhyay |
DATE | 6 |
| 2024 | Driving Autonomy with Event-Based Cameras: Algorithm and Hardware PerspectivesabstractIn high-speed robotics and autonomous vehicles, rapid environmental adaptation is necessary. Traditional cameras often face issues with motion blur and limited dynamic range. Event-based cameras address these by tracking pixel changes continuously and asynchronously, offering higher temporal resolution with minimal blur. In this work, we highlight our recent efforts in solving the challenge of processing event-camera data efficiently from both algorithm and hardware perspective. Specifically, we present how brain-inspired algorithms such as spiking neural networks (SNNs) can efficiently detect and track object motion from event-camera data. Next, we discuss how we can leverage associative memory structures for efficient event-based represen-tation learning. And finally, we show how our developed Application Specific Integrated Circuit (ASIC) architecture for low-latency, energy-efficient processing outperforms typical GPU/CPU solutions, thus enabling real-time event-based processing. With a 100x reduction in latency and a 1000x lower energy per event compared to state-of-the-art GPU/CPU setups, this enhances the front-end camera systems capability in autonomous vehicles to handle higher rates of event generation, improving control. Nael Mizanur Rahman, Uday Kamal, Manish Nagaraj, Shaunak Roy, Saibal Mukhopadhyay |
DATE | 5 |
| 2024 | Efficient Learning of Event-Based Dense Representation Using Hierarchical Memories with Adaptive Update
Uday Kamal, Saibal Mukhopadhyay |
ECCV (82) | 2 |
| 2024 | Multi-Tier 3D SRAM Module Design: Targeting Bit-Line and Word-Line FoldingabstractThis study presents a novel approach to address the challenges of scaling limitations and high read/write latencies inherent in conventional 2D SRAM designs through the development of a 15nm FinFET-based standalone 3D SRAM Array. Utilizing Monolithic Intertier Vias, our two-tiered implementation effectively mitigates these limitations along both X and Y axes. We introduce two distinct design methodologies for 3D SRAM arrays - Wordline and Bitline folding. Through post-layout simulations conducted on arrays with capacities of 2kB (256WLx64BL) and 8kB (512WLx128BL), our 3D designs exhibit superior performance metrics. Notably, the 3D Wordline-Folded design achieves a remarkable 57.54% average reduction in footprint compared to the 2D baseline across both array sizes. Furthermore, the 3D Bitline-Folded configuration demonstrates consistent superiority in speed, with an average read latency improvement of 17.16% and a notable 54.2% enhancement in write latency. Conversely, the 3D Wordline-Folded array emerges as the most energy-efficient option, boasting an average reduction of 15.3% in read energy and over 21% in write energy compared to the 2D baseline. Aditya Iyer 0001, Daehyun Kim 0002, Saibal Mukhopadhyay, Sung Kyu Lim |
ICCAD | 3 |
| 2024 | Sparse Spiking Neural Network: Exploiting Heterogeneity in Timescales for Pruning Recurrent SNNabstractRecurrent Spiking Neural Networks (RSNNs) have emerged as a computationally efficient and brain-inspired machine learning model. The design of sparse RSNNs with fewer neurons and synapses helps reduce the computational complexity of RSNNs. Traditionally, sparse SNNs are obtained by first training a dense and complex SNN for a target task and, next, eliminating neurons with low activity (activity-based pruning) while maintaining task performance. In contrast, this paper presents a task-agnostic methodology for designing sparse RSNNs by pruning an untrained (arbitrarily initialized) large model.
We introduce a novel Lyapunov Noise Pruning (LNP) algorithm that uses graph sparsification methods and utilizes Lyapunov exponents to design a stable sparse RSNN from an untrained RSNN. We show that the LNP can leverage diversity in neuronal timescales to design a sparse Heterogeneous RSNN (HRSNN). Further, we show that the same sparse HRSNN model can be trained for different tasks, such as image classification and time-series prediction. The experimental results show that, in spite of being task-agnostic, LNP increases computational efficiency (fewer neurons and synapses) and prediction performance of RSNNs compared to traditional activity-based pruning of trained dense models. Biswadeep Chakraborty, Beomseok Kang, Saibal Mukhopadhyay |
ICLR | 4 |
| 2024 | Topological Representations of Heterogeneous Learning Dynamics of Recurrent Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) have become an essential paradigm in neuroscience and artificial intelligence, providing brain-inspired computation. Recent advances in literature have studied the network representations of deep neural networks. However, there has been little work that studies representations learned by SNNs, especially using unsupervised local learning methods like spike-timing dependent plasticity (STDP). Recent work by [1] has introduced a novel method to compare topological mappings of learned representations called Representation Topology Divergence (RTD). Though useful, this method is engineered particularly for feedforward deep neural networks and cannot be used for recurrent networks like Recurrent SNNs (RSNNs). This paper introduces a novel methodology to use RTD to measure the difference between distributed representations of RSNN models with different learning methods. We propose a novel reformulation of RSNNs using feedforward autoencoder networks with skip connections to help us compute the RTD for recurrent networks. Thus, we investigate the learning capabilities of RSNN trained using STDP and the role of heterogeneity in the synaptic dynamics in learning such representations. We demonstrate that heterogeneous STDP in RSNNs yield distinct representations than their homogeneous and surrogate gradient-based supervised learning counterparts. Our results provide insights into the potential of heterogeneous SNN models, aiding the development of more efficient and biologically plausible hybrid artificial intelligence systems. Biswadeep Chakraborty, Saibal Mukhopadhyay |
IJCNN | 2 |
| 2024 | Structured Latent Space for Lightweight Prediction in Locally Interacting Discrete Dynamical SystemsabstractModeling the large-scale dynamical systems is a computationally expensive task. This is particularly a problem when the focus is solely on understanding the local behavior or state of the systems. Our primary objective is to determine when the propagation of such local interactions will reach a specific region of interest. Although conventional approaches that reconstruct the states of entire dynamic nodes can be used, they may entail unnecessary computational costs. In this paper, we investigate a Structured Latent space for Localized Prediction (SLLP) for the computationally efficient prediction of local behavior in the dynamical systems. The proposed model comprises a CNN encoder to represent the system in a low-dimensional vector, a LSTM module to learn the dynamics in the vector space, and a MLP decoder to predict the future state of a dynamic node. We evaluate the proposed method in the forest fire and stock market models in the task of predicting the burned state of a tree node and buy state of a investor node in future. We compare the proposed model with general ConvLSTM that reconstructs and predicts the entire systems. The proposed model exhibits similar or slightly worse AUC but significantly reduces computational costs, such as FLOPs (×131) and latency (×4.9), than ConvLSTM when predicting a single dynamic node. Beomseok Kang, Minah Lee, Saibal Mukhopadhyay |
IJCNN | 4 |
| 2024 | Harmonica: Hybrid Accelerator to Overcome Imperfections of Mixed-signal DNN AcceleratorsabstractIn recent years, PIM-based mixed-signal accelerators have been proposed as energy- and area-efficient solutions with ultra-high throughput to accelerate DNN computations. However, PIM designs are sensitive to imperfections such as noise, weight/conductance variations, and cell programming errors that substantially degrade the DNN accuracy. To address this issue, we propose a novel algorithm-hardware co-design framework called Harmonica that simultaneously avoids accuracy degradation due to imperfections, improves area utilization and execution time, and reduces energy consumption. Harmonica proposes to select imperfection-sensitive weights using an input channel-wise method and transfer them to a novel and robust digital accelerator while the main computations are performed in the analog PIM cores. Harmonica is adapted to leverage the preceding weight selection method by reducing ADC precision, employing smaller peripheral circuitry, and a hybrid quantization to optimize the design. Our comprehensive experiments show that even in the presence of imperfections as high as 50%, Harmonica reduces the accuracy degradation from 60% - 90% in designs without a protection solution (e.g., in ISAAC or SRE baselines) to 1% - 2% for different DNNs across diverse datasets. In addition, compared to the ISAAC (SRE), Harmonica improves the execution time, energy, area, power, area-efficiency, and power-efficiency by 26% (14%), 52% (40%), 28% (28%), 57% (45%), 43% (7.5×), and 91% (7.3×), respectively. By employing architecture-based differential cells, where two separated categories of crossbars are used for positive and negative weights, Harmonica outperforms ISAAC (SRE) by 75% (9.2×) and 2.65× (10.2×) in terms of area- and power-efficiency. Payman Behnam, Uday Kamal, Ali Shafiee, Alexey Tumanov, Saibal Mukhopadhyay |
IPDPS | 5 |
| 2024 | Enhancing IoT Security with a Hardware Accelerated Machine Learning Model coupling Autoencoder and Long-Short-Term-Memory for Anomaly DetectionabstractThis paper proposes a hardware accelerator for machine learning-based anomaly detection to enhance IoT security. Our model integrates Multilayer Perceptron (MLP) with Long Short-Term Memory (LSTM), utilizing an MLP-based Autoencoder and Isolation Forest algorithm for data dimensionality reduction and computational complexity reduction. Prototyped on a Zynq UltraScale+ XCZU9EG FPGA, our AE-LSTM model surpasses baseline MLP-only and LSTM-only models in resource utilization efficiency and detection accuracy. Compared to these baselines, it reduces parameters by 79.4% and 98% and LUT usage by 61.4% and 90.8%, respectively, while minimizing other resource utilization. Furthermore, power consumption is lowered to about 40% of the MLP-based model's consumption rate and 36% of the LSTM-based model's rate, with latency reduced to less than one-third from both baselines. Kuchul Jung, Jongseok Woo, Saibal Mukhopadhyay |
ISCAS | 3 |
| 2024 | Passive Lightweight On-chip Sensors for Power Side Channel Attack DetectionabstractHardware implementations of encryption engines are susceptible to Power Side Channel Attacks (PSCA) and countermeasures only make attacks harder without eliminating them. This work presents a temperature tolerant 65nm CMOS passive on-chip PSCA detection sensor for a AES-128 engine. The design uses on-chip sub-threshold oscillator to detect the resistor normally placed by an attacker on the power-line of an AES chip to mount PSCA. The measurement demonstrates 99.9% probability of successful detection, 1.1ms detection time, minimal area overhead compared to the encryption area and 0.83mW power. Nael Mizanur Rahman, Uday Kamal, Venakata Chaitanya Krishna Chekuri, Saibal Mukhopadhyay |
ISCAS | 5 |
| 2024 | Efficient Hardware Design of DNN for RF Signal Modulation RecognitionabstractThis paper presents an efficient deep neural network (DNN) accelerator design for the application of modulation recognition of Radio Frequency (RF) signals. A low complexity DNN model utilizing the ternary weights is demonstrated with co-analysis of the classification accuracy and the hardware design. In order to maximize the benefits of the ternary weight quantization, the dedicated hardware design called merged layer architecture is proposed. Physical design analysis shows that the proposed method can improve the bandwidth of the received signal and reduce hardware costs significantly. The physical design analysis is based on the Application Specific Integrated Circuit (ASIC) to evaluate the dedicated hardware design, and the functionality of the full DNN system is verified on the FPGA platform. Jongseok Woo, Kuchul Jung, Saibal Mukhopadhyay |
ISCAS | 3 |
| 2024 | Hardware-friendly Hessian-driven Row-wise Quantization and FPGA Acceleration for Transformer-based ModelsabstractRecent advancements in using FPGAs as co-processors for language model acceleration, particularly in terms of energy efficiency and flexibility, face challenges due to limited memory capacity. This issue hinders the deployment of transformer-based language models. To address these issues, we propose a novel software-hardware co-optimization approach. Our approach incorporates a hardware-friendly hessian-based row-wise mixed-precision quantization algorithm and an intra-layer mixed-precision compute fabric. The software algorithm, based on Hessian analysis, quantizes important rows with high precision and unimportant rows with low precision to compress the parameters effectively while maintaining accuracy, enabling fine-grained mixed-precision computation on the FPGA accelerator. Moreover, the integration of row-wise mixed-precision quantization and our energy-efficient data flow optimization scheme enables the accommodation of all necessary parameters of the BERT-base model on the FPGA, eliminating the need for off-chip memory access during runtime. The experimental results demonstrate that our FPGA accelerator generally outperforms existing FPGA accelerators, exhibiting energy efficiency improvements ranging from 4.12X to 14.57X compared to existing FPGA accelerators. Woohong Byun, Jongseok Woo, Saibal Mukhopadhyay |
ISLPED | 3 |
| 2024 | Cryogenic Operation of Computing-In-Memory based Spiking Neural NetworkabstractThis paper introduces a Computing-In-Memory based Spiking Neural Network (SNN) architecture for cryogenic operation of CMOS (Cryo-SNN). The paper demonstrates design strategies to improve energy efficiency of Cryo-SNN by coupling low-voltage operation at cryogenic temperature with innovative design of neuron circuits optimized for cryogenic conditions. By exploiting the enhanced device characteristics of 14 nm FinFET transistors at cryogenic temperatures, our architecture outlines critical adaptations to SNN components for optimal functionality in extreme environments. The circuit simulation using measurement calibrated 14nm FinFET models shows that a Cryo-SNN designed for MNIST classification operates with 4.54X improved energy-delay-product (EDP) over room temperature operation while maintaining similar accuracy. Further, the paper designs an optimized SNN architecture for autonomous health monitoring of miniaturized satellites at cryogenic temperature consuming less than 1mW of power. Laith A. Shamieh, Wei-Chun Wang 0001, Shida Zhang, Rakshith Saligram, Amol D. Gaidhane, Yu Cao 0001, Arijit Raychowdhury, Suman Datta, Saibal Mukhopadhyay |
ISLPED | 9 |
| 2024 | Online Relational Inference for Evolving Multi-agent Interacting SystemsabstractWe introduce a novel framework, Online Relational Inference (ORI), designed to efficiently identify hidden interaction graphs in evolving multi-agent interacting systems using streaming data. Unlike traditional offline methods that rely on a fixed training set, ORI employs online backpropagation, updating the model with each new data point, thereby allowing it to adapt to changing environments in real-time. A key innovation is the use of an adjacency matrix as a trainable parameter, optimized through a new adaptive learning rate technique called AdaRelation, which adjusts based on the historical sensitivity of the decoder to changes in the interaction graph. Additionally, a data augmentation method named Trajectory Mirror (TM) is introduced to improve generalization by exposing the model to varied trajectory patterns. Experimental results on both synthetic datasets and real-world data (CMU MoCap for human motion) demonstrate that ORI significantly improves the accuracy and adaptability of relational inference in dynamic settings compared to existing methods. This approach is model-agnostic, enabling seamless integration with various neural relational inference (NRI) architectures, and offers a robust solution for real-time applications in complex, evolving systems. Beomseok Kang, Priyabrata Saha, Sudarshan Sharma, Biswadeep Chakraborty, Saibal Mukhopadhyay |
NeurIPS | 5 |
| 2023 | Brain-Inspired Spatiotemporal Processing Algorithms for Efficient Event-Based PerceptionabstractNeuromorphic event-based cameras can unlock the true potential of bio-plausible sensing systems that mimic our human perception. However, efficient spatiotemporal processing algorithms must enable their low-power, low-latency, real-world application. In this talk, we highlight our recent efforts in this direction. Specifically, we talk about how brain-inspired algorithms such as spiking neural networks (SNNs) can approximate spatiotemporal sequences efficiently without requiring complex recurrent structures. Next, we discuss their event-driven formulation for training and inference that can achieve realtime throughput on existing commercial hardware. We also show how a brain-inspired recurrent SNN can be modeled to perform on event-camera data. Finally, we will talk about the potential application of associative memory structures to efficiently build representation for event-based perception. Biswadeep Chakraborty, Uday Kamal, Xueyuan She, Saurabh Dash, Saibal Mukhopadhyay |
DATE | 5 |
| 2023 | Heterogeneous Neuronal and Synaptic Dynamics for Spike-Efficient Unsupervised Learning: Theory and Design Principles
Biswadeep Chakraborty, Saibal Mukhopadhyay |
ICLR | 2 |
| 2023 | Associative Memory Augmented Asynchronous Spatiotemporal Representation Learning for Event-based Perception
Uday Kamal, Saurabh Dash, Saibal Mukhopadhyay |
ICLR | 3 |
| 2023 | Unsupervised 3D Object Learning through Neuron Activity aware Plasticity
Beomseok Kang, Biswadeep Chakraborty, Saibal Mukhopadhyay |
ICLR | 3 |
| 2023 | Brain-Inspired Spiking Neural Network for Online Unsupervised Time Series PredictionabstractEnergy and data-efficient online time series prediction for predicting evolving dynamical systems are critical in several fields, especially edge AI applications that need to update continuously based on streaming data. However, current Deep Neural Network (DNN)-based supervised online learning models require a large amount of training data and cannot quickly adapt when the underlying system changes. Moreover, these models require continuous retraining with incoming data making them highly inefficient. We present a novel Continuous Learning-based Unsupervised Recurrent Spiking Neural Network Model (CLURSNN), trained with spike timing dependent plasticity (STDP) to solve these issues. CLURSNN makes online predictions by reconstructing the underlying dynamical system using Random Delay Embedding by measuring the membrane potential of neurons in the recurrent layer of the recurrent spiking neural network (RSNN) with the highest betweenness centrality. We also use topological data analysis to propose a novel methodology using the Wasserstein Distance between the persistent homologies of the predicted and observed time series as a loss function. We show that the proposed online time series prediction methodology outperforms state-of-the-art DNN models when predicting an evolving Lorenz63 dynamical system. Biswadeep Chakraborty, Saibal Mukhopadhyay |
IJCNN | 2 |
| 2023 | Forecasting Evolution of Clusters in Game Agents with Hebbian LearningabstractLarge multi-agent systems such as real-time strategy games are often driven by collective behavior of agents. For example, in StarCraft II, human players group spatially near agents into a team and control the team to defeat opponents. In this light, clustering the agents in the game has been used for various purposes such as the efficient control of the agents in multi-agent reinforcement learning and game analytic tools for the game users. However, despite the useful information provided by clustering, learning the dynamics of multi-agent systems at a cluster level has been rarely studied yet. In this paper, we present a hybrid AI model that couples unsupervised and self-supervised learning to forecast evolution of the clusters in StarCraft II. We develop an unsupervised Hebbian learning method in a set-to-cluster module to efficiently create a variable number of the clusters with lower inference time complexity than K-means clustering. Also, a long short-term memory based prediction module is designed to recursively forecast state vectors generated by the set-to-cluster module to define cluster configuration. We experimentally demonstrate the proposed model successfully predicts complex movement of the clusters in the game. Beomseok Kang, Saibal Mukhopadhyay |
IJCNN | 2 |
| 2023 | CLUE: Cross-Layer Uncertainty Estimator for Reliable Neural Perception using Processing-in-Memory AcceleratorsabstractOne of the primary challenges of deploying deep neural networks (DNNs) is ensuring their reliable performance in unpredictable edge environments, which are often disrupted by a variety of uncertainties and variations. Estimating uncertainty is crucial in order to understand the reliability of task predictions and prevent system failures. However, quantifying uncertainty stemming from non-ideal properties of processing hardware has not yet been thoroughly studied. To address this, we present Cross-Layer Uncertainty Estimator (CLUE), which quantifies task uncertainty originating from both sensing/processing hardware variations and DNN algorithm uncertainty. Our experimental results demonstrate that CLUE provides uncertainty with up to 80.4% less calibration error and only 12% of energy overheads compared to using task DNN solely. Furthermore, CLUE is able to detect unreliable tasks that stem from processing hardware variations, which prior uncertainty estimators were unable to achieve. Finally, we demonstrate an adaptive control of processing hardware using CLUE, which allows a dynamic trade-off control between task accuracy and energy consumption. Minah Lee, Anni Lu, Mandovi Mukherjee, Shimeng Yu, Saibal Mukhopadhyay |
IJCNN | 5 |
| 2023 | XMD: An Expansive Hardware-Telemetry-Based Mobile Malware Detector for Endpoint DetectionabstractHardware-based Malware Detectors (HMDs) have shown promise in detecting malicious workloads. However, the current HMDs focus solely on the CPU core of a System-on-Chip (SoC) and, therefore, do not exploit the full potential of the hardware telemetry. In this paper, we propose XMD, an HMD that uses an expansive set of telemetry channels extracted from the different subsystems of SoC. XMD exploits the thread-level profiling power of the CPU-core telemetry, and the global profiling power of non-core telemetry channels, to achieve significantly better detection performance than currently used Hardware Performance Counter (HPC) based detectors. We leverage the concept of manifold hypothesis to analytically prove that adding non-core telemetry channels improves the separability of the benign and malware classes, resulting in performance gains. We train and evaluate XMD using hardware telemetries collected from 723 benign applications and 1033 malware samples on a commodity Android Operating System (OS)-based mobile device. XMD improves over currently used HPC-based detectors by 32.91% for the in-distribution test data. XMD achieves the best detection performance of 86.54% with a false positive rate of 2.9%, compared to the detection rate of 80%, offered by the best performing signature-based Anti-Virus(AV) on VirusTotal, on the same set of malware samples. Biswadeep Chakraborty, Sudarshan Sharma, Saibal Mukhopadhyay |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Sequence Approximation using Feedforward Spiking Neural Network for Spatiotemporal Learning: Theory and Optimization Methods
Xueyuan She, Saurabh Dash, Saibal Mukhopadhyay |
ICLR | 3 |
| 2022 | Learning Point Processes using Recurrent Graph NetworkabstractWe present a novel Recurrent Graph Network (RGN) approach for predicting discrete marked event sequences by learning the underlying complex stochastic process. Using the framework of Point Processes, we interpret a marked discrete event sequence as the superposition of different sequences each of a unique type. The nodes of the Graph Network use LSTM to incorporate past information whereas a Graph Attention Network (GAT Network) introduces strong inductive biases to capture the interaction between these different types of events. By changing the self-attention mechanism from attending over past events to attending over event types, we obtain a reduction in time and space complexity from$\mathcal{O}(N^{2})$(total number of events) to$\mathcal{O}(\vert \mathcal{Y}\vert^{2})$(number of event types). Experiments show that the proposed approach improves performance in log-likelihood, prediction and goodness-of-fit tasks with lower time and space complexity compared to state-of-the art Transformer based architectures. Saurabh Dash, Xueyuan She, Saibal Mukhopadhyay |
IJCNN | 3 |
| 2022 | Unsupervised Hebbian Learning on Point Sets in StarCraft IIabstractLearning the evolution of real-time strategy (RTS) game is a challenging problem in artificial intelligent (AI) system. In this paper, we present a novel Hebbian learning method to extract the global feature of a point set in StarCraft II game units, and its application to predict the movement of the points. Our model includes encoder, LSTM, and decoder, and we train the encoder with the unsupervised learning method. We introduce the concept of neuron activity aware learning combined with k-Winner-Takes-All. The optimal value of neuron activity is mathematically derived, and experiments support the effectiveness of the concept over the downstream task. Our Hebbian learning rule benefits the prediction with lower loss compared to self-supervised learning. Also, our model significantly saves the computational cost such as activations and FLOPs compared to a frame-based approach. Beomseok Kang, Saurabh Dash, Saibal Mukhopadhyay |
IJCNN | 4 |
| 2022 | Radar Guided Dynamic Visual Attention for Resource-Efficient RGB Object DetectionabstractAn autonomous system's perception engine must provide an accurate understanding of the environment for it to make decisions. Deep learning based object detection networks experience degradation in the performance and robustness for small and far away objects due to a reduction in object's feature map as we move to higher layers of the network. In this work, we propose a novel radar-guided spatial attention for RGB images to improve the perception quality of autonomous vehicles operating in a dynamic environment. In particular, our method improves the perception of small and long range objects, which are often not detected by the object detectors in RGB mode. The proposed method consists of two RGB object detectors, namely the Primary detector and a lightweight Secondary detector. The primary detector takes a full RGB image and generates primary detections. Next, the radar proposal framework creates regions of interest (ROIs) for object proposals by projecting the radar point cloud onto the 2D RGB image. These ROIs are cropped and fed to the secondary detector to generate secondary detections which are then fused with the primary detections via non-maximum suppression. This method helps in recovering the small objects by preserving the object's spatial features through an increase in their receptive field. We evaluate our fusion method on the challenging nuScenes dataset and show that our fusion method with SSD-lite as primary and secondary detector improves the baseline primary yolov3 detector's recall by 14 % while requiring three times fewer computational resources. Hemant Kumawat, Saibal Mukhopadhyay |
IJCNN | 2 |
| 2022 | Lightweight Model Uncertainty Estimation for Deep Neural Object DetectionabstractQuantifying model uncertainty of Deep Neural Network (DNN) is important to understand the reliability of the model prediction and avoid risks in safety critical applications. Various approaches, including Bayesian neural networks, Monte-Carlo dropout, and ensembles, are suggested to measure the model uncertainty; but with huge computational cost. We present ModelNet, an Artificial Neural Network (ANN) that can estimate spatial/semantic model uncertainties of a DNN based object detection with less computation overhead. ModelNet is a deterministic ANN that distills the predictive distribution of stochastic DNN. Experimental results show that ModelNet can learn the uncertainty estimation from stochastic DNN in various architectures. ModelNet can perform as a probabilistic object detector with 39x-179x less number of operations, or as an uncertainty assistant to a task network with 1.4x more parameters and 38x less number of operations compared to stochastic DNN. Moreover, a case study of uncertainty driven adaptive sensor using ModelNet is presented. Minah Lee, Burhan Ahmad Mudassar, Saibal Mukhopadhyay |
IJCNN | 3 |
| 2022 | A Methodology for Understanding the Origins of False Negatives in DNN Based Object DetectorsabstractIn this paper we present two novel complimentary methods namely the gradient analysis and the activation discrepancy analysis to analyze the perception failures occurring inside the DNN based object detectors. The gradient analysis localizes the nodes within the network that fail consistently in a scenario, thus creating a ‘signature’ of False Negatives (FNs). This method traces a set of False Negatives through the network and finds sections of the network that contribute to this set. The signatures show the location of the faulty nodes is sensitive to input conditions (such as darkness, glare etc.), network architecture, training hyperparameters, object class etc. Certain nodes of the network fail consistently throughout the training process thus implying that some False Negatives occur due to the global optimization nature of Stochastic Gradient Descent (SGD) based training. This analysis requires the knowledge of False Negatives and therefore can be used for post-hoc diagnostic analysis. On the other hand, the activation discrepancy analysis analyzes the discrepancy in forward activations of a DNN. This method can be conducted online and shows that the pattern of the activation discrepancy is sensitive to input conditions and detection recall. Kruttidipta Samal, Hemant Kumawat, Marilyn Wolf, Saibal Mukhopadhyay |
IJCNN | 4 |
| 2022 | Analysis of the Effect of Hot Carrier Injection in An Integrated Inductive Voltage RegulatorabstractThis paper presents a simulation-based study to evaluate the effect of Hot Carrier Injection (HCI) on the characteristics of an on-chip, digitally-controlled, switched inductor voltage regulator (IVR) architecture. Our methodology integrates device-level aging models, circuit simulations in SPICE, and control loop simulations in Simulink. We characterize the effect of HCI on individual components of an IVR, and their combined effect on the efficiency and transient performance. Our analysis using an IVR designed in 65nm CMOS shows that aging of the power stages has a smaller impact on performance compared to that of the control loop. Further, we perform a comparative analysis to show that, with a 1.8V supply, HCI leads to higher aging-induced degradation of IVR than Negative Bias Temperature Instability (NBTI). Finally, our simulation shows that parasitic inductance near IVR input aggravates NBTI and parasitic capacitance near IVR output aggravates HCI effects on IVR’s performance. Shida Zhang, Nael Mizanur Rahman, Venakata Chaitanya Krishna Chekuri, Carlos Tokunaga, Saibal Mukhopadhyay |
ISLPED | 5 |
| 2022 | A ReRAM Memory Compiler for Monolithic 3D Integrated Circuits in a Carbon Nanotube ProcessabstractWe present a ReRAM memory compiler for monolithic 3D (M3D) integrated circuits (IC). We develop ReRAM architectures for M3D ICs using 1T-1R bit cells and single and multiple tiers of transistors for access and peripheral circuits. The compiler includes an automated flow for generation of subarrays of different dimensions and larger arrays of a target capacity by integrating multiple subarrays. The compiler is demonstrated using an M3D process design kit (PDK) based on a Carbon Nanotube Transistor technology. The PDK includes multiple layers of transistors and back-end-of-the-line integrated ReRAM. Simulations show the compiled ReRAM macros with multiple tiers of transistors reduces footprint and improves performance over the macros with single-tier transistors. The compiler creates layout views that are exported into library exchange format or graphic data system for full-array assembly and schematic/symbol views to extract per-bit read/write energy and read latency. Comparison of the proposed M3D subarray architectures with baseline 2D subarrays, generated with a custom-designed set of bit cells and peripherals, demonstrate up to 48% area reduction and 13% latency improvement. Dae Hyun Kim 0004, Sung Kyu Lim, Saibal Mukhopadhyay |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2022 | Robust Processing-In-Memory With Multibit ReRAM Using Hessian-Driven Mixed-Precision ComputationabstractThis article presents an algorithmic approach to design reliable deep neural networks (DNNs) in the presence of stochastic variations in the network parameters induced by process variations in the bit cells in a processing-in-memory (PIM) architecture. We propose and derive a Hessian-based sensitivity metric that can be computed without computing or storing the full Hessian to identify and protect the “important” network parameters while allowing large variations in unprotected parameters. We also show that this metric can be used to aggressively quantize unprotected network parameters in the PIM for improved inference efficiency and compute density. Experiments on modern DNNs like ResNet, MobileNetv2, and DenseNet on CIFAR10 using measured RRAM device data shows the effectiveness of our approach. Saurabh Dash, Yandong Luo, Anni Lu, Shimeng Yu, Saibal Mukhopadhyay |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Introspective Closed-Loop Perception for Energy-efficient SensorsabstractTask-driven closed-loop perception-sensing systems have shown considerable energy savings over traditional open-loop systems. Prior works on such systems have used simple feedback signals such as object detections and tracking which led to poor perception quality. This paper proposes an improved approach based on perceptual risk. First, a method is proposed to estimate the risk of failure to detect a target of interest. The risk estimate is used as a signal in a feedback system to determine how sensor resources are utilized. Two feedback algorithms are proposed: one based on proportional/integral methods and the other based on 0/1 (bang-bang) methods. These feedback algorithms are compared based on the efficiency with which they use available sensor resources as well as their absolute detection rates. Experiments on two real-world autonomous driving datasets show that the proposed system has better object detection recall and lower marginal cost of prediction than prior work. Kruttidipta Samal, Marilyn Wolf, Saibal Mukhopadhyay |
AVSS | 3 |
| 2021 | Towards Improving the Trustworthiness of Hardware based Malware Detector using Online Uncertainty EstimationabstractHardware-based Malware Detectors (HMDs) using Machine Learning (ML) models have shown promise in detecting malicious workloads. However, the conventional black-box based machine learning (ML) approach used in these HMDs fail to address the uncertain predictions, including those made on zero-day malware. The ML models used in HMDs are agnostic to the uncertainty that determines whether the model “knows what it knows,” severely undermining its trustworthiness. We propose an ensemble-based approach that quantifies uncertainty in predictions made by ML models of an HMD, when it encounters an unknown workload than the ones it was trained on. We test our approach on two different HMDs that have been proposed in the literature. We show that the proposed uncertainty estimator can detect > 90% of unknown workloads for the Power-management based HMD, and conclude that the overlapping benign and malware classes undermine the trustworthiness of the Performance Counter-based HMD. Nikhil Chawla, Saibal Mukhopadhyay |
DAC | 3 |
| 2021 | Reliable Edge Intelligence in Unreliable EnvironmentabstractA key challenge for deployment of artificial intelligence (AI) in real-time safety-critical systems at the edge is to ensure reliable performance even in unreliable environments. This paper will present a broad perspective on how to design AI platforms to achieve this unique goal. First, we will present examples of AI architecture and algorithm that can assist in improving robustness against input perturbations. Next, we will discuss examples of how to make AI platforms robust against hardware induced noise and variation. Finally, we will discuss the concept of using lightweight networks as reliability estimators to generate early warning of potential task failures. Minah Lee, Xueyuan She, Biswadeep Chakraborty, Saurabh Dash, Burhan Ahmad Mudassar, Saibal Mukhopadhyay |
DATE | 6 |
| 2021 | Closed-loop Approach to Perception in Autonomous SystemabstractCurrently, functional tasks within Autonomous Systems are balkanized into several sub-systems such as object detection, tracking, motion planning, multi-sensor fusion etc. which are developed and tested in isolation. In recent times, deep learning is used in the perception systems for improved accuracy, but such algorithms are not adaptive to the transient real-world requirements of an Autonomous System such as latency and energy. These limitations are critical for resource constrained systems such as autonomous drones. Therefore, a holistic closed-loop system design is required for building reliable and efficient perception systems for autonomous drones. The closed-loop perception system creates a focus-of-attention based feedback from end-task such as motion planning to control computation within the deep neural networks (DNNs) used in early perception tasks such as object detection. We observe that this closed-loop perception system improves resource utilization of resource hungry DNNs within perception system with minimal impact on motion planning. Kruttidipta Samal, Marilyn Wolf, Saibal Mukhopadhyay |
DATE | 3 |
| 2021 | Securing IoT Devices Using Dynamic Power Management: Machine Learning ApproachabstractThe shift in paradigm from cloud computing toward edge has resulted in faster response times, a more secure and energy-efficient edge. Internet-of-Things (IoT) devices form a vital part of the edge, but despite legions of benefits it offers, increasing vulnerabilities and escalation in malware generation has rendered them insecure. Software-based approaches are prominent in malware detection, but they fail to meet the requirements for IoT devices. Dynamic power management (DPM) is architecture agnostic and inherently pervasive component existing in all low-power IoT devices. In this article, we demonstrate dynamic voltage and frequency scaling (DVFS) states form a signature pertinent to an application, and its runtime variations comprise of features essential for securing IoT devices against malware attacks. We have demonstrated this proof of concept by performing experimental analysis on a Snapdragon 820 mobile processor, hosting the Android operating system (OS). We developed a supervised machine learning model for application classification and malware identification by extracting features from the DVFS states time series. The experimental results show$> 0.7~F1$score in classifying different android benchmarks and >0.88 in classifying benign and malware applications when evaluated across different DVFS governors. We also performed power measurements under different governors to evaluate power-security aware governor. We have observed higher detection accuracy and lower power dissipation under settings of the ondemand governor. Nikhil Chawla, Monodeep Kar, Saibal Mukhopadhyay |
IEEE Internet Things J. | 5 |
| 2021 | Physics-incorporated convolutional recurrent neural networks for source identification and forecasting of dynamical systems
Priyabrata Saha, Saurabh Dash, Saibal Mukhopadhyay |
Neural Networks | 3 |
| 2021 | ScieNet: Deep learning with spike-assisted contextual information extraction
Xueyuan She, Daehyun Kim 0002, Saibal Mukhopadhyay |
Pattern Recognit. | 4 |
| 2021 | Machine Learning in Wavelet Domain for Electromagnetic Emission Based Malware AnalysisabstractThis paper presents a signal processing and machine learning (ML) based methodology to leverage Electromagnetic (EM) emissions from an embedded device to remotely detect a malicious application running on the device and classify the application into a malware family. We develop Fast Fourier Transform (FFT) based feature extraction followed by Support Vector Machine (SVM) and Random Forest (RF) based ML models to detect a malware. We further propose methods to learn characteristic behavior of different malwares from EM traces to reveal similarities to known malware families and improve efficiency of malware analysis. We propose to use Discrete Wavelet Transform (DWT) based feature extraction from spectrograms of EM side-channel traces and perform ML on the extracted features to learn fine-grained patterns of malware families. The experimental demonstration on Open-Q 820 development platform demonstrate 0.99 F1score in detecting malware and 0.88 F1score in uniquely classifying malwares among 8 malware family evaluated using Support Vector Machines (SVM) and Random Forest (RF) Machine Learning(ML) models. We also demonstrate capability of proposed framework in identifying new unknown applications with 0.99 recall and unknown malware family with 0.87 recall. Nikhil Chawla, Saibal Mukhopadhyay |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | A Fully Spiking Hybrid Neural Network for Energy-Efficient Object DetectionabstractThis paper proposes a Fully Spiking Hybrid Neural Network (FSHNN) for energy-efficient and robust object detection in resource-constrained platforms. The network architecture is based on a Spiking Convolutional Neural Network using leaky-integrate-fire neuron models. The model combines unsupervised Spike Time-Dependent Plasticity (STDP) learning with back-propagation (STBP) learning methods and also uses Monte Carlo Dropout to get an estimate of the uncertainty error. FSHNN provides better accuracy compared to DNN based object detectors while being more energy-efficient. It also outperforms these object detectors, when subjected to noisy input data and less labeled training data with a lower uncertainty error. Biswadeep Chakraborty, Xueyuan She, Saibal Mukhopadhyay |
IEEE Trans. Image Process. | 3 |
| 2020 | WarningNet: A Deep Learning Platform for Early Warning of Task Failures under Input Perturbation for Reliable Autonomous PlatformsabstractThere is a growing interest in deploying complex deep neural networks (DNN) in autonomous systems to extract task-specific information from real-time sensor data and drive critical tasks. The perturbations in sensor data due to noise or environmental conditions can lead to errors in information extraction and degrade reliability of the entire autonomous systems. This paper presents a light-weight deep learning plat-form, WarningNet, that operates on sensor data to estimate potential task failures due to spatiotemporal input perturbations. Experimental results show that WarningNet can provide early warning of the performance degradation of different tasks within a fraction of the time required for the task to complete. As a case-study, we show that the early warning can be leveraged to improve the task reliability under adverse condition using on-demand input pre-processing. Minah Lee, Burhan Ahmad Mudassar, Taesik Na, Saibal Mukhopadhyay |
DAC | 4 |
| 2020 | Q-PIM: A Genetic Algorithm based Flexible DNN Quantization Method and Application to Processing-In-Memory PlatformabstractThis paper presents a genetic algorithm (GA) based training free layer-wise quantization method, named as GAQ, to reduce model complexity of arbitrary DNN architectures. The proposed algorithm formulates an optimization problem to determine the quantization level for each DNN layer under the constrain of maximum accuracy degradation and uses genetic algorithm to solve the problem at the inference stage of any pre-trained DNN models. The experimental results on various DNNs for image classification demonstrate 5x to 17x weight compression rate with insignificant (< 2%) accuracy loss, comparable with existing quantization algorithms which typically require multi-pass retraining and handcrafted tuning. To evaluate the computational benefits of GAQ, we present a SRAM based flexible precision all-digital processing-in-memory (PIM) architecture, named as Q-PIM, that leverages GAQ to optimally control precision for each DNN layer to enhance efficiency. The simulation in 28nm CMOS shows potential for significant energy and latency advantage over fixed-precision PIM architectures. Daehyun Kim 0002, Saibal Mukhopadhyay |
DAC | 4 |
| 2020 | Hessian-Driven Unequal Protection of DNN Parameters for Robust InferenceabstractThis paper presents an algorithmic approach to design reliable deep neural networks (DNN) in the presence of stochastic variations in the network parameters induced by process variations in the bit-cells in a processing-in-memory (PIM) architecture. We propose and derive a Hessian based sensitivity metric that can be computed without computing or storing the full Hessian to identify and protect the "important" network parameters while allowing large variations in unprotected parameters. Experiments on modern DNNs like ResNet, MobileNetv2, DenseNet on CIFAR10 demonstrates that by shielding only a small (1% -- 5%) fraction of parameters one can achieve less than 1% accuracy degradation even under large (50%) stochastic variations in other parameters. Saurabh Dash, Saibal Mukhopadhyay |
ICCAD | 2 |
| 2020 | RTL-to-GDS Design Tools for Monolithic 3D ICsabstractIn this paper, we propose RTL-to-GDS design flow for monolithic 3D ICs (M3D) built with carbon nanotube field-effect transistors and resistive memory. Our tool flow is based on commercial 2D tools and smart ways to extend them to conduct M3D design and simulation. We provide a post-route optimization flow, which exploits the full potential of the underlying M3D process design kit (PDK) for power, performance and area (PPA) optimization. We also conduct IR-drop and thermal analysis on M3D designs to improve the reliability. To enhance the testability of our M3D designs, we develop design-for-test (DFT) methodologies and integrate a low-overhead built-in self-test module into our design for testing inter-layer vias (ILVs) as well as logic circuitries in the individual tiers. Our benchmark design is RISC-V Rocketcore, which is an open source processor. Our experiments show 8.1% of power, 19.6% of wirelength and 55.7% of area savings with M3D designs at iso-performance compared to its 2D counterpart. In addition, our IR-drop and thermal analyses indicate acceptable power and thermal integrity in our M3D design. Gauthaman Murali, Pruek Vanna-Iampikul, Dae Hyun Kim 0004, Arjun Chaudhuri, Sanmitra Banerjee, Krishnendu Chakrabarty, Saibal Mukhopadhyay, Sung Kyu Lim |
ICCAD | 9 |
| 2020 | Silicon vs. Organic Interposer: PPA and Reliability Tradeoffs in Heterogeneous 2.5D Chiplet IntegrationabstractThe optimal selection of an interposer substrate is important in 2.5D systems, because its physical, material and electrical characteristics govern the overall system performance, reliability and cost. Several materials have been proposed that offer various tradeoffs including silicon, organic, glass and etc. In this paper, we conduct a quantitative comparison between two 2.5D IC designs based on silicon vs. liquid crystal polymer (LCP) interposer technologies in the overall system level for the first time. We also investigate tradeoffs in power, performance and area (PPA), signal integrity (SI) and power integrity (PI) depending on the interposer technologies. Through our flow, we generate a large-scale benchmark architecture with commercial-grade GDS layouts of interposer and chiplets using two different interposer substrates. Then, we model transmission lines and power delivery network (PDN) of each 2.5D IC design. Finally, we perform PPA analysis, SI and PI on both 2.5D IC designs to observe the quantitative tradeoffs between two designs. Our experiment shows that silicon interposer-based design has 10.46% less power, 0.25× smaller area and 0.57× shorter average wirelength compared to LCP interposer-based design. However, LCP-based design has 0.59× smaller PDN DC impedance and 0.75× shorter worst delay of interposer wire while maintaining the power delivery efficiency. Lastly, our cost analysis of 2.5D IC design indicates that the overall cost of organic LCP technology, if both the chiplets and their interposer costs are combined, is 2.69× higher than the silicon even the cost of LCP interposer is 1.91% of silicon interposer. This indicates that LCP technology is prohibitive unless the interconnect and bump dimensions are dramatically reduced. Venakata Chaitanya Krishna Chekuri, Nael Mizanur Rahman, Majid Ahadi Dolatsara, Hakki Mert Torun, Madhavan Swaminathan, Saibal Mukhopadhyay, Sung Kyu Lim |
ICCD | 7 |
| 2020 | MagNet: Discovering Multi-agent Interaction Dynamics using Neural NetworkabstractWe present the MagNet, a neural network-based multi-agent interaction model to discover the governing dynamics and predict evolution of a complex multi-agent system from observations. We formulate a multi-agent system as a coupled non-linear network with a generic ordinary differential equation (ODE) based state evolution, and develop a neural network-based realization of its time-discretized model. MagNet is trained to discover the core dynamics of a multi-agent system from observations, and tuned on-line to learn agent-specific parameters of the dynamics to ensure accurate prediction even when physical or relational attributes of agents, or number of agents change. We evaluate MagNet on a point-mass system in two-dimensional space, Kuramoto phase synchronization dynamics and predator-swarm interaction dynamics demonstrating orders of magnitude improvement in prediction accuracy over traditional deep learning models. Priyabrata Saha, Burhan Ahmad Mudassar, Saibal Mukhopadhyay |
ICRA | 5 |
| 2020 | Flex-PIM: A Ferroelectric FET based Vector Matrix Multiplication Engine with Dynamical Bitwidth and Floating Point PrecisionabstractThis paper presents Flex-PIM, a ferroelectric FET (FeFET) based processing-in-memory (PIM) engine for vector-matrix-multiplication (VMM). With FeFET as the basic memory cell, Flex-PIM features low read latency/programming energy, non-volatility and high density. The core of Flex-PIM micro-architecture is an all-digital VMM engine integrated with innovative memory array peripherals to realize dynamically controllable bitwidth and floating point precision. The Flex-PIM architecture is simulated in 28nm CMOS technology and shows multiplication-accumulation (MAC) operations from 32-bit floating point (99 GMACS/W) to 4 bit fixed-point (3.3 TMACS/W). A system level design with specialized instruction set is presented to acclerate training and inference of deep neural networks (DNN) using Flex-PIM. The full-chip simulations show that Flex-PIM can increase computing efficiency of training and inference by 32x and 120x, respectively, over desktop GPUs while maintaining high accuracy over a wide-range of DNN models using flexible precision. Daehyun Kim 0002, Saibal Mukhopadhyay |
IJCNN | 4 |
| 2020 | Hybridization of Data and Model based Object Detection for Tracking in Flash LidarsabstractIn recent times deep neural networks have become very successful in solving traditionally hard problems in Computer Vision such as Object Detection. This is due to their ability to find hidden patterns in high dimensional data such as images. But if there is a known structure within data that can be accurately represented by a pre-defined model, then by merging this model based algorithm and deep neural network, overall system accuracy can be increased. We apply this idea for solving the task of flash lidar object detection and tracking. Flash lidar is an emerging lidar sensing technology which is getting a lot of attention lately due to their lack of moving parts compared to prevalent scanning lidars. Samples from flash lidar suffer from both spatial and temporal noise which coupled with low angular resolution and FoV lead to low accuracy in object detection. In this paper we present a data driven deep learning based flash lidar object detector and tracker. To our knowledge, this is the first work to use deep learning for flash lidar object detection/tracking. Our tracker has two detectors- 1. supervised object detector and 2. unsupervised class agnostic foreground/moving object detector which are merged to achieve multi-object tracking accuracy of 47.9% on CAMEL dataset. Kruttidipta Samal, Marilyn Wolf, Saibal Mukhopadhyay |
IJCNN | 3 |
| 2020 | SAFE-DNN: A Deep Neural Network With Spike Assisted Feature Extraction For Noise Robust InferenceabstractWe present a Deep Neural Network with Spike Assisted Feature Extraction (SAFE-DNN) to improve robustness of classification under stochastic perturbation of inputs. The proposed network augments a DNN with unsupervised learning of low-level features using spiking neural network (SNN) with spike-timing-dependent plasticity (STDP). The complete network learns to ignore local perturbation while performing global feature detection and classification. The experimental results on CIFAR-10 and ImageNet subset demonstrate improved noise robustness for multiple DNN architectures without sacrificing accuracy on clean images. Xueyuan She, Priyabrata Saha, Daehyun Kim 0002, Saibal Mukhopadhyay |
IJCNN | 5 |
| 2020 | BiasP: a DVFS based exploit to undermine resource allocation fairness in linux platformsabstractDynamic Voltage and Frequency Scaling (DVFS) plays an integral role in reducing the energy consumption of mobile devices, meeting the targeted performance requirements at the same time. We examine the security obliviousness of CPUFreq, the DVFS framework in Linux-kernel based systems. Since Linux-kernel based operating systems are present in a wide array of applications, the high-level CPUFreq policies are designed to be platform-independent. Using these policies, we present BiasP exploit, which restricts the allocation of CPU resources to a set of targeted applications, thereby degrading their performance. The exploit involves detecting the execution of instructions on the CPU core pertinent to the targeted applications, thereafter using CPUFreq policies to limit the available CPU resources available to those instructions. We demonstrate the practicality of the exploit by operating it on a commercial smartphone, running Android OS based on Linux-kernel. We can successfully degrade the User Interface (UI) performance of the targeted applications by increasing the frame processing time and the number of dropped frames by up to 200% and 947% for the animations belonging to the targeted-applications. We see a reduction of up to 66% in the number of retired instructions of the targeted-applications. Furthermore, we propose a robust detector which is capable of detecting exploits aimed at undermining resource allocation fairness through malicious use of the DVFS framework. Nikhil Chawla, Saibal Mukhopadhyay |
ISLPED | 3 |
| 2020 | Architecture, Chip, and Package Codesign Flow for Interposer-Based 2.5-D Chiplet Integration Enabling Heterogeneous IP ReuseabstractA new trend in system-on-chip (SoC) design is chiplet-based IP reuse using 2.5-D integration. Complete electronic systems can be created through the integration of chiplets on an interposer, rather than through a monolithic flow. This approach expands access to a large catalog of off-the-shelf intellectual properties (IPs), allows reuse of them, and enables heterogeneous integration of blocks in different technologies. In this article, we present a highly integrated design flow that encompasses architecture, circuit, and package to build and simulate heterogeneous 2.5-D designs. Our target design is 64core architecture based on Reduced Instruction Set Computer (RISC)-V processor. We first chipletize each IP by adding logical protocol translators and physical interface modules. We convert a given register transfer level (RTL) for 64-core processor into chiplets, which are enhanced with our centralized network-onchip. Next, we use our tool to obtain physical layouts, which is subsequently used to synthesize chip-to-chip I/O drivers and these chiplets are placed/routed on a silicon interposer. Our package models are used to calculate power, performance, and area (PPA) and reliability of 2.5-D design. Our design space exploration (DSE) study shows that 2.5-D integration incurs 1.29× power and 2.19× area overheads compared with 2-D counterpart. Moreover, we perform DSE studies for power delivery scheme and interposer technology to investigate the tradeoffs in 2.5-D integrated chip (IC) designs. Gauthaman Murali, Heechun Park, Eric Qin 0001, Hyoukjun Kwon, Venakata Chaitanya Krishna Chekuri, Nael Mizanur Rahman, Nihar Dasari, Minah Lee, Hakki Mert Torun, Kallol Roy, Madhavan Swaminathan, Saibal Mukhopadhyay, Tushar Krishna, Sung Kyu Lim |
IEEE Trans. Very Large Scale Integr. Syst. | 14 |
| 2020 | Low Power Unsupervised Anomaly Detection by Nonparametric Modeling of Sensor StatisticsabstractThis article presents anomaly detection by examining sensor stream statistics (AEGIS), a novel mixed-signal framework for real-time AEGIS. AEGIS utilizes kernel density estimation (KDE)-based nonparametric density estimation to generate a real-time statistical model of the sensor data stream. The likelihood estimate of the sensor data point can be obtained based on the generated statistical model to detect outliers. We present CMOS Gilbert Gaussian cell-based design to realize Gaussian kernels for KDE. For outlier detection, the decision boundary is defined in terms of kernel standard deviation (σKernel) and likelihood threshold (PThres). We adopt a sliding window to update the detection model in real time. We use time-series data set provided from Yahoo to benchmark the performance of AEGIS. A f1-score higher than 0.87 is achieved by optimizing parameters such as length of the sliding window and decision thresholds which are programmable in AEGIS. Discussed architecture is designed using 45-nm technology node and our approach on average consumes ~75-μW power at a sampling rate of 2 MHz while using ten recent inlier samples for density estimation. Ahish Shylendra, Priyesh Shukla, Saibal Mukhopadhyay, Swarup Bhunia, Amit Ranjan Trivedi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | A Spatiotemporal Pre-processing Network for Activity Recognition under Rain
Minah Lee, Burhan Ahmad Mudassar, Taesik Na, Saibal Mukhopadhyay |
BMVC | 4 |
| 2019 | Rethinking Convolutional Feature Extraction for Small Object Detection
Burhan Ahmad Mudassar, Saibal Mukhopadhyay |
BMVC | 2 |
| 2019 | Architecture, Chip, and Package Co-design Flow for 2.5D IC Design Enabling Heterogeneous IP ReuseabstractA new trend in complex SoC design is chiplet-based IP reuse using 2.5D integration. In this paper we present a highly-integrated design flow that encompasses architecture, circuit, and package to build and simulate heterogeneous 2.5D designs. We chipletize each IP by adding logical protocol translators and physical interface modules. These chiplets are placed/routed on a silicon interposer next. Our package models are then used to calculate PPA and signal/power integrity of the overall system. Our design space exploration study using our tool flow shows that 2.5D integration incurs 2.1x PPA overhead compared with 2D SoC counterpart. Gauthaman Murali, Heechun Park, Eric Qin 0001, Hyoukjun Kwon, Venakata Chaitanya Krishna Chekuri, Nihar Dasari, Minah Lee, Hakki Mert Torun, Kallol Roy, Madhavan Swaminathan, Saibal Mukhopadhyay, Tushar Krishna, Sung Kyu Lim |
DAC | 13 |
| 2019 | RTL-to-GDS Tool Flow and Design-for-Test Solutions for Monolithic 3D ICsabstractMonolithic 3D IC overcomes the limitation of the existing through-silicon-via (TSV) based 3D IC by providing denser vertical connections with nano-scale inter-layer vias (ILVs). In this paper, we demonstrate a thorough RTL-to-GDS design flow for monolithic 3D IC, which is based on commercial 2D place-and-route (P&R) tools and clever ways to extend them to handle 3D IC designs and simulations. We also provide a low-cost built-in-self-test (BIST) method to detect various faults that can occur on ILVs. Lastly, we present a resistive random access memory (ReRAM) compiler that generates memory modules that are to be integrated in monolithic 3D ICs. Heechun Park, Kyungwook Chang, Bon Woong Ku, Daehyun Kim 0002, Arjun Chaudhuri, Sanmitra Banerjee, Saibal Mukhopadhyay, Krishnendu Chakrabarty, Sung Kyu Lim |
DAC | 9 |
| 2019 | Design of Reliable DNN Accelerator with Un-reliable ReRAMabstractThis paper presents an algorithmic approach to design reliable ReRAM based Processing-in-Memory (PIM) architecture for Deep Neural Network (DNN) acceleration under intrinsic stochastic behavior of ReRAM devices. We employ the dynamical fixed point (DFP) data representation format to adaptively change the decimal point location based on the data range, minimizing the unused most significant bits (MSBs). Further, we propose a device variability aware (DVA) training methodology where stochastic noise is added to the parameters during training to enhance the robustness of network to the parameter's variation. Simulations indicate that, on average, the proposed algorithms improve the computing accuracy by more than 20% considering various benchmark DNNs (convolutional and recurrent). Moreover, the proposed approach enhances robustness of the DNN to noisy input data. Xueyuan She, Saibal Mukhopadhyay |
DATE | 3 |
| 2019 | A Camera with Brain - Embedding Machine Learning in 3D SensorsabstractThe cameras today are designed to capture signals with highest possible accuracy to most faithfully represent what it sees. However, many mission-critical autonomous applications ranging from traffic monitoring to disaster recovery to defense requires quality of information, where useful information depends on the tasks and is defined using complex features, rather than only changes in captured signal. Such applications require cameras that capture useful information from a scene with highest quality while meeting system constraints such as power, performance, and bandwidth. This paper will discuss the feasibility of a camera that learns how to capture task-dependent information with highest quality, paving the pathway to design a camera with brain. 3D integration of digital pixel sensors with massively parallel computing platform for machine learning creates a hardware architecture for such a camera. The paper will discuss embedded machine learning algorithms that can run on such platform to enhance quality of useful information by real-time control of the sensor parameters. We conclude by identifying critical challenges as well as opportunities for hardware and algorithmic innovations to enable machine learning in the feedback loop of a 3D image sensor based camera. Burhan Ahmad Mudassar, Priyabrata Saha, Mohammad Faisal Amir, Evan Gebhardt, Taesik Na, Jong Hwan Ko, Marilyn Wolf, Saibal Mukhopadhyay |
DATE | 9 |
| 2019 | Fast and Low-Precision Learning in GPU-Accelerated Spiking Neural NetworkabstractSpiking neural network (SNN) uses biologically inspired neuron model coupled with Spike-timing-dependent-plasticity (STDP) to enable unsupervised continuous learning in artificial intelligence (AI) platform. However, current SNN algorithms shows low accuracy in complex problems and are hard to operate at reduced precision. This paper demonstrates a GPU-accelerated SNN architecture that uses stochasticity in the STDP coupled with higher frequency input spike trains. The simulation results demonstrate 2 to 3 times faster learning compared to deterministic SNN architectures while maintaining high accuracy for MNIST (simple) and fashion MNIST (complex) data sets. Further, we show stochastic STDP enables learning even with 2 bits of operation, while deterministic STDP fails. Xueyuan She, Saibal Mukhopadhyay |
DATE | 3 |
| 2019 | Mitigating Power Supply Glitch based Fault Attacks with Fast All-Digital Clock Modulation CircuitabstractThis paper experimentally demonstrates that an on-chip integrated fast all-digital clock modulation (F-ADCM) circuit can be used as a countermeasure against supply glitch and temperature variations-based fault injection attacks (FIA). The F-ADCM circuit modulates clock edges in presence of DC/transient supply glitches and temperature variations to ensure correct operation of the underlying cryptographic circuit. With a testchip manufactured in 130nm CMOS process, we first demonstrate an inexpensive methodology to conduct a fault attack on hardware implementation of a 128-bit advanced encryption standard (AES) engine using externally controlled supply glitches. Next, we show that with F-ADCM circuit, it is no longer possible to inject supply/temperature glitch-based faults even after 10 million encryptions across varying operating conditions. Moreover, in extreme operating conditions, the F-ADCM circuit doesn't generate any clock edges, leading to complete failure of the AES encryption, indicating no exploitable faults are present. Monodeep Kar, Nikhil Chawla, Saibal Mukhopadhyay |
DATE | 4 |
| 2019 | A Spectral Convolutional Net for Co-Optimization of Integrated Voltage Regulators and Embedded InductorsabstractIntegrated voltage regulators (IVR) with embedded inductors is an emerging technology that provides point-of-load voltage regulation to high-performance systems. Conventional two-step approaches to the design of IVRs can suffer from suboptimal design as the optimal inductor depends on the characteristics of the buck converter (BC). Furthermore, inductor-level trade-offs such as AC and DC resistance, inductance and area can not be determined independently from the BC. This co-dependency of the BC and the inductor creates a highly non-linear response surface, which raises the necessity of co-optimization, involving multiple time-consuming electromagnetics (EM) simulations. In this paper, we propose a machine learning based optimization methodology that eliminates EM simulations from the optimization loop to significantly reduce the optimization complexity. A novel technique named as Spectral Transposed Convolutional Neural Network (S-TCNN) is presented to derive an accurate predictive model of the inductor frequency response using a small amount of training data. The derived S-TCNN is then used along with a time-domain model of the BC to perform multi-objective optimization that approximates the Pareto front for 5 objectives, namely inductor area, BC settling time, voltage conversion efficiency, droop and ripple. The resulting methodology provides multiple Pareto optimal inductors in an efficient and fully automated fashion, thereby allows to rapidly determine the optimal trade-offs for possibly contradicting design objectives. We demonstrate the proposed framework on co-optimization of solenoidal inductor with magnetic core and BC that are integrated on silicon interposer. Hakki Mert Torun, Huan Yu 0011, Nihar Dasari, Venakata Chaitanya Krishna Chekuri, Sung Kyu Lim, Saibal Mukhopadhyay, Madhavan Swaminathan |
ICCAD | 8 |
| 2019 | Application Inference using Machine Learning based Side Channel AnalysisabstractThe proliferation of ubiquitous computing requires energy-efficient as well as secure operation of modern processors. Side channel attacks are becoming a critical threat to security and privacy of devices embedded in modern computing infrastructures. Unintended information leakage via physical signatures such as power consumption, electromagnetic emission (EM) and execution time have emerged as a key security consideration for SoCs. Also, information published on purpose at user privilege level accessible through software interfaces results in software-only attacks. In this paper, we used a supervised learning based approach for inferring applications executing on android platform based on features extracted from EM side-channel emissions and software exposed dynamic voltage frequency scaling (DVFS) states. We highlight the importance of machine learning based approach in utilizing these multi-dimensional features on a complex SoC, against profiling-based approaches. We also show that learning the instantaneous frequency states polled from on-board frequency driver (cpufreq) is adequate to identify a known application and flag potentially malicious unknown application. The experimental results on benchmarking applications running on ARMv8 processor in Snapdragon 820 board demonstrates early detection of these apps in 700 ms, and atleast 85% accuracy in detecting unknown applications. Overall, the highlight is to utilize a low-complexity path to application inference attacks through learning instantaneous frequency states pattern of CPU core. Nikhil Chawla, Monodeep Kar, Saibal Mukhopadhyay |
IJCNN | 4 |
| 2019 | FocalNet - Foveal Attention for Post-processing DNN OutputsabstractThis paper presents FocalNet - an iterative information extraction algorithm that uses the concept of foveal attention to post-process the outputs of Deep Neural Networks (DNNs) by performing variable sampling of the input/feature space. FocalNet is integrated into an existing task-driven deep learning model without modifying the weights of the network. Layers, at which to perform foveation are automatically selected using a data-driven approach. We apply FocalNet to the task of object detection using a state of the art convolutional detector, RFCN ResNet-101. On the PASCAL VOC 2007 dataset, we are able to achieve a mAP increase of 3.7%. On the MS COCO 2017 validation dataset we achieve an increase in mAP by 0.3%. Further, a higher increase in mAP is observed for a computationally efficient detector (1.7% for Faster R-CNN with ResNet50). In additon to object detection, we show effectiveness of FocalNet to the problem of single object tracking with a 0.5% increase in average IoU on the MOT17 dataset over a simple tracking by detection approach using DNNs. Burhan Ahmad Mudassar, Saibal Mukhopadhyay |
IJCNN | 2 |
| 2019 | Mixture of Pre-processing Experts Model for Noise Robust Deep Learning on Resource Constrained PlatformsabstractDeep learning on an edge device requires energy efficient operation due to ever diminishing power budget. Intentional low quality data during the data acquisition for longer battery life, and natural noise from the low cost sensor degrade the quality of target output which hinders adoption of deep learning on an edge device. To overcome these problems, we propose simple yet efficient mixture of pre-processing experts (MoPE) model to handle various image distortions including low resolution and noisy images. We also propose to use adversarially trained auto encoder as a pre-processing expert for the noisy images. We evaluate our proposed method for various machine learning tasks including object detection on MS-COCO 2014 dataset, multiple object tracking problem on MOT-Challenge dataset, and human activity classification on UCF 101 dataset. Experimental results show that the proposed method achieves better detection, tracking and activity classification accuracies under noise without sacrificing accuracies for the clean images. The overheads of our proposed MoPE are 0.67% and 0.17% in terms of memory and computation compared to the baseline object detection network. Taesik Na, Minah Lee, Burhan Ahmad Mudassar, Priyabrata Saha, Jong Hwan Ko, Saibal Mukhopadhyay |
IJCNN | 6 |
| 2019 | Improving Robustness of ReRAM-based Spiking Neural Network Accelerator with Stochastic Spike-timing-dependent-plasticityabstractSpike-timing-dependent-plasticity (STDP) is an unsupervised learning algorithm for spiking neural network (SNN), which promises to achieve deeper understanding of human brain and more powerful artificial intelligence. While conventional computing system fails to simulate SNN efficiently, process-in-memory (PIM) based on devices such as ReRAM can be used in designing fast and efficient STDP based SNN accelerators, as it operates in high resemblance with biological neural network. However, the real-life implementation of such design still suffers from impact of input noise and device variation. In this work, we present a novel stochastic STDP algorithm that uses spiking frequency information to dynamically adjust synaptic behavior. The algorithm is tested in pattern recognition task with noisy input and shows accuracy improvement over deterministic STDP. In addition, we show that the new algorithm can be used for designing a robust ReRAM based SNN accelerator that has strong resilience to device variation. Xueyuan She, Saibal Mukhopadhyay |
IJCNN | 3 |
| 2019 | Automatic GDSII Generator for On-Chip Voltage Regulator for Easy Integration in Digital SoCsabstractThis paper demonstrates an electronic design automation (EDA) tool flow for generation of digitally controlled high-bandwidth on-chip voltage regulators for easy integration in a digital SoC. The proposed flow optimizes the control loop and power stage of an integrated voltage regulator (IVR) to achieve desired transient performance and/or efficiency. The GDSII of the optimized IVR is generated by coupling logic synthesis and physical design of digital blocks, automated generation of power stage layout, and top-level integration of all modules. We demonstrate the developed flow for inductive IVRs in 130nm and 65nm CMOS process. We show feasibility of fast design space exploration, optimization, layout generation as well as automated integration of the generated IVR with a RISC-V core. Venakata Chaitanya Krishna Chekuri, Nihar Dasari, Saibal Mukhopadhyay |
ISLPED | 4 |
| 2019 | Multigated Carbon Nanotube Field Effect Transistors-Based Physically Unclonable Functions As Security KeysabstractEnabling data security from unauthorized access is a major challenge for electronics devices. Most of the conventional cryptographic techniques store “keys” in nonvolatile memory, which is vulnerable to external attacks like physical attacks, side-channel attacks, fault attacks, etc. Physically unclonable functions (PUFs) have the potential to overcome these challenges because they do not store keys permanently and are difficult to reproduce. The next generation of electronic and opto-electronic devices may use semiconducting materials like carbon nanotubes (CNTs) or 2-D materials due to their superior electrical, optical, thermal, and mechanical properties. There is a need for PUFs, which are low-cost and more efficient than existing silicon-based PUFs and compatible with future electronic technologies. We propose multigated CNT field effect transistors (CNT-FETs)-based PUFs, where inherent randomness of CNT network and a multigated channel are utilized to generate high-quality random keys. We have shown that while conventional single-gate channel FETs can generate binary keys, multigated CNT-FETs, where different gate voltages are applied in different sections of the channel, can enable the creation of multiple challenges and current levels to produce not only ternary but up to base-17 (heptadecimal) keys. Such keys can create significantly more entropy than binary or ternary keys of the same size generated by typical PUFs. Monodeep Kar, Suresh K. Sitaraman, Saibal Mukhopadhyay |
IEEE Internet Things J. | 5 |
| 2019 | Energy Efficient and Side-Channel Secure Cryptographic Hardware for IoT-Edge NodesabstractDesign of ultralightweight but secure encryption engine is a key challenge for Internet-of-Things edge devices. This paper explores the system level design space for an ultralow power image sensor node for secure communication and proposes an optimized datapath architecture for 128-bit SIMON (SIMON128), a lightweight block cipher, for minimal performance, power, and area overheads with increased level of side-channel security. Various datapath architectures for SIMON are explored for simultaneously increasing energy-efficiency and resistance to power-based side-channel analysis (PSCA) attacks. Alternative datapath architectures are implemented on ASIC (15 nm CMOS) and field programmable gate array (FPGA) (Spartan-6, 45 nm) to perform power, performance, and area analysis. We show that, although a bitserial datapath minimizes area and power, a round unrolled datapath provides 80× higher energy-efficiency and 143× higher performance, compared to the baseline bitserial design. Moreover, the PSCA measurements performed using Sakura-G board with Spartan-6 FPGA, demonstrate that a 6-round unrolled datapath improves minimum-traces-to-disclosure for correlation power analysis (CPA) by at least 384× over baseline bitserial design with no successful CPA even with 500000 measurements. Finally, application to the image-sensor node demonstrates that optimized unrolled SIMON128 can provide equivalent performance to AES128 at lower area, higher energy efficiency, and improved side channel security. Nikhil Chawla, Jong Hwan Ko, Monodeep Kar, Saibal Mukhopadhyay |
IEEE Internet Things J. | 5 |
| 2019 | Design and Analysis of a Neural Network Inference Engine Based on Adaptive Weight CompressionabstractNeural networks generally require significant memory capacity/bandwidth to store/access a large number of synaptic weights. This paper presents design of an energy-efficient neural network inference engine based on adaptive weight compression using a JPEG image encoding algorithm. To maximize compression ratio with minimum accuracy loss, the quality factor of the JPEG encoder is adaptively controlled depending on the accuracy impact of each block. With 1% accuracy loss, the proposed approach achieves 63.4× compression for multilayer perceptron (MLP) and 31.3× for LeNet-5 with the MNIST dataset, and 15.3× for AlexNet and 10.2× for ResNet-50 with ImageNet. The reduced memory requirement leads to higher throughput and lower energy for neural network inference (3× effective memory bandwidth and 22× lower system energy for MLP). Jong Hwan Ko, Duckhwan Kim 0001, Taesik Na, Saibal Mukhopadhyay |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | Autotuning of Integrated Inductive Voltage Regulator Using On-Chip Delay Sensor to Tolerate Process and Passive VariationsabstractThis paper demonstrates autotuning of the coefficients of the feedback loop of an inductive integrated voltage regulator (IVR) using an on-chip delay sensor. The proposed approach improves the effective performance of the digital core under variations in the on-die/package integrated passives and transistor process. A 130-nm CMOS test-chip is designed containing a multisampled 125-MHz IVR with a wirebond inductor, on-die capacitor, and all-digital proportional-integral-differential (PID) controller powering a parallel Advanced Encryption Standard (AES) engine. The autotuning is performed using a Vernier delay line based on-chip delay sensor and an all-digital tuning engine. The measurement results demonstrate up to 5.2% improvement in the maximum operating frequency of the AES core using performance-based autotuning. Venakata Chaitanya Krishna Chekuri, Monodeep Kar, Saibal Mukhopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | Edge-Host Partitioning of Deep Neural Networks with Feature Space Encoding for Resource-Constrained Internet-of-Things PlatformsabstractThis paper introduces partitioning an inference task of a deep neural network between an edge and a host platform in the IoT environment. We present a DNN as an encoding pipeline, and propose to transmit the output feature space of an intermediate layer to the host. Encoding of the feature space is proposed to enhance the maximum input rate supported by the edge platform and/or reduce the energy of the edge platform. Simulation results show that partitioning a DNN coupled with feature space encoding enables significant improvement in the energy-efficiency and throughput over the baseline configurations that perform the entire inference at the edge or at the host. Jong Hwan Ko, Taesik Na, Mohammad Faisal Amir, Saibal Mukhopadhyay |
AVSS | 4 |
| 2018 | Adaptive Control of Camera Modality with Deep Neural Network-Based Feedback for Efficient Object TrackingabstractRound-the-clock surveillance requires robust object detection and tracking independent of lighting conditions. Fusing information from visual-infrared object detection network pair at feature level or decision level shows promising accuracy. However, such fused object detection network is not suitable for edge devices with limited processing power and memory. In this paper, we propose a technique to control spatial modality using feedback from the object detection network and create a mixed-modality image by eliminating the redundancy between visual and infrared information. Mixed-modality image enables object tracking with a single deep neural network as opposed to the decision- level fusion with two separate networks for visual image and infrared image. Proposed approach achieves at least 8% better object tracking accuracy than decision-level fusion while operating at 2X frame-rate and consuming 50% less energy. Priyabrata Saha, Burhan Ahmad Mudassar, Saibal Mukhopadhyay |
AVSS | 3 |
| 2018 | Edge-cloud collaborative processing for intelligent internet of things: a case study on smart surveillanceabstractLimited processing power and memory prevent realization of state of the art algorithms on the edge level. Offloading computations to the cloud comes with tradeoffs as compression techniques employed to conserve transmission bandwidth and energy adversely impact accuracy of the algorithm. In this paper, we propose collaborative processing to actively guide the output of the sensor to improve performance on the end application. We apply this methodology to smart surveillance specifically the task of object detection from video. Perceptual quality and object detection performance is characterized and improved under a variety of channel conditions. Burhan Ahmad Mudassar, Jong Hwan Ko, Saibal Mukhopadhyay |
DAC | 3 |
| 2018 | Performance based tuning of an inductive integrated voltage regulator driving a digital core against process and passive variationsabstractThis paper presents an auto-tuning method for fully integrated voltage regulators (IVRs) driving digital cores against variations in passive as well as process/temperature of the core. The key contribution is to perform auto-tuning of the coefficients of the feedback loop of the IVR based on the performance of the digital cores. Simulations using a high-frequency IVR Simulink model and digital logic in 45nm CMOS process shows that the proposed performance driven auto-tuning demonstrates potential for up to 12% increase in system performance under inductance and threshold variation. Venakata Chaitanya Krishna Chekuri, Monodeep Kar, Saibal Mukhopadhyay |
DATE | 4 |
| 2018 | Accelerating biophysical neural network simulation with region of interest based approximationabstractModeling the dynamics of biophysical neural network (BNN) is essential to understand brain operation and design cognitive systems. Large-scale and biophysically plausible BNN modeling requires solving multiple-terms, coupled and non-linear differential equations, making simulation computationally complex and memory intensive. This paper presents an adaptive simulation methodology in which neurons in the region of interest (ROI) follow high biological accurate models while the other neurons follow computation friendly models. To enable ROI based approximation, we propose a generic template based computing algorithm which unifies the data structure and computing flow for various neuron models. We implement the algorithms on CPU, GPU and embedded platforms, showing llx speedup with insignificant loss of biological details in the region of interest. Xueyuan She, Saibal Mukhopadhyay |
DATE | 3 |
| 2018 | The CAMEL approach to stacked sensor smart camerasabstractStacked image sensor systems combine an image sensor, memory, and processors using 3D technology. Stacking camera components that have traditionally been packaged separately provides several benefits: very high bandwidth out of the image sensor, allowing for higher frame rates; very low latency, providing opportunities for image processing and computer vision algorithms which can adapt at very high rates; and lower power consumption. This paper will review the characteristics of stacked image sensor systems and discuss novel algorithmic and systems concepts that are made possible by these stacked sensors. Saibal Mukhopadhyay, Marilyn Wolf, Mohammed Faisal Amir, Evan Gebhardt, Jong Hwan Ko, Jaeha Kung 0001, Burhan Ahmad Mudassar |
DATE | 1 |
| 2018 | Exploiting on-chip power management for side-channel securityabstractThe high-performance and energy-efficient encryption engines have emerged as a key component for modern System-On-Chip (SoC) in various platforms including servers, desktops, mobile, and IoT edge devices. A key bottleneck to secure operation of encryption engines is leakage of information through various side-channels. For example, an adversary can extract the secret key by performing statistical analysis on measured power and electromagnetic (EM) emission signatures generated by the hardware during encryption. Countermeasures to such side-channel attacks often come at high power, area, or performance overheads. Therefore, design of side-channel secure encryption engines is a critical challenge for high-performance and/or power-/energy efficient operations. This paper reviews that although low-power requirement imposes critical challenge for side-channel security, but circuit techniques traditionally developed for power management also present new opportunities for side-channel resistance. As a case study, we review the feasibility of using integrated voltage regulator and dynamic voltage frequency scaling normally used for efficient power management, for increasing power-side-channel resistance of AES engines. The hardware measurement results from test-chip fabricated in 130nm process are presented to demonstrate the impact of power management circuits on side-channel security. Monodeep Kar, Sanu Mathew, Anand Rajan, Vivek De, Saibal Mukhopadhyay |
DATE | 6 |
| 2018 | An Unsupervised Anomalous Event Detection Framework with Class Aware Source SeparationabstractThis paper presents a novel problem of detection and localization of anomalous events due to a certain class of objects in video data with applications to smart surveillance. A baseline system is proposed that uses a convolutional neural network (CNN) to generate pixel level masks corresponding to objects of a class of interest. A Restricted Boltzmann Machine (RBM) is then trained on the mask to learn patterns of normal behavior. The free energy of the RBM is used to detect the presence of an anomaly while the reconstruction error is used to localize the anomaly. Our approach is scalable to a low power and energy constrained setting with 1930.48 ms of latency and 4826 mJ energy consumed per frame on a mGPU. Burhan Ahmad Mudassar, Jong Hwan Ko, Saibal Mukhopadhyay |
ICASSP | 3 |
| 2018 | A ferroelectric FET based power-efficient architecture for data-intensive computingabstractIn this paper, we present a ferroelectric FET (FeFET) based power-efficient architecture to accelerate data-intensive applications such as deep neural networks (DNNs). We propose a cross-cutting solution combining emerging device technologies, circuit optimizations, and micro-architecture innovations. At device level, FeFET crossbar is utilized to perform vector-matrix multiplication (VMM). As a field effect device, FeFET significantly reduces the read/write energy compared with the resistive random-access memory (ReRAM). At circuit level, we propose an all-digital peripheral design, reducing the large overhead introduced by ADC and DAC in prior works. In terms of micro-architecture innovation, a dedicated hierarchical network-on-chip (H-NoC) is developed for input broadcasting and on-the-fly partial results processing, reducing the data transmission volume and latency. Speed, power, area and computing accuracy are evaluated based on detailed device characterization and system modeling. For DNN computing, our design achieves 254x and 9.7x gain in power efficiency (GOPS/W) compared to GPU and ReRAM based designs, respectively. Taesik Na, Prakshi Rastogi, Karthik Rao, Asif Islam Khan, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
ICCAD | 7 |
| 2018 | Cascade Adversarial Machine Learning Regularized with a Unified Embedding
Taesik Na, Jong Hwan Ko, Saibal Mukhopadhyay |
ICLR (Poster) | 3 |
| 2018 | DeepTrain: A Programmable Embedded Platform for Training Deep Neural NetworksabstractThis paper presents, DeepTrain, an embedded platform for high-performance and energy-efficient training of deep neural network (DNN). The key architectural concept of DeepTrain is to develop a spatially homogeneous computing (and memory) fabric with temporally heterogeneous programmable data flows to optimize memory mapping and data reuse during different phases of training operation.The DeepTrain is demonstrated as an in-memory accelerator integrated in the logic layer of a 3-D memory module. A programming model and supporting architecture utilizes the flexible data flow to efficiently accelerate training of various types of DNNs. The cycle level simulation and synthesized design in 15 nm FinFET shows power efficiency of 500 GFLOPS/W, and almost similar throughput for a wide range of DNNs, including convolutional, recurrent, and mixed (CNN+RNN) networks. Duckhwan Kim 0001, Taesik Na, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Adaptive Precision Cellular Nonlinear Network
Jaeha Kung 0001, Duckhwan Kim 0001, Saibal Mukhopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | ReRAM-Based Processing-in-Memory Architecture for Recurrent Neural Network AccelerationabstractWe present a recurrent neural network (RNN) accelerator design with resistive random-access memory (ReRAM)-based processing-in-memory (PIM) architecture. Distinguished from prior ReRAM-based convolutional neural network accelerators, we redesign the system to make it suitable for RNN acceleration. We measure the system throughput and energy efficiency with the detailed circuit and device characterization. Reprogrammability is enabled with our design, and an RNN friendly pipeline is employed to increase the system throughput. We observe that on average the proposed system achieves 79× improvement of computing efficiency compared with graphics processing unit baseline. Our simulation also indicates that to maintain high accuracy and computing efficiency, the read noise standard deviation should be less than 0.2, the device resistance should be at least 1 MQ, and the device writes latency should be minimized. Taesik Na, Saibal Mukhopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Design of an Energy-Efficient Accelerator for Training of Convolutional Neural Networks using Frequency-Domain ComputationabstractConvolutional neural networks (CNNs) require high computation and memory demand for training. This paper presents the design of a frequency-domain accelerator for energy-efficient CNN training. With Fourier representations of parameters, we replace convolutions with simpler pointwise multiplications. To eliminate the Fourier transforms at every layer, we train the network entirely in the frequency domain using approximate frequency-domain nonlinear operations. We further reduce computation and memory requirements using sinc interpolation and Hermitian symmetry. The accelerator is designed and synthesized in 28nm CMOS, as well as prototyped in an FPGA. The simulation results show that the proposed accelerator significantly reduces training time and energy for a target recognition accuracy. Jong Hwan Ko, Burhan Ahmad Mudassar, Taesik Na, Saibal Mukhopadhyay |
DAC | 4 |
| 2017 | Adaptive weight compression for memory-efficient neural networksabstractNeural networks generally require significant memory capacity/bandwidth to store/access a large number of synaptic weights. This paper presents an application of JPEG image encoding to compress the weights by exploiting the spatial locality and smoothness of the weight matrix. To minimize the loss of accuracy due to JPEG encoding, we propose to adaptively control the quantization factor of the JPEG algorithm depending on the error-sensitivity (gradient) of each weight. With the adaptive compression technique, the weight blocks with higher sensitivity are compressed less for higher accuracy. The adaptive compression reduces memory requirement, which in turn results in higher performance and lower energy of neural network hardware. The simulation for inference hardware for multilayer perceptron with the MNIST dataset shows up to 42X compression with less than 1% loss of recognition accuracy, resulting in 3X higher effective memory bandwidth and ~19X lower system energy. Jong Hwan Ko, Duckhwan Kim 0001, Taesik Na, Jaeha Kung 0001, Saibal Mukhopadhyay |
DATE | 5 |
| 2017 | Clock data compensation aware clock tree synthesis in digital circuits with adaptive clock generationabstractAdaptive clock generation to track critical path delay enables lowering supply voltage with improved timing slack under supply noise. This paper presents how to synthesize clock tree in adaptive clocking to fully exploit the clock data compensation (CDC) effect in digital circuits. The paper first provides analytical proof of ideal CDC effect for ring oscillator based clock generation. Second, the paper analyzes non-ideal CDC effect in a gate dominated critical path and wire dominated clock tree design. The paper shows the delay sensitivity mismatch between clock tree and critical path can degrade CDC effect by analyzing timing slack under power supply noise (PSN). Finally, the paper proposes simple but efficient clock tree synthesis (CTS) technique to maximize timing slack under PSN in digital circuits with adaptive clock generation. Taesik Na, Jong Hwan Ko, Saibal Mukhopadhyay |
DATE | 3 |
| 2017 | On-chip training of recurrent neural networks with limited numerical precisionabstractTraining of neural network can be accelerated by limited numerical precision together with specialized low-precision hardware. This paper studies how low precision can impact on entire training of RNNs. We emulate low precision training for recently proposed gated recurrent unit (GRU) and use dynamic fixed point as a target numeric format. We first show that batch normalization on input sequences can help speed up training with low precision as well as high precision. We also show that the overflow rate should be carefully controlled for dynamic fixed point. We study low precision training with various rounding options including bit truncation, round to nearest, and stochastic rounding. Stochastic rounding shows superior results than the other options. The effect of fully low precision training is also analyzed by comparing partial low precision training. We show that the piecewise linear activation function with stochastic rounding can achieve comparable training results with floating point precision. Low precision multiplier and accumulator (MAC) with linear-feedback shift register (LFSR) is implemented with 28nm Synopsys PDK for energy and performance analysis. Implementation results show low precision hardware is 4.7× faster, and energy per task is up to 4.55× lower than that of floating point hardware. Taesik Na, Jong Hwan Ko, Jaeha Kung 0001, Saibal Mukhopadhyay |
IJCNN | 4 |
| 2017 | A Programmable Hardware Accelerator for Simulating Dynamical SystemsabstractThe fast and energy-efficient simulation of dynamical systems defined by coupled ordinary/partial differential equations has emerged as an important problem. The accelerated simulation of coupled ODE/PDE is critical for analysis of physical systems as well as computing with dynamical systems. This paper presents a fast and programmable accelerator for simulating dynamical systems. The computing model of the proposed platform is based on multilayer cellular nonlinear network (CeNN) augmented with nonlinear function evaluation engines. The platform can be programmed to accelerate wide classes of ODEs/PDEs by modulating the connectivity within the multilayer CeNN engine. An innovative hardware architecture including data reuse, memory hierarchy, and near-memory processing is designed to accelerate the augmented multilayer CeNN. A dataflow model is presented which is supported by optimized memory hierarchy for efficient function evaluation. The proposed solver is designed and synthesized in 15nm technology for the hardware analysis. The performance is evaluated and compared to GPU nodes when solving wide classes of differential equations and the power consumption is analyzed to show orders of magnitude improvement in energy efficiency. Jaeha Kung 0001, Duckhwan Kim 0001, Saibal Mukhopadhyay |
ISCA | 4 |
| 2017 | Invited paper: Low power requirements and side-channel protection of encryption engines: Challenges and opportunitiesabstractPower attack is a critical challenge to security of encryption engines. Countermeasures to side-channel attacks often come at high power, area, or performance overhead. Therefore, design of side-channel secure encryption engines is a critical challenge for power-/resource-constrained platforms. This paper discusses that although low-power need imposes critical challenge for side-channel security, but circuit techniques traditionally developed for power management also present new opportunities for side-channel resistance. As a case-study, we show the feasibility of using integrated voltage regulator, normally used for efficient power management, for increasing side-channel resistance of AES engines. Monodeep Kar, Sanu Mathew, Anand Rajan, Vivek De, Saibal Mukhopadhyay |
ISLPED | 6 |
| 2016 | (Invited paper) energy delivery for self-powered IoT devicesabstractDistributed small-scale electronics for IoT applications are on the rise. Power delivery for such electronics requires innovative design techniques to improve energy efficiency. This paper summarizes energy delivery challenges for IoT devices and discusses several design techniques for efficient power delivery units. Such design solutions cover challenges like energy harvesting from very low input voltage, maximized energy harvesting, energy delivery with multiple voltage domains and design using low voltage devices to sustain higher than breakdown voltages. Zakir K. Ahmed 0001, Monodeep Kar, Saibal Mukhopadhyay |
ASP-DAC | 3 |
| 2016 | An energy-efficient wireless video sensor node with a region-of-interest based multi-parameter rate controller for moving object surveillanceabstractThis paper presents a lightweight video sensor node for moving object surveillance using region-of-interest (ROI) based coding and an on-line multi-parameter rate controller. The proposed ROI-based coding scheme determines ROI blocks, pre-processes non-ROI blocks using bit-truncation, and encodes all blocks using Motion JPEG. The on-line rate controller modulates the parameters of the ROI-based coding scheme to match the encoded data rate and transmission data rate under the variations in channel bandwidth and input video content. The low-complexity hardware of the ROI-based coding scheme reduces computation energy, and the on-line rate controller minimizes buffer requirement. The sensor node is designed in 130nm CMOS and prototyped in a Virtex-V FPGA. Simulations show that, under the same ROI quality, the proposed approach reduces system energy by 61% compared to H.264/AVC. Jong Hwan Ko, Taesik Na, Saibal Mukhopadhyay |
AVSS | 3 |
| 2016 | Behavioral modeling of timing slack variation in digital circuits due to power supply noise
Taesik Na, Saibal Mukhopadhyay |
DATE | 2 |
| 2016 | Amdahl's law for lifetime reliability scaling in heterogeneous multicore processorsabstractHeterogeneous multicore processors have been suggested as alternative microarchitectural designs to enhance performance and energy efficiency. Using Amdahl's Law, heterogeneous models were primarily analyzed in performance and energy efficiency aspects to demonstrate its advantage over conventional homogeneous systems. In this paper, we further extend the study to understand the lifetime reliability consequences of heterogeneous multicore processors, as reliability becomes an increasingly important constraint. We present the lifetime reliability models of multicore processors based on Amdahl's Law, including compact thermal estimation that has strong correlation with device aging. Lifetime reliability is analyzed by varying i) core utilization (Amdahl's scaling factor), ii) processor composition (number of big and small cores), and iii) thread scheduling method. The study shows that the heterogeneous processor may have a serious reliability challenge. If the processor is comprised of only one big core and many small cores, stresses can be biased to the big core especially when workloads spend more time on sequential operations. Our study reveals that incorporating multiple big cores can mitigate reliability bottleneck in big cores and enhance processor lifetime, but adding too many big cores will have an adverse impact on lifetime reliability as well as performance. William J. Song, Saibal Mukhopadhyay, Sudhakar Yalamanchili |
HPCA | 2 |
| 2016 | A single-inductor-cascaded-stage topology for high conversion ratio boost regulatorabstractA single-inductor-cascaded-stage boost regulator topology is presented that time-multiplexes a single inductor using one-nFET-two-pFET power stage and a bias-gated Pulse-Frequency Modulation controller to achieve high conversion ratio. A test-chip in 130nm CMOS demonstrates 120× conversion using a single inductor while consuming 140nA bias current. Zakir K. Ahmed 0001, Saibal Mukhopadhyay |
ICCD | 2 |
| 2016 | What does ultra low power requirements mean for side-channel secure cryptography?abstractThe design of low power and side-channel-attack resistant encryption engine is a key challenge to enhance security of resource-constrained platforms. This paper present case studies to show that the low-power requirement is a challenge as well as an opportunity for improving side-channel resistance. On one hand, low-power encryption architecture can be more vulnerable to power-attack; and the countermeasures comes with significant overhead. However, on the other hand, low-power circuit techniques such as integrated voltage regulation or adaptive clocking can also be exploited to improve power-attack resistance. The analysis shows the need for future research on low-power and side-channel secure cryptography. Monodeep Kar, Anand Rajan, Vivek De, Saibal Mukhopadhyay |
ICCD | 5 |
| 2016 | ReRAM Crossbar based Recurrent Neural Network for human activity detectionabstractWe present a programmable high-efficient Recurrent Neural Network (RNN) with Synapses design using Resistive Random Access Memory (ReRAM). The presented ReRAM-RNN employs crossbar ReRAM arrays as synapses. A fast synapses programming methodology is realized by CMOS-based neuron with in-built programming circuitry. The simulations are performed using experimentally verified physical resistive switching model, instead of only functional models, providing better estimate of system speed and power efficiency. Simulation results show that ReRAM-RNN can provide higher computation efficiency and/or more compact design than software realization of RNN, and dedicated CMOS based digital-and analog-RNN. We show that the efficiency improvement of ReRAM-based neural network design is more significant in feedback networks than in feedforward networks. Eui Min Jung, Jaeha Kung 0001, Saibal Mukhopadhyay |
IJCNN | 4 |
| 2016 | Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D MemoryabstractThis paper presents a programmable and scalable digital neuromorphic architecture based on 3D high-density memory integrated with logic tier for efficient neural computing. The proposed architecture consists of clusters of processing engines, connected by 2D mesh network as a processing tier, which is integrated in 3D with multiple tiers of DRAM. The PE clusters access multiple memory channels (vaults) in parallel. The operating principle, referred to as the memory centric computing, embeds specialized state-machines within the vault controllers of HMC to drive data into the PE clusters. The paper presents the basic architecture of the Neurocube and an analysis of the logic tier synthesized in 28nm and 15nm process technologies. The performance of the Neurocube is evaluated and illustrated through the mapping of a Convolutional Neural Network and estimating the subsequent power and performance for both training and inference. Duckhwan Kim 0001, Jaeha Kung 0001, Sek M. Chai, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
ISCA | 5 |
| 2016 | Exploiting Fully Integrated Inductive Voltage Regulators to Improve Side Channel Resistance of Encryption EnginesabstractThis paper explores fully integrated inductive voltage regulators (FIVR) as a technique to improve the side channel resistance of encryption engines. We propose security aware design modes for low passive FIVR to improve robustness of an encryption-engine against statistical power attacks in time and frequency domain. A Correlation Power Analysis is used to attack a 128-bit AES engine synthesized in 130nm CMOS. The original design requires ~250 Measurements to Disclose (MTD) the 1st byte of key; but with security-aware FIVR, the CPA was unsuccessful even after 20,000 traces. We present a reversibility based threat model for the FIVR-based protection improvement and show the robustness of security aware FIVR against such threat. Monodeep Kar, Sanu Mathew, Anand Rajan, Vivek De, Saibal Mukhopadhyay |
ISLPED | 6 |
| 2016 | An Energy-Aware Approach to Noise-Robust Moving Object Detection for Low-Power Wireless Image Sensor PlatformsabstractThis paper presents an energy-aware approach to moving object detection that requires very low computation and memory while ensuring robust performance under noisy environments. The proposed approach is integrated into a wireless image sensor platform with a block-based processing unit and the motion JPEG encoder. The sensor platform is designed as an ASIC in 130nm CMOS for energy/area analysis, as well as prototyped into Virtex-5 FPGA for functional validation. The sensor platform designed with the proposed approach consumes less energy and area than the platforms with the existing methods such as Gaussian Mixture Model, while maintaining a reliable delivery of region-of-interest. Jong Hwan Ko, Saibal Mukhopadhyay |
ISLPED | 2 |
| 2016 | Dynamic Approximation with Feedback Control for Energy-Efficient Recurrent Neural Network HardwareabstractThis paper presents methodology of feedback-controlled dynamic approximation to enable energy-accuracy trade-off in digital recurrent neural network (RNN). A low-power digital RNN engine is presented that employs the proposed dynamic approximation. The on-chip feedback controller is realized by utilizing hysteretic or proportional controller. The dynamic adaptation of bit-precisions during the RNN computation is selected as approximation approach. Considering various applications, the digital RNN engine designed in 28nm CMOS shows ~36% average energy saving compared to the baseline case, with only ~4% of accuracy degradation on average. Jaeha Kung 0001, Duckhwan Kim 0001, Saibal Mukhopadhyay |
ISLPED | 3 |
| 2016 | Speeding up Convolutional Neural Network Training with Dynamic Precision Scaling and Flexible Multiplier-AccumulatorabstractTraining convolutional neural network is a major bottleneck when developing a new neural network topology. This paper presents a dynamic precision scaling (DPS) algorithm and flexible multiplier-accumulator (MAC) to speed up convolutional neural network training. The DPS algorithm utilizes dynamic fixed point and finds good enough numerical precision for target network while training. The precision information from DPS is used to configure our proposed MAC. The proposed MAC can perform fixed point computation with variable precision mode providing differentiated computation time which enables speeding up training for lower precision computation. Simulation results show that our work can achieve 5.7x speed-up while consuming 31% energy compared to baseline for modified Alexnet on Flickr image style recognition task. Taesik Na, Saibal Mukhopadhyay |
ISLPED | 2 |
| 2016 | Partitioning Methods for Interface Circuit of Heterogeneous 3-D-ICs Under Process VariationabstractThis paper presents the design of tier-to-tier interface circuits for 3-D-ICs, where different tiers may operate at different voltages and/or frequencies. The design and partitioning methodologies for the tier-to-tier interface circuit are discussed. The footprint, power, and performance of the interface are analyzed considering the effects of tier-to-tier process variations in 3-D-ICs and technology scaling. The simulation results show that dividing the interface circuit evenly between two tiers reduces footprint but increases power dissipation. For heterogeneous systems with different voltages for reading and writing tiers, dividing the interface between tiers provides better performance than the worst case scenario. On the other hand, placing the interface circuit in the reading tier maximizes throughput for a homogeneous system where both tiers operate at the same voltage. In advanced CMOS nodes, placing interface circuit in the reading tier is a better option due to high delay of the level shifters. Duckhwan Kim 0001, Saibal Mukhopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Optimization of FinFET-based circuits using a dual gate pitch techniqueabstractSource/drain stressors in FinFET-based circuits lose their effectiveness at smaller contacted gate pitches. To improve circuit performance, a dual gate pitch technique is proposed in this work, where standard cells with double the gate pitch are selectively used on the gates of the circuit critical paths, at minimal area and power costs. A stress-aware library characterization is performed for FinFET-based standard cells by obtaining stress distributions using finite element simulations on a subset of structures. The stresses are then employed to create look-up tables for mobility multipliers and threshold voltage shifts, for subsequent performance characterization of FinFET-based standard cells. Finally, a circuit delay optimizer is applied using the dual gate pitch approach and is compared with an alternative gate sizing approach. Using a combination of gate sizing and the dual gate pitch approach, it is shown that the average power delay product improves by 12.9% and 15.9% in 14nm and 10nm technologies, respectively. Sravan K. Marella, Amit Ranjan Trivedi, Saibal Mukhopadhyay, Sachin S. Sapatnekar |
ICCAD | 3 |
| 2015 | A power-aware digital feedforward neural network platform with backpropagation driven approximate synapsesabstractThis paper proposes a power-aware digital feedforward neural network platform that utilizes the backpropagation algorithm during training to enable energy-quality trade-off. Given a quality constraint, the proposed approach identifies a set of synaptic weights for approximation in a neural network. The approach selects synapses with small impact on output error, estimated by the backpropagation algorithm, for approximation. The approximations are achieved by a coupled software (reduced bit-width) and hardware (approximate multiplication in the processing engine) based design approaches. The full-chip design in 130nm CMOS shows, compared to a baseline accurate design, the proposed approach reduces system power by ~38% with 0.4% lower recognition accuracy in a classification problem. Jaeha Kung 0001, Duckhwan Kim 0001, Saibal Mukhopadhyay |
ISLPED | 3 |
| 2015 | Exploring power attack protection of resource constrained encryption engines using integrated low-drop-out regulatorsabstractThe power attack protection of encryption engines often comes at the expense of area, power, and/or performance overheads making the design of a low-power and compact but secure encryption engine challenging. This paper explores the feasibility of using an on-chip low dropout regulator (LDO) as a countermeasure to power attack of low-power and compact encryption engine. We design an area minimized implementation of Advanced Encryption Standard (AES) using predictive 45nm node and show that lightweight implementations are more susceptible to power attack. Using behavioral modeling, we show that an on-chip LDO can enhance power attack resistance of this compact AES engine; however, the tradeoff between LDO performance and power attack protection is essential. Our analysis shows that LDO can increase power attack resistance of the compact AES by >800X with marginal area (1.4%) and power (5%) overheads. Monodeep Kar, Jong Hwan Ko, Saibal Mukhopadhyay |
ISLPED | 4 |
| 2015 | Experimental characterization of in-package microfluidic cooling on a System-on-ChipabstractThis paper, for the first time, experimentally demonstrated the in-package microfluidic cooling on a commercial System-on-Chip (SoC). The pinfin interposer attached to the commercial SoC achieved energy efficient cooling for the SPLASH-2 benchmark suite in measurement. The low-power piezoelectric pump controlled by the SoC ensures the thermal integrity and reduces the system-level energy consumption through leakage reduction. The measurements demonstrated that the in-package fluidic cooling improves the SoC's energy-efficiency and reduces design footprint compared to the external passive cooling. Wen Yueh, Zhimin Wan, Yogendra Joshi, Saibal Mukhopadhyay |
ISLPED | 4 |
| 2015 | On the Impact of Energy-Accuracy Tradeoff in a Digital Cellular Neural Network for Image ProcessingabstractThis paper studies the opportunities of energy-accuracy tradeoff in cellular neural network (CNN). Algorithmic characteristics of CNN is coupled with hardware-induced error distribution of a digital CNN cell to evaluate energy-accuracy tradeoff for simple image processing tasks as well as a complex application. The analysis shows that errors modulate the cell dynamics and propagate through the network degrading the output quality and increasing the convergence time. The error propagation is determined by the task being performed by the CNN, specifically, the strength of the feedback template. Controlling precision is observed to be a more effective approach for energy-accuracy tradeoff in CNN than voltage over scaling. Jaeha Kung 0001, Duckhwan Kim 0001, Saibal Mukhopadhyay |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2014 | On the Design of Reliable 3D-ICs Considering Charged Device Model ESD Events During Die StackingabstractThis paper studies charged device model electrostatic discharge (CDM-ESD) events in die stacking process of 3D-ICs and investigates CDM-ESD protection circuits for individual TSVs to prevent high voltage stress on transistor connected to TSV. The models for power, area, delay, and signal integrity of TSVs considering ESD protection are presented. The models are used to drive a methodology to design reliable 3D-ICs considering CDM-ESD while minimizing the overheads. We study the impact of ESD protection on die-to-die asynchronous interface circuit. Duckhwan Kim 0001, Saibal Mukhopadhyay |
DAC | 2 |
| 2014 | Ultra-low power electronics with Si/Ge tunnel FETabstractSi/Ge Tunnel FET (TFET) with its subthermal subthreshold swing is attractive for low power analog and digital designs. Greater Ion/Ioffratio of TFET can reduce the dynamic power in digital designs, while higher gm/IDScan lower the bias power of analog amplifier. However, the above benefits of TFET are eclipsed by MOSFET at a higher power/performance point. Ultra low power scalability of the key analog and digital circuits, SRAM and operational transconductance amplifier (OTA), with TFET is demonstrated. Analyzing a TFET based cellular neural network, this work shows the feasibility of ultra-low-power neuromorphic computing with TFET. Amit Ranjan Trivedi, Mohammad Faisal Amir, Saibal Mukhopadhyay |
DATE | 3 |
| 2014 | An on-chip autonomous thermoelectric energy management system for energy-efficient active coolingabstractThis paper presents an on-chip thermoelectric (TE) energy management system for energy-efficient on-demand active cooling of integrated circuits. Embedding a TE module (TEM) within the package has shown potential for on-demand cooling of integrated circuits (ICs); however, the additional cooling energy limits the effectiveness of TE coolers (TEC). The proposed on-chip system monitors the IC temperature and provides cooling during critical thermal events by operating the TEM in the Peltier mode. During normal operation, the TEM is operated in the Seebeck mode to harvest the otherwise wasted heat energy generated by the IC and reduce the net cooling energy. A boost regulator harvests energy in an output capacitor and a programmable current source controls the cooling. The design is implemented in a 130nm CMOS test-chip, and tested with an external thermoelectric device. Borislav Alexandrov, Zakir K. Ahmed 0001, Saibal Mukhopadhyay |
ISLPED | 3 |
| 2014 | Impact of process variation in inductive integrated voltage regulator on delay and power of digital circuitsabstractThis paper analyzes the effect of variations in the parameters of an Integrated Voltage Regulator (IVR) and its impact on the power/performance of a system of IVR driven digital logic circuit. The coupled analysis of IVR and digital logic considering variations in the integrated passives, power train FETs and controller transistors shows, compared to an off-chip VR, variations in IVR induce much larger shifts in the operating frequency of the logic and total system power. Variations in the output filter passives cause most prominent variations in the system power and performance, particularly pronounced at low voltage operation of the core. We also show that the mean performance of the system can be traded-off to reduce the variability by modifying IVR parameters, such as controller zeroes or output capacitors. Monodeep Kar, Sergio Carlo, Harish Krishnamurthy, Saibal Mukhopadhyay |
ISLPED | 4 |
| 2014 | Energy Introspector: A parallel, composable framework for integrated power-reliability-thermal modeling for multicore architecturesabstractSustaining processor performance growth is challenged by physical limitations due to increased power and heat dissipations. Power and thermal management techniques combined with inherent workload dynamics create the spatiotemporal variations of power, temperature, and degradation in processors. As industry moves to smaller feature sizes, the performance will become increasingly dominated by the physics. The challenge is in understanding how the physics is manifested at the microarchitecture level. This requires the modeling and simulation environment that can capture multiple, distinct physical phenomena and their concurrent impact on the microarchitecture. William J. Song, Saibal Mukhopadhyay, Sudhakar Yalamanchili |
ISPASS | 2 |
| 2014 | TSV-Aware Interconnect Distribution Models for Prediction of Delay and Power Consumption of 3-D Stacked ICsabstract3-D integrated circuits (3-D ICs) are expected to have shorter wirelength, better performance, and less power consumption than 2-D ICs. These benefits come from die stacking and use of through-silicon vias (TSVs) fabricated for interconnections across dies. However, the use of TSVs has several negative impacts such as area and capacitance overhead. To predict the quality of 3-D ICs more accurately, TSV-aware 3-D wirelength distribution models considering the negative impacts were developed. In this paper, we apply an optimal buffer insertion algorithm to the TSV-aware 3-D wirelength distribution models and present various prediction results on wirelength, delay, and power consumption of 3-D ICs. We also apply the framework to 2-D and 3-D ICs built with various combinations of process and TSV technologies and predict the quality of today and future 3-D ICs. Dae Hyun Kim 0004, Saibal Mukhopadhyay, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Control Principles and On-Chip Circuits for Active Cooling Using Integrated Superlattice-Based Thin-Film Thermoelectric DevicesabstractSuperlattice thin-film thermoelectric coolers (TECs) are emerging as a promising technology for hot spot mitigation in microprocessors. This paper studies the prospect of on-demand cooling with advanced TECs integrated at the back of the heat spreader inside a package (integrated TEC). Using thermal compact models of the chip and package with integrated TECs, the control principles for TEC-assisted transient cooling are presented. The control principles are implemented in a 130-nm CMOS process and cosimulated with the thermal system to show their feasibility and energy overheads. The simulation results show potential for extending the time for which a chip and package can sustain a high power load. Borislav Alexandrov, Owen Sullivan, William J. Song, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2014 | A Variation-Aware Preferential Design Approach for Memory-Based Reconfigurable ComputingabstractStatic random access memory arrays designed in sub-90-nm technologies are highly vulnerable to process variation-induced read/write/access failures. In memory-based reconfigurable computing frameworks, which use large high-density memory array, such failures lead to incorrect execution of mapped applications. It causes loss in quality of service (QoS) for digital signal processing (DSP) applications. In this paper, we analyze the effect of parameter variations on QoS in a memory-based reconfigurable computing framework. Next, we propose a preferential design approach at both application mapping and circuit level, which can significantly improve QoS and yield under large parameter variations. The proposed application mapping process considers the reliability map of a memory array and maps the important components with respect to QoS to more reliable memory blocks under performance constraint. At circuit level, we exploit the read-dominant memory access pattern to skew the memory cells for better read stability leading to improved QoS. Such a architecture/circuit codesign approach can also tolerate increased failure rate at low operating voltage, thus facilitating low-power operation. The effect of the approach is studied for two common DSP applications, namely discrete cosine transform and finite-impulse response (FIR) filter. The simulation results for FIR application show 45% improvement in power at iso-QoS and 47% in yield for a target peak signal to noise ratio at 45-nm technology. Somnath Paul, Saibal Mukhopadhyay, Swarup Bhunia |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | On the potential of 3D integration of inductive DC-DC converter for high-performance power deliveryabstractThis paper studies the potential and challenges of integrating an inductor based DC-DC converter based voltage regulator module (VRM) as a separate die with processor for high-performance power delivery network (PDN). The frequency domain analysis of PDN considering the converter shows 3D integration of VRM improves PDN impedance but the effectiveness depends on the converter design and whether the LC filter is integrated on-board, on-package, or on-die with the die-stack. The methodologies to co-design the converter with PDN and packaging scenarios are discussed and implications on PDN impedance and power losses are studied to maximally exploit the advantage of 3D stacking. Sergio Carlo, Wen Yueh, Saibal Mukhopadhyay |
DAC | 3 |
| 2013 | Exploring tunnel-FET for ultra low power analog applications: a case study on operational transconductance amplifierabstractThis work studies the potentials and challenges of designing ultra-low-power analog circuits exploiting unique characteristics of Tunnel-FET (TFET). TFET can achieve ultra-low quiescent current (~pA). In the subthreshold operation, TFET exhibit subthreshold swing lower than 60mV/decade, and hence higher transconductance per bias current than the MOSFET. TFET also exhibit very weak temperature dependence, and higher output resistance. Among several challenges, TFET demonstrate higher Shot noise at low biasing current. Through design of TFET based Operational Transconductance Amplifier (OTA) these challenges and opportunities are discussed. For implantable bio-medical applications, TFET OTA based neural amplifier design is studied. Amit Ranjan Trivedi, Sergio Carlo, Saibal Mukhopadhyay |
DAC | 3 |
| 2013 | Role of power grid in side channel attack and power-grid-aware secure designabstractSide-channel attack (SCA) is a method in which an attacker aims at extracting secret information from crypto chips by analyzing physical parameters (e.g. power). SCA has emerged as a serious threat to many mathematically unbreakable cryptography systems. From an attacker's point of view, the difficulty of mounting SCA largely depends on Signal-to-Noise Ratio (SNR) of the side-channel information. It has been shown that SNR primarily depends on algorithmic and circuit-level implementation, measurement noise, as well as device thermal noise. However, to the best of our knowledge, there has not been any study on the effect of power delivery network (PDN) on SCA resistance. We note that the PDN plays a significant role in SNR of measured supply current. Furthermore, SCA resistance strongly depends on the operating frequency due to RLC structure of a power grid. In this paper, we analyze the effect of power grid on SCA and provide quantitative results to demonstrate the frequency-dependent SCA resistance due to PDN-induced noise. This property can potentially be exploited by an attacker to facilitate the attack by operating a device at favorable frequency points. On the other hand, from a designer's perspective, one can explore countermeasures to secure the device at all operating frequencies while minimizing the design overhead. Based on this observation, we propose a frequency-dependent noise-injection based compensation technique to efficiently protect against SCA. Simulation results using realistic PDN model as well as experimental measurements using FPGA test board validate the observations on role of PDN in SCA and the efficacy of the proposed compensation approach. Xinmu Wang, Wen Yueh, Debapriya Basu Roy, Seetharam Narasimhan, Yu Zheng 0011, Saibal Mukhopadhyay, Debdeep Mukhopadhyay, Swarup Bhunia |
DAC | 6 |
| 2013 | Perceptual quality preserving SRAM architecture for color motion picturesabstractThis work proposes a low power methodology for video framebuffers to preserve the perceptual quality while reducing SRAM power. The bank-wise voltage scaling combined with error masking circuitry is proposed where voltage domains are separated according to the importance of luminous and color channels. The implementation may apply to standard embedded memory cores without redesigning specialized hardware within the SRAM bank. The simulation results showed that the proposed channel protection technique produced better energy-quality trade-off than the conventional higher-order-bit protection for the uncompressed as well as compressed motion image frames. Wen Yueh, Minki Cho, Saibal Mukhopadhyay |
DATE | 3 |
| 2013 | Physics of computing as an introduction to computer engineeringabstractThis paper describes a new required course in the Georgia Tech computer engineering curriculum, ECE 3030, Physical Foundations of Computer Systems. Traditional introductory courses take a constructive approach to logic design and computer organization. 3030, in contrast, introduces the major physical concepts underlying computation. It shows how they determine basic properties of computers such as speed and energy consumption. It also explores design trade-offs by showing how changes that improve one type of property inevitably, due to physics, cause another useful property to degrade. The course emphasizes CMOS but many of its principles apply to other logic technologies as well. Students do not directly design logic or learn assembly language-for example, delay and energy consumption are studied for inverter chains. However, they have time in the course to study in detail the basic physical phenomena that underlie design choices in digital systems. Those principles help students absorb material in later classes such as VLSI design. 3030 introduces certain topics to students much earlier in the curriculum than is traditional. We believe that an early introduction to principles is important not just for students who become logic designers but for all computer engineers. Marilyn Wolf, Saibal Mukhopadhyay |
FIE | 2 |
| 2013 | Error resilient logic circuits under dynamic variationsabstractThe design of low power and robust circuits under dynamic variations has emerged as a key challenge for silicon technologies. A particularly challenging problem is to tolerate transient supply noise that can occur in nanoseconds to microseconds time scales. The use of voltage or timing safety margin helps tolerate dynamic variations but at the expense of reduced performance or increased power dissipation. This talk will present adaptive circuit techniques to design resilient pipelines under fast transient variations. The presented techniques will allow a pipeline circuit to operate with minimal safety margin while preventing timing errors by adaptive clocking as well as time-borrowing and clock stretching. The measurement data from test-chips designed in 130nm CMOS technology will be presented to demonstrate the effectiveness of adaptive circuit techniques in designing low-power and resilient pipeline circuits. Kwanyeob Chae, Saibal Mukhopadhyay |
IOLTS | 2 |
| 2013 | Electrothermal analysis of spin-transfer-torque random access memory arraysabstractSpin Transfer Torque RAM (STTRAM) is a promising candidate for fast, scalable, high-density, nonvolatile memory in nanometer technology. However, relatively high write current density and small volume of the memory device indicate the possibility of significant self-heating in the STTRAM structure. This article performs a critical analysis of the self-heating induced temperature variations in STTRAM. We perform a 3D finite volume method based study to characterize self-heating effect in a single cell. The analysis is extended for STTRAM arrays by developing a computationally efficient RC compact model based thermal analyzer. The analysis shows that self-heating can results in considerable increase in both steady-state value and transient change in temperature of individual cells. The effect is less pronounced at the array level and depends on the activity level, that is, number of active cells within an array size. The analysis further illustrates that self-heating negatively impacts electrical reliability metrics namely, read margin and detection accuracy; degrades cell performance; and modulates energy dissipation. Subho Chatterjee, Sayeef S. Salahuddin, Saibal Mukhopadhyay |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2012 | Prospects of active cooling with integrated super-lattice based thin-film thermoelectric devices for mitigating hotspot challenges in microprocessorsabstractSuper-lattice thin-film thermoelectric coolers (TEC) are emerging as a promising technology for hot spot mitigation in microprocessors. This paper studies the prospect of on-demand cooling with advanced TECs integrated at the back of the heat spreader inside a package (integrated TEC). The thermal compact models of the chip and package with integrated TECs are developed and used for steady-state and transient temperature analysis. The control principles for TEC assisted transient cooling are presented and their impact on reducing thermal violations in microprocessors and TEC energy dissipations are discussed. Borislav Alexandrov, Owen Sullivan, Saibal Mukhopadhyay |
ASP-DAC | 4 |
| 2012 | Tier-adaptive-voltage-scaling (TAVS): A methodology for post-silicon tuning of 3D ICsabstractThis paper presents tier-adaptive-voltage-scaling (TAVS) as a post-silicon tuning methodology for improving parametric yield of 3D integrated circuits considering die-to-die and within-die process variations. The TAVS methodology senses process corners of individual tiers using on-tier delay sensors and adapt the supply voltage of each tier. The overall TAVS architecture is presented and the circuit issues associated with design of 3D level shifters are discussed. Circuit level simulation and statistical analysis of the TAVS architecture in predictive 45nm technology show the possibility of 26%-39% reduction in chip delay distribution. Kwanyeob Chae, Saibal Mukhopadhyay |
ASP-DAC | 2 |
| 2012 | Self-adaptive power gating with test circuit for on-line characterization of energy inflection activityabstractA test circuit is presented for post-silicon and on-line characterization of the energy-inflection activity of power-gated circuits (the activity when overhead energy is equal to leakage savings) under static (process) and dynamic (voltage/temperature/input) variations. The test circuit is applied to design self-adaptive power-gating for energy-efficient SRAM. Amit Ranjan Trivedi, Saibal Mukhopadhyay |
VTS | 2 |
| 2012 | On the parametric failures of SRAM in a 3D-die stack considering tier-to-tier supply cross-talkabstractThis paper analyzes the supply crosstalk between logic cores and SRAMs on separate tiers in a 3D die-stack using a distributed RLC based 3D power grid model. The analysis shows that due to the supply cross-talk power variation in cores modulates the performances and parametric failures in SRAM. Wen Yueh, Subho Chatterjee, Amit Ranjan Trivedi, Saibal Mukhopadhyay |
VTS | 4 |
| 2012 | Modeling and Designing for Accuracy and Energy Efficiency in Wireless Electroencephalography SystemsabstractRemote wireless monitoring of physiological signals has emerged as a key enabler for biotelemetry and can significantly improve the delivery of healthcare. Improving the energy efficiency and battery lifetime of the monitoring units without sacrificing the acquired signal quality is a key challenge in large-scale deployment of bioelectronic systems for remote wireless monitoring. In this article, we present a design methodology for accuracy aware, energy efficient wireless monitoring of electroencephalography (EEG) data. The proposed design performs a real-time accuracy energy trade-off by controlling the volume of transmitted data based on the information content in the EEG signal. We consider the effect of different system parameters in order to design an optimal system. We analyze the impact of noise of the wireless channel. Our analysis shows that the proposed system design approach can provide up to 10X energy savings in a 32 channel wireless EEG system with minimal impact on the monitored EEG signal accuracy. Jeremy R. Tolbert, Pratik Kabali, Simeranjit Brar, Saibal Mukhopadhyay |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2012 | Variation-Aware Clock Network Design Methodology for Ultralow Voltage (ULV) CircuitsabstractThis paper presents a design methodology for robust and low-energy clock networks for ultralow voltage (ULV) circuits. We show that both clock slew and skew play important roles in achieving high maximum operating frequency$(F_{\max})$and low clock energy in ULV circuits. In addition, clock networks in ULV circuits are highly sensitive to process variations. We propose a variation-aware methodology that controls both clock skew and slew to maximize$F_{\max}$and minimize clock power. In addition, we implement dynamic programming (DP)-based ULV clock routing and buffering methods (deferred merging and embedding) for deterministic and statistical conditions. Experimental results show that our clock network design method achieves lower energy (more than 20% savings) at comparable or even higher$F_{\max}$compared with the existing methods. Xin Zhao 0001, Jeremy R. Tolbert, Saibal Mukhopadhyay, Sung Kyu Lim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | Variation-aware clock network design methodology for ultra-low voltage (ULV) circuits
Xin Zhao 0001, Jeremy R. Tolbert, Chang Liu 0034, Saibal Mukhopadhyay, Sung Kyu Lim |
ISLPED | 4 |
| 2011 | Modeling and Analysis of Image Dependence and Its Implications for Energy Savings in Error Tolerant Image ProcessingabstractWe present an analysis of the relationship between input images and energy consumption in error tolerant image processing. Under aggressive voltage scaling, the output image quality of image processing depends on input images for two reasons: 1) error tolerance among images is naturally disparate in terms of perceptual image quality assessment, and 2) the error rate under aggressive voltage scaling varies by input image types. Based on both effects, the supply voltage can be optimized for a given quality requirement so as to achieve ultralow power/energy dissipation. Our analysis demonstrates the significance of the accurate delay estimation, which depends on not only combinational inputs but also the previous state of the logic. We present a new sequential model for accurate error estimation. Based on the model, our experimental results demonstrate that different input image types lead to very different output quality. We also present the effect of process variation on the relationship between input image and output quality. The dependence of energy consumption on input images provides a new perspective for low-power multimedia and image processing system design. Se Hun Kim, Saibal Mukhopadhyay, Marilyn Wolf |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2011 | Analysis and Design of Energy and Slew Aware Subthreshold Clock SystemsabstractIn this paper, we analyze the effect of clock slew in subthreshold circuits. Specifically, we address the issue that variations in clock slew at the register control can cause serious timing violations. We show that clock slew variations can cause frequency targets to deviate by as much as 28% from the design goals. Based on these observations, we recognize the importance of clock slew control in subthreshold circuits. We propose a systematic approach to design the clock tree for subthreshold circuits to reduce the clock slew variations while minimizing the energy dissipation in the tree. The combined approach, including the wire sizing and dynamic nodal capacitance control, can achieve better slew control (and better timing control) at lower energy in subthreshold circuits. Jeremy R. Tolbert, Xin Zhao 0001, Sung Kyu Lim, Saibal Mukhopadhyay |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2011 | A Scalable Design Methodology for Energy Minimization of STTRAM: A Circuit and Architecture PerspectiveabstractIn this paper, we analyze the energy dissipation in spin-torque-transfer random access memory array (STTRAM). We present a methodology for exploring the design space to minimize the energy dissipation of the array while maintaining required read and write quality for a given magnetic tunnel junction technology. The proposed method shows the need for proper choice of the silicon transistor width and array operating voltage to minimize the energy dissipation of the STTRAM array. The write energy is found to be 10 × greater than read energy. Hence, read-write ratio becomes a crucial factor that determines energy for STTRAM last level caches (L2). An exploration is performed across several architectural benchmarks including shared and non-shared caches for detailed energy analysis. Subho Chatterjee, Mitchelle Rasquinha, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2011 | Reconfigurable SRAM Architecture With Spatial Voltage Scaling for Low Power Mobile Multimedia ApplicationsabstractThis paper presents a dynamically reconfigurable SRAM array for low-power mobile multimedia application. The proposed structure use a lower voltage for cells storing low-order bits and a nominal voltage for cells storing higher order bits. The architecture allows reconfigure the number of bits in the low-voltage mode to change the error characteristics of the array in run-time. Simulations in predictive 70 nm nodes show that the proposed array can obtain 45% savings in memory power with a marginal (~10%) reduction in image quality. Minki Cho, Jason Schlessman, Marilyn Wolf, Saibal Mukhopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2011 | SRAM Write-Ability Improvement With Transient Negative Bit-Line VoltageabstractIncreasing variations in device parameters significantly degrades the write-ability of SRAM cells in deep sub-100 nm CMOS technology. In this paper, a transient negative bit-line voltage technique is presented to improve write-ability of SRAM cell. Capacitive coupling is used to generate a transient negative voltage at the low-going bit-line during Write operation without using any on-chip or off-chip negative voltage source. Statistical simulations in a 45-nm PD/SOI technology show a 103X reduction in the Write-failure probability with the proposed method. Saibal Mukhopadhyay, Rahul M. Rao, Jae-Joon Kim, Ching-Te Chuang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2010 | Design method and test structure to characterize and repair TSV defect induced signal degradation in 3D systemabstractIn this paper we present a test structure and design methodology for testing, characterization, and self-repair of TSVs in 3D ICs. The proposed structure can detect the signal degradation through TSVs due to resistive shorts and variations in TSV. For TSVs with moderate signal degradations, the proposed structure reconfigures itself as signal recovery circuit to improve signal fidelity. The paper presents the design of the test/recovery structure, the test methodologies, and demonstrates its effectiveness through stand alone simulations as well as in a full-chip physical design of a 3D IC. Minki Cho, Chang Liu 0034, Dae Hyun Kim 0004, Sung Kyu Lim, Saibal Mukhopadhyay |
ICCAD | 5 |
| 2010 | Analysis of thermal behaviors of spin-torque-transfer RAM: a simulation studyabstractWe present an accurate model of the self-heating effect in the Spin-Torque-Transfer RAM (STTRAM) using finite-volume-methods and thermal RC based compact models. We couple device level thermal simulation to the self-heating phenomenon to show that self-heating during write operation can result in significant temperature increase in STTRAM which in turn adversely affect the read disturb, leakage energy and sensing accuracy. Subho Chatterjee, Sayeef S. Salahuddin, Saibal Mukhopadhyay |
ISLPED | 4 |
| 2010 | An energy efficient cache design using spin torque transfer (STT) RAMabstractThe on-chip memory is a dominant source of power and energy consumption in modern and future processors. This paper explores the use of a new emerging non-volatile memory technology as a replacement for SRAM based lower level caches - Spin Torque Transfer(STT) RAM. While STTRAM achieves a reduction in leakage energy of 90% compared to SRAM, the dynamic energy for a write operation is 2X that of SRAM. Consequently, we propose additional microarchitectural optimizations to reduce overall dynamic energy which achieve an average reduction in dynamic energy over the base case of 30% with a range of 16% to 60% across 10 benchmarks. Mitchelle Rasquinha, Dhruv Choudhary, Subho Chatterjee, Saibal Mukhopadhyay, Sudhakar Yalamanchili |
ISLPED | 4 |
| 2010 | Optimization of burn-in test for many-core processors through adaptive spatiotemporal power migrationabstractWe present adaptive spatiotemporal power migration (ASTPM) for burn-in of many core chips. ASTPM adapts the number of simultaneously stressed cores and dynamically varies their location to prevent thermal runaway, improve test-quality, and optimize burn-in time. Minki Cho, Nikhil Sathe, Arijit Raychowdhury, Saibal Mukhopadhyay |
ITC | 4 |
| 2010 | Self-Repairing SRAM Using On-Chip Detection and CompensationabstractIn nanometer scale static-RAM (SRAM) arrays, systematic inter-die and random within-die variations in process parameters can cause significant parametric failures, severely degrading parametric yield. In this paper, we investigate the interaction between the inter-die and intra-dieV tvariations on SRAM read and write failures. To improve the robustness of the SRAM cell, we propose a closed-loop compensation scheme using on-chip monitors that directly sense the global read stability and writability of the cell. Simulations based on 45-nm partially depleted silicon-on-insulator technology demonstrate the viability and the effectiveness of the scheme in SRAM yield enhancement. Niladri Narayan Mojumder, Saibal Mukhopadhyay, Jae-Joon Kim, Ching-Te Chuang, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Accuracy-aware SRAM: a reconfigurable low power SRAM architecture for mobile multimedia applicationsabstractWe propose a dynamically reconfigurable SRAM architecture for low-power mobile multimedia applications. Parametric failures due to manufacturing variations limit the opportunities for power saving in SRAM. We show that, using a lower voltage for cells storing low-order bits and a nominal voltage for cells storing higher order bits, ~45% savings in memory power can be achieved with a marginal (~10%) reduction in image quality. A reconfigurable array structure is developed to dynamically reconfigure the number of bits in different voltage domains. Minki Cho, Jason Schlessman, Marilyn Wolf, Saibal Mukhopadhyay |
ASP-DAC | 4 |
| 2009 | Yield estimation of SRAM circuits using "Virtual SRAM Fab"abstractStatic Random Access Memories (SRAMs) are key components of modern VLSI designs and a major bottleneck to technology scaling as they use the smallest size devices with high sensitivity to manufacturing details. Analysis performed at the "schematic" level can be deceiving as it ignores the interdependence between the implementation layout and the resulting electrical performance. We present a computational framework, referred to as "Virtual SRAM Fab", for analyzing and estimating pre-Si SRAM array manufacturing yield considering both lithographic and electrical variations. The framework is being demonstrated for SRAM design/optimization in 45nm nodes and currently being used for both 32nm and 22nm technology nodes. The application and merit of the framework are illustrated using two different SRAM cells in a 45nm PD/SOI technology, which have been designed for similar stability/performance, but exhibit different parametric yields due to layout/lithographic variations. We also demonstrate the application of Virtual SRAM Fab for prediction of layout-induced imbalance in an 8T cell, which is a popular candidate for SRAM implementation in 32-22nm technology nodes. Aditya Bansal, Rama N. Singh, Rouwaida Kanj, Saibal Mukhopadhyay, Jin-Fuw Lee, Emrah Acar, Amith Singhee, Keunwoo Kim, Ching-Te Chuang, Sani R. Nassif, Fook-Luen Heng, Koushik K. Das |
ICCAD | 4 |
| 2009 | A methodology for robust, energy efficient design of Spin-Torque-Transfer RAM arrays at scaled technologiesabstractIn this paper we propose a methodology for energy efficient Spin-Torque-Transfer Random Access Memory (STTRAM) array design at scaled technology nodes. We present a model to estimate and analyze the energy dissipation of an STTRAM array. The presented model shows the strong dependence of the array energy on the silicon transistor width, word line voltage and row/column organization. Using the array energy model we propose a design methodology for STTRAM arrays which minimizes the energy dissipation while maintaining the required robustness in read and write operations at scaled technologies. Subho Chatterjee, Mitchelle Rasquinha, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
ICCAD | 4 |
| 2009 | A circuit-software co-design approach for improving EDP in reconfigurable frameworksabstractUse of two-dimensional memory array for lookup table (LUT) based reconfigurable computing frameworks has been proposed earlier for improvement in performance and energy-delay product (EDP). In this paper, we propose an integrated solution for achieving significantly higher EDP in these frameworks by leveraging on the read-dominant memory access pattern. First, we propose to employ an asymmetric memory cell design, which provides higher read performance (~2X) and lower read power (~1.6X) in order to improve the overall EDP during operation. Exploiting the fact that the proposed memory cell provides better read power/performance for cells storing logic '0', next we propose a content-aware application mapping approach, which tries to maximize the logic '0' content in the LUTs. We show that the joint circuit and application mapping level optimization approach provides significant improvement in system EDP for a set of benchmark circuits. Somnath Paul, Subho Chatterjee, Saibal Mukhopadhyay, Swarup Bhunia |
ICCAD | 3 |
| 2009 | A variation-aware preferential design approach for memory based reconfigurable computingabstractStatic Random Access Memory (SRAM) arrays designed in sub-90nm technologies are highly vulnerable to process variation induced read/write/access failures. In memory based reconfigurable computing frameworks, which use large high density memory array, such failures lead to incorrect execution of mapped applications. It causes loss in Quality of Service (QoS) for Digital Signal Processing (DSP) applications. We propose a Preferential Design approach at both application mapping and circuit level, which can significantly improve QoS and yield under large parameter variations. Such a architecture/circuit co-design approach can also tolerate increased failure rate at low operating voltage, thus facilitating low-power operation. Simulation results for a common DSP application show 45% improvement in power at iso--QoS and 47% in yield for a target Peak Signal to Noise Ratio (PSNR) at 45nm technology. Somnath Paul, Saibal Mukhopadhyay, Swarup Bhunia |
ICCAD | 2 |
| 2009 | On improving the algorithmic robustness of a low-power FIR filterabstractVoltage scaling is a promising approach to reduce the power consumption in signal processing circuits. However aggressive voltage scaling can introduce errors in the output signal, thus degrading the algorithmic performance of the circuit. We consider the specific case of the finite impulse response (FIR) filter, and identify two different sources of errors occurring due to voltage scaling: (a) errors introduced because of increased delay along the logic path and (b) errors caused by failures in the memory due to process variations. We design a FIR filter by using a simple feedback based approach to reduce the memory errors and a linear predictor structure for correcting the logic errors. The proposed filter is more robust to both logic and memory errors caused by voltage scaling. The results show a considerable improvement in the output Signal to Noise ratio (at least around 10 dB) for a probability of error (Perr) even as high as 0.5. We also utilize the proposed technique for an image filtering application and observe a considerable improvement in the visual quality of the output image along with an improvement of over 10 dB in the Peak Signal to Noise ratio for Perras high as 0.5. Sourabh Khire, Saibal Mukhopadhyay |
ICCD | 2 |
| 2009 | Experimental analysis of sequence dependence on energy saving for error tolerant image processingabstractWe present experimental analysis to exploit the sequence dependence on energy saving in error tolerant image processing. Our analysis shows that the error distributions depend not only on combinational inputs but also on the previous state of the logic. We present a new sequential model for low-power delay faults. Our experimental results demonstrate the importance of considering the state of logic when analyzing errors and its dependence on output quality and energy saving. The results show that different input image type leads to very different output quality. This means that more error tolerant image can be operated at lower voltage. The sequence dependence in output quality and energy saving provide a new perspective to design low power multimedia system. Se Hun Kim, Saibal Mukhopadhyay, Marilyn Wolf |
ISLPED | 2 |
| 2009 | Slew-aware clock tree design for reliable subthreshold circuitsabstractIn the paper, we analyze the effect of clock slew in subthreshold circuits. Specifically, we address the issue that variations in clock slew at the register control can cause serious timing violations. We show that clock slew variations can cause latch timing metrics such as setup, hold and clock-to-q times to deviate by 90% from the design goals. Based on these observations, we recognize the importance of clock slew control in subthreshold circuits. We propose a systematic approach to design the clock tree for subthreshold circuits to reduce the clock slew variations while minimizing the power dissipation in the tree. We show that a tighter nodal capacitance control is necessary to control the slew in a subthreshold clock tree, which can increase the power dissipation. Recognizing that the wire resistances have a negligible effect in subthreshold circuits, we show proper wire sizing is necessary to reduce the clock power. Finally, we propose a dynamic nodal capacitance control technique that allows larger slew at the earlier nets of the tree while controlling it more aggressively near the sink nodes. The combined approach, including the wire sizing and dynamic nodal capacitance control, can achieve better slew control (and better timing control) at lower power in subthreshold circuits. Jeremy R. Tolbert, Xin Zhao 0001, Sung Kyu Lim, Saibal Mukhopadhyay |
ISLPED | 4 |
| 2008 | Hybrid CMOS-STTRAM non-volatile FPGA: design challenges and optimization approachesabstractResearch efforts to develop a novel memory technology that combines the desired traits of non-volatility, high endurance, high speed and low power have resulted in the emergence of Spin Torque Transfer-RAM (STTRAM) as a promising next generation universal memory. However, the prospect of developing a non-volatile FPGA framework with STTRAM exploiting its high integration density remains largely unexplored. In this paper, we propose a novel CMOS-STTRAM hybrid FPGA framework; identify the key design challenges; and propose optimization techniques at circuit, architecture and application mapping levels. Simulation results show that a STTRAM based optimized FPGA framework achieves an average improvement of 48.38% in area, 22.28% in delay and 16.1% in dynamic power for ISCAS benchmark circuits over a conventional CMOS based FPGA design. Somnath Paul, Saibal Mukhopadhyay, Swarup Bhunia |
ICCAD | 2 |
| 2008 | Pre-Si estimation and compensation of SRAM layout deficiencies to achieve target performance and yieldabstractWith technology scaling, process constraints and imperfections result in significant variation of post-Si performance and stability of SRAM from designed/target pre-Si parameters. Modification/ re-optimization of SRAM cell and/or tuning of process parameters to meet target performance and stability are limited by area constraints and involve several technology ramp-up cycles. For reducing access failures, if process is not fine tuned, memory access clock cycle period may need to be increased thereby compromising performance. We propose a design methodology to meet the target performance and reduce access failures by tuning the SRAM array peripherals instead of tuning the SRAM cell and process parameters. Proposed design methodology is supported by numerical framework and validated by simulation results on 45nm PDSOI technology. We further show that our methodology does not impact the READ stability of a cell. Aditya Bansal, Rama N. Singh, Saibal Mukhopadhyay, Geng Han, Fook-Luen Heng, Ching-Te Chuang |
ICCD | 3 |
| 2008 | Capacitive coupling based transient negative bit-line voltage (Tran-NBL) scheme for improving write-ability of SRAM design in nanometer technologiesabstractIncreasing process variation can significantly degrade the write-ability of an SRAM. In this paper, we propose negative bit- line voltage technique to improve cell write-ability without using any on-chip or off-chip negative voltage source. Capacitive coupling is used to generate a transient negative voltage at the low bit-line during write operation. Simulations in 45 nm PD/SOI technology show a 103times reduction in the write-failure probability with the proposed technique. Saibal Mukhopadhyay, Rahul M. Rao, Jae-Joon Kim, Ching-Te Chuang |
ISCAS | 1 |
| 2008 | Design and Analysis of a Self-Repairing SRAM with On-Chip Monitor and Compensation CircuitryabstractIn an SRAM array, the systematic inter-die and the random within-die variations in process parameters cause significant number of parametric failures, to degrade process yield in the nanometer technology regime. In this paper, we investigate the interaction between the inter-die and intra-die Vt variations on SRAM read and write failures. To improve robustness of SRAM cell, we propose a closed-loop compensation scheme using on-chip monitors that directly sense the global read stability and writability of the cell directly. Computer simulations based on 45nm PD/SOI technology demonstrate the viability and effectiveness of the scheme in SRAM yield enhancement. Niladri Narayan Mojumder, Saibal Mukhopadhyay, Jae-Joon Kim, Ching-Te Chuang, Kaushik Roy 0001 |
VTS | 2 |
| 2008 | Reduction of Parametric Failures in Sub-100-nm SRAM Array Using Body BiasabstractIn this paper, we present a postsilicon-tuning technique to improve parametric yield of SRAM array using body bias (BB). First, we show that, although parametric failures in SRAM are due to local random intradie variations, the parametric failures increase at extreme interdie corners. Next, we show that proper BB can reduce different types of parametric failures. Finally, we show that adaptive application of BB to different dies, based on their interdie corners, reduces the total number of parametric failures in those dies. This helps to repair the faulty dies at different interdie corners, thereby improving SRAM yield. We show that postsilicon-tuning using BB can result in significant yield enhancement for SRAM (8%-25% in predictive 70-nm technology). Saibal Mukhopadhyay, Hamid Mahmoodi, Kaushik Roy 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2008 | Profit Aware Circuit Design Under Process Variations Considering Speed BinningabstractIn this paper, a profit-aware design metric is proposed to consider the overall merit of a design in terms of power and performance. A statistical design methodology is then developed to improve the economic merit of a design considering frequency binning and product price profile. A low-complexity sensitivity-based gate sizing algorithm is developed to improve economic gain of a design over its initial yield-optimized design. Finally, we present an integrated design methodology for simultaneous sizing and bin boundary determination to enhance profit under an area constraint. Experiments on a set of ISCAS'85 benchmarks show in average 19% improvement in profit for simultaneous sizing and bin boundary determination, considering both leakage power dissipation and delay bounds compared to a design initially optimized for 90% yield at iso-area in 70-nm bulk CMOS technology. Animesh Datta, Swarup Bhunia, Jung Hwan Choi, Saibal Mukhopadhyay, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2007 | Design and analysis of Thin-BOX FD/SOI devices for low-power and stable SRAM in sub-50nm technologiesabstractThis paper demonstrates viable device design options for low-leakage and robust SRAM in sub-50nm FD/SOI technology. We explore the possibilities of reducing the body-doping of FD/SOI devices with proper tuning of back-gate bias or gate workfunction to achieve a given leakage target. The reduction of body-doping density helps reduce the effect of the random dopant fluctuation (RDF), while the Vt and leakage are controlled using the back-gate bias. Our analysis show that, body-doping reduction combined with back-gate biasing is the most efficient FD/SOI device design for low-leakage and robust SRAM. Saibal Mukhopadhyay, Keunwoo Kim, Ching-Te Chuang |
ISLPED | 1 |
| 2006 | Speed binning aware design methodology to improve profit under parameter variationsabstractDesigning high-performance systems with high yield under parameter variations has raised serious design challenges in nanometer technologies. In this paper, we propose a profit-aware yield model, based on which we present a statistical design methodology to improve profit of a design considering frequency binning and product price profile. A low-complexity sensitivity-based gate sizing algorithm is developed to improve the profitability of design over an initial yield-optimized design. We also propose an algorithm to determine optimal bin boundaries for maximizing profit with frequency binning. Finally, we present an integrated design methodology for simultaneous sizing and bin placement to enhance profit under an area constraint. Experiments on a set of ISCAS85 benchmarks show up to 26% (36%) improvement in profit for fixed bin (for simultaneous sizing and bin placement) with three frequency bins considering both leakage and delay bounds compared to a design optimized for 90% yield at iso-area. Animesh Datta, Swarup Bhunia, Jung Hwan Choi, Saibal Mukhopadhyay, Kaushik Roy 0001 |
ASP-DAC | 4 |
| 2006 | Self-calibration technique for reduction of hold failures in low-power nano-scaled SRAMabstractIncreasing source voltage (Source-Biasing) is an efficient technique for reducing gate and sub-threshold leakage of SRAM arrays. However, due to process variation, a higher source voltage can significantly increase data flipping in standby mode (Hold Failures) resulting in faulty memories. This imposes serious concerns in reducing standby power with source-bias. In this paper, we analyze the effect of source bias on hold failures under both inter-die and intra-die variations. We propose a self-calibrating SRAM for aggressively reducing leakage while maintaining the hold failures under control. Swaroop Ghosh, Saibal Mukhopadhyay, Keejong Kim, Kaushik Roy 0001 |
DAC | 2 |
| 2006 | Circuit-aware device design methodology for nanometer technologies: a case study for low power SRAM designabstractIn this paper, we propose a general Circuit-aware Device Design methodology, which can improve the overall circuit design by taking advantages of the individual circuit characters during the device design phase. The proposed methodology analytically derives the optimal device in terms of the pre-specified circuit quality factor. We applied the proposed methodology to SRAM design and achieved significant reduction in standby leakage and access time (11% and 7%, respectively, for conventional 6T-SRAM). Also, we observed that the optimal devices selected depend considerably on the applied circuit techniques. We believe that the proposed Circuit-aware Device Design methodology will be useful in the sub-90nm technology, where different leakage components (subthreshold, gate, and junction tunneling) are comparable in magnitude. Also, in this work, we have presented a design automation framework for SRAM, which is conventionally custom designed and optimized. Qikai Chen, Saibal Mukhopadhyay, Aditya Bansal, Kaushik Roy 0001 |
DATE | 2 |
| 2006 | Delay Modeling and Statistical Design of Pipelined Circuit Under Process VariationabstractUnder inter-die and intra-die parameter variations, the delay of a pipelined circuit follows a statistical distribution. This paper presents analytical models to estimate yield for a pipelined design based on delay distributions of individual pipe stages. Using the proposed models, it is shown that a change in logic depth and an imbalance between stage yields can improve the design yield and the area of a pipeline a circuit. A novel statistical methodology is developed to enhance yield of a pipelined circuit under an area constraint. Based on the concept of area borrowing, the results show that incorporating a proper imbalance among stage areas in a four-stage pipeline improves design yield up to 15.4% for the same area (and reduces area up to 8.4% under a yield constraint) compared with a balanced design Animesh Datta, Swarup Bhunia, Saibal Mukhopadhyay, Kaushik Roy 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2006 | Modeling and analysis of loading effect on leakage of nanoscaled bulk-CMOS logic circuitsabstractIn nanoscale complementary metal-oxide-semiconductor (CMOS) devices, a significant increase in subthreshold, gate, and reverse-biased junction band-to-band-tunneling (BTBT) leakage results in large leakage power in logic circuits. Leakage components interact with each other at the device level (through device geometry and the doping profile) and at the circuit level (through the node voltages). Due to the circuit-level interaction of the different leakage components, the leakage of a logic gate depends on the circuit topology, i.e., the number and the nature of the other logic gates connected to its input and output. In this paper, the effect of loading on a leakage of a circuit is analyzed for the first time. The authors have also proposed a method to accurately estimate the total leakage in a logic circuit from its logic-level description considering the impact of loading and transistor stacking. Saibal Mukhopadhyay, Swarup Bhunia, Kaushik Roy 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2006 | Modeling and Analysis of Leakage Currents in Double-Gate TechnologiesabstractThis paper models and analyzes subthreshold and gate leakage currents in different double-gate (DG) devices, namely, a doped body symmetric device with polysilicon gates, an intrinsic body symmetric device with metal gates, and an intrinsic body asymmetric device with different front and back gate materials. The effect of variations in device parameters on the leakage components is also analyzed. Using the developed models, digital circuits (logic gates and static random access memory cells) designed with different DG structures are also analyzed. The analysis shows that the use of (near mid-gap) metal gate and intrinsic body devices significantly reduces both the total leakage and its sensitivity to parametric variations in DG devices and circuits Saibal Mukhopadhyay, Keunwoo Kim, Ching-Te Chuang, Kaushik Roy 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2006 | A novel high-performance and robust sense amplifier using independent gate control in sub-50-nm double-gate MOSFETabstractDouble-gate (DG) transistor has emerged as one of the most promising devices for nano-scale circuit design. In this paper, we propose a high-performance and robust sense-amplifier design using independent gate control in symmetric and asymmetric DG devices for sub-50-nm technologies. The proposed sense amplifier has better performance (30%-35% less sensing delay) and robustness (60%-80% less minimum input bit-differential for correct operation considering 10% worst case silicon thickness mismatch) compared to the connected gate design. Hence, the proposed design successfully demonstrates the benefit of using independent gate control in DG devices for efficient circuit design in sub-50-nm regime. Saibal Mukhopadhyay, Hamid Mahmoodi, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2005 | A Statistical Approach to Area-Constrained Yield Enhancement for Pipelined Circuits under Parameter VariationsabstractUnder inter- and intra-die parameter variations, delay of a pipelined circuit follows a statistical distribution. Hence, a pipelined circuit suffers yield loss with respect to violation of target delay constraint unless an overly pessimistic worst-case design approach is followed. We propose a statistical approach for pipeline design to enhance yield with respect to a target delay under an area budget. Right choice of the number of pipeline stages to enhance yield under an area constraint is addressed using simple statistical yield models. Next, individual stages are designed for maximizing yield under area constraint for the stages. Once the independently optimized stages are combined to form a pipeline, we propose a final global optimization step to improve pipeline yield with no area overhead, based on a concept of area borrowing. Optimization results show that, the proposed statistical design approach for pipeline improves the overall yield up to 12% over conventional design for equal area. Animesh Datta, Swarup Bhunia, Saibal Mukhopadhyay, Kaushik Roy 0001 |
Asian Test Symposium | 3 |
| 2005 | Leakage Current Based Stabilization Scheme for Robust Sense-Amplifier Design for Yield Enhancement in Nano-scale SRAMabstractIn this paper, we develop a method to analyze the probability of access failure in SRAM array (due to random Vt variation in transistors) by jointly considering variations in cell and senseamplifiers. Our analysis shows that, improving robustness of senseamplifier is extremely important for reducing memory access failure probability and improving yield. We present a process variation tolerant sense amplifier suitable for SRAM array designed in sub- 100nm CMOS technologies. The proposed technique reduces the failure probability of sense amplifiers by more than 80% with negligible penalty in the sensing delay. Saibal Mukhopadhyay, Arijit Raychowdhury, Hamid Mahmoodi, Kaushik Roy 0001 |
Asian Test Symposium | 1 |
| 2005 | Statistical Modeling of Pipeline Delay and Design of Pipeline under Process Variation to Enhance Yield in sub-100nm TechnologiesabstractOperating frequency of a pipelined circuit is determined by the of the slowest pipeline stage. However, under statistical delay variation in sub-100 nm technology regime, the slowest stage is not readily identifiable and the estimation of the pipeline yield with respect to a target delay is a challenging problem. We have proposed analytical models to estimate yield for a pipelined design based on delay distributions of individual pipe stages. Using the proposed models, we have shown that change in logic depth and imbalance between the stage delays can improve the yield of a pipeline. A statistical methodology has been developed to optimally design a pipeline circuit for enhancing yield. Optimization results show that, proper imbalance among the stage delays in a pipeline improves design yield by 9% for the same area and performance (and area reduction by about 8.4% under a yield constraint) over a balanced design. Animesh Datta, Swarup Bhunia, Saibal Mukhopadhyay, Nilanjan Banerjee, Kaushik Roy 0001 |
DATE | 3 |
| 2005 | Modeling and Analysis of Loading Effect in Leakage of Nano-Scaled Bulk-CMOS Logic CircuitsabstractIn nanometer scaled CMOS devices, a significant increase in the subthreshold, the gate and the reverse biased junction band-to-band-tunneling (BTBT) leakage results in a large increase of the total leakage power in a logic circuit. Leakage components interact with each other at the device level (through device geometry, doping profile) and also at the circuit level (through node voltages). Due to the circuit level interaction of the different leakage components, the leakage of a logic gate strongly depends on the circuit topology, i.e., the number and nature of the other logic gates connected to its input and output. For the first time, we analyze the loading effect on leakage and propose a method to estimate accurately, from its logic level description, the total leakage in a logic circuit, considering the impact of loading and transistor stacking. Saibal Mukhopadhyay, Swarup Bhunia, Kaushik Roy 0001 |
DATE | 1 |
| 2005 | Double-gate SOI devices for low-power and high-performance applicationsabstractDouble-gate (DG) transistors have emerged as promising devices for nano-scale circuits due to their better scalability compared to bulk CMOS. Among the various types of DG devices, quasi-planar SOI FinFETs are easier to manufacture compared to planar double-gate devices. DG devices with independent gates (separate contacts to back and front gates) have recently been developed. DG devices with symmetric and asymmetric gates have also been demonstrated. Such device options have direct implications at the circuit level. Independent control of front and back gate in DG devices can be effectively used to improve performance and reduce power in sub-50nm circuits. Independent gate control can be used to merge parallel transistors in noncritical paths. This results in reduction in the effective switching capacitance and hence power dissipation. We show a variety of circuits in logic and memory that can benefit from independent gate operation of DG devices. As examples, we show the benefit of independent gate operation in circuits such as dynamic logic circuits, Schmitt triggers, sense amplifiers, and SRAM cells. In addition to independent gate option, we also investigate the usefulness of asymmetric devices and the impact of width quantization and process variations on circuit design. Kaushik Roy 0001, Hamid Mahmoodi, Saibal Mukhopadhyay, Hari Ananthan, Aditya Bansal, Tamer Cakici |
ICCAD | 3 |
| 2005 | A Feasibility Study of Subthreshold SRAM Across Technology GenerationsabstractIn this paper, we have explored the feasibility of designing an SRAM array in the subthreshold domain of device operation. We have performed a nominal corner analysis of power and stability and a statistical analysis of the different failure probabilities of the subthreshold SRAM. Our analysis shows that subthreshold SRAM gives significant reduction (/spl sim/100/spl times/) of operating and standby power at iso-performance (/spl sim/100MHz) compared to the superthreshold counterpart. However, with increasing intra-die variation owing to technology scaling, the failure probability of subthreshold SRAM increases thereby masking the power benefits. Arijit Raychowdhury, Saibal Mukhopadhyay, Kaushik Roy 0001 |
ICCD | 2 |
| 2005 | Process Variation Tolerant Online Current Monitor for Robust SystemsabstractLarge inter-die and intra-die process variations result in significant uncertainty in delay of circuits. Large delay variations may lead to parametric/functional failures. In this paper we propose a leakage-variation-tolerant online current monitor, namely leakage canceling current sensor, to detect completion of operations in logic blocks. The current monitor is applied to self timed logic to design process variation tolerant circuits. It is observed that, for self-timed circuits, the probability of functional failures can be reduced by 50% with no performance degradation and with same power consumption. Qikai Chen, Saibal Mukhopadhyay, Hamid Mahmoodi, Kaushik Roy 0001 |
IOLTS | 2 |
| 2005 | Yield Prediction of High Performance Pipelined Circuit with Respect to Delay Failures in Sub-100nm TechnologyabstractIn nanoscale technology, large variations in process parameters produce wide delay spread in high performance circuit. In this paper the authors developed analytical models for yield prediction with respect to delay variation of pipeline design. The converse problem of estimating the design space for individual pipe stages based on a target yield has been addressed. For an example 4 stage pipelined circuit proposed analytical models are verified to predict yield within 2% of results obtained from Monte-Carlo Hspice simulation Animesh Datta, Saibal Mukhopadhyay, Swarup Bhunia, Kaushik Roy 0001 |
IOLTS | 2 |
| 2005 | Modeling and analysis of total leakage currents in nanoscale double gate devices and circuitsabstractIn this paper we model (numerically and analytically) and analyze sub-threshold, gate-to-channel tunneling, and edge direct tunneling leakage in Double Gate (DG) devices. We compare the leakage of different DG structures, namely, doped body symmetric device with polysilicon gates, intrinsic body symmetric device with metal gates and intrinsic body asymmetric device with different front and back gate material. It is observed that, use of (near-mid-gap) metal gate and intrinsic body devices significantly reduces both the total leakage and its sensitivity to parametric variations in DG circuits Saibal Mukhopadhyay, Keunwoo Kim, Ching-Te Chuang, Kaushik Roy 0001 |
ISLPED | 1 |
| 2005 | Reliable and self-repairing SRAM in nano-scale technologies using leakage and delay monitoringabstractThe inter-die and intra-die variations in process parameters result in large number of failures in an SRAM array degrading the design yield. In this paper, we propose an adaptive repairing technique for SRAM based on leakage and delay monitoring. Leakage and delay monitoring is used to effectively separate dies with different inter-die Vts from each other. Using the leakage (or delay) monitoring and adaptive body bias, we propose a reliable and self-repairing SRAM which has reduced number of parametric failures under high inter-die and intra-die Vt variations. The proposed self-repairing SRAM improves the design yield by 5%-40% in predictive 70nm technology from BPTM. Saibal Mukhopadhyay, Kunhyuk Kang, Hamid Mahmoodi, Kaushik Roy 0001 |
ITC | 1 |
| 2005 | Modeling of failure probability and statistical design of SRAM array for yield enhancement in nanoscaled CMOSabstractIn this paper, we have analyzed and modeled failure probabilities (access-time failure, read/write failure, and hold failure) of synchronous random-access memory (SRAM) cells due to process-parameter variations. A method to predict the yield of a memory chip based on the cell-failure probability is proposed. A methodology to statistically design the SRAM cell and the memory organization is proposed using the failure-probability and the yield-prediction models. The developed design strategy statistically sizes different transistors of the SRAM cell and optimizes the number of redundant columns to be used in the SRAM array, to minimize the failure probability of a memory chip under area and leakage constraints. The developed method can be used in an early stage of a design cycle to enhance memory yield in nanometer regime. Saibal Mukhopadhyay, Hamid Mahmoodi, Kaushik Roy 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | Accurate estimation of total leakage in nanometer-scale bulk CMOS circuits based on device geometry and doping profileabstractDramatic increase of subthreshold, gate and reverse biased junction band-to-band-tunneling (BTBT) leakage in scaled devices results in the drastic increase of total leakage power in a logic circuit. In this paper, a methodology for accurate estimation of the total leakage in a logic circuit based on the compact modeling of the different leakage current in nanoscaled bulk CMOS devices has been developed. Current models have been developed based on the device geometry, two-dimensional doping profile, and operating temperature. A circuit-level model of junction BTBT leakage has been developed. Simple models of the subthreshold current and the gate current have been presented. Also, the impact of quantum mechanical behavior of substrate electrons, on the circuit leakage has been analyzed. Using the compact current model, a transistor has been modeled as a sum of current sources (SCS). The SCS transistor model has been used to estimate the total leakage in simple logic gates and complex logic circuits (designed with transistors of 25-nm effective length) at room and elevated temperatures. Saibal Mukhopadhyay, Arijit Raychowdhury, Kaushik Roy 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | Low-power scan design using first-level supply gatingabstractReduction in test power is important to improve battery lifetime in portable electronic devices employing periodic self-test, to increase reliability of testing, and to reduce test cost. In scan-based testing, a significant fraction of total test power is dissipated in the combinational block. In this paper, we present a novel circuit technique to virtually eliminate test power dissipation in combinational logic by masking signal transitions at the logic inputs during scan shifting. We implement the masking effect by inserting an extra supply gating transistor in the supply to ground path for the first-level gates at the outputs of the scan flip-flops. The supply gating transistor is turned off in the scan-in mode, essentially gating the supply. Adding an extra transistor in only one logic level renders significant advantages with respect to area, delay, and power overhead compared to existing methods, which use gating logic at the output of scan flip-flops. Moreover, the proposed gating technique allows a reduction in leakage power by input vector control during scan shifting. Simulation results on ISCAS89 benchmarks show an average improvement of 62% in area overhead, 101% in power overhead (in normal mode), and 94% in delay overhead, compared to the lowest cost existing method. Swarup Bhunia, Hamid Mahmoodi, Debjyoti Ghosh, Saibal Mukhopadhyay, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2005 | A forward body-biased low-leakage SRAM cache: device, circuit and architecture considerationsabstractThis paper presents a forward body-biasing (FBB) technique for active and standby leakage power reduction in cache memories. Unlike previous low-leakage SRAM approaches, we include device level optimization into the design. We utilize super high Vt (threshold voltage) devices to suppress the cache leakage power, while dynamically FBB only the selected SRAM cells for fast operation. In order to build a super high Vt device, the two-dimensional (2-D) halo doping profile was optimized considering various nanoscale leakage mechanisms. The transition latency and energy overhead associated with FBB was minimized by waking up the SRAM cells ahead of the access and exploiting the general cache access pattern. The combined device-circuit-architecture level techniques offer 64% total leakage reduction and 7.3% improvement in bit line delay compared to a previous state-of-the-art low-leakage SRAM technique. Static noise margin of the proposed SRAM cell is comparable to conventional SRAM cells. Chris H. Kim, Jae-Joon Kim, Saibal Mukhopadhyay, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2004 | Leakage in nano-scale technologies: mechanisms, impact and design considerationsabstractThe high leakage current in nano-meter regimes is becoming a significant portion of power dissipation in CMOS circuits as threshold voltage, channel length, and gate oxide thickness are scaled. Consequently, the identification of different leakage components is very important for estimation and reduction of leakage. Moreover, the increasing statistical variation in the process parameters has led to significant variation in the transistor leakage current across and within different dies. Designing with the worst case leakage may cause excessive guard-banding, resulting in a lower performance. This paper explores various intrinsic leakage mechanisms including weak inversion, gate-oxide tunneling and junction leakage etc. Various circuit level techniques to reduce leakage energy and their design trade-off are discussed. We also explore process variation compensating techniques to reduce delay and leakage spread, while meeting power constraint and yield. Amit Agarwal 0001, Chris H. Kim, Saibal Mukhopadhyay, Kaushik Roy 0001 |
DAC | 3 |
| 2004 | Statistical design and optimization of SRAM cell for yield enhancementabstractWe have analyzed and modeled the failure probabilities of SRAM cells due to process parameter variations. A method to predict the yield of a memory chip based on the cell failure probability is proposed. The developed method is used in an early stage of a design cycle to minimize memory failure probability by statistically sizing of SRAM cell. Saibal Mukhopadhyay, Hamid Mahmoodi, Kaushik Roy 0001 |
ICCAD | 1 |
| 2004 | A Novel Low-Power Scan Design Technique Using Supply GatingabstractReduction in test power is important to improve battery life in portable devices employing periodic self-test, to increase reliability of testing and to reduce test-cost. In scan-based testing, about 80% of total test power is dissipated in the combinational block. In this paper, we present a novel circuit technique to virtually eliminate test power dissipation in combinational logic by masking signal transition at the logic inputs during scan shifting. We realize the masking effect by inserting an extra supply gating transistor in the VDD to GND path for the first level cells at output of the scan flops. The supply gating transistor is turned off in the scan-in mode, essentially gating the supply. Adding an extra transistor in only one logic level renders significant advantage with respect to area, delay and power (in normal mode of operation) overhead compared to existing methods, which use gating logic at the output of scan flops. Simulation results on ISCAS89 benchmarks show up to 79% improvement in area, up to 32% in power (in normal mode) and up to 7% in delay compared to lowest-cost known alternative. Swarup Bhunia, Hamid Mahmoodi, Saibal Mukhopadhyay, Debjyoti Ghosh, Kaushik Roy 0001 |
ICCD | 3 |
| 2004 | A circuit-compatible model of ballistic carbon nanotube field-effect transistorsabstractCarbon nanotube field-effect transistors (CNFETs) are being extensively studied as possible successors to CMOS. Novel device structures have been fabricated and device simulators have been developed to estimate their performance in a sub-10-nm transistor era. This paper presents a novel method of circuit-compatible modeling of single-walled semiconducting CNFETs in their ultimate performance limit. For the first time, both the I-V and the C-V characteristics of the device have been efficiently modeled for circuit simulations. The model so developed has been used to simulate arithmetic and logic blocks using HSPICE. Arijit Raychowdhury, Saibal Mukhopadhyay, Kaushik Roy 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2003 | Accurate estimation of total leakage current in scaled CMOS logic circuits based on compact current modelingabstractDramatic increase of subthreshold, gate and reverse biased junction band-to-band-tunneling (BTBT) leakage in scaled devices, results in the drastic increase of total leakage power in a logic circuit. In this paper a methodology for accurate estimation of the total leakage in a logic circuit based on the compact modeling of the different leakage current in scaled devices has been developed. Current models have been developed based on the exact device geometry, 2-D doping profile and operating temperature. A circuit level model of junction BTBT leakage (which is unprecedented) has been developed. Simple models of the subthreshold current and the gate current have been presented. Here, for the first time, the impact of quantum mechanical behavior of substrate electrons, on the circuit leakage has been analyzed. Using the compact current model, a transistor has been modeled as a Sum of Current Sources (SCS). The SCS transistor model has been used to estimate the total leakage in simple logic gates and complex logic circuits (designed with transistors of 25nm effective length) at the room and at the elevated temperatures. Saibal Mukhopadhyay, Arijit Raychowdhury, Kaushik Roy 0001 |
DAC | 1 |
| 2003 | Modeling of Ballistic Carbon Nanotube Field Effect Transistors for Efficient Circuit Simulation
Arijit Raychowdhury, Saibal Mukhopadhyay, Kaushik Roy 0001 |
ICCAD | 2 |
| 2003 | A forward body-biased low-leakage SRAM cache: device and architecture considerationsabstractThis paper presents a forward body-biasing (FBB) scheme for active leakage power reduction in cache memories. We utilize super high VT (threshold voltage) devices to suppress the leakage power in unselected portions of a cache while fast operation is achieve by dynamically forward body-biasing the selected SRAM cells. In order to generate a super high VT device, the 2-D halo doping profile was optimized considering different nanometer regime leakage mechanisms. The transition latency and energy overhead associated with FBB could be minimized by (i) waking up the SRAM cells ahead of the access and (ii) exploiting the cache access pattern. The combined device-circuit-architecture level techniques offer 64% total leakage reduction and 7.3% improvement in bitline delay compared to a previous state-of-the-art low-leakage SRAM technique. Chris H. Kim, Jae-Joon Kim, Saibal Mukhopadhyay, Kaushik Roy 0001 |
ISLPED | 3 |
| 2003 | Modeling and estimation of total leakage current in nano-scaled CMOS devices considering the effect of parameter variationabstractIn this paper we have developed analytical models to estimate the mean and the standard deviation in the gate, the subthreshold, the reverse biased source/drain junction band-to-band-tunneling (BTBT) and the total leakage in scaled CMOS devices considering variation in process parameters like device geometry, doping profile, flat-band voltage and supply voltage. We have verified the model using Monte Carlo simulation using an NMOS device of 50nm effective length and analyzed the results to enumerate the effect of different process parameters on the individual components and the total leakage. Saibal Mukhopadhyay, Kaushik Roy 0001 |
ISLPED | 1 |
| 2003 | Leakage current mechanisms and leakage reduction techniques in deep-submicrometer CMOS circuitsabstractHigh leakage current in deep-submicrometer regimes is becoming a significant contributor to power dissipation of CMOS circuits as threshold voltage, channel length, and gate oxide thickness are reduced. Consequently, the identification and modeling of different leakage components is very important for estimation and reduction of leakage power, especially for low-power applications. This paper reviews various transistor intrinsic leakage mechanisms, including weak inversion, drain-induced barrier lowering, gate-induced drain leakage, and gate oxide tunneling. Channel engineering techniques including retrograde well and halo doping are explained as means to manage short-channel effects for continuous scaling of CMOS devices. Finally, the paper explores different circuit techniques to reduce the leakage power consumption. Kauschick Roy, Saibal Mukhopadhyay, Hamid Mahmoodi |
Proc. IEEE | 2 |
| 2003 | Gate leakage reduction for scaled devices using transistor stackingabstractIn this paper, the effect of gate tunneling current in ultra-thin gate oxide MOS devices of effective length (L/sub eff/) of 25nm (oxide thickness=1.1 nm), 50 nm (oxide thickness=1.5 nm) and 90 nm (oxide thickness=2.5 nm) is studied using device simulation. Overall leakage in a stack of transistors is modeled and the opportunities for leakage reduction in the standby mode of operation are explored for scaled technologies. It is shown that, as the contribution of gate leakage relative to the total leakage increases with technology scaling, traditional techniques become ineffective in reducing overall leakage current in a circuit. A novel technique of input vector selection based on the relative contributions of gate and subthreshold leakage to the overall leakage is proposed for reducing total leakage in a circuit. This technique results in 44% savings in total leakage in 50-nm devices compared to the conventional stacking technique. Saibal Mukhopadhyay, Cassondra Neau, R. T. Cakici, Amit Agarwal 0001, Chris H. Kim, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |