EDBT 2026 Demo / reviewers in the wild / expert
Naresh R. Shanbhag
dblp:72/6310
· DBLP profile ↗
128ranked-venue papers
12as first author
13since 2021 · last 2026
0000-0002-4323-9164ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 81 · 11 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 4 first-authorArtificial intelligence and machine learning · 7 · 3 since 2021Computer networks · 7Software engineering, systems software and programming languages · 6 · 1 first-author · 2 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Energy-Accuracy Trade-Offs in Massive MIMO Signal Detection Using SRAM-Based In-Memory ComputingabstractThis paper investigates the use of SRAM-based in-memory computing (IMC) architectures for designing energy efficient and accurate signal detectors for massive multi-input multi-output (MIMO) systems. SRAM-based IMCs are the state-of-the-art for executing deep learning workloads in terms of energy efficiency and compute density. However, given the more stringent accuracy requirements of massive MIMO signal detection compared to deep learning, it remains an open question whether the energy efficiency benefits of IMCs can be preserved while meeting these accuracy requirements. This paper systematically explores the energy-accuracy trade-off in massive MIMO signal detectors designed using SRAM-based IMCs. Through transistor-level behavioral modeling and simulations in 28 nm CMOS process, we evaluate the energy per information bit ($E_{\mathrm {b}}$), error vector magnitude (EVM) and symbol error rate (SER) of linear detectors designed using SRAM-based IMCs for various wireless channels. We perform extensive design space exploration to determine the design parameters and operating conditions under which IMC-based linear detectors achieve significantly better energy efficiency than conventional digital implementations, while maintaining comparable accuracy. Our results show that IMC-based detectors can meet 3GPP EVM specifications for QPSK, 16-QAM, and 64-QAM across both real-world (Argos) and synthetic (WINNER-II) wireless channels, achieving$7.2\times $to$11.7\times $better energy efficiency than digital detectors while incurring$\lt \mathrm {0.1~dB}$penalty in receiver signal-to-noise ratio (RX SNR). We present extensive simulation results that provide insights into how parameters such as MIMO dimension, modulation scheme, precision of detection matrix, IMC bit-cell capacitance, IMC ADC precision, and ADC thermal noise impact the energy efficiency and accuracy of massive MIMO signal detection. Mihir Kavishwar, Naresh R. Shanbhag |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2026 | Characterizing the Intrinsic Bank-Level Accuracy Versus Energy Trade-Off of SRAM-Based Analog In-Memory Computing Architectures in 28 nm CMOS
Shuo Li 0008, Chihun Song, Hyungyo Kim, Nam Sung Kim, Naresh R. Shanbhag |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | On the Security Vulnerabilities of MRAM-based In-Memory Computing Architectures against Model Extraction AttacksabstractThis paper studies the security vulnerabilities of embedded nonvolatile memory (eNVM)-based in-memory computing (IMC) architectures to model extraction attacks (MEAs). These attacks allow the reconstruction of private training data from trained model parameters thereby leaking sensitive user information. The presence of analog noise in eNVM-based IMC computation suggests that they may be intrinsically robust to MEA. However, we show that this conjecture is false. Specifically, we consider the scenario where an attacker aims to retrieve model parameters via input-output query access, and propose three attacks that exploit the statistics of the IMC computation. We demonstrate the efficacy of these attacks in extracting the model parameters of the last layer of a ResNet-20 network from the bitcell array of an MRAM-based IMC prototype in 22 nm process. Employing the proposed MEAs, the attacker obtains a CIFAR-10 accuracy within 0.1% of that of a N = 64 dimensional, 7 b × 4 b fixed-point digital baseline. To the best of our knowledge, this is the first work to demonstrate MEAs for eNVM-based IMC on a real-life IC prototype. Our results indicate the critical importance of investigating the security vulnerabilities of IMCs in general, and eNVM-based IMCs, in particular. Saion K. Roy, Naresh R. Shanbhag |
ICCAD | 2 |
| 2023 | PRIVE: Efficient RRAM Programming with Chip Verification for RRAM-based In-Memory Computing AccelerationabstractAs deep neural networks (DNNs) have been success-fully developed in many applications with continuously increasing complexity, the number of weights in DNNs surges, leading to consistent demands for denser memories than SRAMs. RRAM-based in-memory computing (IMC) achieves high density and energy-efficiency for DNN inference, but RRAM programming remains to be a bottleneck due to high write latency and energy consumption. In this work, we present the Progressive-wRite In-memory program-VErify (PRIVE) scheme, which we verify with an RRAM testchip for IMC-based hardware acceleration for DNNs. We optimize the progressive write operations on different bit positions of RRAM weights to enable error compensation and reduce programming latency/energy, while achieving high DNN accuracy. For 5-bit precision DNNs, PRIVE reduces the RRAM programming energy by 1.82×, while maintaining high accuracy of 91.91% (VGG-7) and 71.47% (ResNet-18) on CIFAR-10 and CIFAR-100 datasets, respectively. Wangxin He, Jian Meng, Sujan K. Gonugondla, Shimeng Yu, Naresh R. Shanbhag, Jae-sun Seo |
DATE | 5 |
| 2023 | Boosting the Accuracy of SRAM-Based in-Memory Architectures Via Maximum Likelihood-Based Error Compensation MethodabstractSRAM-based analog in-memory computing (IMC) architectures have demonstrated high energy efficiency and compute density over digital accelerators for machine learning. However, their compute SNR and achievable dot product (DP) dimension are limited by the analog nature of computations. We present a Maximum Likelihood (ML)-based statistical Error Compensation (MLEC) method to enhance the accuracy of binary DPs in a 6T SRAM-based IMC. MLEC leverages the IMC architecture to extract multiple observations and implements an approximate ML detection rule. Employing simulations in a 28nm CMOS and behavioral modeling, we show that MLEC enhances the compute SNR by 5dB-to-30dB over a conventional IMC with an energy overhead ranging from 10%-to-30% for DP dimensions of 64-to-256. Hyungyo Kim, Naresh R. Shanbhag |
ICASSP | 2 |
| 2023 | Enhancing the Accuracy of Resistive In-Memory Architectures using Adaptive Signal ProcessingabstractAnalog in-memory computing architectures (IMCs) have exhibited high energy efficiency over conventional digital architectures. The use of resistive memory arrays such as magnetic RAM (MRAM) in IMCs has significant potential due to their high-density and non-volatility. However, the analog nature of computation limits the compute accuracy of IMCs. We present activation scaling compensation (ASC), an adaptive signal processing method to compensate for the effects of bitline (BL) and sourceline (SL) parasitic resistances in MRAM-based IMC matrix-vector multiplication (MVM) architectures. Through behavioral modeling and simulation, we demonstrate that ASC enhances the compute signal-to-noise ratio (SNR) by 8.6dB-to-13.9dB with negligible energy overhead. The effectiveness of ASC-enabled MRAM-based IMC is demonstrated in the context of a signal recovery problem. Han-Mo Ou, Naresh R. Shanbhag |
ICASSP | 2 |
| 2023 | On the Robustness of Randomized Ensembles to Adversarial PerturbationsabstractRandomized ensemble classifiers (RECs), where one classifier is randomly selected during inference, have emerged as an attractive alternative to traditional ensembling methods for realizing adversarially robust classifiers with limited compute requirements. However, recent works have shown that existing methods for constructing RECs are more vulnerable than initially claimed, casting major doubts on their efficacy and prompting fundamental questions such as: "When are RECs useful?", "What are their limits?", and "How do we train them?". In this work, we first demystify RECs as we derive fundamental results regarding their theoretical limits, necessary and sufficient conditions for them to be useful, and more. Leveraging this new understanding, we propose a new boosting algorithm (BARRE) for training robust RECs, and empirically demonstrate its effectiveness at defending against strong $\ell_\infty$ norm-bounded adversaries across various network architectures and datasets. Our code can be found at https://github.com/hsndbk4/BARRE. Hassan Dbouk, Naresh R. Shanbhag |
ICML | 2 |
| 2022 | IMPQ: Reduced Complexity Neural Networks Via Granular Precision AssignmentabstractThe demand for the deployment of deep neural networks (DNN) on resource-constrained Edge platforms is ever increasing. Today’s DNN accelerators support mixed-precision computations to enable reduction of computational and storage costs but require networks with precision at variable granularity, i.e., network, layer or kernel level. However, the problem of granular precision assignment is challenging due to an exponentially large search space and efficient methods for such precision assignment are lacking. To address this problem, we introduce the iterative mixed-precision quantization (IMPQ) framework to allocate precision at variable granularity. IMPQ employs a sensitivity metric to order the weight/activation groups in terms of the likelihood of misclassifying input samples due to its quantization noise. It iteratively reduces the precision of the weights and activations of a pretrained full-precision network starting with the least sensitive group. Compared to state-of-the-art methods, IMPQ reduces computational costs by 2× -to-2.5× for compact networks such as MobileNet-V1 on ImageNet with no accuracy loss. Our experiments reveal that kernel-wise granular precision assignment provides 1.7× higher compression than layer-wise assignment. Sujan K. Gonugondla, Naresh R. Shanbhag |
ICASSP | 2 |
| 2022 | Adversarial Vulnerability of Randomized EnsemblesabstractDespite the tremendous success of deep neural networks across various tasks, their vulnerability to imperceptible adversarial perturbations has hindered their deployment in the real world. Recently, works on randomized ensembles have empirically demonstrated significant improvements in adversarial robustness over standard adversarially trained (AT) models with minimal computational overhead, making them a promising solution for safety-critical resource-constrained applications. However, this impressive performance raises the question: Are these robustness gains provided by randomized ensembles real? In this work we address this question both theoretically and empirically. We first establish theoretically that commonly employed robustness evaluation methods such as adaptive PGD provide a false sense of security in this setting. Subsequently, we propose a theoretically-sound and efficient adversarial attack algorithm (ARC) capable of compromising random ensembles even in cases where adaptive PGD fails to do so. We conduct comprehensive experiments across a variety of network architectures, training schemes, datasets, and norms to support our claims, and empirically establish that randomized ensembles are in fact more vulnerable to $\ell_p$-bounded adversarial perturbations than even standard AT models. Our code can be found at https://github.com/hsndbk4/ARC. Hassan Dbouk, Naresh R. Shanbhag |
ICML | 2 |
| 2022 | Fundamental Limits on the Computational Accuracy of Resistive Crossbar-based In-memory ArchitecturesabstractIn-memory computing (IMC) architectures exhibit an intrinsic trade-off between computational accuracy and energy efficiency. This paper determines the fundamental limits on the compute SNR of MRAM-, ReRAM-, and FeFET-based crossbars by employing statistical signal and noise models. For a specific dot-product dimension N, the maximum compute SNR (SNRmax) is shown to occur at an optimum value of sensing resistance $R_{s}^{*}$ where clipping and quantization noise contributions from the analog-to-digital converter (ADC) are balanced out. SNRmaxcan be further improved by choosing devices with higher resistive contrast Roff/Ron, e.g., FeFET, but only until it attains a value in the range 12-15. Beyond this point, mismatch in the input digital-to-analog converters (DACs) and bitcell variations begin to dominate the compute SNR. Finally, by mapping a ResNet20 (CIFAR-10) network onto resistive crossbars, it is shown that the array-level compute SNR maximizing circuit parameters also maximizes the network-level accuracy. Saion K. Roy, Ameya Patil 0001, Naresh R. Shanbhag |
ISCAS | 3 |
| 2022 | Fundamental Limits on Energy-Delay-Accuracy of In-Memory Architectures in Inference ApplicationsabstractThis article obtains fundamental limits on the computational precision of in-memory computing architectures (IMCs). An IMC noise model and associated signal-to-noise ratio (SNR) metrics are defined and their interrelationships analyzed to show that the accuracy of IMCs is fundamentally limited by the compute SNR (${\mathrm {SNR}}_{a}$) of its analog core, and that activation, weight, and output (ADC) precision needs to be assigned appropriately for the final output SNR (${\mathrm {SNR}}_{T}$) to approach${\mathrm {SNR}}_{a}$. The minimum precision criterion (MPC) is proposed to minimize the analog-to-digital converter (ADC) precision and hence its overhead. Three in-memory compute models—charge summing (QS), current summing (IS), and charge redistribution (QR)—are shown to underlie most known IMCs. Noise, energy, and delay expressions for the compute models are developed and employed to derive expressions for the SNR, ADC precision, energy, and latency of IMCs. The compute SNR expressions are validated via Monte Carlo simulations in a 65 nm CMOS process. For a 512 row SRAM array, it is shown that: 1) IMCs have an upper bound on their maximum achievable${\mathrm {SNR}}_{a}$due to constraints on energy, area and voltage swing, and this upper bound reduces with technology scaling for QS-based architectures; 2) MPC enables${\mathrm {SNR}}_{T}$to approach${\mathrm {SNR}}_{a}$to be realized with minimal ADC precision; and 3) QS-based (QR-based) architectures are preferred for low (high) compute SNR scenarios. Sujan K. Gonugondla, Charbel Sakr, Hassan Dbouk, Naresh R. Shanbhag |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Optimizing Selective Protection for CNN ResilienceabstractAs CNNs are being extensively employed in high performance and safety-critical applications that demand high reliability, it is important to ensure that they are resilient to transient hardware errors. Traditional full redundancy solutions provide high error coverage, but the associated overheads are often prohibitively high for resource-constrained systems. In this work, we propose software-directed selective protection techniques to target the most vulnerable work in a CNN, providing a low-cost solution. We propose and evaluate two domain-specific selective protection techniques for CNNs that target different granularities. First, we develop a feature-map level resilience technique (FLR), which identifies and statically protects the most vulnerable feature maps in a CNN. Second, we develop an inference level resilience technique (ILR), which selectively reruns vulnerable inferences by analyzing their output. Third, we show that the combination of both techniques (FILR) is highly efficient, achieving nearly full error coverage (99.78% on average) for quantized inferences via selective protection. Our tunable approach enables developers to evaluate CNN resilience to hardware errors before deployment using MAC operations as overhead for quicker trade-off analysis. For example, targeting 100% error coverage on ResNet50 with FILR requires 20.8% additional MACs, while measurements on a Jetson Xavier GPU shows 4.6% runtime overhead. Abdulrahman Mahmoud, Siva Kumar Sastry Hari, Christopher W. Fletcher, Sarita V. Adve, Charbel Sakr, Naresh R. Shanbhag, Pavlo Molchanov 0001, Michael B. Sullivan 0001, Timothy Tsai 0002, Stephen W. Keckler |
ISSRE | 6 |
| 2021 | Generalized Depthwise-Separable Convolutions for Adversarially Robust and Efficient Neural NetworksabstractDespite their tremendous successes, convolutional neural networks (CNNs) incur high computational/storage costs and are vulnerable to adversarial perturbations. Recent works on robust model compression address these challenges by combining model compression techniques with adversarial training. But these methods are unable to improve throughput (frames-per-second) on real-life hardware while simultaneously preserving robustness to adversarial perturbations. To overcome this problem, we propose the method of Generalized Depthwise-Separable (GDWS) convolution - an efficient, universal, post-training approximation of a standard 2D convolution. GDWS dramatically improves the throughput of a standard pre-trained network on real-life hardware while preserving its robustness. Lastly, GDWS is scalable to large problem sizes since it operates on pre-trained models and doesn't require any additional training. We establish the optimality of GDWS as a 2D convolution approximator and present exact algorithms for constructing optimal GDWS convolutions under complexity and error constraints. We demonstrate the effectiveness of GDWS via extensive experiments on CIFAR-10, SVHN, and ImageNet datasets. Our code can be found at https://github.com/hsndbk4/GDWS. Hassan Dbouk, Naresh R. Shanbhag |
NeurIPS | 2 |
| 2020 | DBQ: A Differentiable Branch Quantizer for Lightweight Deep Neural Networks
Hassan Dbouk, Hetul Sanghvi, Mahesh Mehendale, Naresh R. Shanbhag |
ECCV (27) | 4 |
| 2020 | Low-Complexity Fixed-Point Convolutional Neural Networks For Automatic Target RecognitionabstractThere has been growing interest in developing neural network based automatic target recognition systems for synthetic aperture radar applications. However, these networks are typically complex in terms of storage and computation which inhibits their deployment in the field, where such resources are heavily constrained. In order to bring the cost of implementing these networks down, we develop a set of compact network architectures and train them in fixed-point. Our proposed method achieves an overall 984 reduction in terms of storage requirements and 71 × reduction in terms of computational complexity compared to state-of-the-art con-volutional neural networks for automatic target recognition (ATR), while maintaining a classification accuracy of > 99% on the MSTAR dataset. Hassan Dbouk, Hanfei Geng, Craig M. Vineyard, Naresh R. Shanbhag |
ICASSP | 4 |
| 2020 | SWIPE: Enhancing Robustness of ReRAM Crossbars for In-memory ComputingabstractCrossbar-based in-memory architectures have emerged as an attractive platform for energy-efficient realization of deep neural networks (DNNs). A key challenge in such architectures is achieving accurate and efficient writes due to the presence of bitcell conductance variations. In this paper, we propose the Single-Write In-memory Program-vErify (SWIPE) method that achieves high accuracy writes for crossbar-based in-memory architectures at 5×-to-10× lower cost than standard program-verify methods. SWIPE leverages the bit-sliced attribute of crossbar-based in-memory architectures and the statistics of conductance variations to compensate for device non-idealities. Using SWIPE to write into ReRAM crossbar allows for a 2× (CIFAR-10) and 3× (MNIST) increase in storage density with < 1% loss in DNN accuracy. In particular, SWIPE compensates for 4.8×-to-7.7× higher conductance variations. Furthermore, SWIPE can be augmented with injection-based training methods in order to achieve even greater enhancements in robustness. Sujan K. Gonugondla, Ameya Patil 0001, Naresh R. Shanbhag |
ICCAD | 3 |
| 2020 | Fundamental Limits on the Precision of In-memory ArchitecturesabstractThis paper obtains the fundamental limits on the computational precision of in-memory computing architectures (IMCs). Various compute SNR metrics for IMCs are defined and their interrelationships analyzed to show that the accuracy of IMCs is fundamentally limited by the compute SNR (SNRa) of its analog core, and that activation, weight and output precision needs to be assigned appropriately for the final output SNR SNRT → SNRa. The minimum precision criterion (MPC) is proposed to minimize the output and hence the column analog-to-digital converter (ADC) precision. The charge summing (QS) compute model and its associated IMC QS-Arch are studied to obtain analytical models for its compute SNR, minimum ADC precision, energy and latency. Compute SNR models of QS-Arch are validated via Monte Carlo simulations in a 65 nm CMOS process. Employing these models, upper bounds on SNRa of a QS-Arch-based IMC employing a 512 row SRAM array are obtained and it is shown that QS-Arch's energy cost reduces by 3.3× for every 6 dB drop in SNRa, and that the maximum achievable SNRa reduces with technology scaling while the energy cost at the same SNRa increases. These models also indicate the existence of an upper bound on the dot product dimension N due to voltage headroom clipping, and this bound can be doubled for every 3 dB drop in SNRa. Sujan K. Gonugondla, Charbel Sakr, Hassan Dbouk, Naresh R. Shanbhag |
ICCAD | 4 |
| 2020 | Deep In-Memory Architectures in SRAM: An Analog Approach to Approximate ComputingabstractThis article provides an overview of recently proposed deep in-memory architectures (DIMAs) in SRAM for energyand latency-efficient hardware realization of machine learning (ML) algorithms. DIMA tackles the data movement problem in von Neumann architectures head-on by deeply embedding mixed-signal computations into a conventional memory array. In doing so, it trades off its computational signal-to-noise ratio (compute SNR) with energy and latency, and therefore, it represents an analog form of approximate computing. DIMA exploits the inherent error immunity of ML algorithms and SNR budgeting methods to operate its analog circuitry in a low-swing/low-compute SNR regime, thereby achieving >100× reduction in the energy-delay product (EDP) over an equivalent von Neumann architecture with no loss in inference accuracy. This article describes DIMA's computational pipeline and provides a Shannon-inspired rationale for its robustness to process, temperature, and voltage variations and design guidelines to manage its analog nonidealities. DIMA's versatility, effectiveness, and practicality demonstrated via multiple silicon IC prototypes in a 65-nm CMOS process are described. A DIMA-based instruction set architecture (ISA) to realize an end-to-end application-toarchitecture mapping for the accelerating diverse ML algorithms is also presented. Finally, DIMA's fundamental tradeoff between energy and accuracy in the low-compute SNR regime is analyzed to determine energy-optimum design parameters. Mingu Kang, Sujan K. Gonugondla, Naresh R. Shanbhag |
Proc. IEEE | 3 |
| 2019 | Per-Tensor Fixed-Point Quantization of the Back-Propagation Algorithm
Charbel Sakr, Naresh R. Shanbhag |
ICLR (Poster) | 2 |
| 2019 | Accumulation Bit-Width Scaling For Ultra-Low Precision Training Of Deep Networks
Charbel Sakr, Naigang Wang, Chia-Yu Chen, Jungwook Choi, Ankur Agrawal, Naresh R. Shanbhag, Kailash Gopalakrishnan |
ICLR (Poster) | 6 |
| 2019 | An MRAM-Based Deep In-Memory Architecture for Deep Neural NetworksabstractThis paper presents an MRAM-based deep in-memory architecture (MRAM-DIMA) to efficiently implement multi-bit matrix vector multiplication for deep neural networks using a standard MRAM bitcell array. The MRAM-DIMA achieves an 4.5 × and 70× lower energy and delay, respectively, compared to a conventional digital MRAM architecture. Behavioral models are developed to estimate the impact of circuit non-idealities, including process variations, on the DNN accuracy. An accuracy drop of ≤ 0.5% (≤ 1%) is observed for LeNet-300-100 on the MNIST dataset (a 9-layer CNN on the CIFAR-10 dataset), while tolerating 24% (12%) variation in cell conductance in a commercial 22 nm CMOS-MRAM process. Ameya Patil 0001, Haocheng Hua, Sujan K. Gonugondla, Mingu Kang, Naresh R. Shanbhag |
ISCAS | 5 |
| 2019 | An Energy-Efficient Classifier via Boosted Spin Channel NetworksabstractWith diminishing energy and delay benefits via CMOS scaling, there is much interest in exploring the use of alternative state variables such as electronic spin. Multiple research efforts are underway exploring both Boolean and non-Boolean design space using spin devices in order to make their energy and delay benefits competitive to CMOS. In this paper, we propose spin channel networks (SCNs) - spin-based circuits that exploit exponential decay of spin current to efficiently realize multi-bit dot product computation. We show that proposed SCNs can be employed with adaptive boosting (AdaBoost) learning algorithm to efficiently realize a binary classifier for breast cancer detection. The proposed SCN implementation achieves 112× and 14× lower energy per decision compared to the conventional all spin logic (ASL) and 20nm CMOS designs, respectively, for identical decision throughput. Ameya Patil 0001, Sasikanth Manipatruni, Dmitri E. Nikonov, Ian A. Young, Naresh R. Shanbhag |
ISCAS | 5 |
| 2019 | Shannon-Inspired Statistical Computing for the Nanoscale EraabstractModern day computing systems are based on the von Neumann architecture proposed in 1945 but face dual challenges of: 1) unique data-centric requirements of emerging applications and 2) increased nondeterminism of nanoscale technologies caused by process variations and failures. This paper presents a Shannon-inspired statistical model of computation (statistical computing) that addresses the statistical attributes of both emerging cognitive workloads and nanoscale fabrics within a common framework. Statistical computing is a principled approach to the design of non-von Neumann architectures. It emphasizes the use of information-based metrics; enables the determination of fundamental limits on energy, latency, and accuracy; guides the exploration of statistical design principles for low signal-to-noise ratio (SNR) circuit fabrics and architectures such as deep in-memory architecture (DIMA) and deep in-sensor architecture (DISA); and thereby provides a framework for the design of computing systems that approach the limits of energy efficiency, latency, and accuracy. From its early origins, Shannon-inspired statistical computing has grown into a concrete design framework validated extensively via both theory and laboratory prototypes in both CMOS and beyond. The framework continues to grow at both of these levels, yielding new ways of connecting systems through architectures, circuits, and devices, for the semiconductor roadmap to march into the nanoscale era. Naresh R. Shanbhag, Naveen Verma, Yongjune Kim 0001, Ameya Patil 0001, Lav R. Varshney |
Proc. IEEE | 1 |
| 2018 | True Gradient-Based Training of Deep Binary Activated Neural Networks Via Continuous BinarizationabstractWith the ever growing popularity of deep learning, the tremendous complexity of deep neural networks is becoming problematic when one considers inference on resource constrained platforms. Binary networks have emerged as a potential solution, however, they exhibit a fundamentallimi-tation in realizing gradient-based learning as their activations are non-differentiable. Current work has so far relied on approximating gradients in order to use the back-propagation algorithm via the straight through estimator (STE). Such approximations harm the quality of the training procedure causing a noticeable gap in accuracy between binary neural networks and their full precision baselines. We present a novel method to train binary activated neural networks using true gradient-based learning. Our idea is motivated by the similarities between clipping and binary activation functions. We show that our method has minimal accuracy degradation with respect to the full precision baseline. Finally, we test our method on three benchmarking datasets: MNIST, CIFAR-10, and SVHN. For each benchmark, we show that continuous binarization using true gradient-based learning achieves an accuracy within 1.5% of the floating-point baseline, as compared to accuracy drops as high as 6% when training the same binary activated network using the STE. Charbel Sakr, Jungwook Choi, Kailash Gopalakrishnan, Naresh R. Shanbhag |
ICASSP | 5 |
| 2018 | An Analytical Method to Determine Minimum Per-Layer Precision of Deep Neural NetworksabstractThere has been growing interest in the deployment of deep learning systems onto resource-constrained platforms for fast and efficient inference. However, typical models are overwhelmingly complex, making such integration very challenging and requiring compression mechanisms such as reduced precision. We present a layer-wise granular precision analysis which allows us to efficiently quantize pre-trained deep neural networks at minimal cost in terms of accuracy degradation. Our results are consistent with recent findings that perturbations in earlier layers are most destructive and hence needing more precision than in later layers. Our approach allows for significant complexity reduction demonstrated by numerical results on the MNIST and CIFAR-10 datasets. Indeed, for an equivalent level of accuracy, our fine-grained approach reduces the minimum precision in the network up to 8 bits over a naive uniform assignment. Furthermore, we match the accuracy level of a state-of-the-art binary network while requiring up to ~3.5× lower complexity. Similarly, when compared to a state-of-the-art fixed-point network, the complexity savings are even higher (up to ~14×) with no loss in accuracy. Charbel Sakr, Naresh R. Shanbhag |
ICASSP | 2 |
| 2018 | PROMISE: An End-to-End Design of a Programmable Mixed-Signal Accelerator for Machine-Learning AlgorithmsabstractAnalog/mixed-signal machine learning (ML) accelerators exploit the unique computing capability of analog/mixed-signal circuits and inherent error tolerance of ML algorithms to obtain higher energy efficiencies than digital ML accelerators. Unfortunately, these analog/mixed-signal ML accelerators lack programmability, and even instruction set interfaces, to support diverse ML algorithms or to enable essential software control over the energy-vs-accuracy tradeoffs. We propose PROMISE, the first end-to-end design of a PROgrammable MIxed-Signal accElerator from Instruction Set Architecture (ISA) to high-level language compiler for acceleration of diverse ML algorithms. We first identify prevalent operations in widely-used ML algorithms and key constraints in supporting these operations for a programmable mixed-signal accelerator. Second, based on that analysis, we propose an ISA with a PROMISE architecture built with silicon-validated components for mixed-signal operations. Third, we develop a compiler that can take a ML algorithm described in a high-level programming language (Julia) and generate PROMISE code, with an IR design that is both language-neutral and abstracts away unnecessary hardware details. Fourth, we show how the compiler can map an application-level error tolerance specification for neural network applications down to low-level hardware parameters (swing voltages for each application Task) to minimize energy consumption. Our experiments show that PROMISE can accelerate diverse ML algorithms with energy efficiency competitive even with fixed-function digital ASICs for specific ML algorithms, and the compiler optimization achieves significant additional energy savings even for only 1% extra errors. Prakalp Srivastava, Mingu Kang, Sujan K. Gonugondla, Sungmin Lim, Jungwook Choi, Vikram S. Adve, Nam Sung Kim, Naresh R. Shanbhag |
ISCA | 8 |
| 2018 | Energy-Efficient Deep In-memory Architecture for NAND Flash MemoriesabstractThis paper proposes an energy-efficient deep in-memory architecture for NAND flash (DIMA-F) to perform machine learning and inference algorithms on NAND flash memory. Algorithms for data analytics, inference, and decision-making require processing of large data volumes and are hence limited by data access costs. DIMA-F achieves energy savings and throughput improvement for such algorithms by reading and processing data in the analog domain at the periphery of NAND flash memory. This paper also provides behavioral models of DIMA-F that can be used for analysis and large scale system simulations in presence of circuit non-idealities and variations. DIMA-F is studied in the context of linear support vector machines and k-nearest neighbor for face detection and recognition, respectively. An estimated 8×-to-23× reduction in energy and 9×-to-15× improvement in throughput resulting in EDP gains up to 345× over the conventional NAND flash architecture incorporating an external digital ASIC for computation. Sujan K. Gonugondla, Mingu Kang, Yongjune Kim 0001, Mark Helm, Sean Eilert, Naresh R. Shanbhag |
ISCAS | 6 |
| 2018 | SRAM Bit-line Swings Optimization using Generalized WaterfillingabstractWe propose an information-theoretic approach to optimize non-uniform bit-line swings for static random access memories (SRAMs). We formulate convex optimization problems whose objectives are to minimize energy (for low-power SRAMs), maximize speed (for high-speed SRAMs), and minimize energy-delay product for a given constraint on mean squared error of retrieved words. We show that these optimization problems can be interpreted as generalized water-filling including classical waterfilling, ground-flattening and water-filling, and sand-pouring and water-filling, respectively. Numerical results show that energy-optimal swing assignment reduces energy consumption by half at a peak signal-to-noise ratio of 30dB for an 8-bit accessed word. Yongjune Kim 0001, Mingu Kang, Lav R. Varshney, Naresh R. Shanbhag |
ISIT | 4 |
| 2018 | Generalized Water-Filling for Source-Aware Energy-Efficient SRAMsabstractConventional low-power static random access memories (SRAMs) reduce read energy by decreasing the bit-line voltage swings uniformly across the bit-line columns. This is because the read energy is proportional to the bit-line swings. On the other hand, bit-line swings are limited by the need to avoid decision errors especially in the most significant bits. We propose a principled approach to determine optimal non-uniform bit-line swings by formulating convex optimization problems. For a given constraint on mean squared error of retrieved words, we consider criteria to minimize energy (for low-power SRAMs), maximize speed (for high-speed SRAMs), and minimize energy-delay product. These optimization problems can be interpreted as classical water-filling, ground-flattening and water-filling, and sand-pouring and water-filling, respectively. By leveraging these interpretations, we also propose greedy algorithms to obtain optimized discrete swings. Numerical results show that energy-optimal swing assignment reduces energy consumption by half at a peak signal-to-noise ratio of 30 dB for an 8-bit accessed word. The energy savings increase to four times for a 16-bit accessed word. Yongjune Kim 0001, Mingu Kang, Lav R. Varshney, Naresh R. Shanbhag |
IEEE Trans. Commun. | 4 |
| 2017 | A Systems Approach to Computing in Beyond CMOS Fabrics: InvitedabstractNo abstract available. Ameya Patil 0001, Naresh R. Shanbhag, Lav R. Varshney, Eric Pop, H.-S. Philip Wong, Subhasish Mitra, Jan M. Rabaey, Jeffrey A. Weldon, Lawrence T. Pileggi, Sasikanth Manipatruni, Dmitri E. Nikonov, Ian A. Young |
DAC | 2 |
| 2017 | Minimum precision requirements for the SVM-SGD learning algorithmabstractIt is well-known that the precision of data, weight vector, and internal representations employed in learning systems directly impacts their energy, throughput, and latency. The precision requirements for the training algorithm are also important for systems that learn on-the-fly. In this paper, we present analytical lower bounds on the precision requirements for the commonly employed stochastic gradient descent (SGD) on-line learning algorithm in the specific context of a support vector machine (SVM). These bounds are obtained subject to desired system performance. These bounds are validated using the UCI breast cancer dataset. Additionally, the impact of these precisions on the energy consumption of a fixed-point SVM with on-line training is studied. Simulation results in 45 nm CMOS process show that operating at the minimum precision as dictated by our bounds improves energy consumption by a factor of 5.3× as compared to conventional precision assignments with no observable loss in accuracy. Charbel Sakr, Ameya Patil 0001, Yongjune Kim 0001, Naresh R. Shanbhag |
ICASSP | 5 |
| 2017 | Analytical Guarantees on Numerical Precision of Deep Neural NetworksabstractThe acclaimed successes of neural networks often overshadow their tremendous complexity. We focus on numerical precision – a key parameter defining the complexity of neural networks. First, we present theoretical bounds on the accuracy in presence of limited precision. Interestingly, these bounds can be computed via the back-propagation algorithm. Hence, by combining our theoretical analysis and the back-propagation algorithm, we are able to readily determine the minimum precision needed to preserve accuracy without having to resort to time-consuming fixed-point simulations. We provide numerical evidence showing how our approach allows us to maintain high accuracy but with lower complexity than state-of-the-art binary networks. Charbel Sakr, Yongjune Kim 0001, Naresh R. Shanbhag |
ICML | 3 |
| 2017 | PredictiveNet: An energy-efficient convolutional neural network via zero predictionabstractConvolutional neural networks (CNNs) have gained considerable interest due to their record-breaking performance in many recognition tasks. However, the computational complexity of CNNs precludes their deployments on power-constrained embedded platforms. In this paper, we propose predictive CNN (PredictiveNet), which predicts the sparse outputs of the non-linear layers thereby bypassing a majority of computations. PredictiveNet skips a large fraction of convolutions in CNNs at runtime without modifying the CNN structure or requiring additional branch networks. Analysis supported by simulations is provided to justify the proposed technique in terms of its capability to preserve the mean square error (MSE) of the nonlinear layer outputs. When applied to a CNN for handwritten digit recognition, simulation results show that PredictiveNet can reduce the computational cost by a factor of 2.9χ compared to a state-of-the-art CNN, while incurring marginal accuracy degradation. Yingyan (Celine) Lin, Charbel Sakr, Yongjune Kim 0001, Naresh R. Shanbhag |
ISCAS | 4 |
| 2017 | Slicer Architectures for Analog-to-Information Conversion in Channel EqualizersabstractThe scaling of analog-to-digital converter (ADC) power consumption with communication bandwidth imposes severe limits on its precision, which significantly impacts receiver performance. In this paper, we consider a “space-time” generalization of the flash architecture by allowing a fixed number of slicers to be dispersed in time (i.e., sampling offset) as well as space (i.e., amplitude), with the goal of investigating its capabilities for analog-to-information conversion (i.e., enabling reliable recovery of digital information, rather than faithful reproduction of the input signal) in the context of channel equalization for binary signaling over a dispersive channel. We first study standard symbol-spaced ADC with severe quantization constraints, estimating the minimum number of slicers needed to avoid error floors. We observe that the performance is sensitive to channel realization and sampling phase, which motivates a more flexible space-time architecture. Using ideas similar to those underlying compressive sensing, we prove that such architectures have no fundamental limitations in theory: randomly dispersing enough one-bit slicers over space and time does provide information sufficient for reliable equalization. We then focus on practical designs for symbol-spaced and fractionally-spaced sampling subject to a constraint on the number of slicers, and propose an algorithm for optimizing slicer thresholds, which significantly improves performance over a standard design. Aseem Wadhwa, Upamanyu Madhow, Naresh R. Shanbhag |
IEEE Trans. Commun. | 3 |
| 2016 | Probabilistic Error Models for machine learning kernels implemented on stochastic nanoscale fabrics
Naresh R. Shanbhag |
DATE | 2 |
| 2016 | Analysis of error resiliency of belief propagation in computer visionabstractProbabilistic inference is a versatile tool to solve a large variety of pixel-labeling problems in computer vision such as stereo matching and image denoising. Belief Propagation (BP) is an effective method for such inference tasks, and has also shown attractive error-resilience properties—the ability to converge to usable solutions in the presence of low-level hardware errors. This is of increasing interest, as the looming end of Moore's Law scaling brings with it a vast increase in the statistical variability of nanoscale circuit fabrics. In this work we seek to understand why certain combinations of BP and error-resilience mechanisms work so well in practice. We focus on Algorithmic Noise Tolerance (ANT) techniques for the resilience mechanisms, and Max-Product BP for inference. We analyze the error characteristics of BP in this hardware context, derive novel asymptotic error bounds, and provide theoretical reasoning to explain why ANT works well in this BP context. Experimental results from detailed resilient-BP simulations for various stereo matching tasks offer empirical support for this analysis. Jungwook Choi, Ameya Patil 0001, Rob A. Rutenbar, Naresh R. Shanbhag |
ICASSP | 4 |
| 2016 | Perfect error compensation via algorithmic error cancellationabstractThis paper presents a novel statistical error compensation (SEC) technique — algorithmic error cancellation (AEC)-for designing robust and energy-efficient signal processing and machine learning kernels on scaled process technologies. AEC exhibits a perfect error compensation (PEC) property, i.e., it is able to achieve a post-compensation error rate equal to zero. AEC generates a maximum likelihood (ML) estimate of the hardware error and employs it for error cancellation. AEC is applied to a voltage overscaled 45-tap, 45nm CMOS finite impulse response (FIR) filter employed in a EEG seizure detection system. AEC is shown to perfectly compensate for errors in the main FIR block and its reduced precision replica when they make errors at a rate of up to 73% and 98%, respectively. The AEC-based FIR is compared with an uncompensated architecture, and a fast architecture. AEC's error compensation capability enables it to achieve a 31.5% (at same supply voltage) and 19.7% (at same energy) speed-up over the uncompensated architecture, and a 8. 9% speed-up over a fast architecture at the same energy consumption. At fd, k = 452.3 MHz, AEC results in a 27.7% and 12.4% energy savings over the uncompensated and fast architectures, respectively. Sujan K. Gonugondla, Byonghyo Shim, Naresh R. Shanbhag |
ICASSP | 3 |
| 2016 | Error Resilient and Energy Efficient MRF Message-Passing-Based Stereo MatchingabstractMessage-passing-based inference algorithms have immense importance in real-world applications. In this paper, error resiliency of a message passing based Markov random field (MRF) stereo matching hardware is explored and enhanced through the application of statistical error compensation. Error resiliency is of particular interest for subnanometer and postsilicon devices. The inherent robustness of iteration-based MRF inference algorithms is explored and shows that small errors are tolerable, while large errors degrade the performance significantly. Based on these error characteristics, algorithmic noise tolerance (ANT) has been applied at the arithmetic, iteration, and system levels. Introducing timing errors via voltage overscaling, at the arithmetic level, results show that the ANT-based hardware can tolerate an error rate of 21.3%, with performance degradation of only 3.5% at an overhead of 97.4%, compared with an error-free hardware with an energy savings of 39.7%. To reduce compensation complexity, iteration and system-level compensation was explored. Results show that, compared with arithmetic level, system-level compensation reduces overhead to 59%, while maintaining stereo matching performance with only 2.5% degradation with 16% additional power savings. These results are verified via FPGA emulation with timing errors induced within the message passing unit via relaxed synthesis. Eric P. Kim, Jungwook Choi, Naresh R. Shanbhag, Rob A. Rutenbar |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | An energy-efficient memory-based high-throughput VLSI architecture for convolutional networksabstractIn this paper, an energy efficient, memory-intensive, and high throughput VLSI architecture is proposed for convolutional networks (C-Net) by employing compute memory (CM) [1], where computation is deeply embedded into the memory (SRAM). Behavioral models incorporating CM's circuit non-idealities and energy models in 45nm SOI CMOS are presented. System-level simulations using these models demonstrate that the probability of handwritten digit recognition Pr> 0.99 can be achieved using the MNIST database [2], along with a 24.5× reduced energy delay product, a 5.0× reduced energy, and a 4.9× higher throughput as compared to the conventional system. Mingu Kang, Sujan K. Gonugondla, Min-Sun Keel, Naresh R. Shanbhag |
ICASSP | 4 |
| 2015 | Reduced Overhead Error Compensation for Energy Efficient Machine Learning KernelsabstractLow overhead error-resiliency techniques such as RAZOR [1] and algorithmic noise-tolerance (ANT) [2] have proven effective in reducing energy consumption. ANT has been shown to be particularly effective for signal processing and machine learning kernels. In ANT, an explicit estimator block compensates for large magnitude errors in a main block. The estimator represents the overhead in ANT and can be as large as 30%. This paper presents a low overhead ANT technique referred to as ALG-ANT. In ALG-ANT, the estimator is embedded inside the main block via algorithmic reformulation and thus completely eliminates the overhead associated with ANT. However, ALG-ANT is algorithm-specific. This paper demonstrates the ALG-ANT concept in the context of a finite impulse response (FIR) filter kernel and a dot product kernel, both of which are commonly employed in signal processing and machine learning applications. The proposed ALG-ANT FIR filter and dot product kernels are applied to the feature extractor (FE) and SVM classification engine (CE) of an EEG seizure classification system. Simulation results in a commercial 45nm CMOS process show that ALG-ANT can compensate for error rates of up to 0.41 (errors in FE only), and up to 0.19 (errors in FE and CE) and maintain the true positive rate ptp > 0.9 and false positive rate pfp ≤ 0.01. This represents a greater than 3-orders-of-magnitude improvement in error tolerance over the conventional architecture. This error tolerance is employed to reduce energy via the use of voltage overscaling (VOS). ALG-ANT is able to achieve 44.3% energy savings when errors are in FE only, and up to 37.1% savings when errors are in both FE and CE. Naresh R. Shanbhag |
ICCAD | 2 |
| 2015 | Energy-efficient and high throughput sparse distributed memory architectureabstractThis paper presents an energy-efficient VLSI implementation of Sparse Distributed Memory (SDM). High throughput and energy-efficient Hamming distance-based address decoder (CM-DEC) is proposed by employing compute memory [1], where computation is deeply embedded into a memory (SRAM). Hierarchical binary decision (HBD) is also proposed to enhance area- and energy-efficiency of read operation by minimizing data transfer. The SDM is employed as an auto-associative memory with four read iterations and 16×16 binary noisy input image with input error rates of 15%, 25%, and 30%. The proposed SDM achieves 39× smaller energy delay product with 14.5× and 2.7× reduced delay and energy, respectively as compared to conventional digital implementation of SDM in 45 nm SOI CMOS process with output error rate degradation less than 0.4%. Mingu Kang, Eric P. Kim, Min-Sun Keel, Naresh R. Shanbhag |
ISCAS | 4 |
| 2015 | Statistical information processing: Computing for the nanoscale eraabstractComputing platforms operating at the limits of energy-efficiency need to contend with the issue of robustness. This energy vs. robustness trade-off is fundamental in such systems. This talk will describe a Shannon-inspired framework referred to as statistical information processing (SIP). SIP navigates the energy vs. robustness trade-off by treating the problem of energy-efficient computing as one of information processing on low-SNR and unreliable nanoscale device/circuit fabrics. In doing do, SIP seeks to transform computing from its von Neumann roots in data processing to a Shannon-inspired foundation for information processing. Key elements of SIP are the use of information-based metrics, a stochastic low-SNR circuit fabric, and statistical error compensation techniques based on estimation and detection theory, and machine learning. SIP has been used for designing energy-efficient and robust computation, communication, storage, and mixed-signal analog front-ends. This talk will conclude with a brief overview of the Systems On Nanoscale Information fabriCs (SONIC) Center, a 5-year multi-university research center, focused on developing a Shannon/brain-inspired foundation for information processing on CMOS and beyond CMOS nanoscale fabrics. Naresh R. Shanbhag |
ISLPED | 1 |
| 2015 | A 3.6-mW 50-MHz PN Code Acquisition Filter via Statistical Error Compensation in 180-nm CMOSabstractIn this brief, we present a novel architecture for pseudorandom (PN) code acquisition based on statistical error compensation (SEC), which achieves significant power savings. SEC treats errors in hardware as noise in communication networks, and employs robust estimation theory to compensate for errors. We apply SEC to a 256-tap PN code acquisition filter in a 180-nm CMOS process. Multiple (five) dies were tested under voltage overscaling to achieve a near constant detection probability (Pdet) above 90%. The minimum energy consumption ranged from 72.89 to 210.59 pJ (ave 122.52 pJ) for supply voltages between 0.69 and 0.70 V. These operating conditions result in raw error rates of 85.83%-91.23% (ave 88.99%). Energy savings over a conventional errorfree design ranges from 2.4× to 5.8× (ave 3.86×). Energy savings over past work ranges from 1.55× to 3.79× (ave 2.52×). Improvement in error-tolerance over existing error-tolerant designs range from 2146× to 2281× (ave 2225×). The large energy savings were found to be due to a combination of voltage scaling and activity factor reduction. The proposed design achieves a 2.5× improvement in the figure of merit [normalized power/(#taps * precision * sample rate)] compared with conventional PN code acquisition filters. Eric P. Kim, Daniel J. Baker, Sriram Narayanan, Naresh R. Shanbhag, Douglas L. Jones |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2014 | A robust message passing based stereo matching kernel via system-level error resiliencyabstractIn this paper, we present an error resilient Markov random field (MRF) message passing based stereo matching hardware (HW) architecture. Previously, algorithmic noise tolerance (ANT) has been applied at the arithmetic level of the reparameterize unit and showed greatly enhanced robustness of message passing inference based architectures. In this work, correction was targeted at the system level to reduce correction overhead while maintaining performance. An erroneous FPGA based accelerator was employed as our emulation platform. Through relaxed synthesis, we show that timing errors occur within the message passing unit, and are successfully compensated. Error correction has been implemented at several hierarchical levels, including end of iteration, and the final depth map output. Significant enhancement in robustness is achieved with minimal correction overhead. Compared to HW error compensation at the arithmetic level, system level error compensation reduces overhead by more than 50 %, while maintaining stereo matching performance with only 3.8 % degradation. Eric P. Kim, Jungwook Choi, Naresh R. Shanbhag, Rob A. Rutenbar |
ICASSP | 3 |
| 2014 | An energy-efficient VLSI architecture for pattern recognition via deep embedding of computation in SRAMabstractIn this paper, we propose the concept of compute memory, where computation is deeply embedded into the memory (SRAM). This deep embedding enables multi-row read access and analog signal processing. Compute memory exploits the relaxed precision and linearity requirements of pattern recognition applications. System-level simulations incorporating various deterministic errors from analog signal chain demonstrates the limited accuracy of analog processing does not significantly degrade the system performance, which means the probability of pattern detection is minimally impacted. The estimated energy saving is 63 % as compared to the conventional system with standard embedded memory and parallel processing architecture, for 256×256 target image. Mingu Kong, Min-Sun Keel, Naresh R. Shanbhag, Sean Eilert, Ken Curewitz |
ICASSP | 3 |
| 2014 | Space-time slicer architectures for analog-to-information conversion in channel equalizersabstractAs modern communication transceivers scale to multi-Gbps speeds, the power consumption and cost of highresolution, high-speed analog-to-digital converters (ADCs) become a crucial bottleneck in realizing “mostly digital” receiver architectures that leverage Moore's law. This bottleneck could potentially be alleviated by designing analog front ends for the more specific goal of analog-to-information conversion (i.e., preserving the digital information residing in the received signal). As one possible approach towards this goal, we consider a generalization of the standard flash ADC: instead of implementing n bit quantization of a sample by passing it through 2n-1 slicers as in a standard ADC, the slicers are dispersed in time as well as space (i.e., amplitude). Considering BPSK over a dispersive channel, we first show, using ideas similar to those underlying compressive sensing, that randomly dispersing enough one-bit slicers over space and time does provide information sufficient for reliable demodulation over a dispersive channel. We then propose an iterative algorithm for optimizing the design of the sampling times and amplitude thresholds, and provide numerical results showing that the number of slicers can be significantly reduced relative to a conventional flash ADC with comparable bit error rate (BER). These system-level results motivate further investigation, in terms of both circuit and system design, into looking beyond conventional ADC architectures when designing analog front-ends for high-speed communication. Aseem Wadhwa, Upamanyu Madhow, Naresh R. Shanbhag |
ICC | 3 |
| 2014 | Energy-efficient dot product computation using a switched analog circuit architectureabstractIn this paper, we present switched analog circuit (SAC), a new circuit architecture, to implement an energy-efficient mixed-signal dot product (DP) kernel for machine learning and signal processing applications. SAC operates by fast switching the analog inputs to output via variable width digital pulses. The output accuracy and energy consumption of SAC is analyzed and verified for an average and Gaussian blur filter. Simulations in a commercial 130 nm process for a 120 x 120 image show energy savings of 19x-to-32x compared to a digital implementation for signal-to-noise ratios (SNRs) of 30 dB-to-24 dB, respectively. Ihab Nahlus, Eric P. Kim, Naresh R. Shanbhag, David T. Blaauw |
ISLPED | 3 |
| 2014 | Reducing Energy at the Minimum Energy Operating Point Via Statistical Error CompensationabstractThis paper demonstrates that statistical error compensation reduces the energy consumption Emin at the minimum energy operating point (MEOP), which is known to occur in the subthreshold regime. In particular, the impact of algorithmic noise-tolerance (ANT) [1], in conjunction with frequency overscaling (FOS) and voltage overscaling, is studied in the context of an eight-tap finite impulse response (FIR) filter in a 45-nm CMOS process. At the nominal process corner and using low-Vtdevices, we show that the ANT-based FIR filter achieves 20%-47% reduction in Emin and a 1.8×-2.25× increase in the frequency of operation over a conventional (error free) filter operating at its MEOP. This result is achieved via the ability of ANT to compensate for a precompensation error rate of 70%-85%. The use of high-Vtdevices reduces Emin by 10%. This is due to the reduced effectiveness of FOS and increased sensitivity of delay to voltage variations. In the presence of process variations, the ANT-based FIR filter reduces Emin by 54% over a transistor up-sized design while meeting a fixed throughput constraint, and a parametric yield of 99.7%. Rami A. Abdallah, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Statistical analysis of algorithmic noise toleranceabstractAlgorithmic noise tolerance (ANT) is an effective statistical error compensation technique for digital signal processing systems. This paper proves a long held hypothesis that ANT has a strong Bayesian foundation, and develops an analytical framework for predicting the performance of, and designing performance-optimal ANT-based systems. ANT is shown to approximate an optimal Bayesian detector and an optimal minimum mean squared error (MMSE) estimator. We show that the theoretically optimum threshold and the optimal threshold obtained via Monte Carlo simulations agree to within 8%, with performance degradation of at most 2.1% for a variety of error probability mass functions. For a 2D-DCT implemented in a 45nm CMOS process, we find similar results where the thresholds have a 7.8% difference. Furthermore, both analysis and simulations indicate that ANT's probability of error detection is robust to the choice of the threshold. Eric P. Kim, Naresh R. Shanbhag |
ICASSP | 2 |
| 2013 | Robust and Energy Efficient Multimedia Systems via Likelihood ProcessingabstractThis paper presents likelihood processing (LP) for designing robust and energy-efficient multimedia systems in the presence of nanoscale non-idealities. LP exploits error statistics of the underlying hardware to compute the probability of a particular bit being a one or a zero. Multiple output observations are generated via either: 1) modular redundancy (MR), 2) estimation, or 3) exploiting spatio-temporal correlation. Energy efficiency and robustness of a 2D discrete-cosine transform (DCT) image codec employing LP is studied. Simulations in a commercial 45-nm CMOS process show that LP can tolerate up to 100×, and 5× greater component error probability as compared to conventional and triple-MR (TMR)-based systems, respectively, while achieving a peak-signal-to-noise ratio (PSNR) of 30 dB at a pre-correction error rate of 20%. Furthermore, LP is able to achieve energy savings of 71% over TMR at a PSNR of 28 dB, while tolerating a pre-correction error rate of 4%. Rami A. Abdallah, Naresh R. Shanbhag |
IEEE Trans. Multim. | 2 |
| 2012 | A fully automated technique for constructing FSM abstractions of non-ideal latches in communication systemsabstractThe design of a communications system is typically most effective only when each of its components can be accurately represented by a discrete, symbolic behavioural abstraction. Such abstractions, in addition to providing valuable design intuition, also enable highly efficient and scalable system-level simulation. However, given a SPICE-level description for a subsystem such as a latch, it is a challenge to come up with a discrete, symbol-level abstraction that accurately captures its continuous-time dynamics. Indeed, the manual construction of such an abstraction requires deep knowledge and understanding of the operation of the module in question; moreover, it is very time-consuming, tedious, error-prone and not easily scalable to larger designs. In recent work [1], we adapted methods from computational learning theory to develop an automated technique, DAE2FSM, that produces binary finite state machine (FSM) abstractions of non-linear analog/mixed-signal (AMS) circuits. In the present paper, we demonstrate the application of the DAE2FSM technique to automatically derive FSM abstractions for a mixed-signal communications circuit component, namely a current mode latch (CML) designed in IBM's 90nm LP process technology. We show that the FSMs learned by DAE2FSM not only capture the essence of the latch's behaviour during normal conditions, but also faithfully mimic its behaviour under adverse operating conditions (e.g., under lowered supply voltages). Moreover, in addition to a stand-alone CML, we also generate FSMs for cascades of two and three latches (such topologies are used in the design of power-efficient, bit-error optimised analog-to-digital converters). In spite of the inherent non-linearity of such systems, and in spite of the pronounced “analog-ness” of the waveforms in question, our FSM abstractions are able to produce discrete-time symbol sequences that closely match the data points obtained by sampling from continuous-time SPICE simulations. Aadithya V. Karthik, Yingyan (Celine) Lin, Chenjie Gu, Aolin Xu 0001, Jaijeet S. Roychowdhury, Naresh R. Shanbhag |
ICASSP | 6 |
| 2012 | System-driven metrics for the design and adaptation of analog to digital convertersabstractIn this paper, we review some recent advances in the design of ADCs that exploit system-driven metrics, such as the bit-error rate in a communication link, or mutual information in a scheme employing forward error correction. We show, for example, that ADCs can be designed that maximize the information rate between the quantized output of the channel and the input to the channel for communication links with intersymbol-interference and additive noise. These ADCs dramatically outper-form (in terms of achievable information rates) traditional ADC design methods that are based on fixed uniform quantization. Architectures are also developed for ADCs such that system-metrics can be used to dynamically adapt the structure of the ADC to optimize application meaningful criteria, such as bit-error rate for communication over intersymbol interference links. Rajan Narasimha, Georg Zeitler, Naresh R. Shanbhag, Andrew C. Singer, Gerhard Kramer |
ICASSP | 3 |
| 2012 | Soft N-Modular RedundancyabstractAchieving robustness and energy efficiency in nanoscale CMOS process technologies is made challenging due to the presence of process, temperature, and voltage variations. Traditional fault-tolerance techniques such as N-modular redundancy (NMR) employ deterministic error detection and correction, e.g., majority voter, and tend to be power hungry. This paper proposes soft NMR that nontrivially extends NMR by consciously exploiting error statistics caused by nanoscale artifacts in order to design robust and energy-efficient systems. In contrast to conventional NMR, soft NMR employs Bayesian detection techniques in the voter. Soft voter algorithms are obtained through optimization of appropriate application aware cost functions. Analysis indicates that, on average, soft NMR outperforms conventional NMR. Furthermore, unlike NMR, in many cases, soft NMR is able to generate a correct output even when all N replicas are in error. This increase in robustness is then traded-off through voltage scaling to achieve energy efficiency. The design of a discrete cosine transform (DCT) image coder is employed to demonstrate the benefits of the proposed technique. Simulations in a commercial 45 nm, 1.2 V, CMOS process show that soft NMR provides up to 10× improvement in robustness, and 35 percent power savings over conventional NMR. Eric P. Kim, Naresh R. Shanbhag |
IEEE Trans. Computers | 2 |
| 2011 | Timing error statistics for energy-efficient robust DSP systemsabstractThis paper makes a case for developing statistical timing error models of DSP kernels implemented in nanoscale circuit fabrics. Recently, stochastic computation techniques have been proposed where the explicit use of error-statistics in system design has been shown to significantly enhance robustness and energy-efficiency. However, obtaining the error statistics at different process, voltage, and temperature (PVT) corners is hard. This paper: 1) proposes a simple additive error model for timing errors in arithmetic computations due PVT variations, 2) analyzes the relationship between error statistics and parameters, specifically the input statistics, and 3) presents a characterization methodology to obtain the proposed model parameters and thus enabling efficient implementations of emerging stochastic computing techniques. Key results include the following observations: 1) the output error statistics is a weak function of input statistics, and 2) the output error statistics depends upon the one's probability profile of the input word. These observations enable a one-time off-line statistical error characterization of DSP kernels similar to delay and power characterization done presently for standard cells and IP cores. The proposed error model is derived for a number of DSP kernels in a commercial 45nm CMOS process. Rami A. Abdallah, Yu-Hung Lee, Naresh R. Shanbhag |
DATE | 3 |
| 2011 | System-assisted analog mixed-signal designabstractIn this paper, we propose a system-assisted analog mixed-signal (SAMS) design paradigm whereby the mixed-signal components of a system are designed in an application-aware manner in order to minimize power and enhance robustness in nanoscale process technologies. In a SAMS-based communication link, the digital and analog blocks from the output of the information source at the transmitter to the input of the decision device in the receiver are treated as part of the composite channel. This comprehensive systems-level view enables us to compensate for impairments of not just the physical communication channel but also the intervening circuit blocks, most notably the analog/mixed-signal blocks. This is in stark contrast to what is done today, which is to treat the analog components in the transmitter and the analog front-end at the receiver as transparent waveform preservers. The benefits of the proposed system-aware mixed-signal design approach are illustrated in the context of analog-to-digital converters (ADCs) for high-speed links. CAD challenges that arise in designing system-assisted mixed-signal circuits are also described. Naresh R. Shanbhag, Andrew C. Singer |
DATE | 1 |
| 2011 | Least squares approximation and polyphase decomposition for pipelining recursive filtersabstractCurrent techniques used in pipelining recursive filters require significant hardware complexity. These techniques attempt to preserve the exact frequency response of the original circuit while seeking to construct a pipelined architecture. We present a technique that relaxes the need to preserve the ex act frequency response and instead considers a least-squares formulation in conjunction with the pipelined architecture. The benefit of this design is that it reduces the complexity of the pipelined circuit immensely, while enabling a simple pipelined architecture based on a polyphase decomposition of the original filter. Andrew C. Singer, Naresh R. Shanbhag |
ICASSP | 3 |
| 2011 | System energy minimization via joint optimization of the DC-DC converter and the core
Rami A. Abdallah, Pradeep S. Shenoy, Naresh R. Shanbhag, Philip T. Krein |
ISLPED | 3 |
| 2011 | VLSI Architectures for Soft-Decision Decoding of Reed-Solomon CodesabstractSoft-decision decoding of Reed-Solomon codes delivers significant coding gains over classical minimum distance decoding. In this paper, we present architectures for polynomial interpolation and factorization, the two main steps of the soft-decoding algorithm. We introduce an algorithmic transformation for reducing the iterations required in generating the interpolation polynomial and present efficient architectures by sharing computations. We also describe algorithmic transformations for further reducing the interpolation and factorization latency. An area efficient, folded-pipelined version of the interpolation architecture is also described. Finally, we present an example of a Reed-Solomon soft decoder utilizing the presented architectures, having a 250 Mbps throughput. Arshad Ahmed, Ralf Koetter, Naresh R. Shanbhag |
IEEE Trans. Inf. Theory | 3 |
| 2010 | Stochastic computationabstractStochastic computation, as presented in this paper, exploits the statistical nature of application-level performance metrics, and matches it to the statistical attributes of the underlying device and circuit fabrics. Nanoscale circuit fabrics are viewed as noisy communication channels/networks. Communications-inspired design techniques based on estimation and detection theory are proposed. Stochastic computation advocates an explicit characterization and exploitation of error statistics at the architectural and system levels. This paper traces the roots of stochastic computing from the Von Neumann era into its current form. Design and CAD challenges are described. Naresh R. Shanbhag, Rami A. Abdallah, Rakesh Kumar 0002, Douglas L. Jones |
DAC | 1 |
| 2010 | Soft NMR: Analysis & application to DSP systemsabstractWe have recently proposed the concept of soft N-modular redundancy (soft NMR) in order to design robust and energy-efficient computing systems in nanoscale processes, where soft NMR was shown to achieve orders-of-magnitude improvement in robustness with significant power savings over NMR. In this paper, we analyze the performance of soft NMR and compare it with that of NMR and algorithmic noise-tolerance (ANT). An 8-b multiplier and a DCT-based still image compression system in a commercial 45nm CMOS process is considered. Two metrics of system performance: system reliability Pe,sys, and signal-to-noise ratio (SN R) are analyzed. We show that soft NMR always outperforms NMR and that our analysis predicts Pe,sysand SNR to within 4.2% and 2.7dB on average, respectively, of the results of Monte Carlo simulations. Eric P. Kim, Naresh R. Shanbhag |
ICASSP | 2 |
| 2010 | Robust and energy-efficient DSP systems via output probability processingabstractThis paper proposes to employ error statistics of nanoscale circuit fabrics to design robust energy-efficient digital signal processing (DSP) systems. Architectural level error statistics are exploited to generate probability or the reliability of each output bit of a DSP kernel. The proposed technique is referred to here as bit-level a posteriori probability processing (BLAPP). Energy efficiency and robustness of a 2D discrete cosine transform (2D-DCT) image codec employing BLAPP is studied. Simulations in a commercial 45 nm CMOS process show that BLAPP provides up to 14X improvement in robustness, and 25% power savings over conventional 2D-DCT codec design. Rami A. Abdallah, Naresh R. Shanbhag |
ICCD | 2 |
| 2010 | BER-optimal analog-to-digital converters for communication linksabstractIn this paper, we propose BER-optimal analog-to-digital converters (ADC) where quantization levels and thresholds are set non-uniformly to minimize the bit-error rate (BER). This is in contrast to present-day ADCs which act as transparent waveform preservers. Simulations for various communication channels show that the BER-optimal ADC achieves shaping gains that range from 2.5dB for channels with low intersymbol interference (ISI) to more than 30dB for channels with high ISI. Moreover, a 3-bit BER-optimal ADC achieves the same or even lower BER than a 4-bit uniform ADC. For flash converters, this corresponds a power reduction by 2×. Look-up table based equalizers compatible with BER-optimal ADCs are shown to reduce the power up to 47% and the area up to 66% in a 45nm CMOS process. The shaping gain due to BER-optimal ADCs can be exploited to lower peak transmit swings at the transmitter or decrease power consumption of the ADC. Minwei Lu, Naresh R. Shanbhag, Andrew C. Singer |
ISCAS | 2 |
| 2010 | Stochastic Networked ComputationabstractIn this paper, the stochastic networked computation (SNC) paradigm for designing robust and energy-efficient systems-on-a-chip in nanoscale process technologies, where robust computation is treated as a statistical estimation problem is presented. The benefits of SNC are demonstrated by employing it to design an energy-efficient and robust pseudonoise-code acquisition system for the wireless CDMA2000 standard (http://www.3gpp2.org). Simulations in IBM's 130-nm CMOS process show that the SNC-based architecture enhances the average probability of detection (PDet) in the presence of process variations by two to three orders of magnitude, reduces power by 31%-39%, and reduces the variation in PDetby one to two orders of magnitude at a typical false-alarm rate of 5% over a conventional architecture. SNC performance in the presence of voltage overscaling and across technology nodes (90, 65, 45, and 32 nm) is also studied. Girish Varatkar, Shri Narayanan, Naresh R. Shanbhag, Douglas L. Jones |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2009 | Impact of DFE Error Propagation on FEC-Based High-Speed I/O LinksabstractModern state-of-the-art I/O links today rely exclusively upon a high SNR channel and an equalization-based inner transceiver to achieve a BER of 10-15. The equalizer typically consists of a transmit pre-emphasis driver for pre-cursor equalization and a receive DFE for post-cursor cancellation. Recently, forward error-correction (FEC) coding has been proposed to improve the BER and reduce power in high-speed I/O links. However, error-propagation in the DFE is a significant issue affecting code performance. The link performance is also tied to FEC implementation parameters like degree of parallelism. This paper presents a framework for analyzing the impact of DFE burst errors and implementation parameters on end-to-end link performance. For 10.3125 Gb/s transmission through a channel with 19 dB loss at Nyquist rate and serial FEC implementation, we find that a code rate r = 0.8 gives the best ISI penalty vs coding gain trade-off, and a codeword length of 750 bits is necessary to meet target performance. Further, it is observed that the performance of burst error correction codes does not necessarily improve with codeword length, i.e., there is an BER-optimal block length at a given code rate. Rajan Narasimha, Nirmal Warke, Naresh R. Shanbhag |
GLOBECOM | 3 |
| 2008 | Trends in energy-efficiency and robustness using stochastic sensor network-on-a-chipabstractThe stochastic sensor network-on-chip (SSNOC) was recently proposed as an effective computational paradigm for jointly achieving energy-efficiency and robustness in nanoscale processes. In this paper, we study the trends in energy-efficiency and robustness exhibited by an SSNOC architecture as the feature size scales from 130nm to 32nm for a PN-code acquisition application. The conventional architecture exhibits a 3 orders-of-magnitude loss in detection probability P_{det} due to process variations in the 130nm and smaller technology nodes. At the 130nm and 90nm nodes, the proposed SSNOC architecture recovers from this performance loss, and exhibits a 2 orders-of-magnitude smaller variation in P_det compared to the conventional architecture. However, for the 65nm and 45nm technology nodes, the SSNOC architecture with assistance from circuit level techniques such as adaptive body bias (ABB) and adaptive supply voltage (ASV) shows a 2-3 order-of-magnitude better detection performance. In addition, the SSNOC architecture with ABB/ASV achieves 22% to 31% energy savings. For the 32nm node, the current version of SSNOC with ABB/ASV is not robust enough and thus motivates the need to explore even more powerful versions of SSNOC. Girish Varatkar, Sriram Narayanan, Naresh R. Shanbhag, Douglas L. Jones |
ACM Great Lakes Symposium on VLSI | 3 |
| 2008 | Computation as estimation: Estimation-theoretic IC design improves robustness and reduces power consumptionabstractModern integrated circuits (ICs) are designed as massively parallel systems as a consequence of diminishing silicon feature sizes. This has adversely impacted reliability because of increased errors due to process and environmental variations, and particle hits. Viewing hardware errors as analogous to measurement or system noise allows us to borrow results from estimation theory and extend Moore's law. The estimation-theoretic framework provides a design optimization formalization that enables power/reliability trade-off in broad classes of applications. Two applications described here show that specific instantiations of the framework yield significant power savings and system reliability. Shri Narayanan, Girish Varatkar, Douglas L. Jones, Naresh R. Shanbhag |
ICASSP | 4 |
| 2008 | Variation-tolerant, low-power PN-code acquisition using stochastic sensor NOCabstractPresented in this paper is an energy-efficient and variation-tolerant PN-code acquisition architecture for the wireless CDMA2000 standard. The architectures is based on the recently proposed stochastic sensor network-on-chip (SSNOC) computational paradigm. The latter employs the principles of statistically similar decomposition and robust estimation theory to compensate for timing errors due to process variations. Performance of the SSNOC-based PN-code acquisition architecture at the slow process corner indicates that the average probability of detection PDetimproves by up to 3 orders-of-magnitude over that of the conventional architecture, while the variation in PDet(sigma / mu) is reduced by up to 2 orders-of-magnitude over that of the conventional architecture while simultaneously achieving a power reduction of 39%. Girish Varatkar, Sriram Narayanan, Naresh R. Shanbhag, Douglas L. Jones |
ISCAS | 3 |
| 2008 | Error-resilient low-power Viterbi decodersabstractTwo low-power Viterbi decoder (VD) architectures are presented in this paper. In the first, limited decision errors are introduced in the add-compare-select units (ACSUs) of a VD to reduce their critical path delays so that they can be operated at lower supply voltages in absence of timing errors. In the second one, we allow data-dependent timing errors which occur whenever a critical path in the ACSU is excited. Algorithmic noise-tolerance (ANT) is then applied at the level of the ACSU to correct for these errors. Power reduction in this design is achieved by either overscaling the supply voltage (voltage overscaling (VOS)) or designing at the nominal process corner and supply voltage (average-case design). Power savings in the first and second design are 58% and 40% at a coding loss of 0:15 dB and 1:1 dB respectively in a IBM 130nm CMOS process. Rami A. Abdallah, Naresh R. Shanbhag |
ISLPED | 2 |
| 2008 | Joint Equalization and Coding for On-Chip Bus CommunicationabstractIn this paper, we propose using joint equalization and coding to improve on-chip communication speeds by signaling at rates beyond the rate governed by resistance-capacitance (RC) delay of the interconnect. Operating beyond the RC limit introduces inter-symbol interference (ISI). We mitigate the effects of ISI by employing equalization. The proposed equalizer employs a variable threshold inverter whose switching threshold is modified as a function of past output of the bus. We demonstrate even higher speedups by combining equalization with crosstalk avoidance coding. Specifically, simulation results for a 10-mm 32-bit bus in 0.13-mum CMOS technology show that 1.28 speedup is achievable by equalization alone and 2.30 speedup is achievable by joint equalization and coding. Srinivasa R. Sridhara, Ganesh Balamurugan, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2008 | Error-Resilient Motion Estimation ArchitectureabstractIn this paper, we propose an energy-efficient motion estimation architecture. The proposed architecture employs the principle of error-resiliency to combat logic level timing errors that may arise in average-case designs in presence of process variations and/or due to overscaling of the supply voltage [voltage overscaling (VOS)] and thereby achieves power reduction. Error-resiliency is incorporated via algorithmic noise-tolerance (ANT). Referred to as input subsampled replica ANT (ISR-ANT), the proposed technique incorporates an input subsampled replica of the main sum-of-absolute-difference (MSAD) block for detecting and correcting errors in the MSAD block. Simulations show that the proposed technique can save up to 60% power over an optimal error-free system in a 130-nm CMOS technology. These power savings increase to 78% in a 45-nm predictive process technology. Performance of the ISR-ANT architecture in the presence of process variations indicates that average peak signal-to-noise ratio (PSNR) of the ISR-ANT architecture increases by up to 1.8 dB over that of the conventional architecture in 130-nm IBM process technology. Furthermore, the PSNR variation (sigma/mu) is also reduced by 7times over that of the conventional architecture at the slow corner while achieving a power reduction of 33%. Girish Varatkar, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2007 | Coding for Reliable On-Chip Buses: A Class of Fundamental Bounds and Practical CodesabstractA reliable high-speed bus employing low-swing signaling can be designed by encoding the bus to prevent crosstalk and provide error correction. Coding for on-chip buses requires additional bus wires and codec circuits. In this paper, fundamental bounds on the number of wires required to provide joint crosstalk avoidance and error correction using memoryless codes are presented. The authors propose a code construction that results in practical codec circuits with the number of wires being within 35% of the fundamental bounds. When applied to a 10-mm 32-bit bus in a 0.13-μm CMOS technology with low-swing signaling, one of the proposed codes provides 2.14× speedup and 27.5% energy savings at the cost of 2.1× area overhead, but without any loss in reliability. Srinivasa R. Sridhara, Naresh R. Shanbhag |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | Energy-efficient motion estimation using error-toleranceabstractPresented is an energy-efficient motion estimation architecture using error-tolerance. The technique employs overscaling of the supply voltage (voltage overscaling (VOS)) to reduce power at the expense of timing errors, which are then corrected using algorithmic noise-tolerance (ANT) techniques. Referred to as input subsampled replica ANT (ISR-ANT), the proposed technique incorporates an input subsampled replica of the main sum of absolute difference (MSAD) block for obtaining the motion vectors in the presence of errors induced by VOS. Simulations show that the proposed technique can save up to 60% power over an optimal error-free present day system in a 130nm CMOS technology. Power savings increase to 79% in a 45nm predictive process technology. Girish Varatkar, Naresh R. Shanbhag |
ISLPED | 2 |
| 2006 | Soft-Error-Rate-Analysis (SERA) MethodologyabstractWe present a soft-error-rate analysis (SERA) methodology for combinational and memory circuits. SERA is based on a modeling and analysis approach that employs a judicious mix of probability theory, circuit simulation, graph theory, and fault simulation. SERA achieves five orders of magnitude speedup over Monte Carlo-based simulation approaches with less than 5% error. Dependence of the soft-error rate (SER) of combinational logic circuits on a supply voltage, clock period, latching window, circuit topology, and input vector is explicitly captured and studied for a typical 0.18-$muhboxm$CMOS process. Results show that the SER of logic is a much stronger function of timing parameters than the supply voltage. Also, an SER peaking phenomenon in multipliers is observed where the center bits have an SER that are orders of magnitude greater than those of the LSBs and the MSBs. An increase of up to 25% in the SER for multiplier circuits of various sizes has been observed as technology scales from 0.18 to 0.13$muhboxm$. Ming Zhang 0017, Naresh R. Shanbhag |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | Energy-efficient soft error-tolerant digital signal processingabstractIn this paper, we present energy-efficient soft error-tolerant techniques for digital signal processing (DSP) systems. The proposed technique, referred to as algorithmic soft error-tolerance (ASET), employs low-complexity estimators of a main DSP block to achieve reliable operation in the presence of soft errors. Three distinct ASET techniques - spatial, temporal and spatio-temporal- are presented. For frequency selective finite-impulse response (FIR) filtering, it is shown that the proposed techniques provide robustness in the presence of soft error rates of up to P/sub er/=10/sup -2/ and P/sub er/=10/sup -3/ in a single-event upset scenario. The power dissipation of the proposed techniques ranges from 1.1 X to 1.7 X (spatial ASET) and 1.05 X to 1.17 X (spatio-temporal and temporal ASET) when the desired signal-to-noise ratio SNR/sub des/=25 dB. In comparison, the power dissipation of the commonly employed triple modular redundancy technique is 2.9 X. Byonghyo Shim, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | Sequential Element Design With Built-In Soft Error ResilienceabstractThis paper presents a built-in soft error resilience (BISER) technique for correcting radiation-induced soft errors in latches and flip-flops. The presented error-correcting latch and flip-flop designs are power efficient, introduce minimal speed penalty, and employ reuse of on-chip scan design-for-testability and design-for-debug resources to minimize area overheads. Circuit simulations using a sub-90-nm technology show that the presented designs achieve more than a 20-fold reduction in cell-level soft error rate (SER). Fault injection experiments conducted on a microprocessor model further demonstrate that chip-level SER improvement is tunable by selective placement of the presented error-correcting designs. When coupled with error correction code to protect in-pipeline memories, the BISER flip-flop design improves chip-level SER by 10 times over an unprotected pipeline with the flip-flops contributing an extra 7-10.5% in power. When only soft errors in flips-flops are considered, the BISER technique improves chip-level SER by 10 times with an increased power of 10.3%. The error correction mechanism is configurable (i.e., can be turned on or off) which enables the use of the presented techniques for designs that can target multiple applications with a wide range of reliability requirements Ming Zhang 0017, Subhasish Mitra, Norbert Seifert, Nicholas J. Wang, Kee Sup Kim, Naresh R. Shanbhag, Sanjay J. Patel |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2005 | A low-power bus design using joint repeater insertion and codingabstractIn this paper, we propose joint repeater insertion and crosstalk avoidance coding as a low-power alternative to repeater insertion for global bus design in nanometer technologies. We develop a methodology to calculate the repeater size and separation that minimize the total power dissipation for joint repeater insertion and coding for a specific delay target. This methodology is employed to obtain power vs. delay trade-offs for 130-nm, 90-nm, 65-nm, and 45-nm technology nodes. Using ITRS technology scaling data, we show that proposed technique provides 54%, 67%, and 69% power savings over optimally repeater-inserted 10-mm 32-bit bus at 90-nm, 65-nm, and 45-nm technology nodes, respectively, while achieving the same delay Srinivasa R. Sridhara, Naresh R. Shanbhag |
ISLPED | 2 |
| 2005 | Area-efficient high-throughput MAP decoder architecturesabstractIterative decoders such as turbo decoders have become integral components of modern broadband communication systems because of their ability to provide substantial coding gains. A key computational kernel in iterative decoders is the maximum a posteriori probability (MAP) decoder. The MAP decoder is recursive and complex, which makes high-speed implementations extremely difficult to realize. In this paper, we present block-interleaved pipelining (BIP) as a new high-throughput technique for MAP decoders. An area-efficient symbol-based BIP MAP decoder architecture is proposed by combining BIP with the well-known look-ahead computation. These architectures are compared with conventional parallel architectures in terms of speed-up, memory and logic complexity, and area. Compared to the parallel architecture, the BIP architecture provides the same speed-up with a reduction in logic complexity by a factor of M, where M is the level of parallelism. The symbol-based architecture provides a speed-up in the range from 1 to 2 with a logic complexity that grows exponentially with M and a state metric storage requirement that is reduced by a factor of M as compared to a parallel architecture. The symbol-based BIP architecture provides speed-up in the range M to 2M with an exponentially higher logic complexity and a reduced memory complexity compared to a parallel architecture. These high-throughput architectures are synthesized in a 2.5-V 0.25-/spl mu/m CMOS standard cell library and post-layout simulations are conducted. For turbo decoder applications, we find that the BIP architecture provides a throughput gain of 1.96 at the cost of 63% area overhead. For turbo equalizer applications, the symbol-based BIP architecture enables us to achieve a throughput gain of 1.79 with an area savings of 25%. Seok-Jun Lee, Naresh R. Shanbhag, Andrew C. Singer |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Coding for system-on-chip networks: a unified frameworkabstractGlobal buses in deep-submicron (DSM) system-on-chip designs consume significant amounts of power, have large propagation delays, and are susceptible to errors due to DSM noise. Coding schemes exist that tackle these problems individually. In this paper, we present a coding framework derived from a communication-theoretic view of a DSM bus to jointly address power, delay, and reliability. In this framework, the data is first passed through a nonlinear source coder that reduces self and coupling transition activity and imposes a constraint on the peak coupling transitions on the bus. Next, a linear error control coder adds redundancy to enable error detection and correction. The framework is employed to efficiently combine existing codes and to derive novel codes that span a wide range of tradeoffs between bus delay, codec latency, power, area, and reliability. Using simulation results in 0.13-/spl mu/m CMOS technology, we show that coding is a better alternative to repeater insertion for delay reduction as it reduces power dissipation at the same time. For a 10-mm 4-bit bus, we show that a bus employing the proposed codes achieves up to 2.17/spl times/ speed-up and 33% energy savings over a bus employing Hamming code. For a 10-mm 32-bit bus, we show that 1.7/spl times/ speed-up and 27% reduction in energy are achievable over an uncoded bus by employing low-swing signaling without any loss in reliability. Srinivasa R. Sridhara, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2004 | A communication-theoretic design paradigm for reliable SOCsabstractPresented is a design paradigm, pioneered at the University of Illinois in 1997, for reliable and energy-efficient system-on-a-chip (SOC) in nanometer process technologies. These technologies are characterized by non-idealities such as coupling, leakage, soft errors, and process variations, which contribute to a reliability problem. Increasing complexity of systems-on-a-chip (SOC) leads to a related power problem. The proposed paradigm provides solutions to both problems by viewing SOCs as communication networks, and employs ideas from error-control coding, communications, and information theory in order to achieve the dual goals of reliability and energy-efficiency. Naresh R. Shanbhag |
DAC | 1 |
| 2004 | Coding for system-on-chip networks: a unified frameworkabstractIn this paper, we present a coding framework derived from a communication-theoretic view of a DSM bus to jointly address power, delay, and reliability. In this framework, the data is first passed through a nonlinear source coder that reduces self and coupling transition activity and imposes a constraint on the peak coupling transitions on the bus. Next, a linear error control coder adds redundancy to enable error detection and correction. The framework is employed to efficiently combine existing codes and to derive novel codes that span a wide range of trade-offs between bus delay, codec latency, power, area, and reliability. Simulation results, for a 1-cm 32-bit bus in a 0.18-$mu$m CMOS technology, show that 31 reduction in energy and 62 reduction in energy-delay product are achievable. Srinivasa R. Sridhara, Naresh R. Shanbhag |
DAC | 2 |
| 2004 | Switching LMS linear turbo equalizationabstractTurbo equalization using linear filters for data detection has been shown to perform nearly as well as those based on the original maximum a posteriori probability (MAP) detection approach. Such linear equalization methods have taken on many forms in the literature, from simple least-mean-square (LMS)-based adaptive filtering approaches, to minimum mean square error (MMSE)-based methods that are recursively computed for each output symbol for each iteration. In this paper, we consider a class of turbo equalization algorithms in which complexity requirements dictate that a fixed set of filter coefficients must be used for all symbols and for all iterations. By computing one such set of coefficients via the LMS algorithm assuming unreliable soft information, and another set assuming highly reliable soft information, we show that a switching strategy can be employed, nearly achieving the performance of recomputing the coefficients at each iteration. Seok-Jun Lee, Andrew C. Singer, Naresh R. Shanbhag |
ICASSP (4) | 3 |
| 2004 | VLSI architectures for soft-decision decoding of Reed-Solomon codesabstractWe present the architectures for bivariate polynomial interpolation and factorization; the two main steps in algebraic soft-decision decoding of Reed-Solomon codes. We present an efficient formulation of the interpolation algorithm in which dependencies among the discrepancy coefficient computations are utilized to reduce interpolation complexity. Interpolation and factorization complexity is also reduced by using an FFT-like formulation for univariate polynomial translation. The modifications required to incorporate the recently proposed algorithm level modifications for efficient interpolation and factorization are also presented. We determine the latency and hardware requirements for soft-decoding a [255,239] Reed-Solomon code using the proposed architectures. Arshad Ahmed, Ralf Koetter, Naresh R. Shanbhag |
ICC | 3 |
| 2004 | A soft error rate analysis (SERA) methodologyabstractWe present a soft error rate analysis (SERA) methodology for combinational and memory circuits. SERA is based on a modeling and analysis-based approach that employs a judicious mix of probability theory, circuit simulation, graph theory and fault simulation. SERA achieves five orders of magnitude speed-up over Monte Carlo based simulation approaches with less than 5% error. Dependence of soft error rate (SER) of combinational circuits on supply voltage, clock period, latching window, circuit topology, and input vector values are explicitly captured and studied for a typical 0.18 /spl mu/m CMOS process. Results show that the SER of logic is a much stronger function of timing parameters than the supply voltage. Also, an "SER peaking" phenomenon in multipliers is observed where the center bits have an SER that is in order of magnitude greater than that of LSBs and MSBs. Ming Zhang 0017, Naresh R. Shanbhag |
ICCAD | 2 |
| 2004 | Area and Energy-Efficient Crosstalk Avoidance Codes for On-Chip BusesabstractCapacitive crosstalk between adjacent wires in long on-chip buses significantly increases propagation delay in the deep submicron regime. A high-speed bus can be designed by eliminating crosstalk delay through bus encoding. In this paper, we present an overview of the existing coding schemes and show that they require either a large wiring overhead or complex encoder-decoder circuits. We propose a family of codes referred to as overlapping codes that reduce both overheads. We construct two codes from this family and demonstrate their superiority over existing schemes in terms of area and energy dissipation. Specifically, for a 1-cm 32-bit bus in 0.13-/spl mu/m CMOS technology, we present a 48-wire solution that has 1.98/spl times/ speed-up, 10% energy savings and requires 20% less area than shielding. Srinivasa R. Sridhara, Arshad Ahmed, Naresh R. Shanbhag |
ICCD | 3 |
| 2004 | Reduced complexity interpolation for soft-decoding of reed-solomon codesabstractThe re-encoding based interpolation algorithm (R. Koetter et al. 2003) is modified such that intermediate interpolation results are also useful towards solving the algebraic soft-decoding problem. By factorization of a chosen subset of intermediate results, desired coding gains are obtained at lower interpolation costs Arshad Ahmed, Ralf Koetter, Naresh R. Shanbhag |
ISIT | 3 |
| 2004 | Reliable low-power digital signal processing via reduced precision redundancyabstractIn this paper, we present a novel algorithmic noise-tolerance (ANT) technique referred to as reduced precision redundancy (RPR). RPR requires a reduced precision replica whose output can be employed as the corrected output in case the original system computes erroneously. When combined with voltage overscaling (VOS), the resulting soft digital signal processing system achieves up to 60% and 44% energy savings with no loss in the signal-to-noise ratio (SNR) for receive filtering in a QPSK system and the butterfly of fast Fourier transform (FFT) in a WLAN OFDM system, respectively. These energy savings are with respect to optimally scaled (i.e., the supply voltage equals the critical voltage V/sub dd-crit/) present day systems. Further, we show that the RPR technique is able to maintain the output SNR for error rates of up to 0.09/sample and 0.06/sample in an finite impulse response filter and a FFT block, respectively. Byonghyo Shim, Srinivasa R. Sridhara, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2003 | Analysis of linear turbo equalizer via EXIT chartabstractWe propose a method for the analysis of linear turbo equalization based on extrinsic information transfer (EXIT) charts. Given channel knowledge, and therefore the optimum linear equalizer coefficients, the evolution of soft information of the soft-input soft-output (SISO) equalizer can be estimated by computing bit error rates (BER) analytically. Compared to conventional analysis methods, the proposed method predicts the linear turbo equalizer performance without running extensive simulations to obtain the SISO equalizer EXIT charts. Using an empirically generated SISO decoder EXIT chart, convergence analysis can be undertaken. Further, the method provides a bound on the achievable BER for given channels and an estimate of equalizer complexity. These approximate analyses are validated via computer simulations. Seok-Jun Lee, Andrew C. Singer, Naresh R. Shanbhag |
GLOBECOM | 3 |
| 2003 | Modeling and Mitigation of Jitter in Multi-Gbps Source-Synchronous I/O LinksabstractJitter significantly limits the maximum achievable data rates (MADR) over high-speed source-synchronous I/O links. In this paper, we present a simple model that comprehends transmitter and receiver jitter in a source-synchronous I/O link. We show that the channel can have a significant impact on transmit jitter at high data rates, resulting in 1.1X-3.8X jitter amplification for typical cases. We quantify the performance degradation of transmit/receive equalization and multilevel modulation schemes, due to jitter in highspeed I/O links. We present two design techniques to mitigate the effect of jitter on performance - transmission of a slower source-synchronous clock, and jitter equalization. Both techniques can improve MADR by 13% when signaling over a 20" FR4 channel. Ganesh Balamurugan, Naresh R. Shanbhag |
ICCD | 2 |
| 2003 | A low-power VLSI architecture for turbo decodingabstractPresented in this paper is a low-power architecture for turbo decodings of parallel concatenated convolutional codes. The proposed architecture is derived via the concept of block-interleaved computation followed by folding, retiming and voltage scaling. Block-interleaved computation can be applied to any data processing unit that operates on data blocks and satisfies the following three properties: 1.) computation between blocks are independent, 2.) a block can be segmented into computationally independent sub-blocks, and 3.) computation within a sub-block is recursive. The application of block-interleaved computation, folding and retiming reduces the critical path delay in the add-compare-select (ACS) kernel of MAP decoders by 50% - 84% with an area overhead of 14% - 70%. Subsequent application of voltage scaling results in up to 65% savings in power for block-interleaving depth of 6. Experimental results obtained by transistor-level timing and power analysis tools demonstrate power savings of 20% - 44% for a block-interleaving depth of 2 in 0.25μm CMOS process. Seok-Jun Lee, Naresh R. Shanbhag, Andrew C. Singer |
ISLPED | 2 |
| 2003 | VLSI architectures for SISO-APP decodersabstractVery large scale integration (VLSI) design methodology and implementation complexities of high-speed, low-power soft-input soft-output (SISO) a posteriori probability (APP) decoders are considered. These decoders are used in iterative algorithms based on turbo codes and related concatenated codes and have shown significant advantage in error correction capability compared to conventional maximum likelihood decoders. This advantage, however, comes at the expense of increased computational complexity, decoding delay, and substantial memory overhead, all of which hinge primarily on the well-known recursion bottleneck of the SISO-APP algorithm. This paper provides a rigorous analysis of the requirements for computational hardware and memory at the architectural level based on a tile-graph approach that models the resource-time scheduling of the recursions of the algorithm. The problem of constructing the decoder architecture and optimizing it for high speed and low power is formulated in terms of the individual recursion patterns which together form a tile graph according to a tiling scheme. Using the tile-graph approach, optimized architectures are derived for the various forms of the sliding-window and parallel-window algorithms known in the literature. A proposed tiling scheme of the recursion patterns, called hybrid tiling, is shown to be particularly effective in reducing memory overhead of high-speed SISO-APP architectures. Simulations demonstrate that the proposed approach achieves savings in area and power in the range of 4.2%-53.1% over state of the art. Mohammad M. Mansour, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | High-throughput LDPC decodersabstractA high-throughput memory-efficient decoder architecture for low-density parity-check (LDPC) codes is proposed based on a novel turbo decoding algorithm. The architecture benefits from various optimizations performed at three levels of abstraction in system design-namely LDPC code design, decoding algorithm, and decoder architecture. First, the interconnect complexity problem of current decoder implementations is mitigated by designing architecture-aware LDPC codes having embedded structural regularity features that result in a regular and scalable message-transport network with reduced control overhead. Second, the memory overhead problem in current day decoders is reduced by more than 75% by employing a new turbo decoding algorithm for LDPC codes that removes the multiple checkto-bit message update bottleneck of the current algorithm. A new merged-schedule merge-passing algorithm is also proposed that reduces the memory overhead of the current algorithm for low to moderate-throughput decoders. Moreover, a parallel soft-input-soft-output (SISO) message update mechanism is proposed that implements the recursions of the Balh-Cocke-Jelinek-Raviv (BCJR) algorithm in terms of simple "max-quartet" operations that do not require lookup-tables and incur negligible loss in performance compared to the ideal case. Finally, an efficient programmable architecture coupled with a scalable and dynamic transport network for storing and routing messages is proposed, and a full-decoder architecture is presented. Simulations demonstrate that the proposed architecture attains a throughput of 1.92 Gb/s for a frame length of 2304 bits, and achieves savings of 89.13% and 69.83% in power consumption and silicon area over state-of-the-art, with a reduction of 60.5% in interconnect length. Mohammad M. Mansour, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Energy-efficiency bounds for deep submicron VLSI systems in the presence of noiseabstractIn this paper, we present an algorithm for computing the bounds on energy-efficiency of digital very large scale integration (VLSI) systems in the presence of deep submicron noise. The proposed algorithm is based on a soft-decision channel model of noisy VLSI systems and employs information-theoretic arguments. Bounds on energy-efficiency are computed for multimodule systems, static gates, dynamic circuits and noise-tolerant dynamic circuits in 0.25-/spl mu/m CMOS technology. As the complexity of the proposed algorithm grows linearly with the size of the system, it is suitable for computing the bounds on energy-efficiency for complex VLSI systems. A key result presented is that noise-tolerant dynamic circuits offer the best trade off between energy-efficiency and noise-immunity when compared to static and domino circuits. Furthermore, employing a 16-bit noise-tolerant Manchester adder in a CDMA receiver, we demonstrate a 31.2%-51.4% energy reduction over conventional systems when operating in the presence of noise. In addition, we compute the lower bounds on energy dissipation for this CDMA receiver and show that these lower bounds are 2.8/spl times/ below the actual energy consumed, and that noise-tolerance reduces the gap between the lower bounds and actual energy dissipation by a factor of 1.9/spl times/. Lei Wang 0003, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Low-power MIMO signal processingabstractIn this paper, we present a new adaptive error-cancellation (AEC) technique, denoted as multi-input-multi-output (MIMO)-AEC, for the design of low-power MIMO signal processing systems. The MIMO-AEC technique builds on the previously proposed AEC technique by employing an algorithm transformation denoted as MIMO decorrelating (MIMO-DECOR) transform. MIMO-DECOR reduces complexity by exploiting correlations inherent in MIMO systems, thereby improving the energy efficiency of AEC. The proposed MIMO-AEC enables energy minimization of MIMO systems by correcting transient/soft errors that arise in very large scale integration signal processing implementations due to inherent process nonidealities and/or aggressive low-power design styles, such as voltage overscaling. We employ the MIMO-AEC in the design of a low-power Gigabit Ethernet 1000Base-T device. Simulation results indicate 69.1%-64.2% energy savings over optimally voltage-scaled present-day systems with no loss in algorithmic performance. Lei Wang 0003, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2002 | Reliable and energy-efficient digital signal processingabstractThis paper provides an overview of algorithmic noise-tolerance (ANT) for designing reliable and energy-efficient digital signal processing systems. Techniques such as prediction-based, error cancellation-based, and reduced precision redundancy based ANT are discussed. Average energy-savings range from 67% to 71% over conventional systems. Fluid IP core generators are proposed as a means of encapsulating the benefits of an ANT-based low-power design methodology. CAD issues resident in such a methodology are also discussed. Naresh R. Shanbhag |
DAC | 1 |
| 2002 | Turbo decoder architectures for low-density parity-check codesabstractTurbo decoding of low-density parity-check (LDPC) and generalized low-density (GLD) codes and the corresponding decoder architectures are considered. A regular (c, r)-LDPC code of length n is viewed as the intersection of c interleaved super-codes where each super-code is the direct sum of n/r independent single parity-check sub-codes. Extensions to GLD codes simply utilize more powerful sub-codes. The turbo decoding schedule is employed to decode LDPC and GLD codes using constituent soft-input soft-output (SISO) decoders that communicate through c interleavers. The proposed schedule exhibits a faster convergence behavior, and hence lower decoding latency, than the commonly employed two-phase schedule, and has a reduced memory requirement that is a function of the number of super-codes. The performance of the turbo decoding schedule is evaluated through simulations over an AWGN channel. Mohammad M. Mansour, Naresh R. Shanbhag |
GLOBECOM | 2 |
| 2002 | Design methodology for high-speed iterative decoder architecturesabstractWe propose a novel approach to the design and analysis of VLSI architectures for the soft-input soft-output a posteriori probability (SISO-APP) decoding algorithm used in iterative decoders such as turbo decoders. The approach is based on a tile-graph composed of recursion patterns that model the resource-time scheduling of the forward-backward recursion equations of the algorithm. The problem of constructing a SISO-APP architecture is formulated as a three-step process of constructing and counting the patterns needed and then tiling them. The problem of optimizing the architecture for high speed and low power reduces to optimizing the individual patterns and the tiling scheme for minimal delay and storage overhead. The various forms of the sliding and parallel-window (PW) architectures in the literature are instances of the proposed tile-graph. Using the tile-graph approach, a new PW architecture controlled by the window width r is proposed that achieves for r = 10 a 45%, a 71 %, a 51%, and a 25% reduction in decoding delay, state, input, and output metrics storage respectively, compared to a conventional architecture with a 10% increase in resources. Mohammad M. Mansour, Naresh R. Shanbhag |
ICASSP | 2 |
| 2002 | Low-power VLSI decoder architectures for LDPC codesabstractIterative decoding of low-density parity check codes (LDPC) using the message-passing algorithm have proved to be extraordinarily effective compared to conventional maximum-likelihood decoding. However, the lack of any structural regularity in these essentially random codes is a major challenge for building a practical low-power LDPC decoder. In this paper, we jointly design the code and the decoder to induce the structural regularity needed for a reduced complexity parallel decoder architecture. This interconnect-driven code design approach eliminates the need for a complex interconnection network while still retaining the algorithmic performance promised by random codes. Moreover, we propose a new approach for computing reliability metrics based on the BCJR algorithm that reduces the message switching activity in the decoder compared to existing approaches. Simulations show that the proposed approach results in power savings of up to 85.64% over conventional implementations. Mohammad M. Mansour, Naresh R. Shanbhag |
ISLPED | 2 |
| 2001 | Low-power AEC-based MIMO signal processing for gigabit ethernet 1000Base-T transceiversabstractArticle Share on Low-power AEC-based MIMO signal processing for gigabit ethernet 1000Base-T transceivers Authors: Lei Wang Coordinated Science Laboratory, Department of Electrical and Computer Engineering, University of Illinois at Urbana-Champaign, 1308 West Main Street, Urbana, IL Coordinated Science Laboratory, Department of Electrical and Computer Engineering, University of Illinois at Urbana-Champaign, 1308 West Main Street, Urbana, ILView Profile , Naresh Shanbhag Coordinated Science Laboratory, Department of Electrical and Computer Engineering, University of Illinois at Urbana-Champaign, 1308 West Main Street, Urbana, IL Coordinated Science Laboratory, Department of Electrical and Computer Engineering, University of Illinois at Urbana-Champaign, 1308 West Main Street, Urbana, ILView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 334–339https://doi.org/10.1145/383082.383179Online:06 August 2001Publication History 3citation1,230DownloadsMetricsTotal Citations3Total Downloads1,230Last 12 Months5Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Lei Wang 0003, Naresh R. Shanbhag |
ISLPED | 2 |
| 2001 | Soft digital signal processingabstractIn this paper, we propose a framework for low-energy digital signal processing (DSP), where the supply voltage is scaled beyond the critical voltage imposed by the requirement to match the critical path delay to the throughput. This deliberate introduction of input-dependent errors leads to degradation in the algorithmic performance, which is compensated for via algorithmic noise-tolerance (ANT) schemes. The resulting setup that comprises of the DSP architecture operating at subcritical voltage and the error control scheme is referred to as soft DSP. The effectiveness of the proposed scheme is enhanced when arithmetic units with a higher "delay imbalance" are employed. A prediction-based error-control scheme is proposed to enhance the performance of the filtering algorithm in the presence of errors due to soft computations. For a frequency selective filter, it is shown that the proposed scheme provides 60-81% reduction in energy dissipation for filter bandwidths up to 0.5 /spl pi/ (where 2 /spl pi/ corresponds to the sampling frequency f/sub s/) over that achieved via conventional architecture and voltage scaling, with a maximum of 0.5-dB degradation in the output signal-to-noise ratio (SNR/sub o/). It is also shown that the proposed algorithmic noise-tolerance schemes can also be used to improve the performance of DSP algorithms in presence of bit-error rates of up to 10/sup -3/ due to deep submicron (DSM) noise. Rajamohana Hegde, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2001 | High-speed architectures for Reed-Solomon decodersabstractNew high-speed VLSI architectures for decoding Reed-Solomon codes with the Berlekamp-Massey algorithm are presented in this paper. The speed bottleneck in the Berlekamp-Massey algorithm is in the iterative computation of discrepancies followed by the updating of the error-locator polynomial. This bottleneck is eliminated via a series of algorithmic transformations that result in a fully systolic architecture in which a single array of processors computes both the error-locator and the error-evaluator polynomials. In contrast to conventional Berlekamp-Massey architectures in which the critical path passes through two multipliers and 1+[log/sub 2/,(t+1)] adders, the critical path in the proposed architecture passes through only one multiplier and one adder, which is comparable to the critical path in architectures based on the extended Euclidean algorithm. More interestingly, the proposed architecture requires approximately 25% fewer multipliers and a simpler control structure than the architectures based on the popular extended Euclidean algorithm. For block-interleaved Reed-Solomon codes, embedding the interleaver memory into the decoder results in a further reduction of the critical path delay to just one XOR gate and one multiplexer, leading to speed-ups of as much as an order of magnitude over conventional architectures. Dilip V. Sarwate, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | Low-power digital filtering via soft DSPabstractWe propose a low-power filtering algorithm developed via the soft DSP framework. Soft DSP refers to scaling the supply voltage of a DSP implementation beyond the voltage required to match its critical path delay to the throughput. This deliberate introduction of input-dependent errors leads to degradation in the algorithmic performance, which is then compensated for via algorithmic error-control schemes. The proposed error-control schemes, based on forward/backward linear prediction, provides improved performance over the ones proposed in the past by exploiting correlation in both leading and trailing samples with a latency penalty. It is shown that (a) the proposed scheme provides 60-80% reduction in energy dissipation over that achieved via conventional voltage scaling and (b) for the same algorithmic performance, the overhead involved in the proposed algorithm is more than 50% smaller than existing schemes for medium bandwidth filters. Rajamohana Hegde, Naresh R. Shanbhag |
ICASSP | 2 |
| 2000 | Coupling-Driven Signal Encoding Scheme for Low-Power Interface DesignabstractCoupling effects between on-chip interconnects must be addressed in ultra deep submicron VLSI and system-on-a-chip (SoC) designs. A new low-power bus encoding scheme is proposed to minimize coupled switchings which dominate the on-chip bus power consumption. The coupling-driven bus invert method use slim encoder and decoder architecture to minimize the hardware overhead. Experimental results indicate that our encoding methods save effective switchings as much as 30% in an 8-bit bus with one-cycle redundancy. Ki-Wook Kim, Kwang-Hyun Baek, Naresh R. Shanbhag, C. L. Liu 0001 |
ICCAD | 3 |
| 2000 | Low-power decimation filters for oversampling ADCs via the decorrelating (DECOR) transformabstractThe area and power consumption of oversampling analog-to-digital converters (ADCs) are governed largely by the associated digital decimation filter. This paper presents a low power, area-efficient digital decimation filter for an oversampling ADC application that employs the decorrelating (DECOR) transform in order to reduce the power dissipation and area. The DECOR transform exploits the correlation in the coefficients and data sequences to reduce the precision. Simulation results indicate that a decorrelated 8192-tap decimation filter with a decimation ratio of 64 results in a reduction of 5 bits in the coefficient and accumulator size. This corresponds to savings in complexity of 25%. In multi-stage decimation filters, it is shown that the decimation ratio of the last stage needs to be greater than 4 for DECOR to be useful. Dongwon Seo, Naresh R. Shanbhag, Milton Feng |
ISCAS | 2 |
| 2000 | Energy-efficiency bounds for noise-tolerant dynamic circuitsabstractPresented in this paper are lower bounds on energy-efficiency of the mirror noise-tolerant dynamic circuit technique. These lower bounds are derived by solving an energy optimization problem subject to an information-theoretic constraint. Design overheads associated with the noise-tolerant circuit techniques are discussed and incorporated into the optimization problem. Simulation results for a 3-input OR gate transferring information at a rate R=150 M bit/s in 0.35 /spl mu/m CMOS indicate that the lower bound on energy consumption of the noise-tolerant circuit is 25 f J/bit, which is 31% below that of the conventional domino circuit. This lower bound is achieved when the noise-immunity of the mirror technique is 1.64/spl times/ more than that of the domino circuit technique. Naresh R. Shanbhag, Lei Wang 0003 |
ISCAS | 1 |
| 2000 | Architecture driven filter transformationsabstractIn this paper, we present the sum of powers-of-two (SPOT) algorithm transformation that results in a high-speed IIR filter architecture by forcing the first few coefficients of lhe denominator polynomial to powers of two or sums of powers of two. The SPOT transform achieves the same result as achieved by conventional pipelining techniques such as scattered look-ahead and minimum order augmentation but with significantly smaller pipelining overhead and similar sensitivity to coefficient quantization. For typical examples, the SPOT transform roughly saves 30% hardware complexity over existing techniques. Architectures for implementation of the transformed filter transfer functions have also been described. Naresh R. Shanbhag |
ISCAS | 2 |
| 2000 | Reliable low-power design in the presence of deep submicron noise (embedded tutorial session)abstractScaling of feature size in semiconductor technology has been responsible for increasingly higher computational capacity of silicon. This has been the driver for the revolution in communications and computing. However, questions regarding the limits of scaling (and hence Moore's Law) have arisen in recent years due to the emergence of deep submicron noise. The tutorial describes noise in deep submicron CMOS and their impact on digital as well as analog circuits. In particular, noise-tolerance is proposed as an effective means for achieving energy and performance efficiency in the presence of DSM noise. Naresh R. Shanbhag, Krishnamurthy Soumyanath, Samuel Martin |
ISLPED | 1 |
| 2000 | Toward achieving energy efficiency in presence of deep submicron noiseabstractPresented in this paper are: 1) information-theoretic lower bounds on energy consumption of noisy digital gates and 2) the concept of noise tolerance via coding for achieving energy efficiency in the presence of noise. In particular, lower bounds on a) circuit speed f/sub c/ and supply voltage V/sub dd/; b) transition activity t in presence of noise; c) dynamic energy dissipation; and d) total (dynamic and static) energy dissipation are derived. A surprising result is that in a scenario where dynamic component of power dissipation dominates, the supply voltage for minimum energy operation (V/sub dd, opt/) is greater than the minimum supply voltage (V/sub dd, min/)for reliable operation. We then propose noise tolerance via coding to approach the lower bounds on energy dissipation. We show that the lower bounds on energy for an off-chip I/O signaling example are a factor of 24/spl times/ below present day systems. A very simple Hamming code can reduce the energy consumption by a factor of 3/spl times/, while Reed-Muller (RM) codes give a 4/spl times/ reduction in energy dissipation. Rajamohana Hegde, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | Low-power channel coding via dynamic reconfigurationabstractPresented in this paper are energy-optimum reconfiguration strategies for channel codecs. These strategies are derived by solving an optimization problem, which has energy consumption as the objective function and a constraint on the bit error-rate (BER). Energy consumption models for a reconfigurable Reed-Solomon (RS) codec are derived via gate-level simulation of the finite field arithmetic modules. These energy models along with the BER expressions are then employed to derive the energy-optimum reconfiguration strategies. The energy savings are computed by comparing the energy consumption of the reconfigurable codec with that of the static codec. The energy savings range from 0%-83% for channel signal-to-noise ratio (SNR) variations from 7 dB-10 dB. On an average 55% energy savings are achieved. Manish Goel, Naresh R. Shanbhag |
ICASSP | 2 |
| 1999 | Energy-efficient dynamic circuit design in the presence of crosstalk noiseabstractThis paper describes the impact of crosstalk noise on low power design techniques based on voltage scaling. It is shown that this power saving strategy aggravates the crosstalk noise problem and reduces circuit noise immunity. A new energy-efficient, noise-tolerant dynamic circuit technique is presented to address this problem. In a 0.35 m CMOS technology and at a given supply voltage, the proposed technique provides an improvement in noise immunity of 1.8X(for an AND gate) and 2.5X(for an adder carry chain) over domino at the same speed. We use this fact to operate the noise-tolerant circuit at a lower supply voltage to obtain energy savings of about 30%, while expending 30 % more area. Also, to achieve a given noise immunity, the proposed technique consumes 40 % less energy compared to existing noise-tolerance techniques. Ganesh Balamurugan, Naresh R. Shanbhag |
ISLPED | 2 |
| 1999 | Energy-efficient signal processing via algorithmic noise-toleranceabstractIn this paper, we propose a framework for low-energy digital signal processing (DSP) where the supply voltage is scaled beyond the critical voltage required to match the critical path delay to the throughput. This deliberate introduction of input-dependent errors leads to degradation in the algorithmic performance, which is compensated for via algorithmic noise-tolerance (ANT) schemes. The resulting setup that comprises of the DSP architecture operating at sub-critical voltage and the error control scheme is referred to as soft DSP. It is shown that technology scaling renders the proposed scheme more effective as the delay penalty suffered due to voltage scaling reduces due to short channel effects. The effectiveness of the proposed scheme is also enhanced when arithmetic units with a higher "delay-imbalance" are employed. A prediction based error-control scheme is proposed to enhance the performance of the filtering algorithm in presence of errors due to soft computations. For a frequ... Rajamohana Hegde, Naresh R. Shanbhag |
ISLPED | 2 |
| 1999 | A low power data-adaptive motion estimation algorithmabstractA new data adaptive fast search motion estimation algorithm is proposed. Motion estimation is one of the major blocks in video coding applications and is highly computationally intensive. Although a number of fast and efficient algorithms have been developed, most of these do not take into account the varying nature of the input data. We present a novel technique in which we dynamically alter the motion estimation algorithm thereby impacting the execution rate and hence the power consumed. An average reduction of 30% in memory accesses and 60% in arithmetic operations, as compared to the full search block matching algorithm, was observed for the test cases. The loss in peak signal-to-error ratio (PSER) was less than 0.3 dB on average. Jayanto Minocha, Naresh R. Shanbhag |
MMSP | 2 |
| 1999 | Dynamic algorithm transformations (DAT)-a systematic approach to low-power reconfigurable signal processingabstractIn this paper, dynamic algorithm transformations (DATs) for designing low-power reconfigurable signal-processing systems are presented. These transformations minimize energy dissipation while maintaining a specified level of mean squared error or signal-to-noise ratio. This is achieved by modeling the nonstationarities in the input as temporal/spatial transitions between states in the input state-space. The reconfigurable hardware fabric is characterized by its configuration state-space. The configurable parameters are taken to be the filter taps, coefficient and data precisions, and supply voltage V/sub dd/. An energy-optimal reconfiguration strategy is derived as a mapping from the input to the configuration state-space. In this strategy, taps are powered down starting with the tap with the smallest value [w/sub k//sup 2///spl Sigma//sub m/(w/sub k/)] (where w/sub k/ and /spl Sigma//sub m/(w/sub k/) are, respectively, the adders, redundant-to-binary conversion, tree adders, coefficient and energy dissipation of the kth tap). Optimal values for precision and supply voltage V/sub dd/ are subsequently computed from the roundoff error and critical path delay requirements, respectively. The DAT-based adaptive filter is employed as a near-end crosstalk (NEXT) canceller in a 155.52-Mb/s asynchronous transfer mode-local area network transceiver over category-3 wiring. Simulation results indicate that the energy savings range from -2% to 87% as the cable length varies from 110 to 40 m, respectively, with an average saving of 69%. An average saving of 62% is achieved for the case where the supply voltage V/sub dd/ is kept fixed. Manish Goel, Naresh R. Shanbhag |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | A coding framework for low-power address and data bussesabstractThis paper presents a source-coding framework for the design of coding schemes to reduce transition activity. These schemes are suited for high-capacitance buses where the extra power dissipation due to the encoder and decoder circuitry is offset by the power savings at the bus. In this framework, a data source (characterized in a probabilistic manner) is first passed through a decorrelating function f/sub 1/. Next, a variant of entropy coding function f/sub 2/ is employed, which reduces the transition activity. The framework is then employed to derive novel encoding schemes whereby practical forms for f/sub 1/ and f/sub 2/ are proposed. Simulation results with an encoding scheme for data buses indicate an average reduction in transition activity of 36%. This translates into a reduction in total power dissipation for bus capacitances greater than 14 pF/b in 1.2 /spl mu/m CMOS technology. For a typical value for bus capacitance of 50 pF/b, there is a 36% reduction in power dissipation and eight times more power savings compared to existing schemes. Simulation results with an encoding scheme for instruction address buses indicate an average reduction in transition activity by a factor of 1.5 times over known coding schemes. Sumant Ramprasad, Naresh R. Shanbhag, Ibrahim N. Hajj |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | Information-theoretic bounds on average signal transition activity [VLSI systems]abstractTransitions on high-capacitance buses in very large scale integration systems result in considerable system power dissipation. Therefore, various coding schemes have been proposed in the literature to encode the input signal in order to reduce the number of transitions. In this paper, we derive lower and upper bounds on the average signal transition activity via an information-theoretic approach, in which symbols generated by a process (possibly correlated) with entropy rate H are coded with an average of R bits per symbol. The bounds are asymptotically achievable if the process is stationary and ergodic. We also present a coding algorithm based on the Lempel-Ziv data-compression algorithm to achieve the bounds. Bounds are also obtained on the expected number of ones (or zeros). These results are applied to determine the activity-reducing efficiency of different coding algorithms such as, entropy coding, transition signaling, and bus-invert coding, and determine the lower bound on the power-delay product given H and R. Two examples are provided where transition activity within 4% and 9% of the lower bound is achieved when blocks of eight symbols and 13 symbols, respectively, are coded at a time. Sumant Ramprasad, Naresh R. Shanbhag, Ibrahim N. Hajj |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1998 | Improving the throughput of flexible-precision DSPS via algorithm transformationabstractIn this paper, we have presented a systematic technique to improve throughput of signal/image processing algorithms when implemented on flexible precision hardware. Many image/signal processing algorithms need 8-16 bit precision while the DSPs available are of much higher precision (32 bit). Significant performance gain can be obtained if multiple low precision computations can be performed in one cycle of a high precision DSP. We have proposed a framework based on algorithm transformation techniques of unfolding and retiming to systematically map low precision algorithms onto high precision DSPs. The improvement in throughput obtained by this framework is linearly related to the ratio of precision used by the processor and that required by the algorithm. The efficacy of this technique has been demonstrated on a IIR filter. We have also established some theoretical bounds on the maximum throughput that can be achieved using the proposed methodology. Manoj Aggarwal, Naresh R. Shanbhag, Narendra Ahuja |
ICASSP | 2 |
| 1998 | Low-power reconfigurable signal processing via dynamic algorithm transformations (DAT)abstractPresented in this paper are dynamic algorithm transformations (DAT) for systematic design of reconfigurable computing engines. These techniques allow dynamic alteration of algorithm properties in response to input non-stationarities. The input is modeled as a set of states with an underlying probability distribution, /spl Pscr//sub S/. For each input state s, a signal monitoring algorithm (SMA) computes a power-optimal configuration for the signal processing algorithm (SPA) block. A fraction /spl alpha/ of the (SPA) block is hardwired and the remaining (1-/spl alpha/) is reconfigurable. Similarly, the (SMA) block computation is partitioned into a fraction /spl beta/ for the memory and the remaining (1-/spl beta/) for the data path. For the given input state distribution, the optimal values of /spl alpha/ (/spl alpha//sub opt/) and /spl beta/ (/spl beta//sub opt/) are determined. It is shown that for frequency selective filtering (FIR filters), the power savings of 35%-45% can be achieved by a DAT-based reconfigurable system as compared to the traditional design based on the worst-case scenario. Manish Goel, Naresh R. Shanbhag |
ICASSP | 2 |
| 1998 | Energy-efficiency in presence of deep submicron noiseabstractPresented in this paper rtr~1.)lower bormds on energy consumption of noisy digitd gates and 2.) the concept of noise tolerance via coding for achieving energy efficiency in the presence of noise.A discrete channel model for noisy digitrd logic in deep submicron technology that captures the manifestation of circuit noise is presented.me lower bounds are derived via an information-theoretic approach whereby a VLSI architecture implemented in a certain technology is viewed as a channel with information transfer capacity C (in bits/see).A computing application is shown to require a minimum information transfer rate R (rdso in bhs/see).Lower bounds are obtained by employing the information theoretic constraint C > R. ~Is constraint ensures refinabilityof computation though in an asymptotic sense.Lower bounds on transition activity at the output of noisy logic gates are rdso obtaind using this constraint.Past work (for noiseless bus coding) is shown to fdl out as a specird case.In addition, lower bounds on energy dissipation is computed by solving an optimization problem where the objective function is the energy subject to the constraint of C > R. A surprising result is that in a scenario where capacitive component of power dissipation dominates: the voltage for minimm energy is greater than the minimm voltage for reliable operation.For an off-chip UO signaling example, we show that the lower boun& are a factor of 24X below present day systems and that a very simple Hamming code can reduce the energy consumption by a factor of 3X.~Is indicates the potential of noise tolerance (via error control coding) in achieving low energy operation in the presence of noise. 1 btroduction me 1997 National Roadmap for Semiconductors [1] describes the ability to continue affordable scaling as one of the Grand Chrdlenges.For future technologies to be affordable, it is essential that high yields be obtained without putting stringent requirements on *t ~is \vork tvw suppotied by DARPA contract DMT63-97-C-MM and NSF CAREER aw,ud hlP-96X737.Pe@sion to tie d~tat or tid copim of aUor part of ti~s ~vorkfor persrmaSor &srmm use is ~ted ~~tithout fee protidd tit copim are not wde or &&b uted for profit or commer~adtmtage and that copim bw this notice ad the fuU dtation on the fit page.To copy otfrmtie, to repubSish,to post on wmem or to ref~tibute to Sis&,rquirfi prior spdc p-ion and/or a fe. Rajamohana Hegde, Naresh R. Shanbhag |
ICCAD | 2 |
| 1998 | Decorrelating (DECOR) transformations for low-power adaptive filtersabstractPresented in this paper are decorrelating transformations (referred to as DECOR transformations) to reduce the power dissipation in adaptive filters. The coefficients generated by the weight update block in an adaptive filter are passed through a decorrelating block such that fewer bits are required to represent the coefficients. Thus, the size of the arithmetic units in the filter (F-block) is reduced thereby reducing the power dissipation. The DECOR transform is well suited for narrow-band filters because there is significant correlation between adjacent coefficients. In addition, the effectiveness of DECOR transforms increases with increase in the order of the filter and decrease in coefficient precision. Simulation results indicate reduction in power dissipation in the F-block ranging from 12% to 38% for filter bandwidths ranging from 0.15 ƒs to 0.025 ƒs (where ƒs is the sample rate). Sumant Ramprasad, Naresh R. Shanbhag, Ibrahim N. Hajj |
ISLPED | 2 |
| 1998 | Efficient wireless image transmission under a total power constraintabstractDue to high data rates and limited bandwidth as well as limited battery power, wireless multimedia communications systems must be optimized in every possible way. We develop a generic matching scheme for wireless image and video communication in which the three most significant components: the source coder, the channel coder, and hardware power consumption, are jointly optimized. That is, we maximize the end-to-end image quality subject to a total power constraint on both the RF transmission power and the power consumption of the digital implementation of the channel coder, which represents a major portion of the total hardware power in short-range applications. Swaroop Appadwedula, Manish Goel, Douglas L. Jones, Kannan Ramchandran, Naresh R. Shanbhag |
MMSP | 5 |
| 1997 | Analytical Estimation of Transition Activity From Word-Level Signal StatisticsabstractPresented here is an analytical methodologyto determine the average signal activity, T, from the high-levelsignal statistics, a statistical signal generation model,and the signal encoding.Simulation results for 16 bit signalsgenerated via AR(1) and MA(1) models indicate anestimation error in T of less than 2%.The applicationof the proposed method to the estimation of T in DSPhardware is also explained. Sumant Ramprasad, Naresh R. Shanbhag, Ibrahim N. Hajj |
DAC | 2 |
| 1997 | Achievable bounds on signal transition activityabstractTransitions on high capacitance busses in VLSI systems result in considerable system power dissipation. Therefore, various coding schemes have been proposed in the literature to encode the input signal in order to reduce the number of transitions. In this paper we derive achievable lower and upper bounds on the expected signal transition activity. These bounds are derived via an information-theoretic approach in which symbols generated by a source (possibly correlated) with entropy rate H are coded with an average of R bits/symbol. These results are applied to, (1) determine the activity reducing efficiency of different coding algorithms such as Entropy coding, Transition coding, and Bus-Invert coding, (2) bound the error in entropy-based power estimation schemes, and (3) determine the lower-bound on the power-delay product. Two examples are provided where transition activity within 4% and 8% of the lower bound is achieved when blocks of 8 and 13 symbols respectively are coded at a time. Sumant Ramprasad, Naresh R. Shanbhag, Ibrahim N. Hajj |
ICCAD | 2 |
| 1997 | Dynamic algorithm transformation (DAT) for low-power adaptive signal processingabstractArticle Free Access Share on Dynamic algorithm transformation (DAT) for low-power adaptive signal processing Authors: Manish Goel Coordinated Science Lab./ECE Department, Univ. of Illinois at Urbana-Champaign, 1308 W. Main Street, Urbana IL Coordinated Science Lab./ECE Department, Univ. of Illinois at Urbana-Champaign, 1308 W. Main Street, Urbana ILView Profile , Naresh R. Shanbhag Coordinated Science Lab./ECE Department, Univ. of Illinois at Urbana-Champaign, 1308 W. Main Street, Urbana IL Coordinated Science Lab./ECE Department, Univ. of Illinois at Urbana-Champaign, 1308 W. Main Street, Urbana ILView Profile Authors Info & Claims ISLPED '97: Proceedings of the 1997 international symposium on Low power electronics and designAugust 1997 Pages 161–166https://doi.org/10.1145/263272.263316Published:01 August 1997Publication History 1citation181DownloadsMetricsTotal Citations1Total Downloads181Last 12 Months2Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Manish Goel, Naresh R. Shanbhag |
ISLPED | 2 |
| 1997 | Analytical estimation of signal transition activity from word-level statisticsabstractPresented in this paper is a novel methodology to determine the average number of transitions in a signal from its word-level statistical description. The proposed methodology employs: (1) high-level signal statistics, (2) a statistical signal generation model, and (3) the signal encoding (or number representation) to estimate the transition activity for that signal. In particular, the signal statistics employed are mean (/spl mu/), variance (/spl sigma//sup 2/), and autocorrelation (/spl rho/). The signal generation models considered are autoregressive moving-average (ARMA) models. The signal encoding includes unsigned, one's complement, two's complement, and sign-magnitude representations. First, the foilowing exact relation between the transition activity (t/sub i/), bit-level probability (p/sub i/), and the bit-level autocorrelation (/spl rho//sub i/) for a single bit signal b/sub i/ is derived: t/sub i/=2p/sub i/(1-p/sub i/)(1-/spl rho//sub i/) (1). Next, two techniques are presented which employ the word-level signal statistics, the signal generation model, and the signal encoding to determine /spl rho//sub i/ (i=0, /spl middot//spl middot//spl middot/, B-1) in (1) for a B-bit signal. The word-level transition activity T is obtained as a summation over t/sub i/ (i=0,/spl middot//spl middot//spl middot/, B-1); where t/sub i/ is obtained from (1). Simulation results for 16-bit signals generated via ARMA models indicate that an error in T of less than 2% can be achieved. Employing AR(1) and MA(10) models for audio and video signals, the proposed method results in errors of less than 10%. Both analysis and simulations indicate the sign-magnitude representation to have lower transition activity than unsigned, ones' complement, or two's complement. Finally, the proposed method is employed in estimation of transition activity in digital signal processing (DSP) hardware. Signal statistics are propagated through various DSP operators such as adders, multipliers, multiplexers, and delays, and then the transition activity T is calculated. Simulation results with ARMA inputs show that errors less than 4% are achievable in the estimation of the total transition activity in the filters. Furthermore, the transpose form structure is shown to have fewer signal transitions as compared to the direct form structure for the same input. Sumant Ramprasad, Naresh R. Shanbhag, Ibrahim N. Hajj |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1996 | Low-power adaptive filter architectures via strength reductionabstractLow-power and high-speed algorithms and architectures for complex adaptive filters are presented in this paper. These architectures have been derived via the application of algebraic and algorithm transformations. The strength reduction transformation when applied at the algorithmic level results in a power reduction by 21% as compared to the traditional cross-coupled structure. A fine-grain pipelined architecture is then developed via the relaxed look-ahead transformation. The pipelined architecture allows high-speed operation with minimum overhead and when combined with power-supply reduction enables additional power-savings of 40-69%. Thus, an overall power-saving of 60-90% over the traditional cross-coupled architecture is achieved. Manish Goel, Naresh R. Shanbhag |
ISLPED | 2 |
| 1996 | Lower bounds on power dissipation for DSP algorithmsabstractThe author presents a fundamental mathematical basis for determining the lower bounds on power dissipation in digital signal processing (DSP) algorithms. This basis is derived from information-theoretic arguments. In particular, a digital signal processing algorithm is viewed as a process of information transfer with an inherent information transfer rate requirement of R bits/sec. Different architectures implementing a given algorithm are equivalent to different communication networks each with a certain capacity C (also in bits/sec). The absolute lower bound on the power dissipation for any given architecture is then obtained by minimizing the signal power such that its channel capacity C is equal to the desired information transfer rate R. The proposed framework is employed to determine the lower bounds for simple digital filters. Furthermore, lower bounds on the power dissipation achievable via adiabatic logic are also presented, thus demonstrating the versatility of the proposed approach. Naresh R. Shanbhag |
ISLPED | 1 |
| 1995 | Pipelined Adaptive IIR Filter ArchitectureabstractThe authors present fine-grain pipelined architectures for adaptive infinite impulse response (AIIR) filters. The AIIR filters are equation error based. The proposed architectures are developed by employing a combination of scattered look-ahead and relaxed look-ahead pipelining techniques. The scattered look-ahead technique is applied do the non-adaptive (but time-varying) recursive section. The relaxed look-ahead technique is applied to the adaptive blocks. It is shown via simulations that speed-ups of up to 8 and more can be achieved with marginal or no degradation in performance. Naresh R. Shanbhag, Gi-Hong Im |
ISCAS | 1 |
| 1993 | Roundoff error analysis of the pipelined ADPCM coder
Naresh R. Shanbhag, Keshab K. Parhi |
ISCAS | 1 |
| 1993 | A Pipelined Adaptive Differential Vector Quantizer for Low-power Speech Coding Applications
Naresh R. Shanbhag, Keshab K. Parhi |
ISCAS | 1 |