EDBT 2026 Demo / reviewers in the wild / expert
Naveen Verma
dblp:90/2252
· DBLP profile ↗
40ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-8208-5030ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | eViper-2D: A Thin Large-Area Soft Robotics PlatformabstractThis paper presents the key principles of eViper-2D - a thin large-area soft robotics platform - as a new development of the previous extendable Vibrating Intelligent Piezo-Electric Robot (eViper) platform. We first introduce the mechanical, electrical, and control framework of eViper-2D, and then develop systematic and scalable methods to study the impact of diverse actuation patterns on robotic motion dynamics and energy efficiency. By integrating power electronics, communication circuits, piezoelectric actuators, and batteries onboard, the eViper-2D platform enables rapid design iteration and quick evaluation of different control strategies for the multi-actuator soft robot. The platform supports data-driven modeling via automated data acquisition. We show that eViper-2D can provide rich insights into optimizing actuation patterns to achieve agile motion and minimal cost of transport (COT). Hsin Cheng, Elias Veilleux, Zhiwu Zheng, Sigurd Wagner, Naveen Verma, James C. Sturm |
ICRA | 5 |
| 2024 | Reshape and Adapt for Output Quantization (RAOQ): Quantization-aware Training for In-memory Computing SystemsabstractIn-memory computing (IMC) has emerged as a promising solution to address both computation and data-movement challenges, by performing computation on data in-place directly in the memory array. IMC typically relies on analog operation, which makes analog-to-digital converters (ADCs) necessary, for converting results back to the digital domain. However, ADCs maintain computational efficiency by having limited precision, leading to substantial quantization errors in compute outputs. This work proposes RAOQ (Reshape and Adapt for Output Quantization) to overcome this issue, which comprises two classes of mechanisms including: 1) mitigating ADC quantization error by adjusting the statistics of activations and weights, through an activation-shifting approach (A-shift) and a weight reshaping technique (W-reshape); 2) adapting AI models to better tolerate ADC quantization through a bit augmentation method (BitAug), complemented by the introduction of ADC-LoRA, a low-rank approximation technique, to reduce the training overhead. RAOQ demonstrates consistently high performance across different scales and domains of neural network models for computer vision and natural language processing (NLP) tasks at various bit precisions, achieving state-of-the-art results with practical IMC implementations. Bonan Zhang, Chia-Yu Chen, Naveen Verma |
ICML | 3 |
| 2024 | Training Neural Networks With In-Memory-Computing Hardware and Multi-Level Radix-4 InputsabstractTraining Deep Neural Networks (DNNs) requires a large number of operations, among which matrix-vector multiplies (MVMs), often of high dimensionality, dominate. In-Memory Computing (IMC) is a promising approach to enhance MVM compute efficiency and throughput, but introduces fundamental tradeoffs with dynamic range of the computed outputs. While IMC has been successful in DNN inference systems, it has not yet shown feasibility for training, which is more sensitive to dynamic range. This work leverages recent work on alternative radix-4 number formats in DNN training on digital architectures, together with recent work on high-precision analog IMC with multi-level inputs, to enable IMC training. Furthermore, we implement a mapping of radix-4 operands to multi-level analog-input IMC in a manner that improves robustness to analog noise effects. The proposed approach is shown in simulations calibrated to silicon-measured IMC noise to be capable of training DNNs on the CIFAR-10 dataset to within 10% of the testing accuracy of standard DNN training approaches, while analysis shows that further reduction of IMC noise to feasible levels results in accuracy within 2% of standard DNN training approaches. Christopher Grimm, Naveen Verma |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | Piezoelectric Soft Robot Inchworm Motion by Tuning Ground Friction Through Robot Shape: Quasi-Static Modeling and Experimental ValidationabstractElectrically-driven soft robots based on piezoelectric actuators may enable compact form factors and maneuverability in complex environments. In most prior work, piezoelectric actuators are used to control a single degree of freedom. In this work, the coordinated activation of five independent piezoelectric actuators, attached to a common metal foil, is used to implement inchworm-inspired crawling motion in a robot that is less than 0.5 mm thick. The motion is based on the control of its friction to the ground through the robot's shape, in which one end of the robot (depending on its shape) is anchored to the ground by static friction, while the rest of its body expands or contracts. A complete analytical model of the robot shape, which includes gravity, is developed to quantify the robot shape, friction, and displacement. After validation of the model by experiments, the robot's five actuators are collectively sequenced for inchworm-like forward and backward motion. Zhiwu Zheng, Prakhar Kumar, Yenan Chen, Hsin Cheng, Sigurd Wagner, Naveen Verma, James C. Sturm |
IEEE Trans. Robotics | 7 |
| 2023 | Wirelessly-Controlled Untethered Piezoelectric Planar Soft Robot Capable of Bidirectional Crawling and RotationabstractElectrostatic actuators provide a promising approach to creating soft robotic sheets, due to their flexible form factor, modular integration, and fast response speed. However, their control requires kilo-Volt signals and understanding of complex dynamics resulting from force interactions by on-board and environmental effects. In this work, we demonstrate an untethered planar five-actuator piezoelectric robot powered by batteries and on-board high-voltage circuitry, and controlled through a wireless link. The scalable fabrication approach is based on bonding different functional layers on top of each other (steel foil substrate, actuators, flexible electronics). The robot exhibits a range of controllable motions, including bidirectional crawling (up to ~0.6 cm/s), turning, and in-place rotation (at ~1 degree/s). High-speed videos and control experiments show that the richness of the motion results from the interaction of an asymmetric mass distribution in the robot and the associated dependence of the dynamics on the driving frequency of the piezoelectrics. The robot's speed can reach 6 cm/s with specific payload distribution. Zhiwu Zheng, Hsin Cheng, Prakhar Kumar, Sigurd Wagner, Naveen Verma, James C. Sturm |
ICRA | 6 |
| 2023 | eViper: A Scalable Platform for Untethered Modular Soft RobotsabstractSoft robots present unique capabilities, but have been limited by the lack of scalable technologies for construction and the complexity of algorithms for efficient control and motion. These depend on soft-body dynamics, high-dimensional actuation patterns, and external/onboard forces. This paper presents scalable methods and platforms to study the impact of weight distribution and actuation patterns on fully untethered modular soft robots. An extendable Vibrating Intelligent Piezo-Electric Robot (eViper), together with an open-source Simulation Framework for Electroactive Robotic Sheet (SFERS) implemented in PyBullet, was developed as a platform to analyze the complex weight-locomotion interaction. By integrating power electronics, sensors, actuators, and batteries onboard, the eViper platform enables rapid design iteration and evaluation of different weight distribution and control strategies for the actuator arrays. The design supports both physics-based modeling and data-driven modeling via onboard automatic data-acquisition capabilities. We show that SFERS can provide useful guidelines for optimizing the weight distribution and actuation patterns of the eViper, thereby achieving maximum speed or minimum cost of transport (COT). Hsin Cheng, Zhiwu Zheng, Prakhar Kumar, Wali Afridi, Ben Kim, Sigurd Wagner, Naveen Verma, James C. Sturm |
IROS | 7 |
| 2023 | Reliable measurement using unreliable binary comparisons
Ryan M. Corey, Sen Tao, Naveen Verma, Andrew C. Singer |
Signal Process. | 3 |
| 2022 | Statistical computing framework and demonstration for in-memory computing systemsabstractWith the increasing importance of data-intensive workloads, such as AI, in-memory computing (IMC) has demonstrated substantial energy/throughput benefits by addressing both compute and data-movement/accessing costs, and holds significant further promise by its ability to leverage emerging forms of highly-scaled memory technologies. However, IMC fundamentally derives its advantages through parallelism, which poses a trade-off with SNR, whereby variations and noise in nanoscaled devices directly limit possible gains. In this work, we propose novel training approaches to improve model tolerance to noise via a contrastive loss function and a progressive training procedure. We further propose a methodology for modeling and calibrating hardware noise, efficiently at the level of a macro operation and through a limited number of hardware measurements. The approaches are demonstrated on a fabricated MRAM-based IMC prototype in 22nm FD-SOI, together with a neural network training framework implemented in PyTorch. For CIFAR-10/100 classifications, model performance is restored to the level of ideal noise-free execution, and generalized performance of the trained model deployed across different chips is demonstrated. Bonan Zhang, Peter Deaville, Naveen Verma |
DAC | 3 |
| 2022 | Scalable Simulation and Demonstration of Jumping Piezoelectric 2-D Soft RobotsabstractSoft robots have drawn great interest due to their ability to take on a rich range of shapes and motions, compared to traditional rigid robots. However, the motions, and underlying statics and dynamics, pose significant challenges to forming well-generalized and robust models necessary for robot design and control. In this work, we demonstrate a five-actuator soft robot capable of complex motions and develop a scalable simulation framework that reliably predicts robot motions. The simulation framework is validated by comparing its predictions to experimental results, based on a robot constructed from piezoelectric layers bonded to a steel-foil substrate. The simulation framework exploits the physics engine PyBullet, and employs discrete rigid-link elements connected by motors to model the actuators. We perform static and AC analyses to validate a single-unit actuator cantilever setup and observe close agreement between simulation and experiments for both the cases. The analyses are extended to the five-actuator robot, where simulations accurately predict the static and AC robot motions, including shapes for applied DC voltage inputs, nearly-static “inchworm” motion, and jumping (in vertical as well as vertical and horizontal directions). These motions exhibit complex non-linear behavior, with forward robot motion reaching ̴1 cm/s. Our open-source code can be found at: https://github.com/zhiwuz/sfers. Zhiwu Zheng, Prakhar Kumar, Yenan Chen, Hsin Cheng, Sigurd Wagner, Naveen Verma, James C. Sturm |
ICRA | 7 |
| 2022 | Neural Network Training on In-Memory-Computing Hardware With Radix-4 GradientsabstractDeep learning training involves a large number of operations, which are dominated by high dimensionality Matrix-Vector Multiplies (MVMs). This has motivated hardware accelerators to enhance compute efficiency, but where data movement and accessing are proving to be key bottlenecks. In-Memory Computing (IMC) is an approach with the potential to overcome this, whereby computations are performed in-place within dense 2-D memory. However, IMC fundamentally trades efficiency and throughput gains for dynamic-range limitations, raising distinct challenges for training, where compute precision requirements are seen to be substantially higher than for inferencing. This paper explores training on IMC hardware by leveraging two recent developments: (1) a training algorithm enabling aggressive quantization through a radix-4 number representation; (2) IMC leveraging compute based on precision capacitors, whereby analog noise effects can be made well below quantization effects. Energy modeling calibrated to a measured silicon prototype implemented in 16 nm CMOS shows that energy savings of over$400\times $can be achieved with full quantizer adaptability, where all training MVMs can be mapped to IMC, and$3\times $can be achieved for two-level quantizer adaptability, where two of the three training MVMs can be mapped to IMC. Christopher Grimm, Naveen Verma |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Neural Network Training With Stochastic Hardware Models and Software AbstractionsabstractMachine learning inference is of broad interest, increasingly in energy-constrained applications. However, platforms are often pushed to their energy limits, especially with deep learning models, which provide state-of-the-art inference performance but are also computationally intensive. This has motivated algorithmic co-design, where flexibility in the model and model parameters, derived from training, is exploited for hardware energy efficiency. This work extends a model-training algorithm referred to as Stochastic Data-Driven Hardware Resilience (S-DDHR) to enable statistical models of computations, amenable for energy/throughput aggressive hardware operating points as well as emerging variation-prone device technologies. S-DDHR itself extends the previous approach of DDHR by incorporating the statistical distribution of hardware variations for model-parameter learning, rather than a sample of the distributions. This is critical to developing accurate and composable abstractions of computations, to enable scalable hardware-generalized training, rather than hardware instanceby-instance training. S-DDHR is demonstrated and evaluated for a bit-scalable MRAM-based in-memory computing architecture, whose energy/throughput trade-offs explicitly motivate statistical computations. Using foundry data to model MRAM device variations, S-DDHR is shown to preserve high inference performance for benchmark datasets (MNIST, CIFAR-10, SVHN) as variation parameters are scaled to high levels, exhibiting less than 3.5% accuracy drop at 10× the nominal variation level. Bonan Zhang, Lung-Yen Chen, Naveen Verma |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2019 | A Programmable Embedded Microprocessor for Bit-scalable In-memory ComputingabstractThis article consists of a collection of slides from the author's conference presentation. Hongyang Jia, Hossein Valavi, Yinqi Tang, Naveen Verma |
Hot Chips Symposium | 5 |
| 2019 | Stochastic Data-driven Hardware Resilience to Efficiently Train Inference Models for Stochastic Hardware ImplementationsabstractMachine-learning algorithms are being employed in an increasing range of applications, spanning high-performance and energy-constrained platforms. It has been noted that the statistical nature of the algorithms can open up new opportunities for throughput and energy efficiency, by moving hardware into design regimes not limited to deterministic models of computation. This work aims to enable high accuracy in machine-learning inference systems, where computations are substantially affected by hardware variability. Previous work has overcome this by training inference model parameters for a particular instance of variation-affected hardware. Here, training is instead performed for the distribution of variation-affected hardware, eliminating the need for instance-by-instance training. The approach is referred to as Stochastic Data-Driven Hardware Resilience (S-DDHR), and it is demonstrated for an in-memory-computing architecture based on magnetoresistive random-access memory (MRAM). S-DDHR successfully address different samples of stochastic hardware, which would otherwise suffer degraded performance due to hardware variability. Bonan Zhang, Lung-Yen Chen, Naveen Verma |
ICASSP | 3 |
| 2019 | Exploiting Emerging Sensing Technologies Toward Structure in Data for Enhancing Perception in Human-Centric ApplicationsabstractStructure in data can be leveraged to enhance learning. In many perception tasks, the embedded signals arising from physical processes of interest naturally have structure of high semantic relevance. However, traditional forms of remote sensing (e.g., vision) preserve such structure only in limited ways. This paper examines how embedded, form-fitting sensing, referred to as physically integrated (PI) sensing, can preserve such structure in richer ways. While the analysis is agnostic to the particular technology for PI sensing, for which a range of options is emerging, especially driven by the Internet of Things, a particular emerging technology called large-area electronics (LAE) is considered. Using synthetic data from 3-D modeling and rendering of human-activity scenes, LAE-based PI sensing and vision-based remote sensing are emulated and perception systems are formed, showing: 1) enhanced data-efficiency of learning models based on PI sensing; 2) potential for selective deployment of PI sensors in new perception tasks, thanks to robust ranking of their value in such tasks; 3) enhanced data-efficiency of learning models based on vision sensing, by integrating PI sensing; and 4) efficient mapping of PI-sensing features across perception tasks to enhance transferability of learning. Murat Ozatay, Naveen Verma |
IEEE Internet Things J. | 2 |
| 2019 | Shannon-Inspired Statistical Computing for the Nanoscale EraabstractModern day computing systems are based on the von Neumann architecture proposed in 1945 but face dual challenges of: 1) unique data-centric requirements of emerging applications and 2) increased nondeterminism of nanoscale technologies caused by process variations and failures. This paper presents a Shannon-inspired statistical model of computation (statistical computing) that addresses the statistical attributes of both emerging cognitive workloads and nanoscale fabrics within a common framework. Statistical computing is a principled approach to the design of non-von Neumann architectures. It emphasizes the use of information-based metrics; enables the determination of fundamental limits on energy, latency, and accuracy; guides the exploration of statistical design principles for low signal-to-noise ratio (SNR) circuit fabrics and architectures such as deep in-memory architecture (DIMA) and deep in-sensor architecture (DISA); and thereby provides a framework for the design of computing systems that approach the limits of energy efficiency, latency, and accuracy. From its early origins, Shannon-inspired statistical computing has grown into a concrete design framework validated extensively via both theory and laboratory prototypes in both CMOS and beyond. The framework continues to grow at both of these levels, yielding new ways of connecting systems through architectures, circuits, and devices, for the semiconductor roadmap to march into the nanoscale era. Naresh R. Shanbhag, Naveen Verma, Yongjune Kim 0001, Ameya Patil 0001, Lav R. Varshney |
Proc. IEEE | 2 |
| 2018 | Genetic Programming for Energy-Efficient and Energy-Scalable Approximate Feature Computation in Embedded Inference SystemsabstractWith the increasing interest in deploying embedded sensors in a range of applications, there is also interest in deploying embedded inference capabilities. Doing so under the strict and often variable energy constraints of the embedded platforms requires algorithmic, in addition to circuit and architectural, approaches to reducing energy. A broad approach that has recently received considerable attention in the context of inference systems is approximate computing. This stems from the observation that many inference systems exhibit various forms of tolerance to data noise. While some systems have demonstrated significant approximation-versus-energy knobs to exploit this, they have been applicable to specific kernels and architectures; the more generally available knobs have been relatively weak, resulting in large data noise for relatively modest energy savings (e.g., voltage overscaling, bit-precision scaling). In this work, we explore the use of genetic programming (GP) to compute approximate features. Further, we leverage a method that enhances tolerance to feature-data noise through directed retraining of the inference stage. Previous work in GP has shown that it generalizes well to enable approximation of a broad range of computations, raising the potential for broad applicability of the proposed approach. The focus on feature extraction is deliberate because they involve diverse, often highly nonlinear, operations, challenging general applicability of energy-reducing approaches. We evaluate the proposed methodologies through two case studies, based on energy modeling of a custom low-power microprocessor with a classification accelerator. The first case study is on electroencephalogram-based seizure detection. We find that the choice of two primitive functions (square root, subtraction) out of seven possible primitive functions (addition, subtraction, multiplication, logarithm, exponential, square root, and square) enables us to approximate features in 0.41$mJ$per feature vector (FV), as compared to 4.79$mJ$per FV required for baseline feature extraction. This represents a feature extraction energy reduction of 11.68$\times$. The important system-level performance metrics for seizure detection are sensitivity, latency, and number of false alarms per hour. Our set of GP models achieves 100 percent sensitivity, 4.37 second latency, and 0.15 false alarms per hour. The baseline performance is 100 percent sensitivity, 3.84 second latency, and 0.06 false alarms per hour. The second case study is on electrocardiogram-based arrhythmia detection. In this case, just one primitive function (multiplication) suffices to approximate features in 1.13$\mu J$per FV, as compared to 11.69$\mu J$per FV required for baseline feature extraction. This represents a feature extraction energy reduction of 10.35$\times$. The important system-level metrics in this case are sensitivity, specificity, and accuracy. Our set of GP models achieves 81.17 percent sensitivity, 80.63 percent specificity, and 81.86 percent accuracy, whereas the baseline achieves 82.05 percent sensitivity, 88.12 percent specificity, and 87.92 percent accuracy. These case studies demonstrate the possibility of a significant reduction in feature extraction energy at the expense of a slight degradation in system performance. Hongyang Jia, Naveen Verma, Niraj K. Jha |
IEEE Trans. Computers | 3 |
| 2018 | Energy-Efficient Pedestrian Detection System: Exploiting Statistical Error Compensation for Lossy Memory Data Compression
Yinqi Tang, Naveen Verma |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Compressive information acquisition with hardware impairments and constraints: A case studyabstractCompressive information acquisition is a natural approach for low-power hardware front ends, since most natural signals are sparse in some basis. Key design questions include the impact of hardware impairments (e.g., nonlinearities) and constraints (e.g., spatially localized computations) on the fidelity of information acquisition. Our goal in this paper is to obtain specific insights into such issues through modeling of a Large Area Electronics (LAE)-based image acquisition system. We show that compressive information acquisition is robust to stochastic nonlinearities, and that appropriately designed spatially localized computations are effective, by evaluating the performance of reconstruction and classification based on the information acquired. Soorya Gopalakrishnan, Tiffany Moy, Upamanyu Madhow, Naveen Verma |
ICASSP | 4 |
| 2017 | Information-processing-driven interfaces in hybrid large-area electronics systemsabstractIn the development of human-centric systems, access to a large number of human information signals is required. Such signals can be acquired from both ambient and on-person (wearable) sensors. Large-area electronics (LAE) provide distinct capabilities for creating the required diverse, distributed and conformal sensors. However, the large volume of and complex correlation to target information within the captured data requires significant processing and inference. This makes an LAE-CMOS hybrid system well-suited to such applications. Interfacing between the two technologies is a challenge in hybrid system design. We demonstrate an emerging solution space based on information-processing-oriented interfaces, through two case studies: 1) an image sensing and compression system based on random projection [1]; 2) an electroencephalogram (EEG) acquisition and biomarker-extraction system using compressive-sensing circuits [2]. Tiffany Moy, Warren Rieutort-Louis, Liechao Huang, Sigurd Wagner, James C. Sturm, Naveen Verma |
ISCAS | 6 |
| 2017 | A 10-b statistical ADC employing pipelining and sub-ranging in 32nm CMOSabstractThis paper presents a 10-b statistical ADC (S-ADC), achieving higher resolution (INL) than any previously reported S-ADC. This resolution requires a large number of statistical observations via comparators (12 k) with offset variation, making code estimation a key challenge. The efficiency of estimation is enhanced by a coarse frontend estimator, employing pipelining and sub-ranging to arrive at a reduced range, which is then provided to a fine backend estimator. The total computations are reduced by 19×, compared to single-stage estimation over the entire analog range. Implemented in a 32 nm process, the S-ADC achieves INLRMS. Designed to run at 20 MHz, excess supply impedance limits comparator speed to 2 MHz. The energy per 10-b conversion for the comparator array (at 2 MHz) is 744 pJ and the energy per 10-b conversion of the digital estimator (at 20 MHz) is 627 pJ. Sen Tao, Naveen Verma, Ryan M. Corey, Andrew C. Singer |
ISCAS | 2 |
| 2017 | Editorial for JETC Special Issue on Alternative Computing SystemsabstractNo abstract available. Rasit Onur Topaloglu, Naveen Verma |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2016 | Robust blind source separation in a reverberant room based on beamforming with a large-aperture microphone arrayabstractLarge-Area Electronics (LAE) technology has enabled the development of physically-expansive sensing systems with a flexible form-factor, including large-aperture microphone arrays. We propose an approach to blind source separation based on leveraging such an array. In our algorithm we carry out delay-sum beamforming, but use frequency-dependent time delays, making it well-suited for a practical reverberant room. This is followed by a binary mask stage for further interference cancellation. A key feature is that it is fully "blind", since it requires no prior information about the location of the speakers or microphones. Instead, we carry out k-means cluster analysis, to estimate time delays in the background from acquired audio signals that represent the mixture of simultaneous sources. We have tested this algorithm in a conference room (T60 = 350 ms), using two linear arrays consisting of: (1) commercial electret capsules, and (2) LAE microphones, fabricated in-house. We have achieved high-quality separation results, obtaining a mean PESQ MOS improvement (relative to the unprocessed signal) for the electret array of 0.7 for two sources and 0.6 for four simultaneous sources, and for the LAE array of 0.5 and 0.3, respectively. Josue Sanz-Robinson, Liechao Huang, Tiffany Moy, Warren Rieutort-Louis, Yingzhe Hu, Sigurd Wagner, James C. Sturm, Naveen Verma |
ICASSP | 8 |
| 2016 | Hybrid large-area systems: Challenges in interfacingabstractHybrid large-area systems aim to leverage the strengths of two complementary technologies: (1) large-area electronics (LAE), which enables dense arrays of diverse transducers on substrates that can be large and flexible; and (2) silicon CMOS ICs, which enable efficient and high-performance instrumentation, computation, and power management. A key challenge in realizing these hybrid systems on a large-scale lies in the interfacing required between the two technologies. We describe methods to ease the interfacing, enabled by device, circuit, and algorithmic advances, thereby suggesting a range of challenges and opportunities that are exposed when thinking about systems. Tiffany Moy, Sigurd Wagner, Warren Rieutort-Louis, Yingzhe Hu, Liechao Huang, Josue Sanz-Robinson, James C. Sturm, Naveen Verma |
ISCAS | 8 |
| 2016 | Hybrid large-area systems and their interconnection backbone (invited paper)abstractHybrid systems combine Large-Area Electronics (LAE) with high-performance technologies (e.g., silicon CMOS) [1]. With architectural concepts for hybrid systems broadening to match the range of emerging applications, this paper examines modular approaches for multi-sheet, multi-technology integration. It identifies the interfaces required as a critical backbone. For interfaces associated with various system functionalities (sensing, processing, powering), specific approaches are surveyed and analyzed, taking from insights derived from several previous experimental demonstrations of complete hybrid systems. Naveen Verma, Levent E. Aygun, Yasmin Afsar, Yingzhe Hu, Liechao Huang, Tiffany Moy, Josue Sanz-Robinson, Warren Rieutort-Louis, Sigurd Wagner, James C. Sturm |
NOCS | 1 |
| 2016 | Strain Sensing Sheets for Structural Health Monitoring Based on Large-Area Electronics and Integrated CircuitsabstractAccurate and reliable damage characterization (i.e., damage detection, localization, and evaluation of extent) in civil structures and infrastructure is an important objective of structural health monitoring (SHM). Highly accurate and reliable characterization of damage at early stages requires continuous or quasi-continuous direct sensing of the critical parameters. Direct sensing requires deploying dense arrays of sensors, to enhance the probability that damage will result in signals that can be directly acquired by the sensors. However, coverage by dense arrays of sensors over the large areas that are of relevance represents an enormous challenge for current technologies. Large area electronics (LAE) is an emerging technology that can enable the formation of dense sensor arrays spanning large areas (several square meters) on flexible substrates. This paper explores the requirements and technology for a sensing sheet for SHM based on LAE and crystalline silicon CMOS integrated circuits (ICs). The sensing sheet contains a dense array of thin-film full-bridge resistive strain sensors, along with the electronics for strain readout, full-system self-powering, and communication. Research on several stages is presented for translating the sensing sheet to practical SHM applications. This includes experimental characterization of an individual sensor's response when exposed to cracks in concrete and steel; theoretical and experimental performance evaluation of various geometrical parameters of the sensing sheet; and development of the electronics necessary for sensor readout, power management, and sensor-data communication. The concept of direct sensing has been experimentally validated, and the potential of a sensing sheet to provide direct sensing and successful damage characterization has been evaluated in the laboratory setting. A prototype of the sensing sheet has also been successfully developed and independently characterized in the laboratory, meeting the required specifications. Thus, a sensing sheet for SHM applications shows promise both in terms of practicality and effectiveness. Branko Glisic, Shue-Ting E. Tung, Sigurd Wagner, James C. Sturm, Naveen Verma |
Proc. IEEE | 6 |
| 2016 | Compressed Signal Processing on Nyquist-Sampled SignalsabstractPattern-recognition algorithms from the domain of machine learning play a prominent role in embedded sensing systems, in order to derive inferences from sensor data. Very often, such systems face severe energy constraints. The focus of this work is to mitigate the computational energy by exploiting a form of compression which preserves a similarity metric widely used for pattern recognition. The form of compression is random projection, and the similarity metric is inner products between source vectors. Given the prominence of random projections within compressive sensing, previous research has explored this idea for application to compressively-sensed signals. In this work, we analyze the error sources faced by such approaches and show that the compressive-sensing setting itself introduces a significant source of feature-computation error ($\sim$30 percent). We show that random projections can be exploited more generally without compressive sensing, enabling significant reduction in computational energy, and avoiding a significant source of error. The approach is referred to ascompressed signal processing (CSP), and it applies to Nyquist-sampled signals. We validate the CSP approach through two case studies. The first focuses on seizure detection using spectral-energy features extracted from electroencephalograms. We show that at a 32$\times$compression ratio, the number of multiply-accumulate (MAC) and operand-access operations required is reduced by 21.2$\times$, while achieving a sensitivity of 100 percent, latency of 4.33 sec, and false alarm rate of 0.22/hr; this compares to a baseline performance of 100 percent, 4.37 sec, and 0.12/hr, respectively. The second case study focuses on neural prosthesis based on extracting wavelet features from a set of detected spikes. We show that at a 32$\times$compression ratio, the number of MAC and operand access computations required is reduced by 3.3$\times$, while spike sorting performance can be maintained within an average error of 4.89 percent for spike count, 3.42 percent for coefficient of variance, and 4.90 percent for firing rate; this compares with a baseline average error of 4.00, 2.75, and 4.00 percent for spike count, coefficient of variance, and firing rate, respectively. Naveen Verma, Niraj K. Jha |
IEEE Trans. Computers | 2 |
| 2015 | Reducing quantization error in low-energy FIR filter acceleratorsabstractComputational energy versus computational precision represents a critical implementation-level tradeoff facing embedded DSP systems. Focusing on multiply-accumulate (MAC) hardware, which is used extensively in DSP implementations (e.g., FIR filtering), this paper proposes an approach that exploits floating-point representation of multipliers to enable optimization of their quantization error. The approach introduces a parameter α for coefficient scaling, and optimizes α to minimize the output error. Applied to FIR filters with coefficient representation of 6 bits, the approach reduces the quantization error by 37×, compared to traditional, linear-quantized fixed-point coefficient representation and by 28×, compared to unoptimized floating-point coefficient representation. Further, the energy and hardware gate-count of a MAC unit is reduced by 1.4× and 1.2×, respectively, compared to an implementation based on fixed-point representation. Zhuo Wang 0001, Naveen Verma |
ICASSP | 3 |
| 2015 | Enabling Scalable Hybrid Systems: Architectures for Exploiting Large-Area Electronics in ApplicationsabstractBy enabling diverse and large-scale transducers, large-area electronics raises the potential for electronic systems to interact much more extensively with the physical world than is possible today. This can substantially expand the scope of applications, both in number and in value. But first, translation into applications requires a base of system functions (instrumentation, computation, power management, communication). These cannot be realized on the desired scale by large-area electronics alone. It is necessary to combine large-area electronics with high-performance, high-efficiency technologies, such as crystalline silicon CMOS, within hybrid systems. Scalable hybrid systems require rethinking the subsystem architectures from the start by considering how the technologies should be interfaced, on both a functional and physical level. To explore platform architectures along with the supporting circuits and devices, we consider as an application driver, a self-powered sheet for high-resolution structural health monitoring (of bridges and buildings). Top-down evaluation of design alternatives within the hybrid design space and pursuit of template architectures exposes circuit functions and device optimizations traditionally overlooked by bottom-up approaches alone. Naveen Verma, Yingzhe Hu, Liechao Huang, Warren Rieutort-Louis, Josue Sanz-Robinson, Tiffany Moy, Branko Glisic, Sigurd Wagner, James C. Sturm |
Proc. IEEE | 1 |
| 2015 | Signal Processing With Direct Computations on Compressively Sensed DataabstractSparsity is characteristic of a signal that potentially allows us to represent information efficiently. We present an approach that enables efficient representations based on sparsity to be utilized throughout a signal processing system, with the aim of reducing the energy and/or resources required for computation, communication, and storage. The representation we focus on is compressive sensing. Its benefit is that compression is achieved with minimal computational cost through the use of random projections; however, a key drawback is that reconstruction is expensive. We focus on inference frameworks for signal analysis. We show that reconstruction can be avoided entirely by transforming signal processing operations (e.g., wavelet transforms, finite impulse response filters, etc.) such that they can be applied directly to the compressed representations. We present a methodology and a mathematical framework that achieve this goal and also enable significant computational-energy savings through operations over fewer input samples. This enables explicit energy-versus-accuracy tradeoffs that are under the control of the designer. We demonstrate the approach through two case studies. First, we consider a system for neural prosthesis that extracts wavelet features directly from compressively sensed spikes. Through simulations, we show that spike sorting can be achieved with 54× fewer samples, providing an accuracy of 98.63% in spike count, 98.56% in firing-rate estimation, and 96.51% in determining the coefficient of variation; this compares with a baseline Nyquist-domain detector with corresponding performance of 98.97%, 99.69%, and 97.09%, respectively. Second, we consider a system for detecting epileptic seizures by extracting spectral-energy features directly from compressively sensed electroencephalogram. Through simulations of the end-to-end algorithm, we show that detection can be achieved with 21× fewer samples, providing a sensitivity of 94.43%, false alarm rate of 0.1543/h, and latency of 4.70 s; this compares with a baseline Nyquist-domain detector with corresponding performance of 96.03%, 0.1471/h, and 4.59 s, respectively. Mohammed Shoaib, Niraj K. Jha, Naveen Verma |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Overcoming Computational Errors in Sensing Platforms Through Embedded Machine-Learning KernelsabstractWe present an approach for overcoming computational errors at run time that originate from static hardware faults in digital processors. The approach is based on embedded machine-learning stages that learn and model the statistics of the computational outputs in the presence of errors, resulting in an error-aware model for embedded analysis. We demonstrate, in hardware, two systems for analyzing sensor data: 1) an EEG-based seizure detector and 2) an ECG-based cardiac arrhythmia detector. The systems use a small kernel of fault-free hardware (constituting <;7.0% and <;31% of the total areas respectively) to construct and apply the error-aware model. The systems construct their own error-aware models with minimal overhead through the use of an embedded active-learning framework. Via an field-programmable gate array implementation for hardware experiments, stuck-at faults are injected at controllable rates within synthesized gate-level netlists to permit characterization. The seizure detector demonstrates restored performance despite faults on 0.018% of the circuit nodes [causing bit error rates (BERs) up to 45%], and the arrhythmia detector demonstrates restored performance despite faults on 2.7% of the circuit nodes (causing BERs up to 50%). Zhuo Wang 0001, Kyong-Ho Lee, Naveen Verma |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Error-adaptive classifier boosting (EACB): Exploiting data-driven training for highly fault-tolerant hardwareabstractTechnological scaling and system-complexity scaling have dramatically increased the prevalence of hardware faults, to the point where traditional approaches based on design margining are becoming un-viable. The challenges are exacerbated in embedded sensing applications due to constraints on system resources (energy, area). Given the importance of classification functions in such applications, this paper presents an architecture for overcoming faults within a classification processor. The approach employs machine learning for modeling not only complex sensor signals but also error manifestations due to hardware faults. Adaptive boosting is exploited in the architecture for performing iterative data-driven training. This enables the effects of faults in preceding iterations to be modeled and overcome during subsequent iterations. We demonstrate a system integrating the proposed classifier, capable of training its model entirely within the architecture by generating estimated training labels. FPGA experiments show that high fault rates (affecting >3% of all circuit nodes) occurring on >80% of the hardware can be overcome, restoring system performance to fault-free levels. Zhuo Wang 0001, Robert E. Schapire, Naveen Verma |
ICASSP | 3 |
| 2013 | Algorithm-Driven Architectural Design Space Exploration of Domain-Specific Medical-Sensor ProcessorsabstractData-driven machine-learning techniques enable the modeling and interpretation of complex physiological signals. The energy consumption of these techniques, however, can be excessive, due to the complexity of the models required. In this paper, we study the tradeoffs and limitations imposed by the energy consumption of high-order detection models implemented in devices designed for intelligent biomedical sensing. Based on the flexibility and efficiency needs at various processing stages in data-driven biomedical algorithms, we explore options for hardware specialization through architectures based on custom instruction and coprocessor computations. We identify the limitations in the former, and propose a coprocessor-based platform that exploits parallelism in computation as well as voltage scaling to operate at a subthreshold minimum-energy point. We present results from post-layout simulation of cardiac arrhythmia detection with patient data from the MIT-BIH database. After wavelet-based feature extraction, which consumes 12.28 μJ, we demonstrate classification computations in the 12.00-120.05 μJ range using 10000-100000 support vectors. This represents 1170× lower energy than that of a low-power processor with custom instructions alone. After morphological feature extraction, which consumes 8.65 μJ of energy, the corresponding energy numbers are 10.24-24.51 μJ, which is 1548× smaller than one based on a custom-instruction design. Results correspond to Vdd=0.4 V and a data precision of 8 b. Mohammed Shoaib, Niraj K. Jha, Naveen Verma |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | Enabling advanced inference on sensor nodes through direct use of compressively-sensed signalsabstractNowadays, sensor networks are being used to monitor increasingly complex physical systems, necessitating advanced signal analysis capabilities as well as the ability to handle large amounts of network data. For the first time, we present a methodology to enable advanced decision support on a low-power sensor node through the direct use of compressively-sensed signals in a supervised-learning framework; such signals provide a highly efficient means of representing data in the network, and their direct use overcomes the need for energy-intensive signal reconstruction. Sensor networks for advanced patient monitoring are representative of the complexities involved. We demonstrate our technique on a patient-specific seizure detection algorithm based on electroencephalograph (EEG) sensing. Using data from 21 patients in the CHB-MIT database, our approach demonstrates an overall detection sensitivity, latency, and false alarm rate of 94.70%, 5.83 seconds, and 0.199 per hour, respectively, while achieving data compression by a factor of 10x. This compares well with the state-of-the-art baseline detector with corresponding results being 96.02%, 4.59 seconds, and 0.145 per hour, respectively. Mohammed Shoaib, Niraj K. Jha, Naveen Verma |
DATE | 3 |
| 2012 | Enabling system-level platform resilience through embedded data-driven inference capabilities in electronic devicesabstractAdvanced devices for embedded and ambient applications represent one of the most compelling classes of electronic systems, but they also impose more severe constraints on system resources than ever before. Although platform non-idealities have always posed a fundamental limitation, the overheads of conventional margining are now reaching intolerable levels. We describe an alternate approach to hardware resilience that applies to applications where advanced modeling and inference capabilities are required, a rapidly increasing emphasis in many applications. We show how a data-driven modeling framework for analyzing application data can also be used to effectively model and overcome a broad range of hardware non-idealities. Specific examples for biomedical sensors are shown that are able to retain performance with minimal on-line overhead despite the presence of severe digital- and analog-circuit non-idealities. Naveen Verma, Kyong-Ho Lee, Kuk Jin Jang, Ali H. Shoeb |
ICASSP | 1 |
| 2011 | A low-energy computation platform for data-driven biomedical monitoring algorithmsabstractA key challenge in closed-loop chronic biomedical systems is the ability to detect complex physiological states from patient signals within a constrained power budget. Data-driven machine-learning techniques are major enablers for the modeling and interpretation of such states. Their computational energy, however, scales with the complexity of the required models. In this paper, we propose a low-energy, biomedical computation platform optimized through the use of an accelerator for data-driven classification. The accelerator retains selective flexibility through hardware reconfiguration and exploits voltage scaling and parallelism to operate at a sub-threshold minimum-energy point. Using cardiac arrhythmia detection algorithms with patient data from the MIT-BIH database, classification is achieved in 2.96 μJ (at Vdd = 0.4 V), over four orders of magnitude smaller than that on a low-power general-purpose processor. The energy of feature extraction is 148 μJ while retaining flexibility for a range of possible biomarkers. Mohammed Shoaib, Niraj K. Jha, Naveen Verma |
DAC | 3 |
| 2011 | Improving kernel-energy trade-offs for machine learning in implantable and wearable biomedical applicationsabstractEmerging biomedical sensors and stimulators offer unprecedented modalities for delivering therapy and acquiring physiological signals (e.g., deep brain stimulators). Exploiting these in intelligent, closed loop systems requires detecting specific physiological states using very low power (i.e., 1-10 mW for wearable devices, 10-100 μW for implantable devices). Machine learning is a powerful tool for modeling correlations in physiological signals, but model complexity in typical biomedical applications makes detection too computationally intensive. We analyze the computational energy trade-offs and propose a method of restructuring the computations to yield more favorable trade-offs, especially for typical biomedical applications. We thus develop a methodology for implementing low-energy classification kernels and demonstrate energy reduction in practical biomedical systems. Two applications, arrhythmia detection using electrocardiographs (ECG) from the MIT-BIH database and seizure detection using electroencephalographs (EEG) from the CHB-MIT database, are used. The proposed computational restructuring can be used with very little performance degradation, and it reduces energy by 2627x and 7.0-36.3x (depending on the patient), respectively. Kyong-Ho Lee, Sun-Yuan Kung, Naveen Verma |
ICASSP | 3 |
| 2011 | Analysis Towards Minimization of Total SRAM Energy Over Active and Idle Operating ModesabstractComputational requirements in highly energy constrained applications are driving the need for ultra-low-power processors. In such devices SRAMs pose a primary energy limitation. This paper analyzes SRAM energy in practical applications using state-of-the-art power-management techniques. The design targets and array biasing for energy minimization are developed. Compared with generic logic, these are characterized by the important difference that SRAMs generally need to retain data. This restricts the use of power-gating for leakage elimination, and thus this paper considers the application of low-leakage data-retention biasing during the idle-mode. The resulting energy tradeoffs have important distinctions, and these are analyzed in the presence of practical variation levels. Naveen Verma |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2010 | Technologies for Ultradynamic Voltage ScalingabstractEnergy efficiency of electronic circuits is a critical concern in a wide range of applications from mobile multi-media to biomedical monitoring. An added challenge is that many of these applications have dynamic workloads. To reduce the energy consumption under these variable computation requirements, the underlying circuits must function efficiently over a wide range of supply voltages. This paper presents voltage-scalable circuits such as logic cells, SRAMs, ADCs, and dc-dc converters. Using these circuits as building blocks, two different applications are highlighted. First, we describe an H.264/AVC video decoder that efficiently scales between QCIF and 1080p resolutions, using a supply voltage varying from 0.5 V to 0.85 V. Second, we describe a 0.3 V 16-bit micro-controller with on-chip SRAM, where the supply voltage is generated efficiently by an integrated dc-dc converter. Anantha P. Chandrakasan, Denis C. Daly, Daniel F. Finchelstein, Joyce Kwong, Yogesh K. Ramadass, Mahmut E. Sinangil, Vivienne Sze, Naveen Verma |
Proc. IEEE | 8 |
| 2006 | Sub-threshold design: the challenges of minimizing circuit energyabstractIn this paper, we identify the key challenges that oppose sub-threshold circuit design and describe fabricated chips that verify techniques for overcoming the challenges. Benton H. Calhoun, Alice Wang 0002, Naveen Verma, Anantha P. Chandrakasan |
ISLPED | 3 |
| 2005 | Design Considerations for Ultra-Low Energy Wireless Microsensor NodesabstractThis tutorial paper examines architectural and circuit design techniques for a microsensor node operating at power levels low enough to enable the use of an energy harvesting source. These requirements place demands on all levels of the design. We propose architecture for achieving the required ultra-low energy operation and discuss the circuit techniques necessary to implement the system. Dedicated hardware implementations improve the efficiency for specific functionality, and modular partitioning permits fine-grained optimization and power-gating. We describe modeling and operating at the minimum energy point in the subthreshold region for digital circuits. We also examine approaches for improving the energy efficiency of analog components like the transmitter and the ADC. A microsensor node using the techniques we describe can function in an energy-harvesting scenario. Benton H. Calhoun, Denis C. Daly, Naveen Verma, Daniel F. Finchelstein, David D. Wentzloff, Alice Wang 0002, Seong-Hwan Cho, Anantha P. Chandrakasan |
IEEE Trans. Computers | 3 |