EDBT 2026 Demo / reviewers in the wild / expert
Boris Murmann
dblp:94/4508
· DBLP profile ↗
41ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-3417-8782ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 35 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Software engineering, systems software and programming languages · 2 · 1 since 2021Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Generative Silicon: The Next Frontier in Open-Source and AI-Driven Analog Design
Mehdi Saligane, Anhang Li, Boris Murmann |
ISCAS | 4 |
| 2024 | TinyForge: A Design Space Exploration to Advance Energy and Silicon Area Trade-offs in tinyML Compute Architectures with Custom Latch ArraysabstractThe proliferation of smart IoT devices has given rise to tinyML, which deploys deep neural networks on resource-constrained systems, benefitting from custom hardware that optimizes for low silicon area and high energy efficiency amidst tinyML's characteristic small model sizes (50-500 KB) and low target frequencies (1-100 MHz). We introduce a novel custom latch array integrated with a compute memory fabric, achieving 8 μm2/B density and 11 fJ/B read energy, surpassing synthesized implementations by 7x in density and 5x in read energy. This advancement enables dataflows that do not require activation buffers, reducing memory overheads. By optimizing systolic vs. combinational scaling in a 2D compute array and using bit-serial instead of bit-parallel compute, we achieve a reduction of 4.8x in area and 2.3x in multiply-accumulate energy. To study the advantages of the proposed architecture and its performance at the system level, we architect tinyForge, a design space exploration to obtain Pareto-optimal architectures and compare the trade-offs with respect to traditional approaches. tinyForge comprises (1) a parameterized template for memory hierarchies and compute fabric, (2) estimations of power, area, and latency for hardware components, (3) a dataflow optimizer for efficient workload scheduling, (4) a genetic algorithm performing multi-objective optimization to find Pareto-optimal architectures. We evaluate the performance of our proposed architecture on all of the MLPerf Tiny Inference Benchmark workloads, and the BERT-Tiny transformer model, demonstrating its effectiveness in lowering the energy per inference while addressing the introduced area overheads. We show the importance of storing all the weights on-chip, reducing the energy per inference by 7.5x vs. utilizing off-chip memories. Finally, we demonstrate the potential of the custom latch arrays and bit-serial digital compute arrays to reduce by up to 1.8x the energy per inference, 2.2x the latency per inference, and 3.7x the silicon area. Massimo Giordano, Rohan Doshi, Qianyun Lu, Boris Murmann |
ASPLOS (3) | 4 |
| 2024 | On Stress: Combining Human Factors and Biosignals to Inform the Placement and Design of a Skin-like Stress SensorabstractWith advances in electronic-skin and wearable technologies, it is possible to continuously measure stress markers from the skin and sweat to monitor and improve wellbeing and health. Understandably, the sensor’s engineering and resolution are important towards its function. However, we find that people looking for an e-skin stress sensor may look beyond measurement precision, demanding a private and stealth design to reduce, for example, social stigmatization. We introduce the idea of a stress sensing "wear index," created from the combination of human-centered design (n=24), physiological (n=10), and biochemical (n=16) data. This wear index can inform the design of stress wearables to fit specific applications, e.g., human factors may be relevant for a wellbeing application, versus a relapse prevention application that may require more sensing precision. Our wear index idea can be further generalized as a method to close gaps between design and engineering practices. Yasser Khan, Matthew Louis Mauriello, Parsa Nowruzi, Akshara Motani, Grace Hon, Nicholas H. Vitale, Jinxing Li 0006, Amir Foudeh, Dalton Duvio, Erika Shols, Megan Chesnut, James A. Landay, Jan T. Liphardt, Leanne M. Williams, Keith D. Sudheimer, Boris Murmann, Zhenan Bao, Pablo Paredes |
CHI | 17 |
| 2024 | Reinforcement Learning-Enhanced Cloud-Based Open Source Analog Circuit Generator for Standard and Cryogenic Temperatures in 130-nm and 180-nm OpenPDKsabstractThis work introduces an open-source, Process Technology-agnostic framework for hierarchical circuit netlist, layout, and Reinforcement Learning (RL) optimization. The layout, netlist, and optimization python API is fully modular and publicly installable. It features a bottom-up hierarchical construction, which allows for complete design reuse across provided PDKs. The modular hierarchy also facilitates parallel circuit design iterations on cloud platforms. To illustrate its capabilities, a two-stage OpAmp with a 5T first-stage, common-source second-stage, and miller compensation is implemented. We instantiate the OpAmp in two different open-source process design kits (OpenPDKs) using both room-temperature models and cryogenic (4K) models. With a human designed version as the baseline, we leveraged the parameterization capabilities of the framework and applied the RL optimizer to adapt to the power consumption limits suitable for cryogenic applications while maintaining gain and bandwidth performance. Using the modular RL optimization framework we achieve a 6x reduction in power consumption compared to manually designed circuits while maintaining gain to within 2%. Ali Hammoud, Anhang Li, Ayushman Tripathi, Harsh Khandeparkar, Ryan Wans, Gregory Kielian, Boris Murmann, Dennis Sylvester, Mehdi Saligane |
ICCAD | 8 |
| 2024 | Practical Aspects of Script-Based Analog Design Using Precomputed Lookup TablesabstractRatio- and lookup-table based sizing methods were created to eliminate the need for SPICE-based iterative tweaking and to achieve a good match between theoretical analysis and simulated performance. Following such approaches can be gratifying for seasoned designers who have experienced the large mismatch between textbook hand calculations and simulation results firsthand. However, the circuit design novice may fall into the trap of tweaking sizing scripts iteratively and abandoning the connection to the analytical underpinnings of the target design. In this paper, we summarize the best practices for script-based analog circuit design toward the development of analog generators. These methods are illustrated using a two-stage amplifier as an example. Boris Murmann |
ISCAS | 1 |
| 2024 | Enhancing the Energy Efficiency and Robustness of tinyML Computer Vision Using Coarsely-quantized Log-gradient Input ImagesabstractThis article studies the merits of applying log-gradient input images to convolutional neural networks (CNNs) for tinyML computer vision (CV). We show that log gradients enable: (i) aggressive 1-bit quantization of first-layer inputs, (ii) potential CNN resource reductions, (iii) inherent insensitivity to illumination changes (1.7% accuracy loss across 2 -5 … 2 3 brightness variation vs. up to 10% for JPEG), and (iv) robustness to adversarial attacks (>10% higher accuracy than JPEG-trained models). We establish these results using the PASCAL RAW image dataset and through a combination of experiments using quantization threshold search, neural architecture search, and a fixed three-layer network. The latter reveals that training on log-gradient images leads to higher filter similarity, making the CNN more prunable. The combined benefits of aggressive first-layer quantization, CNN resource reductions, and operation without tight exposure control and image signal processing (ISP) are helpful for pushing tinyML CV toward its ultimate efficiency limits. Qianyun Lu, Boris Murmann |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2023 | High-Linearity High-Bandwidth (>20GHz) T&H Front Ends Using Active Bootstrapping and Heterogeneous SiGe/CMOS Circuit Co-DesignabstractThis paper explores technology-circuit co-design techniques that leverage the heterogeneous integration of ad-vanced SiGe BiCMOS and deeply-scaled CMOS. We introduce an active bootstrapping concept that extends the bandwidth of a track-and-hold (T&H) front end to several tens of GHz (> 20 GHz), while preserving high spectral purity (> 55 dB across the entire band) and consuming low power. The proposed techniques are demonstrated through circuit simulations of a die stack featuring 130 nm SiGe BiCMOS and 22 nm FDSOI. Using 4x time-interleaving, the T&H front end achieves up to 22 GS/s sample rate, a bandwidth above 20 GHz, an IM3 better than -60 dB (up to 20 GHz), and a power consumption of 150 mW. For comparison, simulation results of a single-die 22 nm FDSOI front end are also provided. Athanasios Ramkaj, Michael Perrott, Baher Haroun, Boris Murmann |
ISCAS | 4 |
| 2021 | TinyML: Current Progress, Research Challenges, and Future RoadmapabstractTinyML: tiny in size, BIG in impact!This paper highlights the current progress, challenges and open research opportunities in the domain of tinyML, benchmarking, and emerging applications for Edge-AI. Muhammad Shafique 0001, Theocharis Theocharides, Vijay Janapa Reddi, Boris Murmann |
DAC | 4 |
| 2021 | A 7-bit 2 GS/s Time-Interleaved SAR ADC With Timing Skew Calibration Based on Current Integrating SamplerabstractThis paper presents a two-way time-interleaved (TI) 7-bit 2-GS/s successive-approximation-register (SAR) analog-to-digital converter (ADC) in 28 nm CMOS. The design achieves wideband operation with an effective resolution bandwidth (ERBW) in the 3rdNyquist zone. The converter's front-end employs current integrating (CI) sampler that provide both buffering and anti-alias (AA) filtering at low power dissipation. Facilitated by the CI-samplers' inherent inter-sample interactions, the timing mismatch among the TI channels can be detected in the amplitude domain, obviating the need for a dedicated reference channel for background calibration. After calibration, the ADC achieves 36.4 dB signal-to-noise-and-distortion ratio (SNDR) near Nyquist and >2.6 GHz ERBW at a sampling rate of 2 GS/s. The ADC's power consumption is 7.62 mW (including the CI buffer) and its Walden figure of merit (FoMw) is 70.8 fJ/conversion-step. Wenning Jiang, Yan Zhu 0001, Chi-Hang Chan, Boris Murmann, Rui Paulo Martins |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | An 800 nW Switched-Capacitor Feature Extraction Filterbank for Sound ClassificationabstractThis paper presents a 32-channel analog filterbank for front-end signal processing in sound classification systems. It employs a passive N-path switched capacitor topology to achieve high power efficiency and reconfigurability. The circuit's unwanted harmonic mixing products are absorbed by the machine learning model during training. To enable a systematic pre-silicon study of this effect, we develop a computationally efficient circuit model that can process large machine learning datasets on practical time scales. Measured results using a 130 nm CMOS prototype IC indicate competitive classification accuracy on datasets for baby cry detection (93.7% AUC) and voice commands (92.4% average precision), while lowering the feature extraction energy compared to digital realizations by approximately 2× and 10×, respectively. The 1.44 mm2chip consumes 800 nW, which corresponds to the lowest normalized power per simultaneously sampled channel in recent literature. Daniel Villamizar, Dante Gabriel Muratore, James B. Wieser, Boris Murmann |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | Mixed-Signal Computing for Deep Neural Network InferenceabstractModern deep neural networks (DNNs) require billions of multiply-accumulate operations per inference. Given that these computations demand relatively low precision, it is feasible to consider analog computing, which can be more efficient than digital in the low-SNR regime. This overview article investigates the potential of mixed analog/digital computing approaches in the context of modern DNN processor architectures, which are typically limited by memory access. We discuss how memory-like and in-memory compute fabrics may help alleviate this bottleneck and derive asymptotic efficiency limits at the processing array level. It is shown that single-digit fJ/op energy efficiencies are feasible for 4-bit mixed-signal arithmetic. In this analysis, special consideration is given to the SNR and amortization requirements of the analog-digital interfaces. In addition, we consider the pros and cons for a variety of implementation styles and highlight the challenge of retaining high compute efficiency for a complete DNN accelerator design. Boris Murmann |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | Analog and Mixed-Signal Layout Automation Using Digital Place-and-Route ToolsabstractToday’s analog and mixed-signal (AMS) layout flow requires long manual iterations and does not leverage computing resources for data-driven optimization. This issue is further compounded by the explosion of design rules and layout-dependent effects (LDEs). We present an AMS layout generation flow that leverages digital place-and-route (PnR) tools, amortizes setup cost with reusable primitives, and prunes layout candidates using a fast evaluation scheme. We also analyze LDEs and parasitics and investigate unique challenges and mitigation strategies associated with using digital PnR tools for AMS circuits. These insights are validated with a generated StrongARM comparator and a voltage-controlled oscillator (VCO). The VCO layout was optimized in 2 h and fabricated in 16-nm FinFET CMOS. Silicon measurement results of the VCO closely track the simulation, verifying the methodology from netlist to silicon. Po-Hsuan Wei, Boris Murmann |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Sensory Particles with Optical TelemetryabstractCurrent retinal prostheses provide electrical stimulation without feedback from the stimulated neurons. Incorporation of multichannel recording electronics would typically require trans-scleral cables for power supply and data transmission. In this work, we explore a wireless, optoelectronic, miniature, modular, and distributed electro-neural interface for recording, which we call Sensory Particles with Optical Telemetry (SPOT). It can be used in an advanced, bi-directional retinal prosthesis and other sensory applications. Emphasis is placed on the novel telemetry stage. SPOTs are powered by near-infrared light and transmit information by light. As a proof of concept, we designed and built a low-power, small-footprint linear transconductance circuit utilizing chopper stabilization in 130nm CMOS. Our design achieved 57 mS transconductance within 3.5 kHz bandwidth, and a near-infrared (NIR) power density of 0.5 mW/mm2, well within the ocular and thermal safety limits. The telemetry circuit consumes 0.015 mm2area, and each SPOT can be powered by a single photovoltaic (PV) supply of area 0.0056 mm2. Electrical spikes transmitted by an 850nm LED were detected with 15 dB SNR, at the output of the optical link. Karthik Ganesan 0001, Thomas A. Flores, Binh Q. Le, Dante Gabriel Muratore, Neal A. Patel, Subhasish Mitra, Boris Murmann, Daniel Palanker |
ISCAS | 7 |
| 2020 | A 32 Gb/s PAM-4 Optical Transceiver with Active Back Termination in 40 nm CMOS TechnologyabstractThis paper describes the design of a 32 Gb/s four-level pulse amplitude modulation (PAM-4) optical transceiver in a 40 nm CMOS technology. At the transmitter side, the laser driver is composed of an asymmetric waveform equalizer, a 3-tap feed-forward equalizer (FFE), and a novel active-back termination (ABT) circuit. The ABT circuit provides a self-tracking, tunable source impedance to match the characteristic impedance of different laser diodes. At the receiver side, the fully integrated optical receiver consists of a transimpedance amplifier, a variable gain amplifier, an automatic threshold tracking circuit (ATC), and a quarter-rate decision feedback equalizer (DFE). By using the adaptive ATC, it reduces the BER induced by the harmonic distortion along the signal path by more than 27X. Both the ATC and DFE are automatically adapted by an on-chip sign-sign LMS (SSLMS) engine. Fabricated in TSMC 40 nm CMOS process, the chip area for the transmitter and receiver are about 0.029 mm2and 0.23 mm2. The power consumptions are about 146.8 mW and 128.8 mW respectively for the PAM-4 transmitter and receiver. Wei-Hsiang Ho, Yi-Hsun Hsieh, Boris Murmann, Wei-Zen Chen |
ISCAS | 3 |
| 2020 | Design Considerations for External Compensation Approaches to OLED Display DegradationabstractThis paper presents design considerations for compensation circuitry that addresses OLED display degradation. It focuses on an external compensation method that utilizes an analog-to-digital converter (ADC) in the column driver IC. External compensation has the advantage of addressing both threshold voltage (VTH) and mobility (μ) shifts in the driving thin-film transistor (TFT). By maintaining pixel-to-pixel luminance uniformity, it addresses not only image sticking issues but also leads to a lifetime extension of the OLED pixels. Especially for large-sized OLED panels, it is important to understand the noise contributions from the analog front-end of the compensation circuitry. Noise contributions for each component are thus analyzed for maximizing the overall compensation performance. The analysis results show that the noise from the driving TFT and the display panel's parasitics dominate the total noise and that reducing the noise bandwidth can be an effective noise mitigation strategy. Furthermore, the presented results provide guidance on the required ADC specifications. Jaewook Kwon, Changuk Lee, Youngcheol Chae, Boris Murmann |
ISCAS | 4 |
| 2020 | Implications of Finite Clock Transition Time for LPTV Circuit AnalysisabstractModeling linear periodically time-varying (LPTV) circuits is challenging due to the presence of frequency translation. Many approaches have been proposed that simplify the analysis and provide intuition into the operation of these circuits. It is critical to select the proper model when designing LPTV systems: too complex, and intuition is lost; too simple, and numerical accuracy degrades. This work shows how a conversion matrix-based model can be used for mixer-first receivers with complex feedback in the presence of finite switch transitions. This model accurately predicts S11below -10 dB for all tested transition times, in contrast with prior models, which are shown to be invalid with transitions beyond 2% of the clock period. As a design tool, this approach models gain, harmonic rejection ratio, and noise figure within 0.1 dB of simulation with switch transitions even at 5% of the clock period. Stephen Weinreich, Dante Gabriel Muratore, Youngcheol Chae, Thomas McKay, Boris Murmann |
ISCAS | 5 |
| 2020 | Wearable System Design using Intrinsically Stretchable Temperature SensorabstractThis paper explores system integration aspects for stretchable electronics. We consider a prototypical wearable device that combines (1) a conformal and stretchable transducer patch, (2) high-fidelity readout electronics based on rigid CMOS integrated circuits (ICs), and (3) flexible and stretchable interconnects that bridge the two mechanical domains. The considered temperature sensing element is based on organic thin-film transistors (OTFTs) and requires bias currents in the sub-microampere regime. We show how to meet this requirement using standard commercial parts. Furthermore, we present an interconnect solution based on encapsulated liquid-phase conductors and show its measured I-V curves under strain. The design guidelines and parameters obtained from this study can be used to aid the development of similar wearable sensing platforms. Chenxin Zhu, Elizabeth Schell, Min-gu Kim, Zhenan Bao, Boris Murmann |
ISCAS | 5 |
| 2019 | Memory-Optimal Direct Convolutions for Maximizing Classification Accuracy in Embedded ApplicationsabstractIn the age of Internet of Things (IoT), embedded devices ranging from ARM Cortex M0s with hundreds of KB of RAM to Arduinos with 2KB RAM are expected to perform increasingly sophisticated classification tasks, such as voice and gesture recognition, activity tracking, and biometric security. While convolutional neural networks (CNNs), together with spectrogram preprocessing, are a natural solution to many of these classification tasks, storage of the network’s activations often exceeds the hard memory constraints of embedded platforms. This paper presents memory-optimal direct convolutions as a way to push classification accuracy as high as possible given strict hardware memory constraints at the expense of extra compute. We therefore explore the opposite end of the compute-memory trade-off curve from standard approaches that minimize latency. We validate the memory-optimal CNN technique with an Arduino implementation of the 10-class MNIST classification task, fitting the network specification, weights, and activations entirely within 2KB SRAM and achieving a state-of-the-art classification accuracy for small-scale embedded systems of 99.15%. Albert Gural, Boris Murmann |
ICML | 2 |
| 2019 | Long-Short Term Memory Neural Network Stability and Stabilization using Linear Matrix InequalitiesabstractA global asymptotic stability condition for Long Short-Term Memory neural networks is presented in this paper. A linear matrix inequality optimization problem is used to describe this global stability condition. The linear matrix inequality formulation can be viewed as a way for stabilization of Long Short-Term Memory neural networks since the networks' weight matrices and biases can be essentially treated as control variables. The condition and how to compute numerical values for the weight matrices and biases are illustrated by some examples. Shankar A. Deka, Dusan M. Stipanovic, Boris Murmann, Claire J. Tomlin |
ISCAS | 3 |
| 2019 | A Data-Compressive Wired-OR Readout for Massively Parallel Neural RecordingabstractThis paper describes an architecture for the massively parallel digitization of neural action potentials. The scheme achieves simultaneous data compression and channel multiplexing through wired-OR interactions within an array of single-slope A/D converters. The achieved compression is lossy but effective at retaining the critical samples belonging to action potential spikes. Simulation results using ex-vivo experimental data from a 512-channel array show compression rates up to ~73x while maintaining ≥90% reconstruction coverage for parasol cells in the primate retina. Dante Gabriel Muratore, Pulkit Tandon, Mary Wootters, E. J. Chichilnisky, Subhasish Mitra, Boris Murmann |
ISCAS | 6 |
| 2019 | Sound Classification using Summary Statistics and N-Path FilteringabstractAlways-on sound classification is a desirable but power-intensive function for a variety of emerging Internet of Everything applications. This work explores the accuracy-complexity tradeoff by using summary statistics for classifying semi-stationary sounds. Compared to contemporary solutions including deep learning, this approach requires one to three orders of magnitude fewer parameters and can therefore be trained over ten times faster. We propose a mixed-signal design using N-path filters for feature extraction to further improve energy efficiency without incurring a large accuracy penalty for a binary classification task (less than 2.5% area reduction under receiver operating characteristic curve). Daniel Villamizar, Daniele Battaglino, Dante Gabriel Muratore, Reza Hoshyar, Boris Murmann |
ISCAS | 5 |
| 2018 | TRIG: hardware accelerator for inference-based applications and experimental demonstration using carbon nanotube FETsabstractThe energy efficiency demands of future abundant-data applications, e.g., those which use inference-based techniques to classify large amounts of data, exceed the capabilities of digital systems today. Field-effect transistors (FETs) built using nanotechnologies, such as carbon nanotubes (CNTs), can improve energy efficiency significantly. However, carbon nanotube FETs (CNFETs) are subject to process variations inherent to CNTs: variations in CNT type (semiconductor or metallic), CNT density, or CNT diameter, to name a few. These CNT variations can degrade CNFET benefits at advanced technology nodes. One path to overcome CNT variations is to co-optimize CNT processing and CNFET circuit design; however, the required CNT process advancements have not been achieved experimentally. We present a new design approach (TRIG, Technique for Reducing errors using Iterative Gray code) to overcome process variations in hardware accelerators targeting inference-based applications that use serial matrix operations (serial: accumulated over at least 2 clock cycles). We demonstrate that TRIG can retain the major energy efficiency benefits (quantified using Energy Delay Product or EDP) of CNFETs despite CNT variations that exist in today's CNFET fabrication - without requiring further CNT processing improvements to overcome CNT variations. As a case study, we analyze the effectiveness of TRIG for a binary neural network hardware accelerator that classifies images. Despite CNT variations that exist today, TRIG can maintain 99% (90%) of projected EDP benefits of CNFET digital circuits for 90% (99%) image classification accuracy target. We also demonstrate experimentally fabricated CNFET circuits to compute scalar product (a common matrix operation, also called dot product), with and without TRIG: TRIG reduces the mean difference between the expected result (no errors) and the experimentally computed result by 30× in the presence of CNT variations, shown experimentally. Gage Hills, Daniel Bankman, Bert Moons, Lita Yang, Jake Hillard, Alex Kahng, Rebecca Park, Marian Verhelst, Boris Murmann, Max M. Shulaker, H.-S. Philip Wong, Subhasish Mitra |
DAC | 9 |
| 2018 | A New Figure of Merit Equation for Analog-to-Digital Converters in CMOS Image SensorsabstractThis paper presents a new figure-of-merit (FoM) equation for column analog-to-digital converters (ADCs) in CMOS image sensors. The proposed FoM incorporates the dynamic range, resulting from the minimum dark-level random noise and the maximum full well capacity, as well as AD conversion time, power consumption, and area. Using this FoM, we have analyzed various ADC architectures including cyclic, delta-sigma, successive approximation register, single-slope (SS), two-step and hybrid topologies. Using the suggested equation and data collected over the past ten years, we elucidate the reasons behind the dominance of the column-based SS ADC in commercial products. The proposed FoM is therefore useful for gauging the potential of new architectures and for architecture selection at an early stage of the design process. Minho Kwon, Boris Murmann |
ISCAS | 2 |
| 2018 | Some Local Stability Properties of an Autonomous Long Short-Term Memory Neural Network ModelabstractIn this paper some local stability results for an autonomous Long Short-Term Memory neural network model with respect to the origin are provided. In particular, it is shown through linearization that the local asymptotic stability conditions with respect to the origin only depend on one of the weight matrices. Simulations indicate that these local stability conditions greatly influence the behavior of the autonomous four-dimensional neural network in the region where each variable's values vary between minus one and one. Finally, some sufficient stability conditions for the nonlinear model are formulated as a convex program involving linear matrix inequalities. Dusan M. Stipanovic, Boris Murmann, Matteo Causo, Aleksandra Lekic, Vicenc Rubies-Royo, Claire J. Tomlin, Edith Beigné, Sébastien Thuries, Mykhailo Zarudniev, Suzanne Lesecq |
ISCAS | 2 |
| 2018 | Bit Error Tolerance of a CIFAR-10 Binarized Convolutional Neural Network ProcessorabstractDeployment of convolutional neural networks (ConvNets) in always-on Internet of Everything (IoE) edge devices is severely constrained by the high memory energy consumption of hardware ConvNet implementations. Leveraging the error resilience of ConvNets by accepting bit errors at reduced voltages presents a viable option for energy savings, but few implementations utilize this due to the limited quantitative understanding of how bit errors affect performance. This paper demonstrates the efficacy of SRAM voltage scaling in a 9-layer CIFAR-10 binarized ConvNet processor, achieving memory energy savings of 3.12× with minimal accuracy degradation (~99% of nominal). Additionally, we quantify the effect of bit error accumulation in a multi-layer network and show that further energy savings are possible by splitting weight and activation voltages. Finally, we compare the measured error rates for the CIFAR-10 binarized ConvNet against MNIST networks to demonstrate the difference in bit error requirements across varying complexity in network topologies and classification tasks. Lita Yang, Daniel Bankman, Bert Moons, Marian Verhelst, Boris Murmann |
ISCAS | 5 |
| 2018 | Toward Always-On Mobile Object Detection: Energy Versus Performance Tradeoffs for Embedded HOG Feature ExtractionabstractThis paper studies the effects of front-end imager parameters on object detection performance and energy consumption. A custom version of histograms of oriented gradient (HOG) features based on 2-b pixel ratios is presented and shown to achieve superior object detection performance for the same estimated energy compared with conventional HOG features. A front-end hardware implementation capable of extracting these features at multiple scales is proposed, and a system-level energy analysis is performed. This energy analysis suggests a potential 19× reduction in I/O energy and a 3.3× reduction in back-end detection energy compared with conventional object detection pipelines. Alex Omid-Zohoor, Christopher Young, David Ta, Boris Murmann |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | LogNet: Energy-efficient neural networks using logarithmic computationabstractWe present the concept of logarithmic computation for neural networks. We explore how logarithmic encoding of non-uniformly distributed weights and activations is preferred over linear encoding at resolutions of 4 bits and less. Logarithmic encoding enables networks to 1) achieve higher classification accuracies than fixed-point at low resolutions and 2) eliminate bulky digital multipliers. We demonstrate our ideas in the hardware realization, LogNet, an inference engine using only bitshift-add convolutions and weights distributed across the computing fabric. The opportunities from hardware work in synergy with those from the algorithm domain. Edward H. Lee, Daisuke Miyashita, Elaina Chai, Boris Murmann, S. Simon Wong |
ICASSP | 4 |
| 2015 | Mixer-based subarray beamforming for sub-Nyquist sampling ultrasound architecturesabstractUltrasound imagers suffer from a large data rate between their analog to digital converter (ADC) front-end and digital beamforming backend. This becomes a limiting factor when the number of elements is increased, such as in modern 2D transducers. To address this issue, prior work considered sub-Nyquist sampling techniques that exploit the disparity between the signal's physical bandwidth and innovation rate. In this work, we extend this framework using an analog-domain subarray beamforming technique that is feasible due to the narrowband nature of the sub-Nyquist signal acquisition. When applied to waveforms taken from a commercial ultrasound machine, this method reduces both the low-rate ADC count by a factor of eight and the total data rate by a factor of 54 with minimal image degradation. Jonathon Spaulding, Yonina C. Eldar, Boris Murmann |
ICASSP | 3 |
| 2015 | Calculation of MOSFET distortion using the transconductance-to-current ratio (gm/ID)abstractWe present analytical expressions for MOSFET distortion as a function of inversion level, represented by gm/IDas a proxy. The expressions are particularly useful for moderate inversion, where the generic textbook equations fail. Unlike previous approaches, the method requires only a small number of technology parameters. For an estimation of gmnonlinearity, only the subthreshold slope is needed. Two additional parameters are required to model gdsnonlinearity. The derived expressions are validated based on SPICE simulations using a well-calibrated 65-nm PSP model set. Paul G. A. Jespers, Boris Murmann |
ISCAS | 2 |
| 2014 | Low-voltage organic transistors for flexible electronicsabstractA process for the fabrication of bottom-gate, top-contact (inverted staggered) organic thin-film transistors (TFTs) with channel lengths as short as 1 μm on flexible plastic substrates has been developed. The TFTs employ vacuum-deposited small-molecule semiconductors and a low-temperature-processed gate dielectric that is sufficiently thin to allow the TFTs to operate with voltages of about 3 V. The p-channel TFTs have an effective field-effect mobility of about 1 cm2/Vs, an on/off ratio of 107, and a signal propagation delay (measured in 11-stage ring oscillators) of 300 ns per stage. For the n-channel TFTs, an effective field-effect mobility of about 0.06 cm2/Vs, an on/off ratio of 106, and a signal propagation delay of 17 μs per stage have been obtained. Ute Zschieschang, Reinhold Rodel, Ulrike Kraft, Kazuo Takimiya, Tarek Zaki, Florian Letzkus, Joerg Butschke, Harald Richter 0001, Joachim N. Burghartz, Boris Murmann, Hagen Klauk |
DATE | 11 |
| 2014 | Low-rate identification of memory polynomialsabstractWe propose a new approach for the low-rate identification of memory polynomials (MPs), which are frequently used to model RF power amplifiers (PAs). Based on ideas from the finite rate of innovation framework, we find the coefficients of the MP in the frequency domain, which requires a relatively small number of measurements (samples), commensurate with the degrees of freedom in the model. By choosing a random set of frequency components, the stability of the identification is ensured. We show that the method can be used directly for special input signals, such as one used for orthogonal frequency division multiplexing, and extend the idea for arbitrary inputs. Experiments using measured data from a class-AB PA demonstrate the effectiveness of the approach. The sampling rate for identifying this PA is reduced by a factor of 1024 (from 107.52MHz to 105 kHz). Nikolaus Hammler, Yonina C. Eldar, Boris Murmann |
ISCAS | 3 |
| 2014 | Design and optimization of continuous-time filters using geometric programmingabstractThis paper is concerned with the design and optimization of continuous-time active-RC and gm-C filters. We demonstrate how we can maximize the filters' dynamic range (DR) for a given voltage swing, area and power consumption. Using closed-form symbolic expressions, the optimization problems are formulated as geometric programs (GPs) and mixed-integer GPs (MIGPs) that can be quickly solved to find the globally-optimal solution. The techniques developed in this paper are applied to the design and optimization of an active-RC filter with 1 MHz bandwidth, and a gm-C filter with 30 MHz bandwidth, both of which implement a 5thorder elliptic transfer function. The proposed design flow is general, and can be extended to any filter topology or any set of specifications in order to find a low-noise, linear, low-area, and low-power circuit solution. Siddharth Seth, Boris Murmann |
ISCAS | 2 |
| 2014 | Teaching an old dog new tricks: Views on the future of mixed-signal IC designabstractIn the past, CMOS feature size scaling has played a big role in overcoming the perceived barriers, routinely enabling cheaper, faster and lower power devices with every new technology node. However, as the benefits of conventional feature size scaling are diminishing, what can we do to meet the needs of next-generation systems? The precise answer to this question is unclear, but most researchers will agree that some of the progress will have to come from innovative re-architecting and looking for better ways to employ the amazing nano-CMOS fabric that we already have. In this talk, I will review opportunities for system-driven architectural innovation and new application areas in mixed-signal IC design. We will discuss a number of examples related to the idea of “fooling Nyquist” and extracting desired analog-domain information using low-rate and low-bandwidth observations. In addition, we will investigate the potential for mixed-signal co-processors in machine learning algorithms, as well as trends in sensor systems. Along with these examples, we will discuss potential challenges and future needs in mixed-signal testing. Boris Murmann |
ITC | 1 |
| 2012 | Towards an integrated circuit design of a compressed sampling wireless receiverabstractIn this paper, we investigate the extension of previous work on compressed sampling receivers from mathematical abstractions and proof of concept work into a system that directly competes with more traditional receivers in standard CMOS integrated circuit technology. As developed in the literature, the Modulated Wideband Converter shows great promise as a compressed sampling receiver due to its flexibility and inherent spectral agility. We propose several modifications to the system that improve its usability and performance in real-world scenarios. Then, using standard LTE receivers as a basis for comparison, we propose a set of target specifications for the Modulated Wideband Converter, discuss the associated circuit challenges, and evaluate potential solutions that build upon prior work in integrated circuit and system design. Douglas Adams 0002, Chester Sungchung Park, Yonina C. Eldar, Boris Murmann |
ICASSP | 4 |
| 2012 | Settling time and noise optimization of a three-stage operational transconductance amplifierabstractThis paper presents the design and optimization of a nested-Miller compensated, three-stage operational transconductance amplifier (OTA) in 90-nm CMOS that is used in a switched-capacitor (SC) gain stage clocked at 200 MHz. Existing design methods for three-stage OTAs lead to sub-optimal solutions because they decouple inter-related metrics like noise and settling performance. In our approach, the problem of finding an optimal design with the best total integrated noise and settling time has been cohesively solved by formulating a nonlinear constrained optimization program. Equality, inequality, and semi-infinite constraints are formed using closed form symbolic expressions obtained by a closed loop analysis of the SC gain stage and the optimization program is solved by using the interior-point algorithm. Simulation results show that the amplifier achieves a ± 0.1% dynamic error settling time of 2.5 ns with a total integrated noise of 240 μVrms, while consuming 5.2 mW from a 1-V power supply. Siddharth Seth, Boris Murmann |
ISCAS | 2 |
| 2010 | Design of analog circuits using organic field-effect transistorsabstractOrganic field-effect transistors (OFETs) can be manufactured at low temperatures, enabling the fabrication of integrated circuits on flexible plastic substrates and the coverage of large areas at potentially low cost. This paper evaluates state-of-the-art OFET technology from the perspective of the analog circuit designer. Specifically, we review important OFET device performance metrics in comparison to generic silicon CMOS transistors. In addition, an overview of recent accomplishments in OFET-based analog design is presented. Boris Murmann |
ICCAD | 1 |
| 2010 | Portable biomarker detection with magnetic nanotagsabstractThis paper presents a hand-held, portable biosensor platform for quantitative biomarker measurement. By combining magnetic nanoparticle (MNP) tags with giant magnetoresistive (GMR) spin-valve sensors, the hand-held platform achieves highly sensitive (picomolar) and specific biomarker detection in less than 20 minutes. The rapid analysis and potential low cost make this technology ideal for point-of-care (POC) diagnostics. Furthermore, this platform is able to detect multiple biomarkers simultaneously in a single assay, creating a promising diagnostic tool for a vast number of applications. Drew A. Hall, Shan X. Wang, Boris Murmann, Richard S. Gaster |
ISCAS | 3 |
| 2008 | Predictive control algorithm for phase-locked loopsabstractPhase-locked loops (PLLs) exhibit a tradeoff between settling time and noise rejection, due to the fact that a low noise PLL requires a narrow bandwidth (BW) loop filter, which degrades settling time. However, the moments when fast settling or good noise rejection is required are clearly identified in a PLL, and this can be used to overcome this tradeoff. A recent technique - PLL gear shifting - exploits this fact by modifying the loop filter BW according to the PLL current objective. In this work, a similar solution based on predictive control techniques is presented. Through a very simple digital loop filter an optimal response is obtained, where a single parameter controls the bandwidth to improve either settling time or noise rejection. Angel Abusleme, Boris Murmann |
ISCAS | 2 |
| 2008 | General analysis on the impact of phase-skew in time-interleaved ADCsabstractTime-interleaved analog-to-digital converters (TIADC) are sensitive to various mismatches that distort the sampled signal. Standard TIADC analysis assumes a sinusoidal input, which may result in pessimistic matching constraints for system-specific ADCs used with wideband input signals. Closed-form expressions bounding the acceptable phase-skew for wideband systems are derived, and are validated through simulations. In one of the examples presented it is shown that standard analysis can overconstrain the bound on acceptable phase-skew variance by a factor of three. Manar El-Chammas, Boris Murmann |
ISCAS | 2 |
| 2008 | Digitally enhanced analog circuits: System aspectsabstractAn overview of digital enhancement techniques for analog circuits is presented. Recent research suggests that the high density and low energy of digital circuits can be leveraged to enable a new generation of interface electronics that is based on minimal precision, low complexity analog blocks. Today, examples of enhancement schemes can be found in diverse applications and include nonlinearity compensation of ADCs, predistortion of power amplifiers and mismatch calibration in radio receivers. Since it is often difficult to identify commonalities among these different, but conceptually related schemes, this tutorial paper aims to provide a unified and system-oriented perspective of the field. Boris Murmann, Christian Vogel 0001, Heinz Koeppl |
ISCAS | 1 |
| 2006 | 4.25 Gb/s laser driver: design challenges and EDA tool limitationsabstractThis paper describes the design methodology, simulation, and tools used to design a 4.25 Gb/s high output swing laser driver (LD) and the electrical to optical interface from the LD to the laser diode. The quality of the optical output of a fiber optic communication channel is mainly determined by the LD and the electrical interface from the LD to the laser diode. Of particular importance in the interface is how well the LD overcomes the impact of the parasitic, resistive, capacitive, and inductive elements associated with the bondpad, bondwires, package, PCB transmission lines, passive components, and laser diode and its bondwires. The EDA tools used to model the electrical parasitics focus on RF and microwave applications and provide high frequency S-parameter models. This environment requires a stable time domain model of the electrical to optical interface. The presented LD integrated circuit operates from 155 Mb/s to 4.25 Gb/s with rise and fall times of 70 ps or less and a wide output voltage range, and a modulation current range of 5 mA to 85 mA. Benjamin Sheahan, John W. Fattaruso, Jennifer Wong, Karlheinz Muth, Boris Murmann |
DAC | 5 |