EDBT 2026 Demo / reviewers in the wild / expert
Benton H. Calhoun
dblp:15/4871
· DBLP profile ↗
72ranked-venue papers
13as first author
16since 2021 · last 2026
0000-0002-3770-5050ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 68 · 10 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 6 first-authorSoftware engineering, systems software and programming languages · 5 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TAuC-SoC: A 0.36 mm2 In-Yarn Integrable Tiny Audio Compressor SoC for E-Textile Based Audio Recording Applications
Suprio Bhattacharya, Charlie D. Hess, Omar Faruqe, J. Keith McElveen, Douglas Cairns, Doug Overland, Jonathan Johnson, Daniel S. Truesdell, Benton H. Calhoun |
ISCAS | 9 |
| 2026 | A High dv/dt-Tolerant Level Shifter with Sub-Nanosecond Delay Using Dummy Pulse Protection
Zane Gunn, Daniel S. Truesdell, Benton H. Calhoun |
ISCAS | 3 |
| 2025 | EveractiveSelf-Powered SoC with Energy Harvesting, Wakeup Receiver, and Energy-Aware Subsystem
Benton H. Calhoun, David D. Wentzloff, Kuo-Ken Huang, Kyle Craig |
HCS | 1 |
| 2025 | Characterization of Stacked PV Cell Configurations in a Deep N-Well 65nm CMOS TechnologyabstractThe selection of a suitable photovoltaic (PV) design for emerging energy harvesting applications remains challenging, as PV cell performance is highly dependent on design variables and environmental setup. These factors vary widely across existing studies, making direct comparisons complex and increasing uncertainty in predicting expected PV cell output characteristics for a given use case. To address this, our paper presents a comparison of various on-chip PV cell configurations, implemented using a 0.6mm x 0.6mm test chip realized in 65nm deep-nwell (DNW) CMOS technology. The experimental results demonstrate that, depending on diode arrangements, separate configurations achieve a maximum open-circuit voltage of 0.54V and a peak power density of 1.18 µW/mm2at 20 klux illumination. This characterization work, therefore, enables designers to select an optimal configuration tailored to specific application scenarios. Anjali Agrawal, Xinjian Liu, Daniel S. Truesdell, Benton H. Calhoun |
ISCAS | 4 |
| 2025 | A 0.36 mm2 On-the-Fly I2C-to-SPI Converter for E-Textile ApplicationsabstractThis paper presents a 0.36-mm2I2C-to-SPI converter chip with on-the-fly conversion for E-textile applications. The on-the-fly operation eliminates the need for on-chip data buffers and clock generation, improving the area and power efficiency versus prior works. The die area is 7.33× smaller than commercially-available off-the-shelf (COTS) I2Cto-SPI converters, enabling it to integrate unobtrusively into E-textile applications. The chip supports conversion at ultra-fast I2C operating frequency up to 5 MHz while consuming only 0.379 mW of power. At the standard I2C speed of 400 kHz, it consumes 0.145 mW, which is 48× more power-efficient than commercially available I2C-to-SPI converters. Omar Faruqe, Zhenghong Chen, Suprio Bhattacharya, Fahim Foysal, Samit Hasan, Daniel S. Truesdell, Benton H. Calhoun |
ISCAS | 7 |
| 2025 | A Fully Integrated, Custom End-to-End PPG Sensing System for Ultra-Low Power WearablesabstractPhotoplethysmography (PPG) is a widely adopted technique for monitoring essential physiological parameters like heart rate and blood oxygen saturation (SpO2). The increasing demand for wearable health monitoring devices necessitates the development of energy-efficient PPG systems. This paper proposes an ultra-low power, end-to-end PPG sensing system with continuous monitoring capabilities. The system consists of three distinct chips: a system-on-chip, an analog-front end chip, and a Bluetooth low-energy transmitter, enabling real-time PPG data transmission to a smartphone. All components are fabricated using the bulk 65 nm low-power CMOS process. This system achieves minimal power consumption of 18.78 µW during idle periods and 148.5 µW during active data transmission when aggressively duty-cycled at 2.4%. Omar Faruqe, Peter Le, Daehyun Lee, Xinjian Liu, Omar Abdelatty, Daniel S. Truesdell, Benton H. Calhoun |
ISCAS | 7 |
| 2025 | A Sub-μW Digital Temperature Compensation Architecture for Arbitrary Voltage and Current Reference GenerationabstractTraditionally, both reference circuits and the components they supply are designed independently to be as immune to temperature change as practical, but this requires power and area overhead to achieve. These overheads can compound in complex systems or consume excessive portions of a low power budget. In contrast, we propose a sub-μwatt digital temperature compensation architecture that generates voltage and current references with a user-defined temperature response rather than a fixed, near-ideal response. This flexible approach allows a single programmable design to be reused easily to produce different profiles over temperature, reducing design time. It also can reduce the temperature non-linearity in the components it supports by providing an input temperature profile that effectively cancels that non-linear response to allow components to operate at their target spec and eliminate excess power consumption. This approach allows designers to prioritize power consumption in their designs and use this compensation strategy to manage performance across temperature. Natalie B. Ownby, Prerana Singaraju, Suprio Bhattacharya, Steven M. Bowers, Benton H. Calhoun |
ISCAS | 5 |
| 2025 | A Compact, Power-Efficient, and On-the-Fly I2C-to-SPI Converter for Distributed E-Textile SystemsabstractThis paper presents a 0.36-mm2I2C-to-SPI converter chip designed for electronic textile (E-textile) applications, featuring an on-the-fly conversion scheme that eliminates the need for on-chip data buffers and internal clock generation. By leveraging the synchronous nature of both I2C and SPI protocols, the proposed design forwards each incoming I2C data bit, SDA (Serial Data Line) directly to the SPI output using the I2C serial clock line (SCL), thereby reducing both area and power consumption. Two versions of the chip are proposed: a ‘full’ die and a ‘compact’ die. The converter enables seamless integration into distributed in-textile architectures by minimizing silicon overhead. The ‘full’ die implementation achieves a$7.33\times $area reduction compared to commercially available I2C-to-SPI converters, while the ‘compact’ version further reduces the footprint by$14.73\times $through the use of a small corner seal-ring and a linear pad ring layout. Measurement results confirm robust operation at ultra-fast I2C frequency (5 MHz), consuming only 0.379 mW for the ‘full’ die and 0.333 mW for the ‘compact’ variant. At standard I2C speeds (400 kHz), the converter demonstrates a$48\times $improvement (‘full’ die) in power efficiency over commercial off-the-shelf (COTS) solutions, making it an effective and unobtrusive solution for next-generation E-textile systems. Omar Faruqe, Zhenghong Chen, Suprio Bhattacharya, Fahim Foysal, Samit Hasan, Jinhua Wang 0007, Daniel S. Truesdell, Benton H. Calhoun |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2024 | Scalable All-Analog LDOs With Reduced Input Offset Variability Using Digital Synthesis Flow in 65-nm CMOSabstractThe design and verification process for analog circuits can be long and tedious, wherein designers rely heavily on manual effort to create circuits and draw layouts, thereby limiting turn-around-time and design scale and increasing costs. Various previous works have tried to solve this issue by leveraging digital automated place-and-route (APR) tools, but they involve replacing analog elements with digital counterparts, thereby dampening performance. In this work, we propose a digital flow-based approach to design all-analog circuits that dramatically speeds up the design and layout process while retaining the benefits of true analog topologies and demonstrate the performance for three low-dropout regulators (LDOs). Fabricated in 65-nm CMOS, measurement results show that the generated LDOs achieve up to 99.95% peak current efficiency, a figure-of-merit (FOM) of 4.6 ps, and up to 63.93% reduction in input offset variability with respect to their manually designed counterparts. Shourya Gupta, Shuo Li 0008, Benton H. Calhoun |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | AuxcellGen: A Framework for Autonomous Generation of Analog and Memory Unit CellsabstractRecent advances in auto-generating analog and mixed-signal (AMS) circuits use standard digital tool flows to compose AMS circuits from a combination of digital standard cells and a set of auxiliary cells (auxcells). Until now, generating auxcell layouts for each new PDK was the last manual step in the flow for auto-generating AMS components, which limited the available auxcells and reduced the optimality of the auto-generated AMS designs. To solve this, we propose AuxcellGen, a framework to auto-generate auxcell layouts and performance models. Aux-cellGen generates a parasitic-aware auxcell performance model using a neural network (NN), auto-sizes and optimizes auxcell schematics for a given design target, and auto-generates auxcell layouts. The framework is demonstrated by auto-generating tristate buffer auxcells for PLLs and sense-amplifier auxcells for SRAM across a range of user specifications that are compatible with standard cell and memory bitcell pitch. Sumanth Kamineni, Arvind K. Sharma, Ramesh Harjani, Sachin S. Sapatnekar, Benton H. Calhoun |
DATE | 5 |
| 2023 | Modeling and Design of Cold-Start Charge Pumps for Photovoltaic Energy HarvestersabstractThis article presents modeling and design optimization of switched-capacitor charge pumps for use as cold-start circuits in photovoltaic energy harvesters. For a successful cold-start from photovoltaic energy, a certain output voltage and current must be generated by the charge pump while jointly minimizing the required input voltage and current from the energy harvesting source. We model the operation of the cold-start charge pump including its oscillator, derive expressions for required input voltage and current based on output requirements, and demonstrate a Pareto-optimal design space that defines the tradeoff between the required input voltage and current to achieve cold-start. We then assess the impacts of process and temperature variation, and suggest a general design methodology for cold-start charge pump circuits. Finally, we validate models with measurements from a design fabricated in 55nm CMOS. Daniel S. Truesdell, James Boley, Atul Wokhlu, Alain Gravel, David D. Wentzloff, Benton H. Calhoun |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | A Photoplethysmography Analog Front-End Model for Rapid Design of Personalized Healthcare HardwareabstractThis paper presents a comprehensive photoplethysmography (PPG) sensing analog front-end model in MATLAB for accelerating the design of personalized healthcare hardware. This model consists of four component modules, 1) the transimpedance amplifier (TIA), 2) the analog-to-digital converter (ADC), 3) the DC offset current cancellation block (DCOC), and 4) the LED driver, along with a system module which interacts with the four component modules for evaluating the system power consumption and signal-to-noise ratio (SNR) metrics. Once the designer inputs the power budget and the SNR target, this model presents the designer with the optimized design parameters, visualized power and noise breakdown plots, and detailed component specifications. Additional variables represent the user variances in the model to evaluate the hardware performance for different user groups. Compared to simulations in Cadence, the proposed model achieves worst-case errors of only 6.9% and 0.4% for estimating power and SNR, respectively. Peng Wang 0166, Benton H. Calhoun |
ISCAS | 2 |
| 2021 | Stacked Transconductance Boosting for Ultra-Low Power 2.4GHz RF Front-End DesignabstractThis paper presents a transconductance boosting method with a stacked current-reused inverter-based gain cell for ultra-low power RF applications. For the same bias current, the proposed method can boost the conventional inverter-based amplifier transconductance by twice, an overall quadruple boost compared to a single MOSFET in sub-threshold region. A 65nm CMOS based post-layout level implementation of the LNA including an on-chip matching network shows a 1.3 dB better noise figure as single inverter implementation for the same power of 45 μW. A cross-coupled stacked-gm VCO implementation achieves 44% higher voltage swing as that of the cross coupled single inverter at 74μA current, owing to the doubled gm value. Anjana Dissanayake, Steven M. Bowers, Benton H. Calhoun |
ISCAS | 3 |
| 2021 | Graph Coloring Using Coupled Oscillator-Based Dynamical SystemsabstractGraph coloring is a NP-hard problem, and computing the solution on a digital computer entails an exponential increase in the computing resources (time, memory) with increasing problem size. This has motivated the search for alternate and more efficient non-Boolean approaches. Here, we experimentally demonstrate the solution to this problem using the phase dynamics of coupled oscillators. Using a 30-oscillator IC platform with reconfigurable all-to-all coupling and minimal post-processing, our approach achieves 98% accuracy in detecting (near-) optimal solutions within 1 color of the optimal solution in comparison to the 77% accuracy achieved with the heuristic Johnson algorithm. Additionally, we propose a new local search-based post-processing scheme to improve the quality of the coloring solution. Finally, using circuit simulations, we demonstrate the scalability and speed up (~ 100×) achievable with the above approach in larger graphs. Antik Mallick, Mohammad Khairul Bashar, Daniel S. Truesdell, Benton H. Calhoun, Nikhil Shukla |
ISCAS | 4 |
| 2021 | Dynamic Read VMIN and Yield Estimation for Nanoscale SRAMsabstractThe design and verification process for SRAMs can be long and tedious due to the very large multi-dimensional design-space and the large computational time of Monte-Carlo (MC) simulations. In this work, we propose a fast analytical model, which takes into account the supply-voltage, temperature, process-variations, and array-design variables to characterize the critical read path and the small signal differential sensing and then evaluates the read-access failure probability and the corresponding VMINand yield. With a low evaluation time of 15 seconds and <; 6% error, the model is used to evaluate ~ 160K different SRAM designs in 20 hours. The results of the dataset are used to analyze the effect of key design-variables on yield and performance, determine inter-variable correlation, and calculate feature importance. In particular, important statistical results about sense-amplifier-enable timing and dynamic behavior of frequency correlation are presented in this work. Thus, the method can be very useful for SRAM designers to quickly calculate design feasibility and analyze the design space to optimize power, area, and speed. Shourya Gupta, Benton H. Calhoun |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Dynamic Write VMIN and Yield Estimation for Nanoscale SRAMsabstractThe dynamic write-access operation in an SRAM has a long-tailed distribution that sets the threshold for write-access failures. The distribution takes a long time to evaluate since rare failure events lie in its long and heavily skewed tail. Moreover, advanced FinFET technologies are becoming increasingly reliant on assist techniques to resolve these rare failures in the tail for improved dynamic performance and stability. In this work, we present various analytical approaches that work well in both super-threshold and subthreshold regions of operation to quickly determine the write-access failure probability. For FinFET based SRAMs, we present a modified sensitivity analysis-based method to evaluate the write-access operation distribution and discuss the evaluation of contention-limited write-access failures. The impact of various write-assist techniques on the performance and stability of FinFET SRAMs is also discussed. All simulations are performed using commercial 65nm bulk planar and 12nm FinFET technologies. Shourya Gupta, Benton H. Calhoun |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2020 | An 85 nW IoT Node-Controlling SoC for MELs Power-Mode Management and Phantom Energy ReductionabstractThis paper presents an ultra-low power (ULP) node-controlling system-on-chip (SoC) used for power-mode management and phantom energy reduction of miscellaneous electric loads (MELs). The SoC is powered from a single 2.5 V voltage supply enabled by the integrated power management unit (PMU) and can control up to 16 MELs due to the on-chip 16-channel correlator and the 32b RISC-V microprocessor. To further reduce the system power consumption, two clock domains have been adopted for the correlator and the processor separately. Fabricated in 65-nm CMOS, the measured minimum power consumption of the proposed SoC is only 85 nW at 0.45 V voltage supply and 1 kHz clock frequency. The measured maximum operating frequency can go up to 148 kHz with a 0.55 V supply. An application experiment successfully demonstrates that the SoC controls the power modes of MELs from wake-up to cut-off to save the average power and phantom energy. Shuo Li 0008, Jacob Breiholz, Sumanth Kamineni, Jaeho Im, David D. Wentzloff, Benton H. Calhoun |
ISCAS | 6 |
| 2020 | Application-Driven Model of a PPG Sensing Modality for the Informed Design of Self-Powered, Wearable Healthcare SystemsabstractThe convergence of self-powered technology with on-body wearable applications creates impactful opportunities for more personalized healthcare. PPG sensing is recognized as a primary method for recovering physiological information but remains relatively high-power compared to the available energy harvesting options in an on-body self-powered context, limiting reliability. This paper introduces a PPG sensing model based on a differential regulated cascode TIA to demonstrate the power, signaling, and circuit tradeoffs that exist for self-powered PPG operation. The model shows that body-worn μW to sub-μW self-powered PPG operation is theoretically achievable and provides insight on challenges and limitations. Henry L. Bishop, Peng Wang 0166, Benton H. Calhoun |
ISCAS | 3 |
| 2020 | An 88.6nW ozone pollutant sensing interface IC with a 159 dB dynamic rangeabstractThis paper presents a low power resistive sensor interface IC designed at 0.6V for ozone pollutant sensing. The large resistance range of gas sensors poses challenges in designing a low power sensor interface. Exiting architectures are insufficient for achieving a high dynamic range while enabling low VDD operation, resulting in high power consumption regardless of the adopted architecture. We present an adaptive architecture that provides baseline resistance cancellation and dynamic current control to enable low VDD operation while maintaining a dynamic range of 159dB across 20kΩ-1MΩ. The sensor interface IC is fabricated in a 65nm bulk CMOS process and consumes 88.6nW of power which is 300x lower than the state-of-art. The full system power ranges between 116 nW - 1.09 μW which includes the proposed sensor interface IC, analog to digital converter and peripheral circuits. The sensor interface's performance was verified using custom resistive metal-oxide sensors for ozone concentrations from 50 ppb to 900 ppb. Rishika Agarwala, Peng Wang 0166, Akhilesh Tanneeru, Bongmook Lee, Veena Misra, Benton H. Calhoun |
ISLPED | 6 |
| 2020 | An Open-source Framework for Autonomous SoC Design with Analog Block GenerationabstractWe present the world's first autonomous mixed-signal SoC framework, driven entirely by user constraints, along with a suite of automated generators for analog blocks. The process-agnostic framework takes high-level user intent as inputs to generate optimized and fully verified analog blocks using a cell-based design methodology. Our approach is highly scalable and silicon-proven by an SoC prototype which includes 2 PLLs, 3 LDOs, 1 SRAM, and 2 temperature sensors fully integrated with a processor in a 65nm CMOS process. The physical design of all blocks, including analog, is achieved using optimized synthesis and APR flows in commercially available tools. The framework is portable across different processes and requires no-human-in-the-Ioop, dramatically accelerating design time. Tutu Ajayi, Sumanth Kamineni, Yaswanth K. Cherivirala, Morteza Fayazi, Kyumin Kwon, Mehdi Saligane, Shourya Gupta, Chien-Hen Chen, Dennis Sylvester, David T. Blaauw, Ronald G. Dreslinski, Benton H. Calhoun, David D. Wentzloff |
VLSI-SOC | 12 |
| 2018 | FGC: A Tool-flow for Generating and Configuring Custom FPGAs(Abstract Only)abstractWe introduce the FGC Toolflow, the only tool providing flexible custom-FPGA generation and configuration to-date. Currently, researchers building custom FPGAs must create for FPGA schematics and bitstreams by hand. Both tasks are prohibitively time intensive and error prone. Additionally, the simulation time for bitcell configuration is very long (often times longer than the functionality), making the verification of FPGA fabrics even more time consuming. Some existing toolflows and software packages designed to help with this process, but they only generate bitcell configurations, leaving schematics to be developed by hand. Others have limitations in circuit-level and architectural parameters, which prevent them from adequately exploring the FPGA design space. The FGC flow is the only flow available that generates a custom full-FPGA schematic from a single parameter text file, and generates the proper configuration bitstream for a target Verilog functionality. The parameter text file can accommodate 100s of different parameters, which include both circuit-level and architectural parameters to fully encompass the FPGA design space. The FGC flow generates both a schematic and a configuration bitstream for an FPGA with 100 CLBs (900,000 transistors) in only 8 minutes. The flow also generates simulation files, allowing the user to quickly set up and perform simulations to verify the FPGA and its configuration at the chip level with SPICE-level accuracy. This flow was used to create, verify, and test a taped-out ultra-low power FPGA. Oluseyi A. Ayorinde, He Qi, Benton H. Calhoun |
FPGA | 3 |
| 2018 | An Ultra-low Power System On Chip Enabling DVS with SR Level Shifting LatchesabstractThis paper presents an ultra-low power flexible on-chip bus with SR level shifting latches to enable fine grained dynamic voltage scaling (DVS) in self-powered systems. This approach allows Ultra-low Power (ULP) systems-on-chip (SoC) to operate in three modes depending on the available energy and harvesting conditions. The proposed SR latch overcomes the area and power limitations of conventional level shifters to allow each block in the system to operate at either its minimum energy or minimum power point. The proposed latch can also level shift up or down allowing users to tune the voltage of each block post fabrication. A test SoC with the proposed SR latches was fabricated in commercial 130nm technology and measured results are reported. The proposed approach allows an average of 45.2% reduction in active power compared to a reference state-of-the-art ULP SoC when operated at 32KHz. An additional 87% savings can be achieved when operating each sub-block at its minimum voltage. Christopher J. Lukas, Farah B. Yahya, Benton H. Calhoun |
ISCAS | 3 |
| 2018 | Multiple Combined Write-Read Peripheral Assists in 6T FinFET SRAMs for Low-VMIN IoT and Cognitive ApplicationsabstractBattery-operated or energy-harvested IoT and cognitive SoCs in modern FinFET processes prefer the use of low-VMIN SRAMs for ultra-low power (ULP) operations. However, the 1:1:1 high-density (HD) FinFET 6T bitcell faces challenges in achieving a lower VMIN across process variation. The 6T bitcell VMIN improves either by increasing the size of the bitcell or by using combinations of peripheral assists (PAs) since a single PA cannot achieve the best VMIN across process variation. State-of-the-art works show some combinations of write and read PAs that lower the VMIN of 6T FinFET SRAMs. However, the better combinations of PA for 14nm HD 6T FinFET SRAMs are unknown. This work compares all the possible dual combinations of PAs and reveals the better ones. We show that in a usual column mux scenario the combination of negative bitline with VDD boosting and VDD collapse with VDD boosting in a proportion of 14% and 6% (total 20%), respectively, maximize the static VMIN improvement close to 191mV for ULP IoT and cognitive applications. We also show that a combination of wordline boosting with negative bitline and wordline boosting with VSS lowering achieve a 150mV and 25mV of dynamic VMIN improvement at the 5GHz frequency for the worst-case write and read corners, respectively, beating other combinations. Arijit Banerjee 0002, Sumanth Kamineni, Benton H. Calhoun |
ISLPED | 3 |
| 2017 | Soft errors: Reliability challenges in energy-constrained ULP body sensor networks applicationsabstractAggressive technology and supply voltage scaling has led to increasing concern for reliability. Optimizing power and energy with sub-threshold (sub-VT) operation exponentially increases the occurrences of both static and dynamic failures. With smaller node capacitances with each technology and supply scaling node, radiation-induced Single Event Upset (SEU) has become a critical design metric for Ultra-Low-Power (ULP) applications. In this paper, we explore the impact of radiation-induced soft errors on sub-threshold SRAM implemented in a Body Sensory Node (BSN) as an ULP application. We also demonstrate an exponential reduction in the critical charge (Qcrit) of a storage node with supply in near- and sub-VTdesign, resulting in a significant design consideration for the low-power applications. The huge process variation in sub-VTresults in 3X Qcritvariation. Finally, we compare the trend of technology scaling and supply voltage scaling on Qcrit. Harsh N. Patel, Benton H. Calhoun, Randy W. Mann |
IOLTS | 2 |
| 2017 | Modeling trans-threshold correlations for reducing functional test time in ultra-low power systemsabstractThis paper presents a methodology for reducing functional test time in subthreshold SoCs targeting ultra-low power (ULP) internet-of-things (IoT) devices. Due to their low operating speed and voltage, subthreshold SoCs require significantly longer time to test than traditional SoCs. The proposed method models trans-threshold correlations to allow high voltage, high speed testing while accurately predicting delay and power at the low, subthreshold operational voltage. This approach is orthogonal to other traditional testing methodologies and can significantly reduce the test time of digital and memory blocks on subthreshold SoCs. Using this process results in 5.4 × savings in test time for sequential and combinational test circuits, and over 2 × savings in test time for memory circuits with no overhead to area or prefabrication design time. Christopher J. Lukas, Farah B. Yahya, Benton H. Calhoun |
ITC | 3 |
| 2016 | An energy-efficient near/sub-threshold FPGA interconnect architecture using dynamic voltage scaling and power-gatingabstractThe rapid development of the Internet-of-Things requires hardware that is both low-energy and flexible, and a near/sub-threshold FPGA is a very promising solution. In the design of near/sub-threshold FPGAs, the biggest challenge is reducing global interconnect energy, which is the most energy-consuming part in the entire FPGA. Dynamic voltage scaling is an effective technique in reducing energy, but it is not widely used in FPGA interconnects because of the high area overhead of separately provisioning the buffers in the switch boxes to support different voltages on different paths. A low-swing interconnect, which removes buffers, allows this technique to be applied to the FPGA interconnects. In this paper, we propose a novel low-swing FPGA interconnect architecture that integrates dynamic voltage scaling and power-gating techniques with custom tool support. While the power-gating technique is widely used in existing designs for reducing leakage energy of idle drivers and buffers, we also apply power-gating to configuration bitcells in switch boxes, because it is a dominant energy consumer in near/sub-threshold. Including the energy overhead of voltage regulators, our work achieves a 10.1% energy saving in active circuits, 27.0% - 91.3% in idle circuits, and 19.0% - 53.1% in the entire FPGA on average, compared to an already optimized base case that only uses low-swing interconnect but no dynamic voltage scaling or power-gating. In addition, our dynamic voltage scaling allows us to adjust the delay of the low-swing FPGA interconnect from 0.14μs to 0.43μs or adjust its energy per operation from 5.5pJ to 35.7pJ when implementing the MCNC benchmarks at 0.6V. He Qi, Oluseyi A. Ayorinde, Benton H. Calhoun |
FPT | 3 |
| 2016 | Exploring circuit robustness to power supply variation in low-voltage latch and register-based digital systemsabstractThis paper compares the impact of power supply variation on the performance of register-based and latch-based digital circuits. A 32-tap, 16-bit FIR filter is fabricated using both flip-flops and latches in a 130nm CMOS process. Measurements show 25-37% improvement in energy-efficiency for the latch-based implementation operating below 0.6V subject to 44-120mV, 1 kHz peak-peak VDD ripple. This paper also presents a low-power, low-frequency droop measurement technique using digital circuits, which consumes 0.9pW at 0.75V. Abhishek Roy 0002, Benton H. Calhoun |
ISCAS | 2 |
| 2015 | Self-powered wearable sensor platforms for wellnessabstractHealth care continues to be one of the biggest challenges facing our society. Factors such as lifestyle choices, genetics, aging, stress and environmental exposures play a critical role in determining health outcomes. Wearable technologies that can enable continuous/long-term personal health monitoring and personal environmental monitoring can empower users to make better lifestyle decisions and improve health outcomes. While wearable devices promise a compelling future of achieving wellness, current wearable products are not addressing the needs of the health space. To achieve this future, key challenges in wearable systems such as battery life, form factor, sensor functionality, configurability and data analysis will have to be carefully addressed to ensure user adoption and effectively manage health. Veena Misra, Benton H. Calhoun, Shekhar Bhansali, John C. Lach, Suman Datta, Mehmet Ozturk, Alper Bozkurt, Ömer Oralkan, Jason Strohmaier |
CASES | 2 |
| 2015 | Using island-style bi-directional intra-CLB routing in low-power FPGAsabstractIncreased clustering in Field Programmable Gate Arrays (FPGAs) has shifted a larger fraction of the overall routing load into the configurable logic blocks (CLBs), reducing usage of the costly global interconnect. However, increases in CLB size introduce additional overheads inside CLBs, which can limit the savings gained by minimizing the global interconnect use, motivating more efficient intra-CLB routing. This paper explores different topologies for the intra-CLB connectivity and identifies how the optimal local-CLB interconnect changes for different FPGA architecture and circuit parameters. This work compares area, delay, and energy for two intra-CLB topologies: multiplexer-based routing and island-style bi-directional routing, similar to the global FPGA interconnect, but used inside the CLB (which we call a mini-FPGA). The mini-FPGA style of local CLB interconnect prove to be favorable for minimum-energy operation, as they can reduce transistor count by as much as 62%, and consume as much as 77.9% less energy. Multipexer-based CLBs have performance benefits by reducing delays by almost 3×. Multiplexer-based CLBs can consume less energy at nominal voltages, but only if additional measures are taken to limit power consumption in the multiplexers. A 130-nm CMOS test chip confirms that simulation results track measured data for mini-FPGA CLBs. Oluseyi A. Ayorinde, He Qi, Yu Huang 0015, Benton H. Calhoun |
FPL | 4 |
| 2015 | Optimizing energy efficient low-swing interconnect for sub-threshold FPGAsabstractFPGA interconnect traditionally dominates energy and delay, and designs such as low-swing interconnect have been proven to reduce the interconnect burden for low energy FPGAs. This paper presents an optimized low-swing interconnect for FPGAs operating in the sub-threshold region. We also address signal degradation along lengthy interconnect paths and examine strategies for inserting low-switching-threshold repeaters. A 130nm test chip implementing low-swing interconnect meshes with different circuit parameters is measured. The results show that optimization of the low-swing interconnect provides up to 60.2% lower energy-delay-product (EDP) than a straightforward, un-optimized low-swing design at VDD= 0.4V. Furthermore, the simulation results show that the optimized low-swing interconnect is 97.7% faster and 42.7% lower energy than a traditional uni-directional interconnect at VDD= 0.4V. He Qi, Oluseyi A. Ayorinde, Yu Huang 0015, Benton H. Calhoun |
FPL | 4 |
| 2015 | Ultra-low power wireless SoCs enabling a batteryless IoT
Benton H. Calhoun, David D. Wentzloff |
Hot Chips Symposium | 1 |
| 2015 | Error-energy analysis of hardware logarithmic approximation methods for low power applicationsabstractThis paper presents an overview of methods for combinational, base-two logarithmic approximation using Mitchell's algorithm, piecewise-linear/quadratic error compensation schemes, and direct approximation. Optimization methods are used for computing linear segments for each compensation scheme including Hamming weight minimization and pattern recognition of segment slopes for multiplier-less error compensation. A novel, near-zero-average error quadratic compensation scheme is also presented. A test chip was fabricated in a commercial 130nm technology including fifteen base-two logarithm approximations and each was evaluated for its standalone accuracy, measured energy, delay, and area in order to determine the best classes of approximation for low-energy and high-accuracy operation. A system-level evaluation of these methods was performed using a software model of a keyword detection speech pipeline to predict the impact of inaccurate logarithmic approximation on the detection accuracy of keywords across various noise levels. Alicia Klinefelter, Joseph F. Ryan 0002, James W. Tschanz, Benton H. Calhoun |
ISCAS | 4 |
| 2015 | A 0.38 pj/bit 1.24 nW chip-to-chip serial link for ultra-low power systemsabstractAs energy-constrained systems continue to reduce their power consumption, Unding an optimal point of operation for the principle components in the energy budget becomes increasingly important. With energy dominant system components like communication circuits, it is important to consider both energy-per-bit and power in the context of the system's use cases. In this paper, we propose optimization of chip-to-chip links considering both energy-per-cycle and energy-per-bit to find the optimal operating voltage and activity factor while minimizing wasted energy and power. A fabricated 130 nm chip was used to verify this finding and resulted in an energy-per-bit of 0.38 pj/bit and power of 1.24 nW. Christopher J. Lukas, Benton H. Calhoun |
ISCAS | 2 |
| 2015 | Flexible Technologies for Self-Powered Wearable Health and Environmental SensingabstractThis article provides the latest advances from the NSF Advanced Self-powered Systems of Integrated sensors and Technologies (ASSIST) center. The work in the center addresses the key challenges in wearable health and environmental systems by exploring technologies that enable ultra-long battery lifetime, user comfort and wearability, robust medically validated sensor data with value added from multimodal sensing, and access to open architecture data streams. The vison of the ASSIST center is to use nanotechnology to build miniature, self-powered, wearable, and wireless sensing devices that can enable monitoring of personal health and personal environmental exposure and enable correlation of multimodal sensors. These devices can empower patients and doctors to transition from managing illness to managing wellness and create a paradigm shift in improving healthcare outcomes. This article presents the latest advances in high-efficiency nanostructured energy harvesters and storage capacitors, new sensing modalities that consume less power, low power computation, and communication strategies, and novel flexible materials that provide form, function, and comfort. These technologies span a spatial scale ranging from underlying materials at the nanoscale to body worn structures, and the challenge is to integrate them into a unified device designed to revolutionize wearable health applications. Veena Misra, Alper Bozkurt, Benton H. Calhoun, Thomas N. Jackson, Jesse Jur, John C. Lach, Bongmook Lee, John Muth, Ömer Oralkan, Mehmet Ozturk, Susan Trolier-McKinstry, Daryoosh Vashaee, David D. Wentzloff, Yong Zhu 0003 |
Proc. IEEE | 3 |
| 2015 | Virtual Prototyper (ViPro): An SRAM Design Tool for Yield Constrained OptimizationabstractThis brief presents a tool for optimizing the energy and delay (E/D) of static RAM designs to meet a specific die yield constraint. This allows the tool to account for the effects of process variation and to trade off yield with performance and energy. To accomplish this, we use a combination of simulation and modeling techniques to determine the minimum wordline (WL) pulsewidth required for both the read and write operations to meet a user-specified die yield. The use of a hierarchical model enables us to calculate the E/D of a full macro that is margined to meet a specific die yield. By sweeping across the possible design space, we are able to identify Pareto optimal designs. The tool structure described in this brief allows comparison across different array topologies, process technologies, and circuit choices including assist methods. Using this tool, we find that adding a WL boosting scheme results in an overall energy savings, despite the overhead of using a charge pump circuit, due to an improved read delay distribution. James Boley, Peter Beshay, Benton H. Calhoun |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Flexibility and Circuit Overheads in Reconfigurable SIMD/MIMD SystemsabstractDynamically reconfigurable SIMD/MIMD architectures made from simple cores have emerged to exploit diverse forms of parallelism in applications [1,2]. In this work, we investigate the circuit-level overhead and flexibility tradeoffs of such architectures through the design of a custom reconfigurable SIMD/MIMD system. Saad Arrabi, Kevin Skadron, Benton H. Calhoun, John C. Lach, Brett H. Meyer |
FCCM | 5 |
| 2014 | A digital dynamic write margin sensor for low power read/write operations in 28nm SRAMabstractThe conventional guard band design approach increases the SRAM Wordline (WL) pulse duration to operate successfully in all the process, voltage and temperature (PVT) corners. This can significantly increase the dynamic energy. This work presents a digital circuit that is able to track and control the WL pulse duration of the SRAM memory across PVT variations, to minimize the dynamic energy while maintaining robust operations. The circuit is applied on a 78kbit SRAM. The results are compared to the worst case margin approach and show a maximum write energy savings of 45% and 49% relative to margining voltage/temperature (VT) and process variations, respectively. Peter Beshay, Vikas Chandra, Robert C. Aitken, Benton H. Calhoun |
ISLPED | 4 |
| 2013 | Flexible on-chip power delivery for energy efficient heterogeneous systemsabstractHeterogeneous systems-on-chip pose a challenge for power delivery given the variety of needs for different components. In this paper, we describe recent work that leverages power switches and conventional EDA toolflows to implement a set of power delivery schemes that provide a flexible, adaptable range of options for power management of SoCs for which energy efficiency is important. We first present an enhanced dynamic voltage scaling (DVS) scheme that uses power switches to provide rapid changes in the energy-speed operating point to match workloads at a component level. To demonstrate this approach, we describe a data flow processor chip in 90nm CMOS that supports flexible operation from 0.25V with super high energy efficiency up to GHz speeds at 1.2V. This chip shows that our low overhead method to scale energy consumption with the performance requirement supports both high performance and ultra low energy (>10X reduction in energy per operation) in the same circuit. We discuss power switch design for this scheme and investigate strategies for optimizing power switches for different operating modes. Finally, we show how segmented power switches offer several advantages for flexibly managing leakage and for modulating local voltages with low overhead. Benton H. Calhoun, Kyle Craig |
DAC | 1 |
| 2013 | Leveraging sensitivity analysis for fast, accurate estimation of SRAM dynamic write VMINabstractCircuit reliability in the presence of variability is a major concern for SRAM designers. With the size of memory ever increasing, Monte Carlo simulations have become too time consuming for margining and yield evaluation. In addition, dynamic write-ability metrics have an advantage over static metrics because they take into account timing constraints. However, these metrics are much more expensive in terms of runtime. Statistical blockade is one method that reduces the number of simulations by filtering out non-tail samples, however the total number of simulations required still remains relatively large. In this paper, we present a method that uses sensitivity analysis to provide a total speedup of ∼112X compared with recursive statistical blockade with only a 3% average loss in accuracy. In addition, we show how this method can be used to calculate dynamic VMIN and to evaluate several write assist methods. James Boley, Vikas Chandra, Robert C. Aitken, Benton H. Calhoun |
DATE | 4 |
| 2013 | Circuit optimizations to minimize energy in the global interconnect of a low-power-FPGA (abstract only)abstractWe compare circuit and architecture choices in the global interconnect of an FPGA in order to find the minimum energy design for low voltage operation. We look at switch box topology, number of repeaters, receiver circuit topology, and dynamic voltage selection, all with the intent of minimizing energy consumption. The results show that using a pass gate switchbox topology with repeaters in the interconnect and a custom receiver lowers delay by up to 63% and energy by up to 87% from the standard FPGA circuit choices. This work also identifies the optimal VDD choices to maximize performance under energy constraints or vice versa. Oluseyi A. Ayorinde, Benton H. Calhoun |
FPGA | 2 |
| 2012 | Optimal power switch design for dynamic voltage scaling from high performance to subthreshold operationabstractThis work explores optimizing power switch design for Dynamic Voltage Scaling schemes that use headers to connect components to voltage supplies ranging from strong inversion to subthreshold values. We propose using NMOS devices with their gate controlled at the nominal voltage as power switches connected to the subthreshold voltage rail. Measured results show that an NMOS can provide the subthreshold voltage with a power switch size >280X smaller than a PMOS. For architectures targeting operation from subthreshold up to nominal voltage, we show that using an asymmetric transmission gate power switch provides a lower overhead way to enable this flexibility. Kyle Craig, Yousef Shakhsheer, Benton H. Calhoun |
ISLPED | 3 |
| 2012 | A programmable resistive power grid for post-fabrication flexibility and energy tradeoffsabstractThis paper explores the benefits of splitting a monolithic power gate transistor into parallel, independently controlled, variable weighted power gates to provide programmable post-fabrication power grid resistance. This power gate topology creates energy saving opportunities by providing adjustable localized voltages during active modes and reducing leakage current in idle blocks while retaining data. Measurements show over 30% active energy savings per operation and 90% savings in idle current with retention. A modeling flow for a resistive power grid was also developed that demonstrates the effectiveness of this approach in a Bulldozer processor core. Kyle Craig, Yousef Shakhsheer, Sudhanshu Khanna, Saad Arrabi, John C. Lach, Benton H. Calhoun, Stephen V. Kosonocky |
ISLPED | 6 |
| 2012 | A charge pump based receiver circuit for voltage scaled interconnectabstractThis paper presents a charge-pump based low swing interconnect receiver circuit. The interconnect circuit is single ended and supports swings of 300mV or lower. A charge pump front end at the receiver boosts the arriving signal before restoring it to the full logic level, improving the performance of the interconnect. For a 10mm long interconnect wire in a 45nm CMOS process, the proposed scheme provides 3X energy reduction at constant speed and 3.5X delay improvement at constant energy relative to prior art. We deploy the interconnect scheme as the data bus between the L1-L2 caches of a 4-core Alpha processor. Over a set of Splash benchmarks, the proposed architecture reduces total energy consumption by 70% while maintaining the same performance. Aatmesh Shrivastava, John C. Lach, Benton H. Calhoun |
ISLPED | 3 |
| 2012 | Body Sensor Networks: A Holistic Approach From Silicon to UsersabstractBody sensor networks (BSNs) are emerging cyber–physical systems that promise to improve quality of life through improved healthcare, augmented sensing and actuation for the disabled, independent living for the elderly, and reduced healthcare costs. However, the physical nature of BSNs introduces new challenges. The human body is a highly dynamic physical environment that creates constantly changing demands on sensing, actuation, and quality of service (QoS). Movement between indoor and outdoor environments and physical movements constantly change the wireless channel characteristics. These dynamic application contexts can also have a dramatic impact on data and resource prioritization. Thus, BSNs must simultaneously deal with rapid changes to both top–down application requirements and bottom–up resource availability. This is made all the more challenging by the wearable nature of BSN devices, which necessitates a vanishingly small size and, therefore, extremely limited hardware resources and power budget. Current research is being performed to develop new principles and techniques for adaptive operation in highly dynamic physical environments, using miniaturized, energy-constrained devices. This paper describes a holistic cross-layer approach that addresses all aspects of the system, from low-level hardware design to higher level communication and data fusion algorithms, to top-level applications. Benton H. Calhoun, John C. Lach, John A. Stankovic, David D. Wentzloff, Kamin Whitehouse, Adam T. Barth, Jonathan K. Brown, Qiang Li 0025, Nathan E. Roberts, Yanqing Zhang 0002 |
Proc. IEEE | 1 |
| 2012 | Nonrandom Device Mismatch Considerations in Nanoscale SRAMabstractCompetitive density, performance, and functional objectives of the SRAM bit cell require design rules which are much more aggressive than those used in base logic designs. Because soft fail yield in SRAM is dependent on the device threshold and threshold mismatch in the bit cell, much research has been directed toward addressing the random contributors to within-cell device threshold variation. We examine four sources of potential nonrandom threshold mismatch that can arise from the use of aggressive design rules in the bit cell: 1) implanted ion straggle in SiO2; 2) polysilicon inter-diffusion driven counter-doping; 3) lateral ion straggle from the photoresist; and 4) photoresist implant shadowing. Using simulation and hardware measurements, we quantify the device parametric impacts and provide a statistical treatment forming the basis for quantification of the functional margin impacts on the bit cell. We examine two lithography-compliant bit-cell layout topologies and quantify the impact of systematic mismatch on the margin limited yield. Randy W. Mann, Terry B. Hook, Phung T. Nguyen, Benton H. Calhoun |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | Tracking On-Chip Age Using Distributed, Embedded SensorsabstractRecent works show bias temperature instability (BTI) is a detrimental hard-aging mechanism in CMOS circuit design. Negative BTI (NBTI) alone degrades circuit speed upwards of 20% over a 10 year life-span. Having the ability to track the actual aging process provides one method to reduce large design margins that are otherwise required to offset circuit aging. This work extends previous research by contributing a sensing scheme that employs on-chip sensors capable of accurately tracking NBTI pMOS current degradations across process, temperature, and varying activity factors. Results show that a 7600$\mu{\hbox {m}}^{2}$sensing area achieves an overall system accuracy of 90% at a voltage threshold precision of 2 mV. We thoroughly describe the sensor design and the underlying statistics used to determine overall accuracy and precision. Furthermore, a novel sensor distribution method is presented that uses an existing scan-chain methodology to mask the overhead of adding the on-chip sensors. Stuart N. Wooters, Adam C. Cabe, Zhenyu Qi 0001, Jiajing Wang, Randy W. Mann, Benton H. Calhoun, Mircea R. Stan, Travis N. Blalock |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2011 | Cost-effective safety and fault localization using distributed temporal redundancyabstractCost pressure is driving vendors of safety-critical systems to integrate previously distributed systems. One natural approach we have previous introduced is On-Demand Redundancy (ODR), which allows safety-critical and non-critical tasks, traditionally isolated to limit interference, to execute on shared resources. Our prior work has shown that relaxed dedication (RD), one ODR strategy which allows non-critical tasks (NCTs) to execute on idle critical task resources (CTRs), significantly increases NCT throughput. Unfortunately, there are circumstances under which, in spite of this opportunity, it is difficult to effectively schedule NCTs. Brett H. Meyer, Benton H. Calhoun, John C. Lach, Kevin Skadron |
CASES | 2 |
| 2011 | Reducing the cost of redundant execution in safety-critical systems using relaxed dedicationabstractWe introduce on-demand redundancy, a set of architectural techniques that leverage the tightly-coupled nature of components in systems-on-chip to reduce the cost of safety-critical systems. On-demand redundancy eases the assumptions that traditionally segregate the execution of critical and non-critical tasks (NCTs), making resources available for critical tasks at potentially arbitrary points in both space and time, and otherwise freeing resources to execute non-critical tasks when critical tasks are not executing. Relaxed dedication is one such technique that allows non-critical tasks to execute on critical task resources. Our results demonstrate that for a wide variety of applications and architectures, relaxed dedication is more cost-effective than a traditional approach that employs dedicated resources executing in lockstep. Applied to dual-modular redundancy (DMR), relaxed dedication exposes 73% more NCT cycles than traditional DMR on average, across a wide variety of usage scenarios. Brett H. Meyer, Nishant George, Benton H. Calhoun, John C. Lach, Kevin Skadron |
DATE | 3 |
| 2011 | Dynamic write limited minimum operating voltage for nanoscale SRAMsabstractDynamic stability analysis for SRAM has been growing in importance with technology scaling. This paper analyzes dynamic writability for designing low voltage SRAM in nanoscale technologies. We propose a definition for dynamic write limited VMIN. To the best of our knowledge, this is the first definition of a VMINbased on dynamic stability. We show how this VMINis affected by the array capacity, the voltage scaling of the word-line pulse, the bitcell parasitics, and the number of cycles prior to the first read access. We observe that the array can be either dynamically or statically write limited depending on the aforementioned factors. Finally, we look at how voltage-bias based write assist techniques affect the dynamic write limited VMIN. Satyanand Nalam, Vikas Chandra, Robert C. Aitken, Benton H. Calhoun |
DATE | 4 |
| 2011 | An analytical model for performance yield of nanoscale SRAM accounting for the sense amplifier strobe signal
Joseph F. Ryan 0002, Sudhanshu Khanna, Benton H. Calhoun |
ISLPED | 3 |
| 2011 | Minimum Supply Voltage and Yield Estimation for Large SRAMs Under Parametric VariationsabstractSRAM cell minimum operation voltage (Vmin) exhibits a skewed distribution in the presence of random parametric variations. Standard Monte Carlo (MC) simulation is prohibitively expensive to estimate the tail of the Vmin distribution for large SRAMs. We propose a fast and accurate method to estimate Vmin based on the statistical trend of static noise margin withVDDscaling. Our preliminary work has shown its efficiency for standby Vmin estimation. In this work, we extend the method to estimate read and write Vmin and yield. We also generalize it for both symmetric and asymmetric types of cells. With comparable accuracy, the proposed model offers a huge speedup over standard MC. Compared with an alternative fast MC method, importance sampling, it shows a good agreement with less complexity. Jiajing Wang, Benton H. Calhoun |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | An Enhanced Canary-Based System With BIST for SRAM Standby Power ReductionabstractTo achieve aggressive standby power reduction for static random access memory (SRAM), we have previously proposed a closed-loop VDDscaling system with canary replicas that can track global variations. In this paper, we propose several techniques to enhance the efficiency of this system for more advanced technologies. Adding dummy cells around the canary cell improves the tracking of systematic variations. A new canary circuit avoids the possibility that a canary cell may never fail because it resets into its more stable data pattern. A built-in self-test (BIST) block incorporates self-calibration of SRAM minimum standbyVDDand the initial failure threshold due to intrinsic mismatch. Measurements from a new 45 nm test chip further demonstrate the function of the canary cells in smaller technology and show that adding dummy cells reduces the variation of the canary cell. Jiajing Wang, A. Hoefler, Benton H. Calhoun |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | Virtual prototyper (ViPro): an early design space exploration and optimization tool for SRAM designersabstractSRAM design in scaled technologies requires knowledge of phenomena at the process, circuit, and architecture level. Decisions made at various levels of the design hierarchy affect the global figures of merit (FoMs) of an SRAM, such as, performance, power, area, and yield. However, the lack of a quick mechanism to understand the impact of changes at various levels of the hierarchy on global FoMs makes an accurate assessment of SRAM design innovations difficult. Thus, we introduce Virtual Prototyper (ViPro), a tool that helps SRAM designers explore the large design space by rapidly generating optimized virtual prototypes of complete SRAM macros. It does so by allowing designers to describe the SRAM components with varying levels of detail and by incorporating them into a hierarchical model that captures circuit and architectural features of the SRAM to optimize a complete prototype. It generates base-case prototypes that provide starting points for design space exploration, and assesses the impact of process, circuit, and architectural changes on the overall SRAM macro design. Satyanand Nalam, Mudit Bhargava, Ken Mai, Benton H. Calhoun |
DAC | 4 |
| 2010 | SRAM-based NBTI/PBTI sensor system designabstractNBTI has been a major aging mechanism for advanced CMOS technology and PBTI is also looming as a big concern. This work first proposes a compact on-chip sensor design that tracks both NBTI and PBTI for both logic and SRAM circuits. Embedded in an SRAM array the sensor takes the form of a 6T SRAM cell and is at least 30x smaller than previous designs. Extensively reusing the SRAM peripheral circuitry minimizes control logic overhead. Sensing overhead is further amortized as the sensors can be both reconfigured and recycled as functional SRAM cells, potentially increasing SRAM yield when other bit cells fail due to initial process variation or long time aging effects. The paper also proposes a variation-aware sensor system design methodology by quantifying and leveraging the tradeoff between the size and number of sensors and the system sensing precision. Design examples show that a system of 500 sensors can achieve 4mV precision with 98.8% confidence, and a system of 1K sensors designed for 1M SRAM bit cells achieves 2000x area overhead reduction compared to a worst-case based approach. Zhenyu Qi 0001, Jiajing Wang, Adam C. Cabe, Stuart N. Wooters, Travis N. Blalock, Benton H. Calhoun, Mircea R. Stan |
DAC | 6 |
| 2010 | System design principles combining sub-threshold circuit and architectures with energy scavenging mechanismsabstractUltra low power (ULP) circuits and energy scavenging mechanisms, though conceptually appealing, have been mainly studied in isolation to date. In this paper, we observe energy harvesting prototypes to derive system-level models and to reveal practical issues specific to various types of energy harvesting systems. Our models predict how design decisions affect overall lifetime. We use our model to derive system driven principles for optimizing architecture, voltage selection, and sub-threshold circuit designs across different types of power harvesting systems. Benton H. Calhoun, Sudhanshu Khanna, Yanqing Zhang 0002, Joseph F. Ryan 0002, Brian P. Otis |
ISCAS | 1 |
| 2010 | Flexible Circuits and Architectures for Ultralow PowerabstractSubthreshold digital circuits minimize energy per operation and are thus ideal for ultralow-power (ULP) applications with low performance requirements. However, a large range of ULP applications continue to face performance constraints at certain times that exceed the capabilities of subthreshold operation. In this paper, we give two different examples to show that designing flexibility into ULP systems across the architecture and circuit levels can meet both the ULP requirements and the performance demands. Specifically, we first present a method that expands on ultradynamic voltage scaling (UDVS) to combine multiple supply voltages with component level power switches to provide more efficient operation at any energy-delay point and low overhead switching between points. This system supports operation across the space from maximum performance, when necessary, to minimum energy, when possible. It thus combines the benefits of single-VDD, multi-VDD, and dynamic voltage scaling (DVS) while improving on them all. Second, we propose that reconfigurable subthreshold circuits can increase applicability for ULP embedded systems. Since ULP devices conventionally require custom circuit design but the manufacturing volume for many ULP applications is low, a subthreshold field programmable gate array (FPGA) offers a cost-effective custom solution with hardware flexibility that makes it applicable across a wide range of applications. We describe the design of a subthreshold FPGA to support ULP operation and identify key challenges to this effort. Benton H. Calhoun, Joseph F. Ryan 0002, Sudhanshu Khanna, Mateja Putic, John C. Lach |
Proc. IEEE | 1 |
| 2010 | Two Fast Methods for Estimating the Minimum Standby Supply Voltage for Large SRAMsabstractThe data retention voltage (DRV) defines the minimum supply voltage for an SRAM cell to hold its state. Intra-die variation causes a statistical distribution of DRV for individual cells in a memory array. We present two fast and accurate methods to estimate the tail of the DRV distribution. The first method uses a new analytical model based on the relationship between DRV and static noise margin. The second method extends the statistical blockade technique to a recursive formulation. It uses conditional sampling for rapid statistical simulation and fits the results to a generalized Pareto distribution (GPD) model. Both the analytical DRV model and the generic GPD model show a good match with Monte Carlo simulation results and offer speedups of up to four or five orders of magnitude over Monte Carlo at the 6σ point. In addition, the two models show a very close agreement with each other at the tail up to 8σ. For error within 5% with a confidence of 95%, the analytical DRV model and the GPD model can predict DRV quantiles out to 8σ and 6.6σ respectively; and for the mean of the estimate, both models offer within 1% error relative to Monte Carlo at the 4σ point. Jiajing Wang, Amith Singhee, Rob A. Rutenbar, Benton H. Calhoun |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2009 | A Technology-Agnostic Simulation Environment (TASE) for iterative custom IC design across processesabstractA designer's intent and knowledge about the critical issues and trade-offs underlying a custom circuit design are implicit in the simulations she sets up for design creation and verification. However, this knowledge is tightly conjoined with technology-specific features and decoupled from the final schematic in traditional design flows. As a result, this knowledge is easily lost when the technology specifics change. This paper presents a technology agnostic simulation environment (TASE), which is a tool that uses simulation templates to capture the designer's knowledge and separate it from the technology-specific components of a simulation. TASE also allows the designer to form groups of related simulations and port them as a unit to a new technology. This allows an actual design schematic to remain tied to the analyses that illuminate the underlying trade-offs and design issues, unlike the case where schematics are ported alone. Giving the designer immediate access to the trade-offs, which are likely to change in new technologies, accelerates the re-design that often must accompany porting of complicated custom circuits. We demonstrate the usefulness of TASE by investigating Read and Write noise margins for a 6T SRAM in predictive technologies down to 16 nm. Satyanand Nalam, Mudit Bhargava, Kyle Ringgenberg, Ken Mai, Benton H. Calhoun |
ICCD | 5 |
| 2009 | Panoptic DVS: A fine-grained dynamic voltage scaling framework for energy scalable CMOS designabstractThe energy efficiency of a CMOS architecture processing dynamic workloads directly affects its ability to provide long battery lifetimes while maintaining required application performance. Existing scalable architecture design approaches are often limited in scope, focusing either only on circuit-level optimizations or architectural adaptations individually. In this paper, we propose a circuit/architecture co-design methodology called Panoptic Dynamic Voltage Scaling (PDVS) that makes more efficient use of common circuit structures and algorithm-level processing rate control. PDVS expands upon prior work by using multiple component-level PMOS header switches to enable fine-grained rate control, allowing efficient dithering among statically scheduled algorithms with sub-block energy savings. This way, PDVS is able to achieve a wide variety of processing rates to match incoming workload as closely as possible, while each iteration takes less energy to process than on architectures with coarser levels of rate control. Measurements taken from a fabricated 90 nm test chip characterize both savings and overheads and are used to inform PDVS synthesis decisions. Results show that PDVS consumes up to 34% and 44% less energy than Multi-VDD and Single-VDD systems, respectively. Mateja Putic, Liang Di, Benton H. Calhoun, John C. Lach |
ICCD | 3 |
| 2009 | Sub-threshold Operation and Cross-hierarchy Design for Ultra Low Power Wearable SensorsabstractThis paper examines the requirements of wearable sensing applications and their implications for designing the next generation of body area sensors. We define key metrics for wearable sensors and discuss how body area sensors differ from generic wireless sensors. To explore the system level issues with a wearable node, we show measurements from a wearable electrocardiogram (ECG) sensor prototype. Using heart rate monitoring as an example, we show how ultra low power (ULP) circuit design must be applied to support the stringent energy and/or power demands of long life wearable sensors. Specifically, sub-threshold operation of digital circuits creates opportunities for re-thinking the entire system. We conclude that we can only reach the lower limits of power consumption through cross-hierarchy design of the entire sensor node that leverages ULP digital circuits. Benton H. Calhoun, Jonathan F. Bolus, Sudhanshu Khanna, Andrew D. Jurik, Alfred C. Weaver, Travis N. Blalock |
ISCAS | 1 |
| 2009 | Sub-threshold Circuit Design with Shrinking CMOS DevicesabstractThis paper examines the impact of technology scaling to 22 nm on sub-threshold circuit design and proposes several solutions for sub-threshold circuits in new processes. To maintain energy-efficient sub-threshold operation, we must reduce variation and suppress leakage current. To combat random variation and minimize energy for nodes below 45 nm, we show that special strategies are needed for different categories of sub-threshold circuits. Benton H. Calhoun, Sudhanshu Khanna, Randy W. Mann, Jiajing Wang |
ISCAS | 1 |
| 2009 | A 2.6 µW sub-threshold mixed-signal ECG SoCabstractThis paper describes a 130 nm CMOS sub-threshold (sub-VT) mixed-signal system-on-chip (SoC) that acquires and processes an electrocardiogram (ECG) signal for wireless ECG monitoring. Low power ECG devices permit continuous monitoring for longer durations between power recharging. Co-designing the digital and analog blocks creates opportunities for reducing power at the system level. The SoC uses a sub-threshold digital microcontroller (μC) for adaptive control of the sub-VT biased analog components and for processing the ECG data. The μC operates from 0.24 V to 1.2 V and consumes as little as 1.51 pJ per instruction. The SoC consumes only 2.6 μW while providing raw ECG data or processed heart rate data, which can lower the wireless data rate by 500X. Steven C. Jocke, Jonathan F. Bolus, Stuart N. Wooters, Travis N. Blalock, Benton H. Calhoun |
ISLPED | 5 |
| 2009 | Serial sub-threshold circuits for ultra-low-power systemsabstractThis paper explores the use of serial circuits for ultra-low-power sub-threshold systems. A serial system leads to a smaller design and higher utilization, yielding 40% active energy, 15x active power, and 32x leakage power benefits. Further, we show that using a serial system in the sub-threshold regime decreases both active energy and leakage power even at the same speed as a parallel system. This is in sharp contrast to strong inversion, where larger bit widths give lower energy and power for the same delay. We identify the unique properties of sub-threshold operation that creates these differences. Sudhanshu Khanna, Benton H. Calhoun |
ISLPED | 2 |
| 2008 | Power switch characterization for fine-grained dynamic voltage scalingabstractDynamic voltage scaling (DVS) provides power savings for systems with varying performance requirements. One low overhead implementation of DVS uses PMOS power switches to connect DVS blocks to one of the available VDDsupplies. While power switches have been analyzed extensively for leakage power gating, proper design of power switches for DVS is less well understood. This paper characterizes power switches for DVS in terms of VDD-switching delay and VDD-switching energy. We show the impact of these switching overheads on a novel fine-grained DVS architecture and present an RC model that allows fast estimation of the overhead. Measurements of a DVS multiplier and adder on a 90 nm CMOS test chip validate the model. Our model and measurements confirm that power switched DVS can provide sufficiently low overhead to give energy savings with only one clock cycle spent at a lower voltage, making this approach a flexible and enticing option for embedded portable systems. Liang Di, Mateja Putic, John C. Lach, Benton H. Calhoun |
ICCD | 4 |
| 2008 | Analyzing static and dynamic write margin for nanometer SRAMsabstractThis paper analyzes write ability for SRAM cells in deeply scaled technologies, focusing on the relationship between static and dynamic write margin metrics. Reliability has become a major concern for SRAM designs in modern technologies. Both local mismatch and scaled VDD degrade read stability and write ability. Several static approaches, including traditional SNM, BL margin, and the N-curve method, can be used to measure static write margin. However, static approaches cannot indicate the impact of dynamic dependencies on cell stability. We propose to analyze dynamic write ability by considering the write operation as a noise event that we analyze using dynamic stability criteria. We also define dynamic write ability as the critical pulse width for a write. By using this dynamic criterion, we evaluate the existing static write margin metrics at normal and scaled supply voltages and assess their limitations. The dynamic write time metric can also be used to improve the accuracy of VCCmin estimation for active VDD scaling designs. Jiajing Wang, Satyanand Nalam, Benton H. Calhoun |
ISLPED | 3 |
| 2008 | Digital Circuit Design Challenges and Opportunities in the Era of Nanoscale CMOSabstractWell-designed circuits are one key ldquoinsulatingrdquo layer between the increasingly unruly behavior of scaled complementary metal-oxide-semiconductor devices and the systems we seek to construct from them. As we move forward into the nanoscale regime, circuit design is burdened to ldquohiderdquo more of the problems intrinsic to deeply scaled devices. How this is being accomplished is the subject of this paper. We discuss new techniques for logic circuits and interconnect, for memory, and for clock and power distribution. We survey work to build accurate simulation models for nanoscale devices. We discuss the unique problems posed by nanoscale lithography and the role of geometrically regular circuits as one promising solution. Finally, we look at recent computer-aided design efforts in modeling, analysis, and optimization for nanoscale designs with ever increasing amounts of statistical variation. Benton H. Calhoun, Yu Cao 0001, Xin Li 0001, Ken Mai, Lawrence T. Pileggi, Rob A. Rutenbar, Kenneth L. Shepard |
Proc. IEEE | 1 |
| 2007 | Analyzing and modeling process balance for sub-threshold circuit designabstractThis paper describes the strong effects on sub-threshold digital circuit operation of the ratio of PMOS and NMOS current in a given process. We define the concept of process balance/imbalance as describing this ratio and explain the impact ofdifferent circuit and environmental parameters on processbalance. Many of these characteristics are best understood by the degree to which they increase or further decrease process balance. We also propose a model that provides accurate estimation of the effects of process balance that is useful for understanding the impact of process variations and the appropriate types of circuits to use for sub-threshold operation in a given process. Joseph F. Ryan 0002, Jiajing Wang, Benton H. Calhoun |
ACM Great Lakes Symposium on VLSI | 3 |
| 2006 | Sub-threshold design: the challenges of minimizing circuit energyabstractIn this paper, we identify the key challenges that oppose sub-threshold circuit design and describe fabricated chips that verify techniques for overcoming the challenges. Benton H. Calhoun, Alice Wang 0002, Naveen Verma, Anantha P. Chandrakasan |
ISLPED | 1 |
| 2005 | Design Considerations for Ultra-Low Energy Wireless Microsensor NodesabstractThis tutorial paper examines architectural and circuit design techniques for a microsensor node operating at power levels low enough to enable the use of an energy harvesting source. These requirements place demands on all levels of the design. We propose architecture for achieving the required ultra-low energy operation and discuss the circuit techniques necessary to implement the system. Dedicated hardware implementations improve the efficiency for specific functionality, and modular partitioning permits fine-grained optimization and power-gating. We describe modeling and operating at the minimum energy point in the subthreshold region for digital circuits. We also examine approaches for improving the energy efficiency of analog components like the transmitter and the ADC. A microsensor node using the techniques we describe can function in an energy-harvesting scenario. Benton H. Calhoun, Denis C. Daly, Naveen Verma, Daniel F. Finchelstein, David D. Wentzloff, Alice Wang 0002, Seong-Hwan Cho, Anantha P. Chandrakasan |
IEEE Trans. Computers | 1 |
| 2004 | Characterizing and modeling minimum energy operation for subthreshold circuitsabstractSubthreshold operation is emerging as an energy-saving approach to many new applications. This paper examines energy minimization for circuits operating in the subthreshold region. We show the dependence of the optimum V DD for a given technology on design characteristics and operating conditions. Solving equations for total energy provides an analytical solution for the optimum V DD and V T to minimize energy for a given frequency in subthreshold operation. SPICE simulations of a 200K transistor FIR filter confirm the analytical solution and the dependence of the minimum energy operating point on important parameters. Benton H. Calhoun, Anantha P. Chandrakasan |
ISLPED | 1 |
| 2003 | Power-aware architectures and circuits for FPGA-based signal processingabstractThis work showcases a power-aware system design methodology for DSP applications on reconfigurable hardware platforms. In particular, an enhanced FPGA architecture is proposed and analyzed for a deep submicron process technology. These enhancements reduce Configurable Logic Block (CLB) usage for distributed arithmetic implementations of signal processing applications by 50% or more thereby reducing the load on interconnect resources. Multi-Threshold CMOS (MTCMOS) circuit design techniques are aggressively applied to reduce subthreshold leakage using an auto power-down feature for unused logic. Results show a 14x reduction in leakage current for unused CLBs or CLBs in deep sleep mode. CLBs in active mode see up to 2.8x steady-state power reduction. A testchip demonstrating these techniques in 0.13 micron technology has been sent out for fabrication. Frank Honoré, Benton H. Calhoun, Anantha P. Chandrakasan |
FPGA | 2 |
| 2003 | Design methodology for fine-grained leakage control in MTCMOSabstractMulti-threshold CMOS is a popular technique for reducing standby leakage power with low delay overhead. MTCMOS designs typically use large sleep devices to reduce standby leakage at the block level. We provide a formal examination of sneak leakage paths and a design methodology that enables gate-level insertion of sleep devices for sequential and combinational circuits. A fabricated 0.13 /spl mu/m, dual V/sub T/ test chip employs this methodology to implement a low-power FPGA core with gate-level sleep FETs and over 8/spl times/ measured standby current reduction. The methodology allows local sleep regions that reduce leakage in active configurable logic blocks (CLBs) by up to 2.2/spl times/ (measured) for some CLB configurations. Benton H. Calhoun, Frank Honoré, Anantha P. Chandrakasan |
ISLPED | 1 |