Tony Tae-Hyoung Kim

dblp:125/7918 · also Tae Hyoung Kim 0001, Tae-Hyoung Kim 0001, Tony T. Kim · DBLP profile ↗
← Back
70ranked-venue papers
5as first author
32since 2021 · last 2026
0000-0002-1779-1799ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 70 · 5 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 A 0.74 pJ/bit 10 Gbps PAM-4 Transceiver with Triple TIA Termination and Power-saving Addition-only Driving for Parallel Memory Interface Channels
Junsen He, Kiho Seong, Dong-Hyun Yoon, Seung-Myeong Yu, Junyoung Song, Jung-Hwan Choi, Tony Tae-Hyoung Kim
ISCAS8
2026 Time-Based Sensing With Linear Current-to-Time Conversion for Multi-Level Resistive Memory
abstract
Resistive Random Access Memory (RRAM) is a promising low-power memory candidate because of a large R-ratio (RHRS/RLRS). Multi-level RRAM cells have been investigated to improve memory density and cost-per-bit. However, sensing multi-level becomes challenging due to the smaller R-ratios between stored digital values. This paper introduces a novel time-based sensing (TBS) scheme for enhancing the sensing speed and robustness for single-level cells (SLC) to multi-level cells (MLC). The proposed time-based sensing scheme converts the bit line (BL) current into a time delay using a novel current-to-time converter (CTC). The BL-current-dependent time delays are utilized to generate digital data. In addition, the proposed sensing scheme executes sensing without need for reference current or reference voltage. Comprehensive simulation in 40nm CMOS technology shows that the proposed TBS scheme achieves better linearity and higher read speed by precise cell current replication in CTC compared to the prior TBS schemes. As a result, the proposed TBS reduces sensing latency by 230%~340% compared to the prior TBS schemes. Furthermore, the proposed TBS improves the variation tolerance of read operation by 5%~33% at 1.1 V.
Byung-Kwon An, Xueyong Zhang, Anh-Tuan Do, Tony Tae-Hyoung Kim
IEEE Trans. Circuits Syst. I Regul. Pap.4
2026 Highly Dense Capacitor SRAM Computation-In-Memory With Dynamic Range Calibrated Column-by-Column ADCs
abstract
This paper presents a highly dense SRAM-based computation-in-memory (CIM) architecture designed for area-efficient AI acceleration. The proposed architecture leverages charge redistribution in a capacitor-based CIM design to enhance linearity, effectively mitigating the impact of variable parasitic capacitance inherent in FETs. To address the dynamic range mismatch between the CIM array and the ADC, a dynamic range calibration scheme is introduced. Recognizing the significant area overhead of ADCs in CIM, this work also proposes an area-efficient SR-latch-based cap-DAC driver. Furthermore, this work compares and analyzes various ADC comparator reference types, ultimately proposing individual reference columns with global drivers to achieve both ADC symmetry and improved area/power efficiency. Fabricated using a 65nm LP process, the prototype occupies a core area of 0.265 mm2, comprising a 108kb ($432\times 256$) SRAM-CIM array and 256 column-by-column ADCs with peripherals. This design achieves a weight density of 415.4 kb/mm2. By employing column-by-column ADCs, which avoid the throughput limitations of ADC sharing, the prototype achieves an energy efficiency of 95.4 TOPS/W and a throughput of 813.6 GOPS. The resulting area and energy-efficient architecture achieves a SWaP (space, watt, and performance) figure-of-merit of 39.63 TOPS/W$\times $Mb/mm2.
Chufeng Yang, Sen Cao, Dong-Hyun Yoon, Yuanjin Zheng, Tony Tae-Hyoung Kim
IEEE Trans. Circuits Syst. I Regul. Pap.6
2026 A Real-Time End-to-End Event-Based Tactile Sensing System
abstract
Event-based tactile sensing offers a promising alternative to frame-based approaches by reducing data redundancy, yet existing systems often lack end-to-end hardware support and remain power-inefficient. This work presents an energy-efficient event-based tactile sensing system that codesigns algorithms and hardware for real-time perception. At the front end, a leakage-compensated event-driven readout circuit integrates multiple tactile sensors into a single node, minimizing wiring and static power. At the algorithm level, a 2-D convolutional neural network (CNN) reconstructs event frames for handwritten digit recognition with strong robustness to varying interaction durations, while a graph neural network (GNN) processes irregular tactile layouts for object classification. A customized event-based tactile processor (ETP) supports end-to-end tactile processing, incorporating a dual-mode event frame builder (EFB) and a neural network processing unit (NPU) with an optimized data-reuse scheme. Evaluated using 28-nm CMOS, the ASIC performs handwritten digit recognition and object classification in 0.44 and 0.37 ms, respectively. The system consumes the total power of 5.2 mW with only 0.11-mW dynamic power, demonstrating an efficient solution for always-on tactile perception.
Yuncheng Lu, Kiho Seong, Si En Timothy Ng, Shibi Varku, Arindam Basu, Nripan Mathews, Tony Tae-Hyoung Kim
IEEE Trans. Very Large Scale Integr. Syst.7
2026 A 65-nm 3T2R ReRAM Computing-in-Memory Macro Using a 4-bit 1-GS/s Two-Stage Pipelined 2-bit Output Comparator With a Successive Approximation-Like Architecture
Xueyong Zhang, Weifeng Sun 0001, Tony Tae-Hyoung Kim
IEEE Trans. Very Large Scale Integr. Syst.5
2025 A 1Mb RRAM Macro with Bipolar Forming for Improved Programming Yield and Cell-by-cell Write Verification Scheme
abstract
Resistive RAM(RRAM) has emerged as a promising candidate for the next generation non-volatile memories (NVMs) due to its low write voltage, compact area, and CMOS compatibility. In this work, we propose a 1Mb RRAM macro with bipolar forming to reduce the forming voltage and improve the programming yield. Additionally, a cell-by-cell write verification scheme is introduced to protect RRAM cells from overstress and improve RRAM yield. The test chip, fabricated using 40nm CMOS technology, occupies a core area of 2.34 mm2.
Byung-Kwon An, Junjie Mu, Putu Andhita Dananjaya, Weng Hong Lai, Wen Siang Lew, Tony Tae-Hyoung Kim
ISCAS6
2025 FlexDCIM: A 400 MHz 249.1 TOPS/W 64 Kb Flexible Digital Compute-in-Memory SRAM Macro for CNN Acceleration
abstract
This work proposes a 64Kb fully reconfigurable SRAM compute-in-memory (CIM) macro for convolutional neural network (CNN) acceleration using a 65nm node. It supports operation up to 400 MHz. The fully digital operation of the proposed macro effectively removes the analog CIM design issues related to process variations, noise susceptibility, and data-conversion overhead. Hence, it offers no accuracy loss, high energy efficiency, and large area saving for computation. To support the digital computation, a new area-efficient Digital Processing Unit (DPU) is proposed which is equivalent to 8.75T per bit storage. Moreover, the proposed macro features full precision reconfigurability (1b to 8b) for both input and weight, and fully flexible input activation ranging from 1 to 64 parallel inputs. It makes the proposed macro feasible for different neural network topologies. Removing sense amplifiers (SAs) for the memory mode of the proposed design suggests additional area and power savings. The proposed CIM macro achieves an energy efficiency of 249.1TOPS/W and a throughput of 819.2 GOPS.
Vishal Sharma 0004, Xin Zhang 0025, Narendra Singh Dhakad, Tony Tae-Hyoung Kim
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 A 2.793 μW Near-Threshold Neuronal Population Dynamics Trajectory Filter for Reliable Simultaneous Localization and Mapping
abstract
This work presents an algorithm hardware co-design implementing a digital neuronal population dynamics simulator intended for the trajectory error correction task within a simultaneous localization and mapping workflow. A custom discretized procedural algorithm approximating a neuronal population dynamics-based inference operation is developed for mapping onto an ultra-lightweight digital macro featuring massively parallel in-situ processing techniques. Fabricated using a 40nm technology, the test chip features a$22\times 22$neuron array with 0.1358mm2 core area and provides a 12-bit computing precision. A time-multiplexed processing element design prevents the use of excessive silicon area. Accomplished via extensive data reuse through massively parallel processing-in-memory architecture attached to a custom I/O interface, a single inference operation is completed within 3277 clock cycles, providing 200 inferences per second operating at a low frequency of 0.667Mhz with a 0.5V core supply and consuming sub-10-$\mu $W power.
Zhengzhe Wei, Boyi Dong, Yuqi Su, Yi Estelle Wang, Chuanshi Yang, Yuncheng Lu, Chao Wang 0096, Tony Tae-Hyoung Kim, Yuanjin Zheng
IEEE Trans. Circuits Syst. I Regul. Pap.8
2025 A Graph-Based Accelerator of Retinex Model With Bit-Serial Computing for Image Enhancements
abstract
This work proposes the Poisson equation formulation of the Retinex model for image enhancements using a low-power graph hardware accelerator performing finite difference updates on a lattice graph processing element (PE) array. By encapsulating the underlying algorithm in a graph hardware structure, a highly localized dataflow that takes advantage of the physical placement of the PEs is enabled to minimize data movement and maximize data reuse. The on-chip dataflow that achieves data sharing, and reuse among neighboring PEs during massively parallel updates is generated in each PE driven by two external control signals. Using a custom accumulator design intended for bit-serial computing, this work enables precision on demand and extensive on-chip data reuse with minimal area overhead, accommodating a non-overlap image mapping scheme in which a$20\times 20$image tile can be processed without external memory access at a time. With increasing user-configurable update count, image noise and shadow can be progressively removed with the inevitable loss of image details. Fabricated using a 65nm technology, the test chip occupies 0.2955mm2 core area and consumes 2.191mW operating at 1V, 25.6MHz, and a reconfigurable 10- or 14-bit precision.
Zhengzhe Wei, Junjie Mu, Yuanjin Zheng, Tony Tae-Hyoung Kim, Bongjin Kim
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 A 65-nm 55.8-TOPS/W Compact 2T eDRAM-Based Compute-in-Memory Macro With Linear Calibration
abstract
Implementing parallel computing inside memory units, compute-in-memory (CIM) has shown significant energy and latency reduction, which are suitable for neural network accelerators, especially for low-power edge devices. This brief presents a compact 2T-eDRAM CIM structure to support signed 4b/4b/6b input/weight/output precision multiply-accumulate (MAC) operation, exploring a near-zero-skipping (NZS) technique to improve energy efficiency further and reduce weight update time. The center weight first (CWF) update method is proposed to extend the overall weight retention time. Furthermore, the analog multiplication and accumulation nonlinear compensation techniques are employed to improve the accuracy and linear range. Fabricated in 65-nm CMOS technology, this chip achieves the weight bit storage density of 3.7 Mb/mm$^{2}$and SWaP figure of merit of 210 TOPS/W Mb/mm$^{2}$. The measured energy efficiency shows an average of 55.8 TOPS/W with the 4b/4b/6b input/weight/output precision at 1.2 V and 100 MHz.
Xueyong Zhang, Tony Tae-Hyoung Kim
IEEE Trans. Very Large Scale Integr. Syst.3
2024 Time-based Sensing with Linear Current-to-Time Conversion for Multi-level Resistive Memory
abstract
Resistive Random Access Memory (RRAM) is a promising low-power memory candidate because of a large R-ratio (RHRS/RLRS). Multi-level RRAM cells have been investigated to improve memory density and cost-per-bit. However, sensing multi-level becomes challenging due to the smaller R-ratios between stored digital values. This paper introduces a novel time-based sensing (TBS) scheme for enhancing the robustness and speed of the read operation, supporting from single-level cells (SLC) to multi-level cells (MLC). In the proposed time-based sensing scheme, the current-to-time converter (CTC) converts the bit line (BL) current into a time delay based on the cell states. Different time delays from HRS and LRS values are compared to generate digital data. Unlike conventional current-based sense amplifiers (CSA) or voltage-based sense amplifiers (VSA), the proposed sensing scheme does not require a reference array or generator. Comprehensive simulation in 40nm CMOS technology shows enhanced linearity and higher read speed because of the precise replication of cell current to CTC. As a result, the proposed TBS achieves a sensing latency of < 2ns with an energy consumption of 61fJ/bit for read operations at 1.1 V and also supports MLC sensing through a single read operation.
Byung-Kwon An, Xueyong Zhang, Anh-Tuan Do, Tony Tae-Hyoung Kim
ISCAS4
2024 Live Demonstration: Real-Time Object Detection & Classification System in IoT with Dynamic Neuromorphic Vision Sensors
abstract
In this paper, we demonstrate an energy-efficient real-time object detection and classification system featuring a hybrid event-based frame generation pipeline and a background-removal region proposal algorithm. The event-based frame is generated by aggregating active events within a programmable time interval, generating an event-based binary image (EBBI). This approach enables the utilization of low-complexity algorithms for denoising and object detection. The background-removal region proposal algorithm reduces memory requirements and removes dynamic backgrounds, leading to better detection performance. The proposed system is demonstrated on Zynq-7000 FPGA device with a DAVIS346 sensor. Experimental results show that the proposed system achieves comparable detection accuracy while requiring significantly less computation than existing event-based trackers.
Wenhao Lu, Yuncheng Lu, Junying Li, Yucen Shi, Yuanjin Zheng, Tony Tae-Hyoung Kim
ISCAS7
2024 An Energy-Efficient Object Detection System in IoT with Dynamic Neuromorphic Vision Sensors
abstract
Neuromorphic vision sensors (NVSs) mimic the function of the human visual system, with significant energy-saving potential in IoT-based object detection systems. Unlike conventional sensors, NVSs only generate asynchronous spiking events in response to changes in light intensity. However, the inherent noise generated by NVSs causes a degradation of detection performance. Moreover, an interested object usually occupies only a portion of the entire image frame. Therefore, a real-time, accurate event-based object detection system is needed to identify the region of interest (Rol) and leverage this spatial redundancy to reduce computational load in subsequent recognition modules. In this article, we present an energy-efficient real-time object detection system featuring a hybrid event-based frame generation pipeline and a background-removal region proposal algorithm. The event-based frame is generated by aggregating active events within a programmable time interval, generating an event-based binary image (EBBI). This approach enables the utilization of low-complexity algorithms for denoising and object detection. The background-removal region proposal algorithm reduces memory requirements and removes dynamic backgrounds, leading to better detection performance. The proposed system is demonstrated on Zynq-7000 FPGA device with a DAVIS346 sensor. Experimental results show that the proposed system achieves comparable accuracy while requiring significantly less computation than existing event-based trackers.
Wenhao Lu, Yuncheng Lu, Junying Li, Yucen Shi, Yuanjin Zheng, Tony Tae-Hyoung Kim
ISCAS7
2024 A Memory-Efficient High-Speed Event-based Object Tracking System
abstract
Dynamic vision sensors (DVS) have become prevalent in edge vision applications due to their low power and short latency attributes. However, current DVS-based object tracking systems suffer from high power consumption or long processing latency due to high computing intensity of the object detection algorithms. This paper proposes an energy-efficient object detection system through algorithm and hardware co-optimization. We design hardware-efficient denoising and region proposal (RP) algorithms to reduce on-chip memory usage and power consumption. Besides, the processing latency is dramatically reduced thanks to the less computing complexity. The devised algorithm is executed on a heterogeneous platform, with segments particularly sensitive to latency being accelerated via FPGA. An RP processor, supporting both parallel and systolic computing modes, is developed to facilitate the computing-intensive RP generation. Remarkably, the proposed system reduces the on-chip memory by 95.3% in contrast to traditional methods that employ connected component labeling. Moreover, the processing time per frame stands at 92.2 ms, marking a reduction of 82.4% compared to CPU-only operations.
Yuncheng Lu, Kaixiang Cui, Yucen Shi, Junying Li, Wenhao Lu, Yuanjin Zheng, Tony Tae-Hyoung Kim
ISCAS8
2024 A 2.793µW Near-Threshold Neuronal Population Dynamics Simulator for Reliable Simultaneous Localization and Mapping
abstract
This work presents an algorithm hardware co-design implementing a digital neuronal population dynamics simulator intended for a component within the back-end of simultaneous localization and mapping. A custom discretized procedural algorithm including injection, finite difference update, activation, and inhibition to approximate neuronal population dynamics is developed for digital implementation. Fabricated using a 40nm technology, the test chip features a scalable neuron 22 × 22 array with 0.1358mm2core area and provides a 12-bit computing precision. A time-multiplexed processing element design prevents the use of excessive silicon area. Accomplished via extensive data reuse through massively parallel processing-in-memory architecture attached to a custom I/O interface, a single inference operation is completed within 3277 clock cycles, providing 200 inferences per second operating at a low frequency of 0.667Mhz with a 0.5V core supply and consuming 2.793µW of power.
Zhengzhe Wei, Boyi Dong, Yuqi Su, Yi Estelle Wang, Chuanshi Yang, Yuncheng Lu, Chao Wang 0016, Tony Tae-Hyoung Kim, Yuanjin Zheng
ISCAS8
2024 A 0.6-to-1.2 V Scaling-Friendly Discrete-Time OTA-Free Linear VCO-Based ΔΣ ADC Suitable for DVFS
abstract
This work presents a scaling-friendly Discrete-Time (DT) 2nd order$\Delta \Sigma $ADC using a passive lossy integrator and a novel VCO-Frequency Delta-Sigma Modulator (FDSM). The proposed ADC has a 1st order$\Delta \Sigma $ADC structure with a 1st-order loop filter and a quantizer. The loop filter consists of a passive lossy integrator, which is OTA-free, energy-efficient, and provides a 1st order noise shaping. The proposed VCO-FDSM as a quantizer also realizes additional 1st order noise shaping. Therefore, the overall structure has 2nd order noise shaping property. The proposed VCO-FDSM employs a linear VCO with negative feedback to improve the linearity. The VCO output frequency is quantized into a multi-bit thermometer digital output code with inherent dynamic element matching (DEM) property by the fully digital FDSM. The FDSM output having DEM property mitigates the mismatch error of the capacitive digital-to-analog converter (CDAC). This highly digital architecture facilitates a trade-off between the supply voltage and the operational bandwidth (BW), which makes the ADC suitable for energy-efficient IoT applications. The test chip fabricated in 65nm standard CMOS technology operates at wide supply voltages ranging from 0.6 V to 1.2 V with a bandwidth range from 1.25 kHz to 5 kHz. The proposed ADC achieves a 75.7 dB peak SNR at 3.5 kHz BW, consuming$7.78~\mu \text{W}$with a Schreier FoMs of 164.7 dB at 1 V.
Kyung-Chan An, Neelakantan Narasimman, Tony Tae-Hyoung Kim
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 A Dual 7T SRAM-Based Zero-Skipping Compute- In-Memory Macro With 1-6b Binary Searching ADCs for Processing Quantized Neural Networks
abstract
This article presents a novel dual 7T static random-access memory (SRAM)-based compute-in-memory (CIM) macro for processing quantized neural networks. The proposed SRAM-based CIM macro decouples read/write operations and employs a zero-input/weight skipping scheme. A 65nm test chip with$528\times 128$integrated dual 7T bitcells demonstrated reconfigurable precision multiply and accumulate operations with$384\times $binary inputs (0/1) and$384\times 128$programmable multi-bit weights (3/7/15-levels). Each column comprises$384\times $bitcells for a dot product,$48\times $bitcells for offset calibration, and$96\times $bitcells for binary-searching analog-to-digital conversion. The analog-to-digital converter (ADC) converts a voltage difference between two read bitlines (i.e., an analog dot-product result) to a 1-6b digital output code using binary searching in 1-6 conversion cycles using replica bitcells. The test chip with 66Kb embedded dual SRAM bitcells was evaluated for processing neural networks, including the MNIST image classifications using a multi-layer perceptron (MLP) model with its layer configuration of 784-256-256-256-10. The measured classification accuracies are 97.62%, 97.65%, and 97.72% for the 3, 7, and 15 level weights, respectively. The accuracy degradations are only 0.58 to 0.74% off the baseline with software simulations. For the VGG6 model using the CIFAR-10 image dataset, the accuracies are 88.59%, 88.21%, and 89.07% for the 3, 7, and 15 level weights, with degradations of only 0.6 to 1.32% off the software baseline. The measured energy efficiencies are 258.5, 67.9, and 23.9 TOPS/W for the 3, 7, and 15 level weights, respectively, measured at 0.45/0.8V supplies.
Chengshuo Yu, Haoge Jiang, Junjie Mu, Kevin Tshun Chuan Chai, Tony Tae-Hyoung Kim, Bongjin Kim
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 An Effective Faulty TSV Detection Scheme for TSVs in High Bandwidth Memory
abstract
Through-Silicon Vias (TSVs) are essential parts of interconnects in 3-D ICs such as high bandwidth memory (HBM). TSV defects from fabrication or electromigration can increase the resistance of defective TSVs and degrade signal integrity. As many TSVs are implemented in HBMs, effective testing of TSVs and identifying defective TSVs become evermore significant. In this paper, we propose a faulty TSV detection scheme for identifying the locations of faulty TSVs without time-consuming manual probing. TSV conditions are sensed by transmission signal between transceivers. Such signal will also operate in transition-time-to-voltage converter and sensor for TSV condition checking. The location of faulty TSV indication signals from the sensors in each core die will transmit back to the base logic die for further operations. Moreover, the proposed detector can detect faulty TSVs while the transceivers are operating.
Junsen He, Dong-Hyun Yoon, Tony Tae-Hyoung Kim
ISCAS3
2023 A Graph-Based Accelerator of Retinex Model with Bit-Serial Computing for Image Processing
abstract
This work implements the Poisson equation formulation of the Retinex model for image enhancements using a graph hardware accelerator performing finite difference updates on a 2D lattice graph PE array. A single clock gating control signal manages the data flow, data sharing, and reuse pattern among neighboring PEs during massively parallel updates. With increasing user-configurable update count, image noise and shadow can be progressively removed with the inevitable loss of image details. Accommodating a non-overlap image mapping scheme in which a$20\times 20$image tile can be processed without external memory access at a time, the proposed accelerator consists of 18$\times 18$regular PEs surrounded by$4\times 20$boundary PEs with reconfigurable data flow and 4 boundary cache registers. Fabricated using a 65nm technology, the test chip occupies 0.2955mm2core area, and consumes 2.191mW operating at 1V, 25.6MHz, and a reconfigurable 10- or 14-bit precision.
Zhengzhe Wei, Junjie Mu, Zhongzhiguang Lu, Yuanjin Zheng, Tony Tae-Hyoung Kim, Bongjin Kim
ISCAS5
2023 282-to-607 TOPS/W, 7T-SRAM Based CiM with Reconfigurable Column SAR ADC for Neural Network Processing
abstract
Compute in memory ($C$iM) is a promising solution for solving the bottleneck of frequent data interface between memory and processor in Von-Neumann architecture. In this work, a hybrid current/charge domain 7T-SRAM based CiM architecture is proposed to mitigate the PVT-induced RBL variation during computation and thus offer a better linearity without significant impact on the operating frequency and area efficiency. Additionally, a column-referenced 1b to 5b reconfigurable SAR ADC is proposed to support multi-bit output. The proposed design is verified by the Monte-Carlo simulations using 40nm CMOS technology. The 5b mode ADC transferred MAC curve's DNL (LSB) ranges from −0.025 to 0.02 and INL (LSB) ranges from −0.13 to 0.25. The largest RBL variation$(\sigma)$from MAC value −64 to MAC value +64 is 2.08 mV, resulting in a MNIST classification accuracy of 97.5%, which is only 0.1% degradation and Google Speech Command classification accuracy of 80.5%, which is only 0.5% degradation compared to the software baseline, respectively. The whole architecture offers energy efficiency of 282-to-607 TOPS/W for 1-5b output in the MAC operation, which is competitive when compared to other state-of-art$C$iM architectures.
Qibang Zang, Wang Ling Goh, Lu Lu 0013, Chengshuo Yu, Junjie Mu, Tony Tae-Hyoung Kim, Bongjin Kim, Dongrui Li, Anh-Tuan Do
ISCAS6
2023 BP-SCIM: A Reconfigurable 8T SRAM Macro for Bit-Parallel Searching and Computing In-Memory
abstract
This work presents BP-SCIM: a reconfigurable 8T static random access memory (SRAM) macro for bit-parallel searching and computing in-memory (CIM). The decoupled read/write ports of the employed 8T SRAM bit-cell eliminate read disturbance during search and CIM operations. BP-SCIM can support both in-memory Boolean logic and arithmetic operations. Novel CIM-friendly algorithms and peripheral circuits are proposed to reduce the latency of complex arithmetic operations such as multiplication and division. In addition, BP-SCIM can be configured as either a binary content-addressable memory (CAM) or a ternary CAM for fast searching. A$256\times64$BP-SCIM test chip was implemented in 65-nm CMOS technology. The 8-bit addition and 8-bit multiplication operations can achieve the maximum energy efficiency of 3.11 TOPS/W and 0.17 TOPS/W, respectively at 0.7 V supply. For the binary CAM search operation, BP-SCIM can achieve the minimum energy consumption of 0.91 fJ/bit/search at 87 MHz and 0.8 V supply.
Yuzong Chen 0001, Junjie Mu, Lu Lu 0013, Tony Tae-Hyoung Kim
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 A 1-16b Reconfigurable 80Kb 7T SRAM-Based Digital Near-Memory Computing Macro for Processing Neural Networks
abstract
This work introduces a digital SRAM-based near-memory compute macro for DNN inference, improving on-chip weight memory capacity and area efficiency compared to state-of-the-art digital computing-in-memory (CIM) macros. A$20\times 256.1$-16b reconfigurable digital computing near-memory (NM) macro is proposed, supporting a reconfigurable 1-16b precision through the bit-serial computing scheme and the weight and input gating architecture for sparsity-aware operations. Each reconfigurable column MAC comprises$16\times $custom-designed 7T SRAM bitcells to store 1-16b weights, a conventional 6T SRAM for zero weight skip control, a bitwise multiplier, and a full adder with a register for partial-sum accumulations.$20\times $parallel partial-sum outputs are post-accumulated to generate a sub-partitioned output feature map, which will be concatenated to produce the final convolution result. Besides, pipelined array structure improves the throughput of the proposed macro. The proposed near-memory computing macro implements an 80Kb binary weight storage in a 0.473mm2 die area using 65nm. It presents the area/energy efficiency of 4329-270.6 GOPS/mm2 and 315.07-1.23TOPS/W at 1-16b precision.
Junjie Mu, Chengshuo Yu, Tony Tae-Hyoung Kim, Bongjin Kim
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 A Robust Time-Based Multi-Level Sensing Circuit for Resistive Memory
abstract
Resistive random access memory (RRAM) is a promising emerging nonvolatile memory (NVM) due to its large resistance ratio in different switching states. To improve memory density and reduce cost-per-bit, multi-level cell (MLC) RRAM stores multiple bits in a single cell, compared to a single-level cell (SLC). However, random mismatch, process variation, and resistance shift lead to reliability issues, degrade the probability of correct read, and increase the bit error rate (BER). This paper presents a time-based sensing scheme for robust read operation and extends to multi-level sensing for SLC and MLC RRAM arrays. Bit line (BL) voltage is converted into time delay by a voltage-to-time converter (VTC) and compared with the implicit timing reference generated by a delay line. By detecting different states in the time domain, the proposed time-mode sense amplifier (TSA) requires no analog reference voltage or current, which is used in the conventional voltage-mode sense amplifiers (VSA) or current-mode sense amplifiers (CSA). Power gating is employed to enable the time sampling only at the sensing points to suppress the short-circuit current. A charge sharing-induced error compensation (CSEC) circuit is used to eliminate the charge sharing-induced voltage drop and expand the sense margin by$1.56\times $. Monte Carlo simulations in 40nm technology show that the proposed TSA improves read reliability and reduces BER by 3–4 orders of magnitude compared to conventional VSA and CSA. The proposed time-based sensing scheme operates from 0.7-1.2 V supply and consumes 49 fJ/bit for read operation under a nominal 1.2 V supply.
Xueyong Zhang, Byung-Kwon An, Tony Tae-Hyoung Kim
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 A Reconfigurable 8T SRAM Macro for Bit-Parallel Searching and Computing In-Memory
abstract
This work presents BP-SCIM: a reconfigurable 8T SRAM macro for bit-parallel searching and computing in-memory (CIM). BP-SCIM can perform in-memory Boolean logic, arithmetic, and content-addressable memory (CAM) operations. Novel peripheral circuits and algorithms are proposed to support complex arithmetic operations such as multiplication and division. A $256\times 64$ BP-SCIM test chip was implemented in 65-nm CMOS technology. The 8-bit addition and 8-bit multiplication operations can achieve the best energy efficiency of 3.11 TOPS/W and 0.17 TOPS/W, respectively at 0.7 V supply. For the binary CAM search operation, BP-SCIM can achieve the minimum energy consumption of 0.91 fJ/bit/search at 87 MHz and 0.8 V supply.
Yuzong Chen 0001, Junjie Mu, Lu Lu 0013, Tony Tae-Hyoung Kim
ISCAS5
2022 A 6T SRAM Based Two-Dimensional Configurable Challenge-Response PUF for Portable Devices
abstract
This work proposes a 2-dimensional programable SRAM-based PUF. The selection of challenge groups, orders, and sequence lengths dominates the responses with challenge-response pairs (CRPs) by order of rows$^{\mathrm {(sequence~\textrm {}length- 1)}} \times $columns$^{\mathrm {(sequence~\textrm {}length - 1)}}$. The PUF bit cell has split word-lines with vertical and horizontal connections, the bit-lines are placed orthogonally to generate one-bit data with four cells, the entropy source is enriched to 24 transistors. The proposed PUF supports multiple data maps from a single chip. A test chip was fabricated in 65 nm CMOS technology. Under 0.8V and 20 °C (nominal point), the bit error rate reaches 3%. In a single chip, the hamming distance achieves 42.49% within the same group and different orders of challenges, and 47.32% within the different groups of challenges (when the sequence length is 5). The measured inter-hamming distance between chips is improved to 49.47%.
Lu Lu 0013, Taegeun Yoo, Tony Tae-Hyoung Kim
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 A 64 Kb Reconfigurable Full-Precision Digital ReRAM-Based Compute-In-Memory for Artificial Intelligence Applications
abstract
This work presents a fully-digital 64 Kb non-volatile ReRAM based compute-in-memory (CIM) macro for the modern artificial intelligence (AI) edge devices, using 65 nm technology. This digital CIM architecture effectively removes the analog-design issues, related to process variations, noise susceptibility, and data-conversion overhead. Hence, it offers no accuracy loss and high energy-efficiency for the computation. To incorporate the digital computation, a novel NAND logic based 3.25T1R bitcell is proposed. The digital behaviour of this cell makes it superior to the conventional 1T1R based analog bitcell. Also, with the inherent non-volatility of ReRAM, the proposed cell can be a good substitute for SRAM-based CIM architectures with$4.62\times $,$1.96\times $,$3.96\times $, and$5.12\times $lower area than the XNOR-based 12T, Twin-8T, 8T, and 6T SRAM cell respectively. Moreover, the proposed CIM architecture allows full reconfigurabiliy from 1 to 16b precision for both input and weight. It also allows activating any number of parallel inputs, ranging from 1 to 128. According to simulation results, the proposed macro successfully operates up to 166.6 MHz for 1/8/15b input/weight/output precision and achieves 27.28 TOPS/W without any accuracy loss. Removing sense amplifiers for the ReRAM mode of the proposed work claims additional area and power savings.
Vishal Sharma 0004, Tony Tae-Hyoung Kim
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 A 3.2-GHz 178-fsrms Jitter Subsampling PLL/DLL-Based Injection-Locked Clock Multiplier
abstract
This article proposes a 3.2-GHz subsampling phase-locked loop (SSPLL)-based injection-locked clock multiplier (ILCM) using a subsampling delay-locked loop (SSDLL). The proposed ILCM achieves a superior noise reduction effect at low offset frequency because of the high feedback gain of SSPLL. Also, SSDLL is proposed for background phase calibration between SSPLL and injection. Since SSPLL and SSDLL use identical SSCP, the power consumption of the frequency and phase calibration is only 1.32 mW. This work is fabricated in a 65-nm CMOS process, and the 10-kHz phase noise is improved by 8.6 dB. The rms jitter from 10 kHz to 30 MHz is 178 fs. The chip-to-chip variations (20 chips) are 17 and 169 fs with/without SSDLL, respectively. The measured results show that FoM1, FoM2, and total power consumption are −245.1, −260.1 dB, and 9.85 mW, respectively.
Dong-Hyun Yoon, Dong-Kyu Jung, Kiho Seong, Jae-Soub Han, Keun-Yong Chung, Ju Eon Kim, Tony Tae-Hyoung Kim, Kwang-Hyun Baek
IEEE Trans. Very Large Scale Integr. Syst.7
2021 A Multi-Functional 4T2R ReRAM Macro Enabling 2-Dimensional Access and Computing In-Memory
abstract
This paper presents a multi-functional resistive random access memory (ReRAM) macro using a novel 4T2R bit-cell. The proposed 4T2R ReRAM enables 2-dimensional (2D) memory access, which offers significant latency and energy reductions for many applications such as matrix operations. Besides non-volatile storage, the proposed 4T2R ReRAM can support two types of computing in-memory operations: ternary content-addressable memory (TCAM) and logic in-memory (LIM). Evaluations on various matrix operations show that the proposed 4T2R ReRAM with 2-D access capability can reduce memory access latency and energy by up to 88% and 82%, respectively compared with conventional 1T1R ReRAM. For TCAM, the proposed 4T2R bit-cell takes a smaller area than SRAM-based TCAM cell, while achieves a comparable search speed. For LIM, we propose an optimized LIM full adder (LIM- FA) that improves the delay and the power by 3.2* and 1.6*, respectively compared with prior LIM-FAs.
Yuzong Chen 0001, Lu Lu 0013, Yuncheng Lu, Tony Tae-Hyoung Kim
ISCAS4
2021 A Configurable Randomness Enhanced RRAM PUF with Biased Current Sensing Scheme
abstract
This paper explains a resistive nonvolatile memory (RRAM)-based physical unclonable function (PUF). The proposed PUF involves more variations and enhances the randomness by using four 1T1R cells to generate one-bit random data. Moreover, we propose a configurable replica column scheme and a biased current sensing amplifier to further improve the randomness. The proposed RRAM PUF is designed in 40nm CMOS technology. The simulated RRAM PUF with configurable replica columns achieves the randomness of 0.5001 with σ of 0.005. It also has a Hamming distance of 0.486 over various responses for a single chip and 0.498 between different chips.
Lu Lu 0013, Yuzong Chen 0001, Tony Tae-Hyoung Kim
ISCAS3
2021 AND8T SRAM Macro with Improved Linearity for Multi-Bit In-Memory Computing
abstract
In this work, we propose a multi-bit precision (4b input, 4b weight and 4b output) in-memory computing (IMC) architecture, based on the voltage scaling and charge sharing scheme, for the artificial intelligence (AI) edge devices. To achieve the efficient computation, a new AND logic based 8T SRAM cell (AND8T) has been used which employs the charge-domain based computation. For such computation, AND8T incorporates an overlaying metal-oxide-metal capacitor (MOM cap) with no bit-cell area overhead. The proposed cell mitigates the linearity issue of multiply and accumulate (MAC) operation for the IMC unit which is highly desirable for the reliable operation of complex neural networks (CNN). Moreover, our high precision AND8T based IMC architecture allows 128 parallel MAC operations avoiding the need of serial multi-bits input implementation through multiple cycles. The proposed design has been successfully verified by the monte carlo simulation results while working at 50MHz clock frequency and 1V supply using standard 65nm node.
Vishal Sharma 0004, Ju Eon Kim, Yuzong Chen 0001, Tony Tae-Hyoung Kim
ISCAS5
2021 A Logic-Compatible eDRAM Compute-In-Memory With Embedded ADCs for Processing Neural Networks
abstract
A novel 4T2C ternary embedded DRAM (eDRAM) cell is proposed for computing a vector-matrix multiplication in the memory array. The proposed eDRAM-based compute-in-memory (CIM) architecture addresses a well-known Von Neumann bottle-neck in the traditional computer architecture and improves both latency and energy in processing neural networks. The proposed ternary eDRAM cell takes a smaller area than prior SRAM-based bitcells using 6-12 transistors. Nevertheless, the compact eDRAM cell stores a ternary state (-1, 0, or +1), while the SRAM bitcells can only store a binary state. We also present a method to mitigate the compute accuracy degradation issue due to device mismatches and variations. Besides, we extend the eDRAM cell retention time to 200μs by adding a custom metal capacitor at the storage node. With the improved retention time, the overall energy consumption of eDRAM macro, including a regular refresh operation, is lower than most of prior SRAM-based CIM macros. A 128×128 ternary eDRAM macro computes a vector-matrix multiplication between a vector with 64 binary inputs and a matrix with 64 × 128 ternary weights. Hence, 128 outputs are generated in parallel. Note that both weight and input bit-precisions are programmable for supporting a wide range of edge computing applications with different performance requirements. The bit-precisions are readily tunable by assigning a variable number of eDRAM cells per weight or adding multiple pulses to input. An embedded column ADC based on replica cells sweeps the reference level for 2N-1 cycles and converts the analog accumulated bitline voltage to a 1-5bit digital output. A critical bitline accumulate operation is simulated (Monte-Carlo, 3K runs). It shows the standard deviation of 2.84% that could degrade the classification accuracy of the MNIST dataset by 0.6% and the CIFAR-10 dataset by 1.3% versus a baseline with no variation. The simulated energy is 1.81fJ/operation, and the energy efficiency is 552.5-17.8TOPS/W (for 1-5bit ADC) at 200MHz using 65nm technology.
Chengshuo Yu, Taegeun Yoo, Tony Tae-Hyoung Kim, Kevin Tshun Chuan Chai, Bongjin Kim
IEEE Trans. Circuits Syst. I Regul. Pap.4
2021 A 16-kb 9T Ultralow-Voltage SRAM With Column-Based Split Cell-VSS, Data-Aware Write-Assist, and Enhanced Read Sensing Margin in 28-nm FDSOI
abstract
This work proposes an static random access memory (SRAM) with column-based split cell-VSS (CS-CVSS), data-aware write-assist (DAWA), and enhanced read sensing margin in 28-nm FDSOI technology. The proposed CS-CVSS and DAWA techniques improve both half-selected (HS) static noise margin (SNM) and write margin. They also improve HS dynamic noise margin (HS-DNM) by leveraging write through virtual ground with reduced load in the proposed write port. The proposed 3T read port enhances sensing margin by minimizing read bitline leakage through negative gate-to-source voltage regardless of cell data. A 16-kb 9T SRAM test chip demonstrated the minimum operating voltage for write and read operations as 0.47 and 0.25 V, respectively. The minimum energy of 6.72 pJ is achieved at 0.5 V.
M. Sultan M. Siddiqui, Zhao Chuan Lee, Tony Tae-Hyoung Kim
IEEE Trans. Very Large Scale Integr. Syst.3
2020 A 16×128 Stochastic-Binary Processing Element Array for Accelerating Stochastic Dot-Product Computation Using 1-16 Bit-Stream Length
abstract
This work presents 16×128 stochastic-binary processing elements for energy/area efficient processing of artificial neural networks. A processing element (PE) with all-digital components consists of an XNOR gate as a bipolar stochastic multiplier and an 8bit binary adder with 8× registers for accumulating partialsums. The PE array comprises 16× dot-product units, each with 128 PEs cascaded in a single row. The latency and energy of the proposed dot-product unit is minimized by reducing the number of bit-streams required for minimizing the accuracy degradation induced by the approximate stochastic computing. A 128-input dot-product operation requires the bit-stream length (N) of 1-to16, which is two orders of magnitude smaller than the baseline stochastic computation using MUX-based adders. The simulated dot-product error is 6.9-to-1.5% for N=1-to-16, while the error from the baseline stochastic method is 5.9-to-1.7% with N=128to-2048. A mean MNIST classification accuracy is 96.11% (which is 1.19% lower than 8b binary) using a three-layer MLP at N=16. The measured energy from a 65nm test-chip is 10.04pJ per dotproduct, and the energy efficiency is 25.5TOPS/W at N=16.
Qian Chen 0027, Yuqi Su, Taegeun Yoo, Tony Tae-Hyoung Kim, Bongjin Kim
DATE5
2020 Reconfigurable 2T2R ReRAM with Split Word-Lines for TCAM Operation and In-Memory Computing
abstract
The increased latency and power consumption due to data movement between memory and ALU have become the major obstacle in modern big-data and machine learning applications. Beyond von-Neumann architectures, particularly in-memory computing, is under intensive research to overcome this memory access bottleneck. In this work, we propose a 2T2R ReRAM structure that supports ternary content addressable memory (TCAM), logic in-memory operations, and in-memory dot product for Deep Neural Networks (DNNs) besides the normal non-volatile memory (NVM) functionality. This is achieved by employing reconfigurable sense amplifiers and novel word-line drivers. The proposed architecture can serve as a high-density storage system as well as an accelerator for data-intensive applications. Simulation results verify that the proposed 2T2R structure functions correctly for TCAM search, logic in-memory operations and in-memory dot product.
Yuzong Chen 0001, Lu Lu 0013, Bongjin Kim, Tony Tae-Hyoung Kim
ISCAS4
2020 Design and Characterization of Radiation-Hardened MCU for Space Application using Error Correction SRAM and Glitch Removal Clock Buffer Cell
abstract
High-energy environmental radiation particles may affect the System on Chip (SoC) which causes errors in the combinational logic, sequential logic, and even in the clock network. In the latter, they appear as clock glitches that propagate, and eventually, incorrectly latch all the sequential circuits, such as flip-flops and latches attached to the clock node concerned. The impact on-chip functionality is usually fatal. In this work, we propose a low-cost adaptive clock glitch removal circuitry for radiation-resilient clock networks. The proposed technique can be adjusted based on the actual clock glitch profile, to ensure that the error in the clock network is removed, while the impact on the clock signal itself is minimized. It is fully synthesizable, and can thus be incorporated in a conventional digital flow. Using the proposed clock buffer cell, the implemented ARM Cortex M0 shows 85% error reduction when exposed to radiation with minimum area overhead.
Anh-Tuan Do, Tony Tae-Hyoung Kim, Xin Liu 0015, Jun Zhou 0017
ISCAS2
2020 A Secure Data-Toggling SRAM for Confidential Data Protection
abstract
We study the security feature of static random access memory (SRAM) against the data imprinting attack and provide a solution to protect the SRAM from this attack. There are four main contributions in this paper. First, the negative-bias temperature-instability (NBTI) degradation of PMOS transistors in the conventional SRAM cell that causes the data imprinting effect is explained. Second, the data imprinting effect that leaks the stored information in the conventional SRAM cell is investigated. Third, a novel low transistor-count transmission-gate-based master-slave SRAM cell is proposed to periodically toggle the stored data for reducing the data imprinting effect. Fourth, an efficient imprinting analysis flow is proposed to evaluate the proposed data-toggling SRAM for quantifying the data imprinting effect. Based on a 65-nm CMOS process, we implement and prototype the proposed 1k-byte data-toggling SRAM design. We perform our imprinting analysis flow on various SRAM ICs and benchmark our proposed data-toggling SRAM IC against the non-toggling SRAM IC and a commercial Lyontek SRAM IC. From the measurement results, the non-toggling SRAM and Lyontek SRAM suffer from 60% and 81% data imprinting effects, respectively, whereas our data-toggling SRAM has only 11% data imprinting effect (at 160-kHz toggling frequency). The data-toggling SRAM could switch between high security (<; 5% data imprinting effect) high power mode for hardware security applications and low power (<; 0.1mW) low security mode for power-saving applications. Particularly, our data-toggling SRAM could feature as low as ~1% data imprinting effect when increasing the toggling frequency to 1.6 MHz by compromising the power dissipation. Using the image analysis flow, the stored information is revealed in both the non-toggling and Lyontek SRAM ICs but is well protected in the proposed data-toggling SRAM IC.
Weng-Geng Ho, Kwen-Siong Chong, Tony Tae-Hyoung Kim, Bah-Hwee Gwee
ISCAS3
2020 Reconfigurable 2T2R ReRAM Architecture for Versatile Data Storage and Computing In-Memory
abstract
Nonvolatile memory (NVM)-based computing in-memory (CIM) is a promising solution to data-intensive applications. This work proposes a 2T2R resistive random access memory (ReRAM) architecture that supports three types of CIM operations: 1) ternary content addressable memory (TCAM); 2) logic in-memory (LiM) primitives and arithmetic blocks such as full adder (FA) and full subtractor; and 3) in-memory dot-product for neural networks. The proposed architecture allows the NVM operations in both 2T2R and conventional 1T1R configurations. The proposed LiM full adder (LiM-FA) improves the delay, the static power, and the dynamic power by$3.2\times $,$1.2\times $, and$1.6\times $, respectively, compared with state-of-the-art LiM-FAs. Furthermore, based on different optimization techniques and robustness analysis, a lower precharge voltage is set for each mode. This reduces the TCAM search energy and 1T1R ReRAM access energy by$1.6\times $and$1.14\times $, respectively, compared with the case without optimizations.
Yuzong Chen 0001, Lu Lu 0013, Bongjin Kim, Tony Tae-Hyoung Kim
IEEE Trans. Very Large Scale Integr. Syst.4
2020 A 137-μW 1.78-mm2 30-Frames/s Real-Time Gesture Recognition SoC for Smart Devices
abstract
Gesture recognition has increasingly become one of the most popular human-machine interaction techniques for smart devices. Existing gesture recognition systems suffer from either excessive power consumption or large size, limiting their applications for ultralow-power wearable devices. This article presents an accurate area-efficient and low-power real-time gesture recognition system for smart wearable devices. The proposed work utilizes an accurate peak-based gesture classification engine with less memory and a low-resolution and low-power on-chip image sensor for achieving high area efficiency and low power. In addition, the feature extraction architecture removes fixed-pattern noises from the low-power on-chip image sensor for accuracy improvement and employs parallelism for recognition speed enhancement. Thus, the proposed system accomplishes accurate real-time gesture recognition for eight motion hand gestures with an average recognition accuracy of 90.6% and latency of 4.228 ms. Measurement results of a test chip fabricated in 65-nm CMOS demonstrate that the proposed system consumes 137.0 μW at 30 frames/s while occupying only 1.78 mm2, which achieves the lowest power and smallest area among the recently reported gesture recognition systems.
Van Loi Le, Taegeun Yoo, Ju Eon Kim, Ngoc Le Ba, Kwang-Hyun Baek, Tony Tae-Hyoung Kim
IEEE Trans. Very Large Scale Integr. Syst.6
2020 A 0.506-pJ 16-kb 8T SRAM With Vertical Read Wordlines and Selective Dual Split Power Lines
abstract
This article presents an 8T static random access memory (SRAM) macro with vertical read wordline (RWL) and selective dual split power (SDSP) lines techniques. The proposed vertical RWL reduces dynamic energy consumption during read operation by charging and discharging only selected read bitlines (RBLs). The data-aware SDSP technique combined with vertical write bitlines enhances both the write margin (WM) and the static noise margin (SNM). A 16-kb SRAM test chip fabricated in 65-nm CMOS technology demonstrates the minimum energy consumption of 0.506 pJ at 0.4 V and the minimum operating voltage of 0.26 V.
Lu Lu 0013, Taegeun Yoo, Van Loi Le, Tony Tae-Hyoung Kim
IEEE Trans. Very Large Scale Integr. Syst.4
2019 A Sequence-Dependent Configurable PUF Based on 6T SRAM for Enhanced Challenge Response Space
abstract
This work proposes a 2D sequence-dependent PUF based on SRAM. It expands challenge to response pairs (CRPs) by the order of rows(sequence length − 1) × columns(sequence length − 1) for reliable authentication. This is achieved by configuring the sequence of SRAM cell selection. Each bit cell has a vertical word-line and a horizontal word-line to utilize the orthogonal word-lines to connect four cells simultaneously to generate one bit data. The proposed technique allows us to generate multiple data maps from one chip. This non-linear behavior also makes the chip more secure. A test chip was fabricated in 65 nm CMOS technology with the area of area is 12580 µm2. The bit error rate is 3% at the nominal point (0.8 V/20°C) and the inter-hamming distance between chips is 0.497. The hamming distance of 0.427 was measured when using the same sequence length with different orders.
Lu Lu 0013, Tony Tae-Hyoung Kim
ISCAS2
2019 A Logic Compatible 4T Dual Embedded DRAM Array for In-Memory Computation of Deep Neural Networks
abstract
Modern deep neural network (DNN) systems evolved under the ever-increasing demands of handling more complex and computation-heavy tasks. Traditional hardware designed for such tasks had larger size memory and power consumption issue due to extensive on/off-chip memory access. In-memory computing, one of the promising solutions to resolve the issue, dramatically reduced memory access and improved energy efficiency by utilizing the memory cell to function as both a data storage and a computing element. Embedded DRAM (eDRAM) is one of the potential candidates for in-memory computation. Its minimal use of circuit components and low static power consumption provided design advantage while its relatively short retention time made eDRAM unsuitable for certain applications. This work introduces a dot-product processing macro using eDRAM array and explores its capability as an in-memory computing processing element. The proposed architecture implemented a pair of 2T eDRAM cells as a processing unit that can store and operate with ternary weights using only four transistors. Besides, we investigated a method to maximize the retention time in conjunction with analyzing the device mismatch. An input/weight bit-precision reconfigurable 4T eDRAM processing array shows the energy efficiency of 1.81fJ/OP (including refresh energy) when it operates with binary inputs and ternary weights.
Taegeun Yoo, Qian Chen 0027, Tony Tae-Hyoung Kim, Bongjin Kim
ISLPED4
2019 Editorial TVLSI Positioning - Continuing and Accelerating an Upward Trajectory
abstract
I. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5].
Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber
IEEE Trans. Very Large Scale Integr. Syst.26
2019 An Area-Efficient 128-Channel Spike Sorting Processor for Real-Time Neural Recording With 0.175µW/Channel in 65-nm CMOS
abstract
This paper presents a power- and area-efficient spike sorting processor (SSP) for real-time neural recordings. The proposed SSP includes novel detection, feature extraction, and improved K-means algorithms for better clustering accuracy, online clustering performance, and lower power and smaller area per channel. Time-multiplexed registers are utilized in the detector for dynamic power reduction. Finally, an ultra-low-voltage 8T static random access memory (SRAM) is developed to reduce area and leakage consumption when compared to D flip-flop-based memory. The proposed SSP, fabricated in 65-nm CMOS process technology, consumes only 0.175 μW/channel when processing 128 input channels at 3.2 MHz and 0.54 V, which is the lowest among the compared state-of-the-art SSPs. The proposed SSP also occupies 0.003 mm2/channel, which allows 333 channels/mm2.
Anh-Tuan Do, Seyed Mohammad Ali Zeinolabedin, Dongsuk Jeon, Dennis Sylvester, Tony Tae-Hyoung Kim
IEEE Trans. Very Large Scale Integr. Syst.5
2017 Library pruning and sigma corner libraries for power efficient variation tolerant processor pipelines
abstract
Error tolerance techniques are widely used to protect processor pipelines from variation induced timing errors. In this paper, we propose two standard cell library tuning techniques to optimize error tolerant processor pipelines for power and area savings. The design utilizes positive slack available in the pipeline stages and re-distributes it to the preceding error-prone critical paths using slack balancing flip-flops. Library pruning analyses the power and area metrics of the flip-flop cells to derive a power efficient subset of the original library. We use statistical sigma corner libraries to replace the critical flip-flop fan-in cone for further power optimization. Results show that the proposed library tuning techniques provide power reductions of 47% and area reductions of 2.8% in the employed execute stage module of a processor pipeline.
Mini Jayakrishnan, Alan Chang, Tony Tae-Hyoung Kim
VLSI-SoC3
2017 Yield Enhancement of Face-to-Face Cu-Cu Bonding With Dual-Mode Transceivers in 3DICs
abstract
When more than one dies are stacked vertically in 3-D integrated circuits (3DICs), the overall system yield degrades significantly. While each die can be tested before stacking, failures in 3DIC interconnects could jeopardize the entire system. In this paper, a dual-mode transceiver is proposed as a built-in-self-test/repair solution to improve the yield of direct face-to-face copper thermocompression bonding (Cu-Cu bonding). The proposed transceiver could improve the Cu-Cu bonding-based interconnect reliability with the introduction of two operation modes: the ohmic mode when Cu-Cu bonding presents low resistance and the capacitive coupling mode when Cu-Cu bonding is showing any sign of failure with high resistance at the bonding interface. Such mode sensing is self-contained in the transceiver itself with the help of the proposed resistance sensor. In this paper, we discuss the modeling of Cu-Cu bonding and the proposed transceiver design with power, latency, jitter, and crosstalk simulations followed by the design guideline for the practical implementation with yield analysis.
Myat Thu Linn Aung, Takefumi Yoshikawa, Chuan Seng Tan, Tony Tae-Hyoung Kim
IEEE Trans. Very Large Scale Integr. Syst.4
2017 Design of Temperature-Aware Low-Voltage 8T SRAM in SOI Technology for High-Temperature Operation (25 %C-300 %C)
abstract
A temperature-aware low-voltage 8T static random access memory (SRAM) for high-temperature operations is presented. A dedicated read port with virtual ground and optimal body bias improves sensing margin under very high temperature (up to 300 °C). Bitline offset voltage for data “0” caused by the virtual ground scheme is also compensated by a replica bitline. The independent body bias control feature of the employed silicon-on-insulator (SOI) technology allows the write margin to be enhanced significantly without using any write-assist circuitry. Test chips were fabricated in a 1-μm SO! technology with tungsten interconnect for reliability at high temperature and lesser process variation. Measurement results demonstrate that the proposed SRAM operates successfully up to 300 °C with the supply voltage range of 2-5 V. At the minimum performance variation point (VDD = 2.5 V), the SRAM consumes 1.48 mW and shows the access time of 156 ns and the maximum clock frequency of 14.38 MHz at 300 °C.
Ngoc Le Ba, Tony Tae-Hyoung Kim
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Area-efficient and low stand-by power 1k-byte transmission-gate-based non-imprinting high-speed erase (TNIHE) SRAM
abstract
We propose a novel 15-T Transmission-gate-based Non-Imprinting High-speed Erase (TNIHE) SRAM cell with emphases on low area overhead and low stand-by power attributes for highly secured data storage applications. We benchmark our proposed 15-T TNIHE SRAM cell against the reported 22-T Non-Imprinting High-speed Erase (NIHE) SRAM cell, and demonstrated three key features of reducing 7 transistors. First, we adopt the transmission gates (as opposed the cross-couple inverters) in the slave circuitry, saving 4 transistors. Second, we eliminate a transistor which uses to reset the slave circuitry, hence saving 1 transistor. Third, we apply the global inverse transistors (as opposed to the local inverse transistors) in the read /write circuit for each SRAM cell, hence further reduce 2 more transistors. As a result, our proposed TNIHE SRAM cell @ 65nm CMOS features ~17% smaller layout area. We design a 1k-byte memory based on the proposed TNIHE SRAM cells. On the basis of simulations, we show that our 1k-byte SRAM memory features overall ~13% smaller area, and dissipates on average, ~30% lower stand-by power than the reported NIHE counterpart.
Weng-Geng Ho, Ne Kyaw Zwa Lwin, N. Prashanth Srinivas, Kwen-Siong Chong, Tony Tae-Hyoung Kim, Bah-Hwee Gwee
ISCAS5
2016 An ultra-low voltage, VCO-based ADC with digital background calibration
abstract
This paper introduces an ultra-low voltage open loop VCO-based ADC with background calibration for ultra-low power applications. A novel calibration scheme is proposed to calibrate the nonlinear voltage-to-frequency tuning curve of the VCO. A replica VCO is used to compute the correction coefficients and the corrected values are stored in a lookup table. The proposed calibration method is at least 64 times faster than other state-of-the-art ones. A test chip was implemented in commercial 65nm CMOS technology. Measurement results confirm the effectiveness of the calibration scheme at 0.4 V. The proposed VCO-based ADC achieves a resolution of 8.8 bits at 10 KHz bandwidth with the power consumption of 1.15 μW in the open loop architecture.
Neelakantan Narasimman, Tony Tae-Hyoung Kim
ISCAS2
2016 Power and area efficient clock stretching and critical path reshaping for error resilience
abstract
Energy efficient semiconductor chips are in high demand to cater the needs of today's smart products.Advanced technology nodes insert high design margins to deal with rising variations at the cost of power, area and performance.Existing run time resilience techniques are not cost effective due to the additional circuits involved.In this paper, we propose a design time resilience technique using a clock stretched flip-flop to redistribute the available slack in the processor pipeline to the critical paths.We use the opportunistic slack to redesign the critical fan in logic using logic reshaping, better than worst case sigma corner libraries and multi-bit flip-flops to achieve power and area savings.Experimental results prove that we can tune the logic and the library to get significant power and area savings of 69% and 15% in the execute pipeline stage of the processor compared to the traditional worst-case design.Whereas, existing run time resilience hardware results in 36% and 2% power and area overhead respectively.
Mini Jayakrishnan, Alan Chang, Tony Tae-Hyoung Kim
VLSI-SoC3
2016 2.31-Gb/s/ch Area-Efficient Crosstalk Canceled Hybrid Capacitive Coupling Interconnect for 3-D Integration
abstract
This paper introduces a hybrid capacitive coupling interconnects (CCIs) array suitable for bumpless flip-chip 3-D integration. Inside the hybrid array, both single-ended and common-centroid differential CCIs are interleaved together to cancel the crosstalk among them. The crosstalk cancellation capability of its own allows CCIs to be placed closer and thus improves the area efficiency. A high gain and high common-mode-rejection ratio receiver is also presented to minimize the jitter caused by the common-mode noise. The process variation track biasing circuit is also proposed for the receiver. The measurement verifies that the proposed transceiver in a 3 × 3 pseudohybrid CCIs array produces only 84 ps or 0.2 unit interval crosstalk related jitter under the worst case crosstalk condition. A total of nine transceivers in the array achieve the data rate of 20.79 Gb/s and consume only 53 μW/Gb/s. The chip was fabricated in 65-nm CMOS technology.
Myat Thu Linn Aung, Eric Teck Heng Lim, Takefumi Yoshikawa, Tony Tae-Hyoung Kim
IEEE Trans. Very Large Scale Integr. Syst.4
2016 Read Bitline Sensing and Fast Local Write-Back Techniques in Hierarchical Bitline Architecture for Ultralow-Voltage SRAMs
abstract
Voltage scalable decoupled SRAMs operating at a subthreshold region have various challenges, such as deteriorated read bitline (RBL) swing resulting in read sensing failure and degraded cell stability due to the half-select write. This paper proposes an equalized bitline scheme to eliminate the leakage dependence on data pattern and thus improves RBL sensing and its resilience against process, voltage, and temperature variations. In addition, we propose a fast local write-back (WB) technique to implement a half-select-free write operation. With hierarchical bitline architecture, it facilitates a local read and a subsequent fast WB action to secure the original data without performance degradation. A 16-kb SRAM test chip has been fabricated in a 65-nm CMOS technology and achieved the minimum operating voltage of 0.24 V with a read access time of 4.88 μs.
Bo Wang 0020, Tony Tae-Hyoung Kim
IEEE Trans. Very Large Scale Integr. Syst.3
2015 An output feedback-based start-up technique with automatic disabling for battery-less energy harvesters
abstract
This paper presents a start-up circuit based upon output voltage feedback for thermal energy harvesting. The feedback loop consisting of an output load, a boost converter core, a charge-pump-based doubler and a CMOS inverter switch to track the output voltage and generate a reset signal to kickstart the boost converter. The proposed start-up circuit shows a remarkably improved start-up time of 35μs at the input voltage of 290mV. The start-up circuit is automatically disabled after the boost converter reaches steady-state. Only one off-chip component is required, improving the cost efficiency. The proposed start-up circuit consumes 3.2nW during the steady state, aiding the boost converter to achieve an efficiency of 77% at input voltage of 290mV. It was designed in standard 65-nm CMOS process technology and occupies the area of 0.072mm2.
Abhik Das, Yuan Gao 0011, Tony Tae-Hyoung Kim
ISCAS3
2015 A 32kb 9T SRAM with PVT-tracking read margin enhancement for ultra-low voltage operation
abstract
Diminishing bitline sensing margin at low voltage condition is one of the most challenging design obstacles for reliable SRAM implementation in nano-scale CMOS technologies. This paper presents a self-biased design technique that improves the bitline sensing margin during the read operation by sourcing a current which is the same as the total leakage along each bitline. It is able to automatically track changes in supply voltage, operating temperature and die-to-die process variations. Furthermore, a 9T SRAM cell is utilized to ensure that bitline leakage is data-independent. Simulation and measurement results using 65 nm CMOS process show that the proposed technique enlarge the bitline swing over a wide range operating temperature and operates successfully down to the supply voltage of 0.18 V.
Anh-Tuan Do, Kiat Seng Yeo, Tony Tae-Hyoung Kim
ISCAS3
2015 Design of a hybrid neural spike detection algorithm for implantable integrated brain circuits
abstract
Real time spike detection is the first critical step to develop spike-sorting for integrated brain circuits interface applications. Nonlinear Energy Operator (NEO) and absolute thresholding have been widely used as the spike detection algorithms where NEO has a better performance measured by the probability of detection and false alarm. This paper proposes a hybrid spike detection algorithm incorporating both spike detection algorithms to reduce the power and to keep the detection rate the same as that of NEO. In the proposed algorithm, the absolute thresholding is performed first to detect a potential spike. Once a potential spike is detected, NEO is executed to check whether the detected spike by absolute thresholding is valid. Since NEO is conditionally conducted, this reduces the overall power consumption. The simulation shows that the proposed hybrid method improves the power consumption by 54.48% compared to NEO in 65 nm CMOS technology.
Seyed Mohammad Ali Zeinolabedin, Anh-Tuan Do, Kiat Seng Yeo, Tony Tae-Hyoung Kim
ISCAS4
2015 A Ring-Oscillator-Based Reliability Monitor for Isolated Measurement of NBTI and PBTI in High-k/Metal Gate Technology
abstract
Ring-oscillator-based test structures that can separately measure the negative bias temperature instability (NBTI) and positive bias temperature instability (PBTI) degradation effects in digital circuits are presented for high-k metal gate devices. The mathematical derivation also shows that the structure for frequency degradation measurement can directly be used for estimating the portion of the NBTI and PBTI in the conventional ring oscillator. The proposed test structures including frequency degradation sensing circuitry have been implemented in an experimental high-k/metal gate SoI process.
Tony Tae-Hyoung Kim, Pong-Fei Lu, Keith A. Jenkins, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.1
2015 An Area- and Energy-Efficient FIFO Design Using Error-Reduced Data Compression and Near-Threshold Operation for Image/Video Applications
abstract
Many image/video processing algorithms require FIFO for filtering. The FIFO size is proportional to the length of the filters and input data width, causing large area and power consumption. We have proposed an energy- and area-efficient FIFO design for image/video applications through FIFO with error-reduced data compression (FERDC) and near-threshold operation. On architecture level, FERDC technique is proposed to reduce the size and power consumption of the FIFO by utilizing the spatial correlation between neighboring pixels and performing error-reduced data compression together with quantization to minimize the mean square error (MSE). On circuit level, near-threshold operation is adopted to achieve further power reduction while maintaining the required performance. To demonstrate the proposed FIFO, it has been implemented using a 0.18-μm CMOS process technology. The implementation covers different FIFO length, including 128, 256, 512, and 1024. The experimental results show that the proposed FIFO operating at 0.5 V and 28.57 MHz achieves up to 99%, 65%, and 34.91% reduction in dynamic power, leakage power, and area, respectively, with a small MSE of 2.76, compared with the conventional FIFO design. The proposed FIFO can be applied to a wide range of image/video signal processing applications to achieve high area and energy efficiency.
Seyed Mohammad Ali Zeinolabedin, Jun Zhou 0017, Xin Liu 0015, Tony Tae-Hyoung Kim
IEEE Trans. Very Large Scale Integr. Syst.4
2014 Design of SRAM PUF with improved uniformity and reliability utilizing device aging effect
abstract
SRAM Physical Unclonable Function (PUF) makes use of efficient silicon fabrication process where duplication of exact replica devices is difficult. One of the major issues with SRAM-PUF is the reliability and uniformity of the start-up pattern with environmental fluctuations. This paper presents a technique for improving uniformity (distribution of 1's & 0's) and reliability (variations in power-up patterns) of SRAM-PUF utilizing aging effects (mainly NBTI). The proposed technique maintains the uniformity of SRAM-PUF by controlling the polarity of the aging in SRAM arrays. The reliability is controlled by further injecting aging to the SRAM arrays after achieving target uniformity.
Achiranshu Garg, Tony Tae-Hyoung Kim
ISCAS2
2014 An area- and power-efficient FIFO with error-reduced data compression for image/video processing
abstract
Filtering is a key component of many digital image/video processing algorithms. It often requires FIFO to temporarily buffer the pixels data for later usage. The FIFO size is proportional to the length of the filters and input data width, causing large area and power consumption. This paper presents a technique named FIFO with error-reduced data compression (FERDC) to reduce the FIFO size for various filters. The proposed FERDC significantly reduces the area and power consumption while keeping the error metrics such as mean square error (MSE) and peak signal to noise ratio (PSNR) in the acceptable range. Simulation results of a two dimensional wavelet filter shows that the proposed FERDC technique achieves the FIFO size reduction of up to 44.44% with PSNR values larger than 39 dB, which leads to the reduction of at least 31.6% in the dynamic power and 44.44% in the leakage power.
Seyed Mohammad Ali Zeinolabedin, Jun Zhou 0017, Xin Liu 0015, Tony Tae-Hyoung Kim
ISCAS4
2013 Design of self-biased fully differential receiver and crosstalk cancellation for capacitive coupled vertical interconnects in 3DICs
abstract
Interconnect density in traditional capacitive coupling electrodes array is limited by capacitive crosstalk between electrodes. In this work, we propose an array structure where single-ended and differential pair electrodes (designed in the common-centroid structure) are alternatively placed horizontally and vertically. The proposed structure not only cancels out the crosstalk noise but also reduces the spacing requirement between electrodes. A novel self-biased fully differential receiver suppresses the common-mode coupled crosstalk in the differential electrodes. The receiver provides CMRR of 25dB and can recover wide band (10 kHz ~ 1 GHz) signals with 32dB gain. It consumes 25 μW at 1 GHz. It is designed and simulated in a 1.5V 0.13μm CMOS technology.
Myat Thu Linn Aung, Eric Teck Heng Lim, Takefumi Yoshikawa, Tony Tae-Hyoung Kim
ISCAS4
2013 An improved read/write scheme for anchorless NEMS-CMOS non-volatile memory
abstract
We proposed a NEMS-based anchorless memory structure with two stable mechanical states (Up and Down) for retaining data even at high operating temperate (>200°C) [5]. Compared to the conventional anchored devices, the anchorless structure offers better scalability and can operate at lower supply voltage, which is more desirable for integration with CMOS processes. This work addresses several issues in the previous NEM memory array implementation such as shuttle oscillation due to electrostatic pendulum and non-polarity write operation. We propose a 128 × 128 NEM memory array consisting of NEM memory cells and CMOS read/write control circuits. Our proposed read/write scheme with the two-level gate control eliminates the non-polarity write issue and achieves 32% power and 14% delay improvements when compared to the previous control scheme.
Anh-Tuan Do, Karthik G. Jayaraman, Vincent Pott, Chua Geng Li, Pushpapraj Singh, Kiat Seng Yeo, Tony Tae-Hyoung Kim
ISCAS7
2013 A 0.4V 7T SRAM with write through virtual ground and ultra-fine grain power gating switches
abstract
This paper presents a 7T near-threshold SRAM with design techniques for improving cell stability and energy efficiency. The proposed write through virtual ground (WTVG) scheme decreases the period of write disturbance by 6.1×. A PVT tracking sensing scheme is presented to track variation and sense small RBL swing. The ultra-fine grain power gating switches are implemented to minimize the redundant leakage caused by the storage of garbage data. The leakage suppression of 52% is achieved after the initial power-up. A 16 kb SRAM test chip was fabricated in a 65nm CMOS technology and showed the minimum energy of 2.01 pJ at 0.4 V.
Yuan Lin Yeoh, Bo Wang 0020, Xiangyao Yu, Tony Tae-Hyoung Kim
ISCAS4
2012 Design of ring oscillator structures for measuring isolated NBTI and PBTI
abstract
Ring oscillator based test structures that can separately measure the NBTI and PBTI degradation effects in digital circuits are presented for high-k metal-gate devices. The proposed test structures enable simultaneous stress of all devices under test in either NBTI or PBTI mode and measure frequency or threshold voltage shifts. The mathematical derivation also shows that the structure for frequency degradation measurement can directly be used for estimating the portion of the NBTI and PBTI in the conventional ring oscillator. The proposed test structures including beat frequency sensing circuitry have been designed in a 0.9V, 45nm SOI technology.
Tony Tae-Hyoung Kim, Pong-Fei Lu, Chris H. Kim
ISCAS1
2010 An On-Chip NBTI Sensor for Measuring pMOS Threshold Voltage Degradation
abstract
Negative bias temperature instability (NBTI) is one of the most critical device reliability issues in sub-130 nm CMOS processes. In order to better understand the characteristics of this mechanism, accurate and efficient means of measuring its effects must be explored. In this work, we describe an on-chip NBTI degradation sensor using a delay-locked loop (DLL), in which the increase in pMOS threshold voltage due to NBTI stress is translated into a control voltage shift in the DLL for high sensing gain. The proposed sensor is capable of supporting both DC and AC stress modes. Measurements from a test chip fabricated in a 130 nm bulk CMOS process show an average gain of 10 in the operating range of interest, with measurement times in tens of microseconds possible for minimal unwanted threshold voltage recovery. NBTI degradation readings across a range of operating conditions are presented to demonstrate the flexibility of this system.
John Keane 0001, Tony Tae-Hyoung Kim, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.2
2008 Circuit techniques for ultra-low power subthreshold SRAMs
abstract
Subthreshold operation has become an important area in applications where minimal power consumption and energy efficiency are the critical constraints. In particular, ultra-low power SRAM designs are critical for implementing such applications due to the large portion of the systems that they account for. However, sub-threshold SRAMs have many design issues such as cell stability, readability, and writability. In this paper, we give an overview of sub-threshold SRAM design issues and discuss several circuit techniques. We will focus on SRAM cell stability during read and write operation, improved writability, and read port circuits for the design of an ultra-low power sub-threshold SRAMs.
Tony Tae-Hyoung Kim, Jason Liu 0004, John Keane 0001, Chris H. Kim
ISCAS1
2008 A multi-story power delivery technique for 3D integrated circuits
abstract
Integrating circuits in the vertical direction can alleviate interconnect related problems and enable heterogeneous chips to be stacked in a single package with a small form factor. This paper addresses the power delivery issues in 3D chips revealing some interesting facts and design challenges. A multi-story power delivery technique that can reduce the worst case DC noise by 45% and lower the overhead power consumed in the power supply network by 65% is proposed. A test chip layout in an SOI process, showing a 5.3% area overhead, demonstrates the feasibility of the scheme.
Pulkit Jain, Tony Tae-Hyoung Kim, John Keane 0001, Chris H. Kim
ISLPED2
2008 Stack Sizing for Optimal Current Drivability in Subthreshold Circuits
abstract
Subthreshold circuit designs have been demonstrated to be a successful alternative when ultra-low power consumption is paramount. However, the characteristics of MOS transistors in the subthreshold region are significantly different from those in strong inversion. This presents new challenges in design optimization, particularly in complex gates with stacks of transistors. In this paper, we present a framework for choosing the optimal transistor stack sizing factors in terms of current drivability for subthreshold designs. We derive a closed-form solution for the correct sizing of transistors in a stack, both in relation to other transistors in the stack, and to a single device with equivalent current drivability. Simulation results show that our framework provides a performance benefit ranging up to more than 10% in certain critical paths.
John Keane 0001, Hanyong Eom, Tony Tae-Hyoung Kim, Sachin S. Sapatnekar, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.3
2007 An on-chip NBTI sensor for measuring PMOS threshold voltage degradation
abstract
Negative Bias Temperature Instability (NBTI) is one of the most critical device reliability issues facing scaled CMOS technology. In order to better understand the characteristics of this mechanism, accurate and efficient means of measuring its effects must be explored. In this work, we describe an on-chip NBTI degradation sensor using two delay-locked loops (DLL). The increase in PMOS transistor threshold due to NBTI stress is translated into the control voltage of a DLL for high sensing gain. Measurements from a 0.13μm test chip show a maximum gain of 16X in the operating range of interest, with microsecond order measurement times for minimal unwanted recovery. The proposed NBTI sensor also supports various DC and AC stress modes.
John Keane 0001, Tony Tae-Hyoung Kim, Chris H. Kim
ISLPED2
2007 Utilizing Reverse Short-Channel Effect for Optimal Subthreshold Circuit Design
abstract
The impact of the reverse short-channel effect (RSCE) on device current is stronger in the subthreshold region due to reduced drain-induced barrier lowering (DIBL) and the exponential dependency of current on threshold voltage. This paper describes a device-size optimization method for subthreshold circuits utilizing RSCE to achieve high drive current, low device capacitance, less sensitivity to random dopant fluctuations, better subthreshold swing, and improved energy dissipation. Simulation results using ISCAS benchmark circuits show that the critical path delay, power consumption, and energy consumption can be improved by up to 10.4%, 34.4%, and 41.2%, respectively.
Tony Tae-Hyoung Kim, John Keane 0001, Hanyong Eom, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.1
2006 Subthreshold logical effort: a systematic framework for optimal subthreshold device sizing
abstract
Subthreshold circuit designs have been demonstrated to be a successful alternative when ultra-low power consumption is paramount. However, the characteristics of MOS transistors in the subthreshold regime are significantly different from those in strong-inversion. This presents new challenges in design optimization, particularly in complex gates with stacks of transistors. In this paper, we demonstrate a new optimal sizing scheme for subthreshold designs which takes these issues into account. We derive a closed-form solution for the correct sizing of transistors in a stack, both in relation to other transistors in the stack, and to a single transistor with equivalent current drivability. Experimental results show that our framework provides a performance improvement of up to 13.5% over the conventional logical effort method on ISCAS benchmark circuits, while one component circuit demonstrated an improvement of 33.1%.
John Keane 0001, Hanyong Eom, Tony Tae-Hyoung Kim, Sachin S. Sapatnekar, Chris H. Kim
DAC3
2006 Utilizing reverse short channel effect for optimal subthreshold circuit design
abstract
The impact of the Reverse Short Channel Effect (RSCE) on device current is stronger in the subthreshold region due to the reduced Drain-Induced-Barrier-Lowering (DIBL) and the exponential dependency of current on threshold voltage. This paper describes a device size optimization method for subthreshold circuits utilizing RSCE to achieve high drive current, low device capacitance, less sensitivity to random dopant fluctuations, and better subthreshold swing. Simulation results using ISCAS benchmark circuits show that the critical path delay and power consumption can be improved by up to 10.4% and 34.4%, respectively.
Tony Tae-Hyoung Kim, Hanyong Eom, John Keane 0001, Chris H. Kim
ISLPED1