Chris H. Kim

dblp:31/4424 · DBLP profile ↗
← Back
74ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-4194-1347ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 71 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 A Hybrid Ising FPGA-COBI Architecture with Hardware-Based Problem Decomposition
abstract
Many combinatorial optimization problems map naturally to Ising Hamiltonians, $H({\text{s}}) = - \sum\nolimits_{i,j} {{J_{ij}}} {s_i}{s_j} - \sum\nolimits_i {{h_i}} {s_i}$ , and CMOS ring-oscillator Ising machines solve them in microseconds at milliwatts [1] , [2] . Their key limitation is capacity : the number of spins one solver core can process in a single solve. Because each hardware spin represents one binary Ising variable, capacity directly sets the largest problem solvable in one shot. Our 28 nm five-core COBI chip solves a 45-spin all-to-all subproblem per core in 77.5 µ s, so larger instances require iterative decomposition. This shifts the bottleneck from analog solving to digital orchestration: a CPU-based decomposer needs ∼321 µ s/iter over PCIe, 4× the core solve time, leaving the solver idle 84.9% of the time. We instead co-locate an FPGA decomposer with the chip and derive sizing laws for the required parallelism, achieving 1.93× geomean speedup and > 40× energy reduction vs. an optimized C++ baseline.
Ruihong Yin, Chaohui Li, Ahmet Efe, Abhimanyu Kumar, Ziqing Zeng, Ulya R. Karpuzcu, Sachin S. Sapatnekar, Chris H. Kim
FCCM9
2026 MIBID: Model Based Fault Diagnosis on Ising Machines
abstract
Model-Based Diagnosis (MBD) identifies faulty components in complex systems by reasoning over a model of expected behavior and observations. Computing minimal-cardinality diagnoses—those involving the smallest number of faulty components—is NP-hard and becomes challenging for large systems due to the combinatorial growth of possible fault combinations. SAT-based formulations provide a compact representation of diagnostic constraints, allowing the diagnosis problem to be expressed as a combinatorial optimization problem. Since Ising Machines are well suited for solving such problems, this representation offers a natural pathway for mapping MBD to Ising-based computation. In this work, we present MIBID, a framework that maps SAT-based MBD formulations to an Ising model for computing minimal-cardinality diagnoses. The framework also incorporates hardware-aware pre-processing and decomposition to adapt the formulation to the capabilities of an Ising Machine. Experimental results on a manufactured Ising Machine show competitive performance with SAT-based methods for single minimal diagnoses. For multiple-diagnosis tasks, MIBID enumerates up to \(34\%\) and \(43.59\%\) more diagnoses than state-of-the-art under the weak and strong fault models, respectively, thereby providing broader coverage of plausible fault explanations.
Nafisa Sadaf Prova, Ahmet Efe, Abhimanyu Kumar, Chris H. Kim, Sachin S. Sapatnekar, Ulya R. Karpuzcu
ACM Great Lakes Symposium on VLSI4
2026 SATIC: An Optimizing Ising Compiler for SAT(isfiability)
Ahmet Efe, M. Hüsrev Cilasun, Abhimanyu Kumar, Nafisa Sadaf Prova, Ziqing Zeng, Tahmida Islam, Ruihong Yin, Chaohui Li, Peter Kreye, Chris H. Kim, Sachin S. Sapatnekar, Ulya R. Karpuzcu
ISCA10
2023 Electromigration Assessment in Power Grids with Account of Redundancy and Non-Uniform Temperature Distribution
abstract
A recently proposed methodology for electromigration (EM) assessment in on-chip power/ground grid of integrated circuits has been validated by means of measurements, performed on dedicated test grids. IR drop degradation in the grid is used for defining the EM failure criteria. Physics-based models are involved for simulation of EM-induced stress evolution in interconnect structures, void formation and evolution, resistance increase of the voided segments, and consequent re-distribution of electric current in the redundant grid paths. A grid-like test structure, fabricated with a 65 nm technology and consisting of two metal layers, allowed to calibrate the voiding models by tracking voltage evolution in all grid nodes in experiment and in simulation. Good fit of the measured and simulated time-to-failure (TTF) probability distribution was obtained in both cases of uniform and non-uniform temperature distribution across the grid. The second test grid was fabricated with a 28 nm technology, consisted of 4 metal layers, and contained power and ground nets connected to "quasi-cells" with poly-resistors, which were specially designed for operating at elevated temperatures ~350°C. The existing current distributions resulted in different behavior of EM-induced failures in these nets: a gradual voltage evolution in power net, and sharp changes in ground net were observed in experiment, and successfully reproduced in simulations.
Armen Kteyan, Valeriy Sukharev, Alexander Volkov, Jun-Ho Choy, Farid N. Najm, Yong Hyeon Yi, Chris H. Kim, Stéphane Moreau
ISPD7
2022 Experimental Validation of a Novel Methodology for Electromigration Assessment in On-Chip Power Grids
abstract
A recently proposed theoretical methodology for the assessment of the electromigration (EM) induced IR-drop degradation in on-chip power/ground grids has been validated by means of measurements performed on real silicon. A voltage tapping technique was employed for the direct measurement of voltage variations at 162 nodes of the power net, stressed with 10 mA constant source current at an elevated temperature of 350 °C. A voltage drop between cathode and anode pads exceeding a specified threshold was considered as a failure. Times-to-failure (TTF) was measured on 19 packaged test grids and used for computing the mean TTF (MTTF). The EM-induced voltage degradation in this grid was also analyzed with an assessment methodology based on a simulation of stress evolution everywhere in the grid, resulting in a voiding in some of grid branches and corresponding resistance increase. A set of voiding compact models for different grid segments was developed and used in the simulations. The stochastic nature of the EM phenomenon was captured by introducing random distributions of atomic diffusivities and critical stresses across the grid and iterating them with Monte Carlo loops. A good fit between the measured voltage evolution kinetics at different grid nodes and that predicted by simulation, and the good agreement between measured and simulated failure distributions can be considered as the ever first experimental validation of this EM assessment methodology for on-chip power/ground (p/g) grids.
Valeriy Sukharev, Armen Kteyan, Farid N. Najm, Yong Hyeon Yi, Chris H. Kim, Jun-Ho Choy, Sofya Torosyan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2022 A 32Gb/s Time-Based PAM-4 Transceiver for High-Speed DRAM Interfaces With In-Situ Channel Loss and Bit-Error-Rate Monitors
abstract
A digital-intensive four-level pulse amplitude (PAM-4) transceiver featuring a 2-tap time-based decision feedback equalization (TB-DFE) circuit was demonstrated in a 65 nm GP CMOS process. A novel inverter-based differential voltage-to-time converter (DVTC) increases the linearity and dynamic range compared to a prior time-based DFE approach enabling reliable PAM-4 operation. The four-level signal comparison and DFE operation were performed entirely in the time domain using programmable delays and a phase detector (PD). Using an on-chip bit error rate (BER) monitor, we verified a BER less than 10−12while achieving an energy-efficiency of 0.97pJ/b at a 32Gb/s data rate. The transmitter (TX) and receiver (RX) circuits occupy an area of 0.009 mm2.
Po-Wei Chiu, Chris H. Kim
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 GeNVoM: Read Mapping Near Non-Volatile Memory
abstract
DNA sequencing is the physical/biochemical process of identifying the location of the four bases (Adenine, Guanine, Cytosine, Thymine) in a DNA strand. As semiconductor technology revolutionized computing, modern DNA sequencing technology (termed Next Generation Sequencing, NGS) revolutionized genomic research. As a result, modern NGS platforms can sequence hundreds of millions of short DNA fragments in parallel. The sequenced DNA fragments, representing the output of NGS platforms, are termed reads. Besides genomic variations, NGS imperfections induce noise in reads. Mapping each read to (the most similar portion of) a reference genome of the same species, i.e., read mapping, is a common critical first step in a diverse set of emerging bioinformatics applications. Mapping represents a search-heavy memory-intensive similarity matching problem, therefore, can greatly benefit from near-memory processing. Intuition suggests using fast associative search enabled by Ternary Content Addressable Memory (TCAM) by construction. However, the excessive energy consumption and lack of support for similarity matching (under NGS and genomic variation induced noise) renders direct application of TCAM infeasible, irrespective of volatility, where only non-volatile TCAM can accommodate the large memory footprint in an area-efficient way. This paper introduces GeNVoM, a scalable, energy-efficient and high-throughput solution. Instead of optimizing an algorithm developed for general-purpose computers or GPUs, GeNVoM rethinks the algorithm and non-volatile TCAM-based accelerator design together from the ground up. Thereby GeNVoM can improve the throughput by up to 3.67×; the energy consumption, by up to 1.36×, when compared to an ASIC baseline, which represents one of the highest-throughput implementations known.
S. Karen Khatamifard, Zamshed I. Chowdhury, Nakul Pande, Meisam Razaviyayn, Chris H. Kim, Ulya R. Karpuzcu
IEEE ACM Trans. Comput. Biol. Bioinform.5
2021 A Back-Sampling Chain Technique for Accelerated Detection, Characterization, and Reconstruction of Radiation-Induced Transient Pulses
abstract
Accurate characterization of radiation-induced soft errors is a critical step toward understanding the impact of these glitches on circuit and system reliability. With process scaling, there has been exponential increase in number of transistors that can be packed on a die which, in turn, results in higher sensitive node count and persistent soft error susceptibilities. In this work, a novel circuit technique employing higher sensitivity toward soft errors is proposed. The circuit makes use of current-starved gates with bias knobs to fine-tune both measurement resolution and strike sensitivity enabling accelerated and efficient induction of errors in a limited-time irradiation test environment. The back-sampling chain (BSC) circuit can measure individual radiation-induced transient pulse with as low amplitude as$0.3\times $VDD while maintaining a high measurement resolution for pulsewidth characterization. The bias knobs allowing tuning of sensitivity and resolution enable, for the first time, a strike pulse waveform reconstruction methodology that can be used to calibrate current pulse models for assessing soft error rate (SER) sensitivity of standard logic gates.
Saurabh Kumar 0003, Minki Cho, Luke R. Everson, Andres Malavasi, Dan Lake, Carlos Tokunaga, Muhammad M. Khellah, James W. Tschanz, Vivek De, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.10
2021 Wide-Range Many-Core SoC Design in Scaled CMOS: Challenges and Opportunities
abstract
The system-on-chip (SoC) designs for future Internet of Things (IoT) systems, spanning client platforms to cloud datacenters, need to deliver uncompromising and scalable performance with extreme energy efficiency for diverse workloads and applications, while satisfying a wide range of energy budgets, as well as platform cooling and power delivery constraints. Low-latency, burst-mode responsiveness, and scalable high-throughput performance must be delivered on demand for a range of thread-parallel, task-parallel, and data-parallel workloads covering traditional and emerging applications. This article discusses the challenges and opportunities for many-core SoC design in scaled CMOS process operating over a wide voltage-frequency range including near-threshold-voltage (NTV) that can meet the compute demands of the future at scale, flexibly, and efficiently. This article covers: 1) circuit design techniques for NTV cores; 2) mitigation techniques for within-die parameter variations via multivoltage frequency schemes; 3) digital integrated voltage regulators (VRs) for fine-grain and wide-range voltage modulation; and 4) radiation-induced soft error rate (SER) characterization and mitigation techniques to enable reliable operation at NTV. Silicon prototype examples will be used to illustrate the different techniques and highlight future research directions.
Sriram R. Vangal, Somnath Paul, Steven Hsu, Amit Agarwal 0001, Saurabh Kumar 0003, Ram Krishnamurthy 0001, Harish Krishnamurthy, James W. Tschanz, Vivek De, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.10
2018 Effect of aging on linear and nonlinear MUX PUFs by statistical modeling
abstract
This paper addresses the effect of aging on linear and non-linear MUX physical unclonable functions (PUFs). It is well known that a PUF response can be modeled in terms of the delay difference of MUX stages. In this paper, we show that the aging effects can be modeled in terms of variations in delay-difference and arbiter delay. Specifically, with aging, the percent delay-difference variation of each MUX stage can be modeled as a ratio of two correlated Gaussian random variables. This ratio distribution is shown to be approximately Gaussian with zero mean and variance increasing with time. In case of the arbiter, the ratio distribution is modeled as a Gaussian with positive mean. The paper makes three contributions: modeling the effect of aging in terms of percent variations in delay-difference of the MUX stages and arbiter delay, analysis of authentication accuracy with aging, and approaches to increase the PUF's lifetime by either recalibrating it to obtain new delay-difference parameters, or by tuning a threshold based on the total delay-difference. A general approach for selecting the threshold values is described in the paper. It is shown that the authentication accuracy of a PUF is significantly affected due to aging effects of the arbiter itself. Therefore, under the assumption that the variations in arbiter delay are considerably more than in delay-differences, the performance degradation in the case of aging alone is prominent compared to noise alone. We show that the authentication accuracy of a feed-forward PUF is more degraded compared to linear or modified feed-forward PUF. Metrics like Jenson-Shannon and Henze-Penrose divergence are also used to analyze the effect of aging.
Anoop Koyily, S. V. Sandeep Avvaru, Chris H. Kim, Keshab K. Parhi
ASP-DAC4
2018 Low-Energy Deep Belief Networks Using Intrinsic Sigmoidal Spintronic-based Probabilistic Neurons
abstract
A low-energy hardware implementation of deep belief network (DBN) architecture is developed using near-zero energy barrier probabilistic spin logic devices (p-bits), which are modeled to realize an intrinsic sigmoidal activation function. A CMOS/spin based weighted array structure is designed to implement a restricted Boltzmann machine (RBM). Device-level simulations based on precise physics relations are used to validate the sigmoidal relation between the output probability of a p-bit and its input currents. Characteristics of the resistive networks and p-bits are modeled in SPICE to perform a circuit-level simulation investigating the performance, area, and power consumption tradeoffs of the weighted array. In the application-level simulation, a DBN is implemented in MATLAB for digit recognition using the extracted device and circuit behavioral models. The MNIST data set is used to assess the accuracy of the DBN using 5,000 training images for five distinct network topologies. The results indicate that a baseline error rate of 36.8% for a 784x10 DBN trained by 100 samples can be reduced to only 3.7% using a 784x800x800x10 DBN trained by 5,000 input samples. Finally, Power dissipation and accuracy tradeoffs for probabilistic computing mechanisms using resistive devices are identified.
Ramtin Zand, Kerem Yunus Çamsari, Steven D. Pyle, Ibrahim Ahmed 0002, Chris H. Kim, Ronald F. DeMara
ACM Great Lakes Symposium on VLSI5
2018 BiometricNet: Deep Learning based Biometric Identification using Wrist-Worn PPG
abstract
Rapid advances in semiconductor fabrication technology have enabled the proliferation of miniaturized body-worn sensors capable of long term pervasive biomedical signal monitoring. In this paper, we present a novel deep learning-based framework (BiometricNET) on biometric identification using data collected from wrist-worn Photoplethysmography (PPG) signals in ambulatory environments. We have formulated a completely personalized data-driven approach, using a four-layer deep neural network - employing two convolution neural network (CNN) layers in conjunction with two long short-term memory (LSTM) layers, followed by a dense output layer for modelling the temporal sequence inherent within the pulsatile signal representative of cardiac activity. The proposed network configuration was evaluated on the TROIKA dataset collected from 12 subjects involved in physical activity, achieved an average five-fold cross-validation accuracy of 96%.
Luke R. Everson, Dwaipayan Biswas, Madhuri Panwar, Dimitrios Rodopoulos, Amit Acharyya, Chris H. Kim, Chris Van Hoof, Mario Konijnenburg, Nick Van Helleputte
ISCAS6
2018 Predicting Soft-Response of MUX PUFs via Logistic Regression of Total Delay Difference
abstract
This paper presents a logistic regression based approach to predict the soft-response for a challenge using the total delay-difference as an input. This approach enables us to determine whether a challenge is stable or not. Soft-response is the probability of response bit corresponding to the challenge being 1. The total delay-difference is computed from the input challenge by assuming that the delay-difference of the stages are known. The approach learns a logistic function based on the total delay-difference which has just 3 parameters. Therefore, this is a simple approach which gives comparable performance against a more complex approach based on artificial neural network (ANN) models. The model demonstrates good sensitivity and precision but poor specificity. Furthermore, we use scaling parameter of the logistic function to study its relation to the arbiter's timing parameters like setup and hold time.
Anoop Koyily, Chris H. Kim, Keshab K. Parhi
ISCAS3
2018 A Physical Unclonable Function based on Capacitor Mismatch in a Charge-Redistribution SAR-ADC
abstract
A Physical Unclonable Function (PUF) using capacitor mismatch in a standard successive approximation register analog-to-digital converter (SAR-ADC) as the entropy source is demonstrated in 65nm CMOS. SAR-ADCs are readily available in many system-on-chips, making the hardware overhead of the proposed PUF almost negligible. The inherent process variation of metal-oxide-metal (MOM) capacitors is harnessed through a charge redistribution operation which is sampled by the voltage comparator. To enhance the stability of the PUF output, soft response generation and dynamic thresholding techniques were adopted. Finally, we verify that performing the enrollment operation at a lower operating voltage can ensure that PUF responses are stable at the nominal supply voltage used during authentication.
Qianying Tang, Won Ho Choi, Luke R. Everson, Keshab K. Parhi, Chris H. Kim
ISCAS5
2018 Key-Based Dynamic Functional Obfuscation of Integrated Circuits Using Sequentially Triggered Mode-Based Design
abstract
This paper proposes a novel technique for hardware obfuscation termed dynamic functional obfuscation. Hardware obfuscation refers to a set of countermeasures used against IC counterfeiting and illegal overproduction. Traditionally, obfuscation encrypts semiconductor circuits using key inputs which must be set to a correct value to operate the circuit correctly. By keeping the key values secret during the manufacturing process, any attempt by unauthorized parties to overproduce chips or pirate designs is thwarted. The proposed dynamic technique differs from existing fixed obfuscation schemes as the obfuscating signals change over time. This results in inconsistent circuit behavior upon input of incorrect key, where the chip operates correctly sometimes and fails sometimes. The advantage of dynamic obfuscation is that it results in stronger obfuscation by increasing the time complexity of deciphering the correct key using brute-force attack, even with shorter keys. Moreover, the dynamic nature of these circuits also makes them resistant to reverse engineering and SAT solver-based attacks. To achieve dynamic obfuscation, ideas from hardware Trojan literature and sequentially triggered counters are utilized. A demonstration of obfuscation on sequential circuits implementing fast Fourier transform (FFT) algorithm and Ethernet IP shows low overall area and power overheads of less than 1%. Security in terms of time to attack for the FFT circuit (for a key size of 30 bits and a system operating at 100 MHz) is increased to 1021,055 years using dynamic obfuscation compared with only 5.36 s using fixed obfuscation schemes. For the Ethernet IP core, time to attack of dynamic obfuscation with a key size of 32 bits is 1046,423,135 years compared with 21.47s with fixed obfuscation. It is also shown that for a key size of K bits, the lower bound for time to attack using brute-force is proportional to K2Kand K22Kfor the proposed design using one and two random number generators, respectively.
Sandhya Koteshwara, Chris H. Kim, Keshab K. Parhi
IEEE Trans. Inf. Forensics Secur.2
2017 A Pathway to Enable Exponential Scaling for the Beyond-CMOS Era: Invited
abstract
Many key technologies of our society, including so-called artificial intelligence (AI) and big data, have been enabled by the invention of transistor and its ever-decreasing size and ever-increasing integration at a large scale. However, conventional technologies are confronted with a clear scaling limit. Many recently proposed advanced transistor concepts are also facing an uphill battle in the lab because of necessary performance tradeoffs and limited scaling potential. We argue for a new pathway that could enable exponential scaling for multiple generations. This pathway involves layering multiple technologies that enable new functions beyond those available from conventional and newly proposed transistors. The key principles for this new pathway have been demonstrated through an interdisciplinary team effort at C-SPIN (a STARnet center), where systems designers, device builders, materials scientists and physicists have all worked under one umbrella to overcome key technology barriers. This paper reviews several successful outcomes from this effort on topics such as the spin memory, logic-in-memory, cognitive computing, stochastic and probabilistic computing and reconfigurable information processing.
Jianping Wang 0006, Sachin S. Sapatnekar, Chris H. Kim, Paul A. Crowell, Steven J. Koester, Supriyo Datta, Kaushik Roy 0001, Anand Raghunathan, Xiaobo Sharon Hu, Michael T. Niemier, Azad Naeemi, Chia-Ling Chien, Caroline A. Ross, Roland Kawakami
DAC3
2017 Secure and Reliable XOR Arbiter PUF Design: An Experimental Study based on 1 Trillion Challenge Response Pair Measurements
abstract
This paper shows that performing an XOR operation between the outputs of parallel arbiter PUFs generates a more secure output at the expense of reduced stability. In this work, we evaluate the security and stability of XOR PUFs using 1,000,000 randomly chosen challenges, applied to 10 custom-designed PUF chips, tested for 100,000 cycles per challenge, under different voltage and temperature conditions. Based on extensive hardware data, we propose a practical method for selecting challenges that will produce stable responses. A linear regression approach based on soft responses collected during enrollment phase was used to build accurate models for each individual arbiter PUF. Hardware data from fabricated chips verify that the approach is highly effective.
Keshab K. Parhi, Chris H. Kim
DAC3
2017 Advanced spintronic memory and logic for non-volatile processors
abstract
Many ultra-low power Internet of things (IoT) systems may be powered by energy harvested from ambient sources (e.g., solar radiation, thermal gradients, and WiFi). However, these energy sources can vary significantly in terms of their strengths and on/off patterns. For volatile systems, the intermittent nature of the energy sources necessitates the use of backup/recovery schemes to guarantee computational correctness and forward progress, which incur performance, area and energy overhead. Non-volatile (NV) processors based on spintronic devices, such as Spin-Transfer Torque (STT) memory and All-Spin-Logic (ASL), are more attractive alternatives. These NV devices are capable of achieving forward progress without relying on backup/recovery schemes. This work establishes a general framework for evaluating NV device-based processors for energy harvesting applications. Results demonstrate that NV spintronic processors can achieve significant energy savings (up to 83 x) versus a hybrid CMOS (computation) and STT-RAM (backup) implementation.
Robert Perricone, Ibrahim Ahmed 0002, Zhaoxin Liang, Meghna G. Mankalale, Xiaobo Sharon Hu, Chris H. Kim, Michael T. Niemier, Sachin S. Sapatnekar, Jianping Wang 0006
DATE6
2017 Hierarchical functional obfuscation of integratec circuits using a mode-based approach
abstract
Hardware obfuscation has been proposed as a hardware security measure against reverse engineering, intellectual property (IP) piracy and integrated circuits (IC) overbuilding. In this paper, we present a novel method of obfuscation using a hierarchical approach. In the design flow, IP vendors obfuscate their designs using a set of keys and provide these keys to the design house. The design house then integrates all the IPs and adds its own keys to create a complete obfuscated system. This prevents both misuse of IPs and illegal use of ICs since only secure parties have access to the correct keys. The obfuscation at each level is performed using a mode-based approach in which the design can operate in meaningful and non-meaningful modes. The design is functionally correct in only one mode. An attacker needs to work through different levels of the design to correctly decipher its operation and correct working mode. Since each of the IPs can work in multiple meaningful modes, the attack becomes more difficult as the number of IPs increases. These ideas are demonstrated using a convolution architecture with fast Fourier transform (FFT) blocks. With only about 13% area and 15% power overhead over an unobfuscated design, it is shown that the proposed design has the flexibility to be obfuscated with different key sizes and overheads depending on the level of security.
Sandhya Koteshwara, Chris H. Kim, Keshab K. Parhi
ISCAS2
2017 An entropy test for determining whether a MUX PUF is linear or nonlinear
abstract
This paper proposes a novel entropy test to determine whether a MUX PUF is linear or not. Three MUX PUF configurations are considered, namely linear, feed-forward and modified feed-forward. In addition to these, we also consider feed-forward structures like overlap, cascade and separate configurations. The approach is focused on computing the conditional entropy of responses to a set of predefined challenges. The challenge set consists of randomly chosen challenges and their 1-bit neighbors. The entropy is computed across the responses of two 1-bit neighboring challenges. For non-linear MUX PUFs like feed-forward, the method determines the MUX stages which are controlled by internally generated challenge bits as opposed to external challenge bits. This is based on the observation that the conditional entropy for each of these stages is zero. Also, the number of zero conditional entropy values across the MUX stages provide an upper bound on the number of internal arbiters present in the PUF. With the proposed approach, we observe 100% sensitivity and 100% specificity for identifying non-linearity. Furthermore, we show that the proposed approach requires very less number of stable random challenges (about 50) for successfully determining whether a PUF is linear or not for real chips.
Anoop Koyily, Chris H. Kim, Keshab K. Parhi
ISCAS3
2017 A multi-phase VCO quantizer based adaptive digital LDO in 65nm CMOS technology
abstract
A digital low-dropout (DLDO) voltage regulator circuit is proposed utilizing a multi-phase VCO based time quantizer. This high-resolution quantizer requires much lower sampling clock frequency compared to the previously proposed 1-bit comparator based architectures and thereby ensures stability over a wide operating condition, while reducing the dynamic power consumption at the same time. The DLDO operates at an input voltage range of 0.6V to 1.2V and delivers a maximum 115mA current with a 50mV dropout, simulated in a 65nm LP CMOS technology. A dynamically adaptive sampling clock can reduce the output voltage droop by 40-60% and provides 3.5-6.5 times faster settling compared to a baseline DLDO design that uses a fixed sampling clock frequency. The FOM calculated at 0.9V output is 0.53ps. The maximum current efficiency is 99.3%.
Somnath Kundu, Chris H. Kim
ISCAS2
2017 A data remanence based approach to generate 100% stable keys from an SRAM physical unclonable function
abstract
The start-up value of an SRAM cell is unique, random, and unclonable as it is determined by the inherent process mismatch between transistors. These properties make SRAM an attractive circuit for generating encryption keys. The primary challenge for SRAM based key generation, however, is the poor stability when the circuit is subject to random noise, temperature and voltage changes, and device aging. Temporal majority voting (TMV) and bit masking were used in previous works to identify and store the location of unstable or marginally stable SRAM cells. However, TMV requires a long test time and significant hardware resources. In addition, the number of repetitive power-ups required to find the most stable cells is prohibitively high. To overcome the shortcomings of TMV, we propose a novel data remanence based technique to detect SRAM cells with the highest stability for reliable key generation. This approach requires only two remanence tests: writing `1' (or `0') to the entire array and momentarily shutting down the power until a few cells flip. We exploit the fact that the cells that are easily flipped are the most robust cells when written with the opposite data. The proposed method is more effective in finding the most stable cells in a large SRAM array than a TMV scheme with 1,000 power-up tests. Experimental studies show that the 256-bit key generated from a 512 kbit SRAM using the proposed data remanence method is 100% stable under different temperatures, power ramp up times, and device aging.
Muqing Liu 0001, Qianying Tang, Keshab K. Parhi, Chris H. Kim
ISLPED5
2017 Reliable PUF-Based Local Authentication With Self-Correction
abstract
Physical unclonable functions (PUFs) can extract chip-unique signatures from integrated circuits (ICs) by exploiting the uncontrollable randomness due to manufacturing process variations. These signatures can then be used for many hardware security applications including authentication, anti-counterfeiting, IC metering, signature generation, and obfuscation. However, most of these applications require error correcting methods to produce consistent PUF responses across different environmental conditions. This paper presents a novel method to enable lightweight, secure, and reliable PUF-based authentication. A two-level finite-state machine (FSM) is proposed to correct erroneous bits generated by environmental variations (e.g., temperature, voltage, and aging variations). In the proposed method, each PUF response is mapped to a key during design phase. The actual key can be determined from the PUF response only after the chip is fabricated. Because the key is not known to the foundry, the proposed approach prevents counterfeiting. The performance of the proposed method and other applications are also discussed. Our experimental results show that the cost of the proposed self-correcting two-level FSM is significantly less than that of the commonly used error correcting codes. It is shown that the proposed self-correcting FSM consumes about 2× to 10× less area and about 20× to 100× less power than the Bose-Chaudhuri-Hochquenghem codes.
Yingjie Lao, Bo Yuan 0001, Chris H. Kim, Keshab K. Parhi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2016 Estimating delay differences of arbiter PUFs using silicon data
S. V. Sandeep Avvaru, Saroj Satapathy, Yingjie Lao, Chris H. Kim, Keshab K. Parhi
DATE5
2016 Soft Response Generation and Thresholding Strategies for Linear and Feed-Forward MUX PUFs
abstract
In this work, we present probability based response generation schemes for MUX based Physical Unclonable Functions (PUFs). Compared to previous implementations where temporal majority voting (TMV) based on limited samples and coarse criteria was utilized to determine final responses, our design can collect soft responses with detailed probability information using simple on-chip circuits. Thresholds with fine accuracy are applied to efficiently distinguish stable and unstable challenge response pairs (CRPs). A 32nm test chip including both linear and feed-forward MUX PUFs was implemented for concept verification. Based on a detailed analysis of the hardware data, we propose several enhanced thresholding strategies for determining stable CRPs. For instance, a stringent threshold can be imposed in enrollment phase for selecting good CRPs, while a relaxed threshold can be used during normal authentication phase. Experimental data shows a high degree of uniqueness and randomness in the PUF responses which can be attributed to the carefully optimized circuit layout. Finally, output characteristic of a feed-forward MUX PUF was compared to that of a standard linear MUX PUF from the same 32nm chip.
Saroj Satapathy, Yingjie Lao, Keshab K. Parhi, Chris H. Kim
ISLPED5
2016 Beat Frequency Detector-Based High-Speed True Random Number Generators: Statistical Modeling and Analysis
abstract
True random number generators (TRNGs) are crucial components for the security of cryptographic systems. In contrast to pseudo--random number generators (PRNGs), TRNGs provide higher security by extracting randomness from physical phenomena. To evaluate a TRNG, statistical properties of the circuit model and raw bitstream should be studied. In this article, a model for the beat frequency detector--based high-speed TRNG (BFD-TRNG) is proposed. The parameters of the model are extracted from the experimental data of a test chip. A statistical analysis of the proposed model is carried out to derive mean and variance of the counter values of the TRNG. Our statistical analysis results show that mean of the counter values is inversely proportional to the frequency difference of the two ring oscillators (ROSCs), whereas the dynamic range of the counter values increases linearly with standard deviation of environmental noise and decreases with increase of the frequency difference. Without the measurements from the test data, a model cannot be created; similarly, without a model, performance of a TRNG cannot be predicted. The key contribution of the proposed approach lies in fitting the model to measured data and the ability to use the model to predict performance of BFD-TRNGs that have not been fabricated. Several novel alternate BFD-TRNG architectures are also proposed; these include parallel BFD, cascade BFD, and parallel-cascade BFD. These TRNGs are analyzed using the proposed model, and it is shown that the parallel BFD structure requires less area per bit, whereas the cascade BFD structure has a larger dynamic range while maintaining the same mean of the counter values as the original BFD-TRNG. It is shown that 3.25 M and 4 M random bits can be obtained per counter value from parallel BFD and parallel-cascade BFD, respectively, where M counter values are computed in parallel. Furthermore, the statistical analysis results illustrate that BFD-TRNGs have better randomness and less cost per bit than other existing ROSC-TRNG designs. For example, it is shown that BFD-TRNGs accumulate 150% more jitter than the original two-oscillator TRNG and that parallel BFD-TRNGs require one-third power and one-half area for same number of random bits for a specified period.
Yingjie Lao, Qianying Tang, Chris H. Kim, Keshab K. Parhi
ACM J. Emerg. Technol. Comput. Syst.3
2016 System-Level Power Analysis of a Multicore Multipower Domain Processor With ON-Chip Voltage Regulators
abstract
In this paper, we study two different ON-chip power delivery schemes, namely, fully integrated voltage regulator (FIVR) and low-dropout regulator (LDO), and analyze their effect on total system power under process variation, assuming a realistic dynamic voltage-frequency scaling (DVFS) system. The impact of different task scheduling algorithms on the overall system power was also analyzed. We find that in a hypothetical 256-core processor, under a per-core DVFS assumption, the FIVR-based power delivery consumes 20% less power than the LDO-based one for a 50% throughput. However, as the number of cores in the processor reduces, the difference in power consumption between the FIVR-based and LDO-based power delivery schemes becomes smaller. For example, in the case of a 16-core processor with per-core DVFS capability, FIVR-based design was found to consume about the same power as the LDO-based design.
Ayan Paul, Sang Phill Park, Dinesh Somasekhar, Young Moon Kim, Nitin Borkar, Ulya R. Karpuzcu, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.7
2015 Fault-tolerant ripple-carry binary adder using partial triple modular redundancy (PTMR)
abstract
Integrated circuit chips fabricated using nano-scale CMOS technologies will be prone to errors caused by fluctuations in threshold voltage, supply voltage, electromigration, random dopant fluctuations, aging, timing errors and soft errors. Design of nano-scale failure-resistant systems has drawn significant interest in past few years. One common approach to reducing errors is the use of triple modular redundancy (TMR). The hardware overhead associated with TMR is significantly high. This paper presents a novel partial triple modular redundancy (PTMR) approach that achieves the same or better fault-tolerance as that of TMR but with significantly less hardware overhead. In a weighted number system, the most significant bits carry greater weight and preserving these bits is more critical than the lower significant bits. In PTMR, only the P most significant bits of the result are computed using TMR as opposed to all the W bits, where W represents the word-length of the operands. The proposed PTMR approach is illustrated in the context of a ripple-carry adder. It is shown that the hardware overhead can be reduced by 75% to 87.5% with P = 4 as the word-length varies from 16 to 32, with average error power equal to or less than that of TMR. It is shown that P = 3 or 4 is sufficient for word-lengths varying from 16 to 32.
Rahul Parhi, Chris H. Kim, Keshab K. Parhi
ISCAS2
2015 Spin-Based Computing: Device Concepts, Current Status, and a Case Study on a High-Performance Microprocessor
abstract
As the end draws near for Moore's law, the search for low-power alternatives to complementary metal-oxide-semiconductor (CMOS) technology is intensifying. Among the various post-CMOS candidates, spintronic devices have gained special attention for their potential to overcome the power and performance limitations of CMOS. In particular, all spin logic (ASL) technology, which performs Boolean operations and transfers the output in the spin domain, has been proposed for enabling new capabilities-such as high density, low device count, and nonvolatility-that were previously impossible with CMOS technology. In this paper, first we provide an overview of the history and the current status of the various spintronic devices being pursued by the research community. Then, we describe how spin-based components are integrated into a computing system and the advantages that result. We use a hypothetical spintronic-based Intel Core i7 as a test vehicle to compare the system-level power requirements of ASL- and CMOS-based systems, taking into consideration the unique demands of spin-based interconnects. We conclude with a brief analysis of current limitations and future directions of spintronic research.
Jongyeon Kim, Ayan Paul, Paul A. Crowell, Steven J. Koester, Sachin S. Sapatnekar, Jianping Wang 0006, Chris H. Kim
Proc. IEEE7
2015 A Ring-Oscillator-Based Reliability Monitor for Isolated Measurement of NBTI and PBTI in High-k/Metal Gate Technology
abstract
Ring-oscillator-based test structures that can separately measure the negative bias temperature instability (NBTI) and positive bias temperature instability (PBTI) degradation effects in digital circuits are presented for high-k metal gate devices. The mathematical derivation also shows that the structure for frequency degradation measurement can directly be used for estimating the portion of the NBTI and PBTI in the conventional ring oscillator. The proposed test structures including frequency degradation sensing circuitry have been implemented in an experimental high-k/metal gate SoI process.
Tony Tae-Hyoung Kim, Pong-Fei Lu, Keith A. Jenkins, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.4
2015 The Dependence of BTI and HCI-Induced Frequency Degradation on Interconnect Length and Its Circuit Level Implications
abstract
The dependence of bias temperature instability (BTI) and hot carrier injection (HCI)-induced frequency degradation on interconnect length has been examined for the first time. Experimental data from 65-nm test chips show that frequency degradation due to BTI decreases monotonically for longer wires because of the shorter effective stress time, while the HCI-induced component has a nonmonotonic relationship with interconnect length due to the combined effect of increased effective stress time and decreased effective stress voltage. Simple aging models are proposed to capture the unique BTI and HCI behavior in global interconnect drivers. A closed-loop simulation methodology that takes into consideration the interplay between the frequency degradation and the stress parameters (such as stress duration and stress voltage) is used to determine the optimal repeater count and sizing in practical interconnect circuits.
Qianying Tang, Pulkit Jain, Dong Jiao, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.5
2014 Improving STT-MRAM density through multibit error correction
abstract
STT-MRAMs are prone to data corruption due to inadvertent bit flips. Traditional methods enhance robustness at the cost of area/energy by using larger cell sizes to improve the thermal stability of the MTJ cells. This paper employs multibit error correction with DRAM-style refreshing to mitigate errors and provides a methodology for determining the optimal level of correction. A detailed analysis demonstrates that the reduction in nonvolatility requirements afforded by strong error correction translates to significantly lower area for the memory array compared to simpler ECC schemes, even when accounting for the increased overhead of error correction.
Brandon Del Bel, Jongyeon Kim, Chris H. Kim, Sachin S. Sapatnekar
DATE3
2014 Distributed On-Chip Switched-Capacitor DC-DC Converters Supporting DVFS in Multicore Systems
abstract
Dynamic voltage and frequency scaling (DVFS) is a powerful technique to reduce power consumption in a chip multiprocessor. To support DVFS in the multicore power delivery network, we integrate on-chip switched-capacitor (SC) dc-dc converters that can work with multiple conversion ratios to provide varying levels of Vdd supplies. We study the application of such SC converters in multicore chips by simulation. Our results show that distributed SC converters can significantly reduce the voltage droop seen by the local core loads by providing better localized power regulation. Considering the fact that the current distribution in a multicore chip is unbalanced, we further develop computer-aided design techniques to automate the design (size) and distribution (number and location) of these SC converters, using the efficiency of the whole power delivery system as the optimization metric. This is a major concern, but has not been addressed at the system level in prior research. We develop models for the power loss of such a system as a function of size and distribution of the SC converters, then propose an approach to optimize the SC converters to maximize the efficiency of the system, while considering all the possible conversion ratios an SC converter can work with. We verify the accuracy of our models for the power loss in the power delivery system, and demonstrate the efficiency of our techniques to optimize the SC converters on both homogenous and heterogenous multicore chips.
Pingqiang Zhou, Ayan Paul, Chris H. Kim, Sachin S. Sapatnekar
IEEE Trans. Very Large Scale Integr. Syst.3
2012 Optimization of on-chip switched-capacitor DC-DC converters for high-performance applications
abstract
On-chip switched-capacitor (SC) DC-DC converters have recently been demonstrated in silicon for high-performance applications such as multicore processors. The efficiency of the power delivery system using SC converters is a major concern, but this has not been addressed at the system level in prior research. This work develops models for the efficiency of such a system as a function of size and layout of the SC converters, and proposes an approach to optimize the size and layout of the SC converter to minimize power loss. The efficiency of these techniques is demonstrated on both homogenous and heterogenous multicore chips.
Pingqiang Zhou, Won Ho Choi, Bongjin Kim, Chris H. Kim, Sachin S. Sapatnekar
ICCAD4
2012 Design of ring oscillator structures for measuring isolated NBTI and PBTI
abstract
Ring oscillator based test structures that can separately measure the NBTI and PBTI degradation effects in digital circuits are presented for high-k metal-gate devices. The proposed test structures enable simultaneous stress of all devices under test in either NBTI or PBTI mode and measure frequency or threshold voltage shifts. The mathematical derivation also shows that the structure for frequency degradation measurement can directly be used for estimating the portion of the NBTI and PBTI in the conventional ring oscillator. The proposed test structures including beat frequency sensing circuitry have been designed in a 0.9V, 45nm SOI technology.
Tony Tae-Hyoung Kim, Pong-Fei Lu, Chris H. Kim
ISCAS3
2011 An Array-Based Test Circuit for Fully Automated Gate Dielectric Breakdown Characterization
abstract
We propose an array-based test circuit for efficiently characterizing gate dielectric breakdown. Such a design is highly beneficial when studying this statistical process, where up to thousands of samples are needed to create an accurate time to breakdown Weibull distribution. The proposed circuit also facilitates investigations of any spatial correlation of dielectric failures, and can monitor a progressive decrease in gate resistance. Measurement results are presented from a 32 × 32 test array implemented in a 130-nm bulk CMOS process. Results show that this system is capable of taking accurate measurements across a range of voltages and temperatures, which is critical for extrapolating accelerated stress experiment results to expected device lifetimes under realistic operating conditions.
John Keane 0001, Shrinivas Venkatraman, Paulo F. Butzen, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.4
2011 Adaptive Techniques for Overcoming Performance Degradation Due to Aging in CMOS Circuits
abstract
Negative bias temperature instability (NBTI) in pMOS transistors has become a major reliability concern in present-day digital circuit design. Further, with the recent introduction of Hf-based high-k dielectrics for gate leakage reduction, positive bias temperature instability (PBTI), the dual effect in nMOS transistors, has also reached significant levels. Consequently, designs are required to build in substantial guardbands in order to guarantee reliable operation over the lifetime of a chip, and these involve large area and power overheads. In this paper, we begin by proposing the use of adaptive body bias (ABB) and adaptive supply voltage (ASV) to maintain optimal performance of an aged circuit, and demonstrate its advantages over a guard banding technique such as synthesis. We then present a hybrid approach, utilizing the merits of both ABB and synthesis, to ensure that the resultant circuit meets the performance constraints over its lifetime, and has a minimal area and power overhead, as compared with a nominally designed circuit.
Sanjay V. Kumar, Chris H. Kim, Sachin S. Sapatnekar
IEEE Trans. Very Large Scale Integr. Syst.2
2010 Logic-compatible embedded DRAM design for memory intensive low power systems
abstract
Circuit techniques for enabling a low power logic-compatible embedded DRAM (eDRAM) are presented. A boosted 3T gain cell utilizes preferential storage node boosting to improve data retention time and increase read margin. A regulated bit-line write scheme is equipped with a steady-state storage node voltage monitor to overcome the data `1' write disturbance problem. An adaptive and die-to-die adjustable read reference bias generator is proposed to cope with PVT variations. Measurement data from 65nm test chips demonstrate a >1.0msec retention time at 0.9V, 85°C and a <;100μW per Mb refresh power at 1.0V, 85°C which translates into a 50% reduction in static power compared to a power gated SRAM.
Ki Chul Chun, Pulkit Jain, Chris H. Kim
ISCAS3
2010 Variation aware performance analysis of gain cell embedded DRAMs
abstract
Gain cell embedded DRAMs are twice as dense as 6T SRAMs, are logic compatible, have decoupled read and write paths providing good low voltage margin, and can drive long bitlines with gain. In this work, we present a variation study of gain cell eDRAM performance using an industrial 1.2V, 65nm low power CMOS process. Two methods are proposed to analyze eDRAM performance which can be used for designing variation tolerant eDRAM circuits, developing redundancy techniques, and guiding the device optimization procedure.
Wei Zhang 0032, Ki Chul Chun, Chris H. Kim
ISLPED3
2010 An On-Chip NBTI Sensor for Measuring pMOS Threshold Voltage Degradation
abstract
Negative bias temperature instability (NBTI) is one of the most critical device reliability issues in sub-130 nm CMOS processes. In order to better understand the characteristics of this mechanism, accurate and efficient means of measuring its effects must be explored. In this work, we describe an on-chip NBTI degradation sensor using a delay-locked loop (DLL), in which the increase in pMOS threshold voltage due to NBTI stress is translated into a control voltage shift in the DLL for high sensing gain. The proposed sensor is capable of supporting both DC and AC stress modes. Measurements from a test chip fabricated in a 130 nm bulk CMOS process show an average gain of 10 in the operating range of interest, with measurement times in tens of microseconds possible for minimal unwanted threshold voltage recovery. NBTI degradation readings across a range of operating conditions are presented to demonstrate the flexibility of this system.
John Keane 0001, Tony Tae-Hyoung Kim, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.3
2009 Adaptive techniques for overcoming performance degradation due to aging in digital circuits
abstract
Negative bias temperature instability (NBTI) in PMOS transistors has become a major reliability concern in present-day digital circuit design. Further, with the recent usage of Hf-based high-k dielectrics for gate leakage reduction, positive bias temperature instability (PBTI), the dual effect in NMOS transistors has also reached significant levels. Consequently, designers are required to build in substantial guard-bands into their designs, leading to large area and power overheads, in order to guarantee reliable operation over the lifetime of a chip. We propose a guard-banding technique based on adaptive body bias (ABB) and adaptive supply voltage (ASV), to recover the performance of an aged circuit, and compare its merits over previous approaches.
Sanjay V. Kumar, Chris H. Kim, Sachin S. Sapatnekar
ASP-DAC2
2009 A 0.9V, 65nm logic-compatible embedded DRAM with > 1ms data retention time and 53% less static power than a power-gated SRAM
abstract
A logic-compatible low power eDRAM is demonstrated in 65nm CMOS achieving a retention time of 1.25msec and a static power dissipation of 91.3µW/Mb at 0.9V, 85ºC. A boosted 3T gain cell enhances data retention time and read speed. A regulated bit-line write scheme and a read reference bias generator mitigate write disturbance issues and improve tolerance to PVT variations.
Ki Chul Chun, Pulkit Jain, Chris H. Kim
ISLPED3
2009 Sleep Transistor Sizing and Adaptive Control for Supply Noise Minimization Considering Resonance
abstract
The conventional sleep transistor sizing schemes do not consider the resonant supply noise which represents the worst-case supply disturbance. This paper investigates the impact of sleep transistor sizing on different on-chip noise components and shows that, contrary to the conventional wisdom, a larger sleep transistor is not always favored in term of performance when the resonant supply noise is taken into account. To minimize the worst-case supply noise, an optimal sizing scheme using an explicit noise and impedance model is developed and verified by benchmark circuits. Employing the proposed technique results in a reduction of the worst-case noise by 19%, as well as a saving of standby leakage and area overhead by 60% in comparison with conventional sizing scheme. In order to deal with the sporadic nature of the resonant, we propose an adaptive sleep transistor circuit which adjusts the size of sleep transistor on the fly to remove the DC noise penalty of the fixed sizing scheme. Simulation results on 32-nm CMOS technology are used to demonstrate the functionality and effectiveness of the proposed adaptive sizing circuits.
Jie Gu 0003, Hanyong Eom, John Keane 0001, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.4
2009 Design and Implementation of Active Decoupling Capacitor Circuits for Power Supply Regulation in Digital ICs
abstract
Control of on-chip power supply noise has become a major challenge for continuous scaling of CMOS technology. Conventional passive decoupling capacitors (decaps) exhibit significant area and leakage penalties. To improve the efficiency of power supply regulation, this paper proposes a distributed active decap circuit for use in digital integrated circuits (ICs). The proposed design uses an operational amplifier to boost the performance of conventional decaps. Simulations proved its enhanced decoupling effect in comparison with passive decaps. The proposed active decap also shows advantages in providing additional damping to the on-chip resonant noise. To verify the performance from the proposed circuit, a 0.18-mu m test chip with on-chip noise generators and sensors has been fabricated. Measurements show a 4-11times boost in decap value over conventional passive decaps for frequencies up to 1 GHz with a total area saving of 40%. Local supply noise distribution and decap gating capability were also examined from the test chip.
Jie Gu 0003, Ramesh Harjani, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.3
2009 Fuer Chris H. Kim 2 Eintraege in Db, Chris H. Kim und Chris Kim. Identisch. Siehe EE-Links: Univ. of Minnesota. Modeling, Analysis, and Application of Leakage Induced Damping Effect for Power Supply Integrity
abstract
Leakage power is becoming the dominant component of chip power consumption with continued CMOS scaling. An important but commonly unnoticed fact is that leaky transistors act as resistors that help dampen the mid-frequency power supply noise. This paper focuses on the damping effect of various on-chip current components including the leakage current which becomes significant in scaled technologies. By developing physics-based damping models for active and leakage currents, we show that leakage, particularly gate tunneling leakage, provides more damping than strong-inversion current. The proposed models were validated in a 32-nm predictive CMOS technology under process-voltage-temperature (PVT) variations. Examples on large circuits such as SRAM caches are shown to illustrate the application of the proposed model. Simulation results show that the leakage induced damping effect can compensate the speed degradation at high temperatures by 7% or offer 61% saving in decap area and leakage power.
Jie Gu 0003, John Keane 0001, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.3
2008 Circuit techniques for ultra-low power subthreshold SRAMs
abstract
Subthreshold operation has become an important area in applications where minimal power consumption and energy efficiency are the critical constraints. In particular, ultra-low power SRAM designs are critical for implementing such applications due to the large portion of the systems that they account for. However, sub-threshold SRAMs have many design issues such as cell stability, readability, and writability. In this paper, we give an overview of sub-threshold SRAM design issues and discuss several circuit techniques. We will focus on SRAM cell stability during read and write operation, improved writability, and read port circuits for the design of an ultra-low power sub-threshold SRAMs.
Tony Tae-Hyoung Kim, Jason Liu 0004, John Keane 0001, Chris H. Kim
ISCAS4
2008 A multi-story power delivery technique for 3D integrated circuits
abstract
Integrating circuits in the vertical direction can alleviate interconnect related problems and enable heterogeneous chips to be stacked in a single package with a small form factor. This paper addresses the power delivery issues in 3D chips revealing some interesting facts and design challenges. A multi-story power delivery technique that can reduce the worst case DC noise by 45% and lower the overhead power consumed in the power supply network by 65% is proposed. A test chip layout in an SOI process, showing a 5.3% area overhead, demonstrates the feasibility of the scheme.
Pulkit Jain, Tony Tae-Hyoung Kim, John Keane 0001, Chris H. Kim
ISLPED4
2008 Enhancing beneficial jitter using phase-shifted clock distribution
abstract
Clock jitter is generally considered undesirable but recent publications have shown that it can actually improve the timing margin. This paper investigates the "beneficial jitter" effect and presents an accurate analytical model which is verified with HSPICE. Based on our model, a phase-shifted clock distribution technique is proposed to enhance the beneficial jitter effect. By having an optimal phase shift between the supply noise and the clock period, the timing margin can be improved by 2.5X to 15% of the clock period. The benefit of the proposed technique is equivalent to that of having a 5X larger decoupling capacitor.
Dong Jiao, Jie Gu 0003, Pulkit Jain, Chris H. Kim
ISLPED4
2008 Statistical Leakage Estimation of Double Gate FinFET Devices Considering the Width Quantization Property
abstract
This paper presents a statistical leakage estimation method for FinFET devices considering the unique width quantization property. Monte Carlo simulations show that the conventional approach underestimates the average leakage current of FinFET devices by as much as 43% while the proposed approach gives a precise estimation with an error less than 5%. Design example on subthreshold circuits shows the effectiveness of the proposed method.
Jie Gu 0003, John Keane 0001, Sachin S. Sapatnekar, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.4
2008 Stack Sizing for Optimal Current Drivability in Subthreshold Circuits
abstract
Subthreshold circuit designs have been demonstrated to be a successful alternative when ultra-low power consumption is paramount. However, the characteristics of MOS transistors in the subthreshold region are significantly different from those in strong inversion. This presents new challenges in design optimization, particularly in complex gates with stacks of transistors. In this paper, we present a framework for choosing the optimal transistor stack sizing factors in terms of current drivability for subthreshold designs. We derive a closed-form solution for the correct sizing of transistors in a stack, both in relation to other transistors in the stack, and to a single device with equivalent current drivability. Simulation results show that our framework provides a performance benefit ranging up to more than 10% in certain critical paths.
John Keane 0001, Hanyong Eom, Tony Tae-Hyoung Kim, Sachin S. Sapatnekar, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.5
2008 A High-Speed Variation-Tolerant Interconnect Technique for Sub-Threshold Circuits Using Capacitive Boosting
abstract
This paper describes an interconnect technique for subthreshold circuits to improve global wire delay and reduce the delay variation due to process-voltage-temperature (PVT) fluctuations. By internally boosting the gate voltage of the driver transistors, operating region is shifted from subthreshold region to super-threshold region enhancing performance and improving tolerance to PVT variations. Simulations of a clock distribution network using the proposed driver shows a 66%-76% reduction in 3sigma clock skew value and 84%-88% reduction in clock tree delay compared to using conventional drivers. A 0.4-V test chip has been fabricated in a 0.18-mum 6-metal CMOS process to demonstrate the effectiveness of the proposed scheme. Measurement results show 2.6times faster switching speed and 2.4times less delay sensitivity under temperature variations.
Jonggab Kil, Jie Gu 0003, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.3
2008 Body Bias Voltage Computations for Process and Temperature Compensation
abstract
With continued scaling into the sub-90-nm regime, the role of process, voltage, and temperature (PVT) variations on the performance of VLSI circuits has become extremely important. These variations can cause the delay and the leakage of the chip to vary significantly from their expected values, thereby affecting the yield. Circuit designers have proposed the use of threshold voltage modulation techniques to pull back the chip to the nominal operational region. One such scheme, known as adaptive body bias (ABB), has become extremely effective in ensuring optimal performance or leakage savings. Our work provides a means to efficiently compute the body bias voltages required for ensuring high performance operation in gigascale systems. We provide a computer-aided design (CAD) perspective for determining the exact amount of bias voltages that can compensate both temperature and process variations. Mathematical models for delay and leakage based on minimal tester measurements are built, and a nonlinear optimization problem is formulated to ensure highest frequency operation under all conditions, and thereby minimize the overall circuit leakage. Three different algorithms are presented and their accuracies and runtimes are compared. The algorithms have been applied to a wide range of process and temperature corners, for a 65- and 45-nm technology node-based process. A suitable implementation mechanism has also been outlined.
Sanjay V. Kumar, Chris H. Kim, Sachin S. Sapatnekar
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Width-dependent Statistical Leakage Modeling for Random Dopant Induced Threshold Voltage Shift
abstract
Statistical behavior of device leakage and threshold voltage shows a strong width dependency under microscopic random dopant fluctuation. Leakage estimation using the conventional square-root method shows a discrepancy as large as 45% compared to the real case because it fails to model the effective VT shift in the subthreshold region. This paper presents a width-dependent statistical leakage model with an estimation error less than 5%. Design examples on SRAMs and domino circuits demonstrate the significance of the proposed model.
Jie Gu 0003, Sachin S. Sapatnekar, Chris H. Kim
DAC3
2007 NBTI-Aware Synthesis of Digital Circuits
abstract
Negative Bias Temperature Instability (NBTI) in PMOS transistors has become a major reliability concern in nanometer scale design, causing the temporal degradation of the threshold voltage of the PMOS transistors, and the delay of digital circuits. A novel method to characterize the delay of every gate in the standard cell library, as a function of the signal probability of each of its inputs, is developed. Accordingly, a technology mapping technique that incorporates the NBTI stress and recovery effects, in order to ensure optimal performance of the circuit, during its entire life-time, is presented. Our technique, demonstrated over 65nm benchmarks shows an average of 10% area recovery, and 12% power savings, as against a pessimistic method that assumes constant stress on all PMOS transistors in the design.
Sanjay V. Kumar, Chris H. Kim, Sachin S. Sapatnekar
DAC2
2007 Modeling and estimating leakage current in series-parallel CMOS networks
abstract
This paper reviews the modeling of subthreshold leakage current and proposes an improved model for general series-parallel CMOS networks. The presence of on-switches in off-networks, ignored by previous works, is considered in static current analysis. Both contributions present significant influence in the logic circuit leakage prediction when CMOS complex gates are extensively used. The proposed leakage model has been validated through electrical simulations, taking into account a 130nm CMOS technology, with good correlation of the results.
Paulo F. Butzen, André Inácio Reis, Chris H. Kim, Renato P. Ribas
ACM Great Lakes Symposium on VLSI3
2007 Sleep transistor sizing and control for resonant supply noise damping
abstract
A fact that has generally been unnoticed is that sleep transistors for leakage reduction can significantly damp the resonant supply noise due to their series resistance. This paper describes an optimal sleep transistor sizing method considering the dominant resonant supply noise. We show that a smaller sleep transistor can offer a smaller worst case supply noise due to the increased damping. We also propose an adaptive sleep transistor technique which automatically dampens the resonant noise only when it is detected. Simulations in 32nm CMOS show that the resonant noise is reduced by 32% using the proposed technique.
Jie Gu 0003, Hanyong Eom, Chris H. Kim
ISLPED3
2007 An on-chip NBTI sensor for measuring PMOS threshold voltage degradation
abstract
Negative Bias Temperature Instability (NBTI) is one of the most critical device reliability issues facing scaled CMOS technology. In order to better understand the characteristics of this mechanism, accurate and efficient means of measuring its effects must be explored. In this work, we describe an on-chip NBTI degradation sensor using two delay-locked loops (DLL). The increase in PMOS transistor threshold due to NBTI stress is translated into the control voltage of a DLL for high sensing gain. Measurements from a 0.13μm test chip show a maximum gain of 16X in the operating range of interest, with microsecond order measurement times for minimal unwanted recovery. The proposed NBTI sensor also supports various DC and AC stress modes.
John Keane 0001, Tony Tae-Hyoung Kim, Chris H. Kim
ISLPED3
2007 Utilizing Reverse Short-Channel Effect for Optimal Subthreshold Circuit Design
abstract
The impact of the reverse short-channel effect (RSCE) on device current is stronger in the subthreshold region due to reduced drain-induced barrier lowering (DIBL) and the exponential dependency of current on threshold voltage. This paper describes a device-size optimization method for subthreshold circuits utilizing RSCE to achieve high drive current, low device capacitance, less sensitivity to random dopant fluctuations, better subthreshold swing, and improved energy dissipation. Simulation results using ISCAS benchmark circuits show that the critical path delay, power consumption, and energy consumption can be improved by up to 10.4%, 34.4%, and 41.2%, respectively.
Tony Tae-Hyoung Kim, John Keane 0001, Hanyong Eom, Chris H. Kim
IEEE Trans. Very Large Scale Integr. Syst.4
2006 Mathematically assisted adaptive body bias (ABB) for temperature compensation in gigascale LSI systems
abstract
Process variations and temperature variations can cause both the frequency and the leakage of the chip to vary significantly from their expected values, thereby decreasing the yield. Adaptive body bias (ABB) can be used to pull back the chip to the nominal operational region. We propose the use of this technique to counter temperature variations along with process variations. We present a CAD perspective for achieving process and temperature compensation using bidirectional ABB. Mathematical models are used to determine the exact amount of body bias required optimizing the delay and leakage, and an algorithmic flow that can be adopted for gigascale LSI systems is provided.
Sanjay V. Kumar, Chris H. Kim, Sachin S. Sapatnekar
ASP-DAC2
2006 Subthreshold logical effort: a systematic framework for optimal subthreshold device sizing
abstract
Subthreshold circuit designs have been demonstrated to be a successful alternative when ultra-low power consumption is paramount. However, the characteristics of MOS transistors in the subthreshold regime are significantly different from those in strong-inversion. This presents new challenges in design optimization, particularly in complex gates with stacks of transistors. In this paper, we demonstrate a new optimal sizing scheme for subthreshold designs which takes these issues into account. We derive a closed-form solution for the correct sizing of transistors in a stack, both in relation to other transistors in the stack, and to a single transistor with equivalent current drivability. Experimental results show that our framework provides a performance improvement of up to 13.5% over the conventional logical effort method on ISCAS benchmark circuits, while one component circuit demonstrated an improvement of 33.1%.
John Keane 0001, Hanyong Eom, Tony Tae-Hyoung Kim, Sachin S. Sapatnekar, Chris H. Kim
DAC5
2006 An analytical model for negative bias temperature instability
abstract
transistors has become a significant reliability concern in present day digital circuit design. With continued scaling, the effect of NBTI has rapidly grown in prominence, forcing designers to resort to a pessimistic design style using guard-banding. Since NBTI is strongly dependent on the time for which the PMOS device is stressed, different gates in a combinational circuit experience varying extents of delay degradation. This has necessitated a mechanism of quantizing the gate-delay degradation, to pave the way for improved design strategies. Our work addresses this issue by providing a procedure for determining the amount of delay degradation of a circuit due to NBTI. An analytical model for NBTI is derived using the framework of the Reaction-Diffusion model, and a mathematical proof for the widely observed phenomenon of frequency independence is provided. Simulations on ISCAS benchmarks under a 70nm technology show that NBTI causes a delay degradation of about 8 % in combinational logic based circuits after 10 years (� s). I.
Sanjay V. Kumar, Chris H. Kim, Sachin S. Sapatnekar
ICCAD2
2006 Modeling and analysis of leakage induced damping effect in low voltage LSIs
abstract
Although there has been extensive research on controlling leakage power, the fact that leaky transistors can act as a damping element for supply noise has been long ignored or unnoticed in the design community. This paper investigates the leakage induced damping effect that helps suppress the supply noise. By developing physics-based impedance models for active and leakage currents, we show that leakage, particularly gate tunneling leakage, provides more damping than strong-inversion current. Simulations were performed in a 32nm CMOS technology to validate our models under PVT variations and to explore the voltage dependent behavior of this phenomenon. Design example utilizing leakage induced damping such as decap assignment is discussed with results showing 15.6% saving in decap area.
Jie Gu 0003, John Keane 0001, Chris H. Kim
ISLPED3
2006 A high-speed variation-tolerant interconnect technique for sub threshold circuits using capacitive boosting
abstract
This paper describes an interconnect technique for sub-threshold circuits to improve global wire delay and reduce the delay variation due to PVT fluctuations. By internally boosting the gate voltage of the driver transistors, operating region is shifted from sub-threshold region to super-threshold region enhancing performance and improving tolerance to PVT variations. A clock distribution network using the proposed drivers shows an 89% reduction in 3σ clock skew value. A 0.4V test chip has been fabricated in a 0.18μm 6-metal CMOS process to demonstrate the effectiveness of the proposed scheme. Measurement results show 2.6X faster switching speed and 2.4X less delay sensitivity under temperature variations.
Jonggab Kil, Jie Gu 0003, Chris H. Kim
ISLPED3
2006 Utilizing reverse short channel effect for optimal subthreshold circuit design
abstract
The impact of the Reverse Short Channel Effect (RSCE) on device current is stronger in the subthreshold region due to the reduced Drain-Induced-Barrier-Lowering (DIBL) and the exponential dependency of current on threshold voltage. This paper describes a device size optimization method for subthreshold circuits utilizing RSCE to achieve high drive current, low device capacitance, less sensitivity to random dopant fluctuations, and better subthreshold swing. Simulation results using ISCAS benchmark circuits show that the critical path delay and power consumption can be improved by up to 10.4% and 34.4%, respectively.
Tony Tae-Hyoung Kim, Hanyong Eom, John Keane 0001, Chris H. Kim
ISLPED4
2006 A process variation compensating technique with an on-die leakage current sensor for nanometer scale dynamic circuits
abstract
This paper describes a process compensating dynamic (PCD) circuit technique for maintaining the performance benefit of dynamic circuits and reducing the variation in delay and robustness. A variable strength keeper that is optimally programmed based on the die leakage, enables 10% faster performance, 35% reduction in delay variation, and 5times reduction in the number of robustness failing dies, compared to conventional designs. A new leakage current sensor design is also presented that can detect leakage variation and generate the keeper control signals for the PCD technique. Results based on measured leakage data show 1.9-10.2times higher signal-to-noise ratio (SNR) and reduced sensitivity to supply and p-n skew variations compared to prior leakage sensor designs
Chris H. Kim, Kaushik Roy 0001, Steven Hsu, Ram Krishnamurthy 0001, Shekhar Borkar
IEEE Trans. Very Large Scale Integr. Syst.1
2005 Self Calibrating Circuit Design for Variation Tolerant VLSI Systems
abstract
Increasing leakage current and aggravating process variations are showing impact on dynamic circuit performance and robustness as technology scales into the nanometer regime. This paper describes a self-calibrating process compensating dynamic (PCD) circuit technique for maintaining the performance benefit of dynamic circuits and reducing the variation in delay and robustness. A variable strength keeper that is optimally programmed based on the die leakage enables 10% faster performance, 35% reduction in delay variation, and 5/spl times/ reduction in the number of robustness failing dies compared to conventional designs. A new leakage current sensor design is also presented that can detect leakage variation and generate the keeper control signals for the PCD technique. The proposed 6-channel leakage current sensor enables high-resolution on-chip leakage measurements from multiple locations of a die, saving testing cost and realizing both die-to-die and within-die process compensation. Results based on measured leakage data show 1.9-10.2/spl times/ higher signal-to-noise ratio and reduced sensitivity to supply and P/N skew variations compared to prior leakage sensor designs. The PCD technique with the on-die leakage current sensor is applied to a 2-read, 2-write ported 128 /spl times/ 32b register file and a test chip is fabricated in 1.2V, 90nm dual-V, CMOS process.
Chris H. Kim, Steven Hsu, Ram Krishnamurthy 0001, Shekhar Borkar, Kaushik Roy 0001
IOLTS1
2005 Multi-story power delivery for supply noise reduction and low voltage operation
abstract
This paper presents a multi-story power delivery scheme which shows significant reduction of supply noise and power consumption compared to conventional power delivery scheme. To maximize the effectiveness of the proposed scheme, a digital voltage regulator is designed to balance the current dissipation of circuits in different voltage domains. Data transfer circuits based on capacitive coupling are developed for efficient inter-story data communication. Simulation results show 66% and 67% reduction of IR noise and Ldi/dt noise, respectively, while the total power consumption was reduced by 5% compared to a conventional power delivery scheme
Jie Gu 0003, Chris H. Kim
ISLPED2
2005 A forward body-biased low-leakage SRAM cache: device, circuit and architecture considerations
abstract
This paper presents a forward body-biasing (FBB) technique for active and standby leakage power reduction in cache memories. Unlike previous low-leakage SRAM approaches, we include device level optimization into the design. We utilize super high Vt (threshold voltage) devices to suppress the cache leakage power, while dynamically FBB only the selected SRAM cells for fast operation. In order to build a super high Vt device, the two-dimensional (2-D) halo doping profile was optimized considering various nanoscale leakage mechanisms. The transition latency and energy overhead associated with FBB was minimized by waking up the SRAM cells ahead of the access and exploiting the general cache access pattern. The combined device-circuit-architecture level techniques offer 64% total leakage reduction and 7.3% improvement in bit line delay compared to a previous state-of-the-art low-leakage SRAM technique. Static noise margin of the proposed SRAM cell is comparable to conventional SRAM cells.
Chris H. Kim, Jae-Joon Kim, Saibal Mukhopadhyay, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2004 Leakage in nano-scale technologies: mechanisms, impact and design considerations
abstract
The high leakage current in nano-meter regimes is becoming a significant portion of power dissipation in CMOS circuits as threshold voltage, channel length, and gate oxide thickness are scaled. Consequently, the identification of different leakage components is very important for estimation and reduction of leakage. Moreover, the increasing statistical variation in the process parameters has led to significant variation in the transistor leakage current across and within different dies. Designing with the worst case leakage may cause excessive guard-banding, resulting in a lower performance. This paper explores various intrinsic leakage mechanisms including weak inversion, gate-oxide tunneling and junction leakage etc. Various circuit level techniques to reduce leakage energy and their design trade-off are discussed. We also explore process variation compensating techniques to reduce delay and leakage spread, while meeting power constraint and yield.
Amit Agarwal 0001, Chris H. Kim, Saibal Mukhopadhyay, Kaushik Roy 0001
DAC2
2004 Larger-than-vdd forward body bias in sub-0.5V nanoscale CMOS
abstract
This paper examines the effectiveness of larger-than-Vdd forward body bias (FBB) in nanoscale bulk CMOS circuits where Vdd is expected to scale below 0.5V. Equal-to and larger-than Vdd FBB schemes offer unique advantages over conventional FBB such as simple design overhead and reverse body bias capability respectively. Compared to zero body bias, they improve process-variation immunity and achieve 71% and 78% standby leakage savings at iso performance and iso active power at room temperature. We also suggest a novel temperature-adaptive body bias scheme to control active leakage and achieve 22% and 40% active power savings at higher temperatures.
Hari Ananthan, Chris H. Kim, Kaushik Roy 0001
ISLPED2
2003 A forward body-biased low-leakage SRAM cache: device and architecture considerations
abstract
This paper presents a forward body-biasing (FBB) scheme for active leakage power reduction in cache memories. We utilize super high VT (threshold voltage) devices to suppress the leakage power in unselected portions of a cache while fast operation is achieve by dynamically forward body-biasing the selected SRAM cells. In order to generate a super high VT device, the 2-D halo doping profile was optimized considering different nanometer regime leakage mechanisms. The transition latency and energy overhead associated with FBB could be minimized by (i) waking up the SRAM cells ahead of the access and (ii) exploiting the cache access pattern. The combined device-circuit-architecture level techniques offer 64% total leakage reduction and 7.3% improvement in bitline delay compared to a previous state-of-the-art low-leakage SRAM technique.
Chris H. Kim, Jae-Joon Kim, Saibal Mukhopadhyay, Kaushik Roy 0001
ISLPED1
2003 Gate leakage reduction for scaled devices using transistor stacking
abstract
In this paper, the effect of gate tunneling current in ultra-thin gate oxide MOS devices of effective length (L/sub eff/) of 25nm (oxide thickness=1.1 nm), 50 nm (oxide thickness=1.5 nm) and 90 nm (oxide thickness=2.5 nm) is studied using device simulation. Overall leakage in a stack of transistors is modeled and the opportunities for leakage reduction in the standby mode of operation are explored for scaled technologies. It is shown that, as the contribution of gate leakage relative to the total leakage increases with technology scaling, traditional techniques become ineffective in reducing overall leakage current in a circuit. A novel technique of input vector selection based on the relative contributions of gate and subthreshold leakage to the overall leakage is proposed for reducing total leakage in a circuit. This technique results in 44% savings in total leakage in 50-nm devices compared to the conventional stacking technique.
Saibal Mukhopadhyay, Cassondra Neau, R. T. Cakici, Amit Agarwal 0001, Chris H. Kim, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2002 Dynamic VTH Scaling Scheme for Active Leakage Power Reduction
abstract
We present a Dynamic V/sub TH/ Scaling (DVTS) scheme to save the leakage power during active mode of the circuit. The power saving strategy of DVTS is similar to that of the Dynamic V/sub DD/ Scaling (DVS) scheme, which adaptively changes the supply voltage depending on the current workload of the system. Instead of adjusting the supply voltage, DVTS controls the threshold voltage by means of body bias control, in order to reduce the leakage power. The power saving potential of DVTS and its impact on dynamic and leakage power when applied to future technologies are discussed. Pros and cons of the DVTS system are dealt with in detail. Finally, a feedback loop hardware for the DVTS which tracks the optimal V/sub TH/ for a given clock frequency, is proposed. Simulation results show that 92% energy savings can be achieved with DVTS for 70 nm circuits.
Chris H. Kim, Kaushik Roy 0001
DATE1
2002 Dynamic Vt SRAM: a leakage tolerant cache memory for low voltage microprocessors
abstract
This paper presents a Dynamic Vt SRAM (DTSRAM) architecture to reduce the subthreshold leakage in cache memories. The Vt of each cache line is controlled separately by means of body biasing. In order to minimize the energy and delay overhead, a cache line is switched to high Vt only when it is not likely to be accessed anymore. Simulation results from SimpleScalar framework show that even after considering the energy overhead, the DTSRAM can save 72% of the cache leakage with a performance loss less than 1%. Layout of the DTSRAM shows that the area penalty is minimal.
Chris H. Kim, Kaushik Roy 0001
ISLPED1