Seong-Ook Jung

dblp:67/1704 · DBLP profile ↗
← Back
61ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0003-0757-2581ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 58 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3Security and privacy · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2025 Asymmetric Voltage Latch Type and Ultra-Low Swing Bitline Sense Amplifiers for Low-Power High-Density 1R1W 8T SRAM in 14 nm FinFET
abstract
We propose two innovative sense amplifiers, the asymmetric voltage latched-type sense amplifier (A-VLSA) and the ultra-low swing bitline sense amplifier (ULS-SA), to enhance read performance and reduce power consumption in the non-hierarchical BL 2-port 8-transistor SRAM (2P-SRAM). A-VLSA minimizes offset voltage through a MOS capacitor-based asymmetric operation, while ULS-SA achieves reduced read power by adopting clipped precharging and a charge-sharing read mechanism without incurring delay penalties. Measurement results demonstrate that A-VLSA (ULS-SA) achieves 63% (37%) and 61% (63%) lower energy-delay product (EDP) at VDD= 0.8V and 0.6V, respectively, with a 24% (43%) smaller area compared to the previous pseudo-differential asymmetric current latched-type sense amplifier (A-CLSA). These improvements address key limitations of prior approaches, delivering significant advancements in both power efficiency and read performance, especially under high cell density conditions.
Keon-hee Cho, Ji Sang Oh, Younmee Bae, Mijung Kim, Sangyeop Baeck, Taejoong Song, Seong-Ook Jung
IEEE Trans. Circuits Syst. I Regul. Pap.8
2025 SRAM BL Predriven Write Operation With Row and Voltage Auto-Tracking Replica BL in Resistance-Dominated Technology Nodes
abstract
In this article, we analyze the effect of the bitline (BL) predriven write operation in alleviating static random access memory (SRAM) writability degradation caused by BL resistance ($R_{\text {BL}}$). In BL predriven write operation, BL is fully driven to the ground voltage regardless of$R_{\text {BL}}$and the cell is written by a strong instantaneous peak write current ($I_{\text {write,peak}}$) between the cell and BL. The writability yield of BL predriven write operation in the resistance-dominated technology nodes can, thus, be significantly improved. In addition, the row and voltage auto-tracking replica BL (RVAT-RepBL) is proposed to generate BL predriven time ($T_{\text {pre}}$) for BL predriven write operation. In the proposed RVAT-RepBL,$T_{\text {pre}}$is generated by automatically tracking the variation in the number of rows per BL,$R_{\text {BL}}$, and the supply voltage ($V_{\text {DD}}$). In order to verify the effect of BL predriven write operation, the test chip was fabricated on 28-nm CMOS technology, and the poly resistor arrays were inserted to the cell array to reflect the interconnect resistance in the advanced technology nodes. BL predriven write operation has a higher writability yield and a wider operating$V_{\text {DD}}$than the conventional write operation. In addition, when the word line (WL) repeater is applied, the results of BL predriven write operation show that the writability yield of BL predriven write operation is further improved as$I_{\text {write,peak}}$increases with the improvement of WL rising slope.
Keon-hee Cho, Minjune Yeo, Seungjae Yei, Sangyeop Baeck, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.6
2024 Design of Physically Unclonable Function Using Ferroelectric FET With Auto Write-Back Technique for Resource-Limited IoT Security
abstract
Physically unclonable function (PUF) is a lightweight encryption technique that generates random digital keys (responses) using intrinsic process variations of devices, which is a promising solution for Internet of Things (IoT) security due to its compatibility with constrained resources. Recent attempts to adopt nonvolatile memory (NVM) into PUFs have enhanced stability through a write-back technique that maintains consistent responses from the enrollment phase even under wide environmental variations by storing the response in the NVM device. However, the stability of the previous NVM PUFs is limited by the low on/off ratio of the NVMs. In addition, the circuit required to implement the write-back technique poses challenges of increased area and energy consumption. Considering the hardware limitations and power constraints of IoT devices, this paper proposes a ferroelectric field-effect transistor (FeFET) PUF as a suitable security solution. The high on/off ratio of FeFET and the proposed auto write-back technique that does not require additional circuitry realize the stability improvement (a bit error rate of <0.0001%) under wide environmental variations without incurring area and energy overheads. The negligible off current of FeFET prevents static power consumption, which leads to the lowest energy consumption of 6.70e-15 J during the response generation of the FeFET PUF. In addition, the compact PUF cell composed of two FeFETs achieves a high density of 87.37 F2.
Sehee Lim, Junghyeon Hwang, Dong Han Ko, Sekeon Kim, Tae Woo Oh, Sanghun Jeon, Seong-Ook Jung
IEEE Internet Things J.7
2024 A CNN-Based Super-Resolution Processor With Short-Term Caching for Real-Time UHD Upscaling
abstract
Super-resolution (SR) tasks, involving the restoration of low-resolution images to high-resolution images, are expected to handle larger images in the near future. This paper proposes the short-term caching (STC) layer fusion to address the increase in cache memory size for image size expansion. The proposed STC layer fusion requires only 0.2% of the memory size compared to the previous method by discarding the overlap data that must be maintained for a long time and recalculating the discarded data later. In addition, the selective-operated dual accumulator (SODA) is proposed to apply a vertically long output patch that minimizes the additional recalculation of the STC layer fusion with resolving the memory access problem. Thus, the number of processing elements and memory bandwidth are reduced by 33.6% and 35.6%, respectively. The proposed SR processor including the proposed STC layer fusion and SODA is fabricated on a 28nm CMOS process, and the die and core area occupy 12.96mm2 and 3.45mm2, respectively. The SR task achieves$\times 2$upscaling to UHD images at 60 fps, achieving a maximum SR throughput of 497.7 Mpixels/s. Compared to previous approaches, the proposed SR processor reduces the on-chip memory size by more than 95% and achieves 4.3 times the SR throughput.
Hong Keun Ahn, Seong-Ook Jung
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 Ferroelectric FET Nonvolatile Sense-Amplifier-Based Flip-Flops for Low Voltage Operation
abstract
Nonvolatile processors (NVPs) are promising for energy-constrained internet-of-things applications in which frequent switch to standby mode occurs due to their fast and energy-efficient backup and restore operations of locally embedded nonvolatile flip-flops (NV-FFs) with zero leakage current. In addition, the most effective method to reduce dynamic energy consumption is to lower the supply voltage (${V} _{DD}$). The sense-amplifier-based flip-flop (SAFF) is considered a suitable choice for the low${V} _{DD}$operation, since it does not suffer from setup time degradation as${V} _{DD}$lowers. This study presents two ferroelectric FET (FeFET) nonvolatile SAFFs (FeFET NV-SAFF-1 and -2) that exhibit significantly low sequencing overhead (setup time + clock-to-Q time) at low${V} _{DD}$and consume low operating energy in the range of femto joules with compact layout area. The FeFET NV-SAFF-1 can operate robustly at low${V} _{DD}$even with a low resistance ratio. The FeFET NV-SAFF-2 has no area overhead and achieves the best power-performance-area at low${V} _{DD}$among state-of-the-art and proposed FeFET NV-FFs.
Sekeon Kim, Sehee Lim, Dong Han Ko, Tae Woo Oh, Seong-Ook Jung
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 Cross-Coupled Ferroelectric FET-Based Ternary Content Addressable Memory With Energy-Efficient Match Line Scheme
abstract
Fast communication between networking devices increases the importance of the ternary content addressable memory (TCAM). The demands for low energy in networking devices have accelerated the research on nonvolatile TCAMs that store data without the power supply. Recently, ferroelectric field-effect transistors (FeFETs), nonvolatile three-terminal devices with a high on/off ratio, have been adopted for TCAMs. Although the previous FeFET TCAMs consume low write energy with a significantly compact TCAM cell area, they suffer from write problems: 1) the write scheme cannot afford the saturated polarization switching in FeFETs and 2) the states stored in the unselected TCAM cells are changed during the write operation. In addition, search yield is deteriorated by wide process variations of FeFETs. This paper proposes a novel FeFET TCAM that is free from the write problems of the previous FeFET TCAMs and tolerant to process variations. This paper also proposes a novel match line scheme to improve search energy and time by reducing match line capacitance and the amount of discharged voltage in the match evaluation phase. Industrial-compatible 28 nm technology-based simulation results with the Preisach FeFET model show that the proposed FeFET TCAM achieves the highest search yield.
Sehee Lim, Dong Han Ko, Sekeon Kim, Seong-Ook Jung
IEEE Trans. Circuits Syst. I Regul. Pap.4
2022 A Sneak Current Compensation Scheme With Offset Cancellation Sensing Circuit for ReRAM-Based Cross-Point Memory Array
abstract
A sneak current compensation scheme with offset cancellation sensing circuit (SCC-OCSC) adopting dummy bitline and wordline is proposed to resolve sneak current problem in the cross-point memory array. The sneak current degrades sensing yield because it disturbs read operation for high resistance cell, especially, by contaminating read current. Since the sneak current increases proportionally to the array size, the array size is limited to achieve the target sensing yield, which makes it hard to implement high density memory. The proposed SCC-OCSC cancels-out the sneak current through two phases. In first phase, the sneak current is sampled by connecting dummy BL or WL to the sensing circuit. In the second phase, the sneak current is compensated by subtracting the current sampled in the first phase. Furthermore, sensing yield is enhanced by applying the offset cancellation technique, sharing the same sensing circuit in the two phases. The proposed SCC-OCSCs with dummy BL and WL effectively improve read margin with generally used biasing schemes, floating and half-VDDschemes, respectively. In Monte-carlo simulation including post-layout sensing circuit with 65nm CMOS technology, it is verified that sensing yield is significantly improved. Thus, the array size assigned to single sensing circuit is extended up to 8-times with SCC-OCSC, leading 88% area reduction with reduced number of required sensing circuit. Performance is improved as 66% and 25% while power consumption is improved as 47% and 37%, by SCC-OCSC with dummy BL in floating scheme and SCC-OCSC with dummy WL in half-VDDscheme, respectively.
Byungkyu Song, In-Jun Jung, Seong-Ook Jung
IEEE Trans. Circuits Syst. I Regul. Pap.4
2021 Self-Referenced Single-Ended Resistance Monitoring Write Termination Scheme for STT-RAM Write Energy Reduction
abstract
Essential design requirements for a sense amplifier (SA) used in the resistance monitoring write termination (RM-WT) scheme are suggested to reduce the write energy of spin-transfer-torque random access memory (STT-RAM) while achieving a write pass yield comparable to that of a conventional write operation. In addition, a self-referenced single-ended RM-WT (SS-RM-WT) scheme is proposed. To reduce the offset voltage, a single-ended sensing circuit (SE-SC) is used in the SA. A data-aware input voltage-transfer method is also adopted in the SE-SC to maximize the input voltage difference. By adopting a capacitor between the output of the SE-SC and the input of an inverter generating a logical output used for the write termination, the conflict between maintaining and changing the output of the SE-SC is resolved. The simulation results using the industry-compatible 65-nm technology HSPICE model parameters show that the proposed SS-RM-WT scheme achieves a 44% write energy saving on average without increasing the write error rate. Area overhead is only 11.8% for a 256-kb STT-RAM array, whereas that of the previous self-referenced RM-WT schemes is up to 42.5%.
Sara Choi, Hong Keun Ahn, Byungkyu Song, Seung-Hyuk Kang, Seong-Ook Jung
IEEE Trans. Circuits Syst. I Regul. Pap.5
2021 Imbalance-Tolerant Bit-Line Sense Amplifier for Dummy-Less Open Bit-Line Scheme in DRAM
abstract
In a conventional open bit-line scheme of DRAM, the edge subarrays (MATs) located at both ends of the cell array block contain alternated real and dummy bit-lines, unavoidably leading to an additional area overhead. To reduce the area overhead, one edge MAT can be eliminated by converting the dummy bit-lines of the other edge MAT into real bit-lines. This strategy causes the conventional bit-line sense amplifiers (BLSAs) in the MATs located at both ends of the cell array block to have a much smaller complementary bit-line capacitance than a true bit-line capacitance. Thus, the sensing operation of a conventional BLSA with this unbalanced bit-line capacitance experiences various problems: sensing voltage decrease, data flipping, and asymmetric equalization. To solve these problems, we propose a novel sensing circuit that can operate effectively even under unbalanced bit-line capacitance, thus suggesting the possibility of an open bit-line scheme without dummy bit-lines. Our proposed dummy-less open bit-line scheme can save approximately 4% of the array height. Compared with the conventional unbalanced BLSA, the proposed BLSA increases the sensing voltage by more than 100%, reduces the voltage peaks by 30% during the data transfer, and reduces equalization time by 1.2 ns in HSPICE Monte Carlo simulation.
Suk Min Kim, Byungkyu Song, Seong-Ook Jung
IEEE Trans. Circuits Syst. I Regul. Pap.3
2021 Environmental-Variation-Tolerant Magnetic Tunnel Junction-Based Physical Unclonable Function Cell With Auto Write-Back Technique
abstract
Recently, with the increase in popularity of Internet of Things (IoT) devices, cryptographic protection techniques have become necessary for high-security applications. In general, IoT devices have strict power and area constraints. Thus, use of a physical unclonable function (PUF), which can generate a secret key at low cost, can be advantageous for high-security IoT devices. This paper presents a novel environmental-variation-tolerant (EVT) magnetic tunnel junction (MTJ)-based PUF that has a small area, high randomness, and low bit error rate (BER) compared to previous PUFs. The simulation results obtained using industry-compatible 65-nm model parameters indicate that the proposed PUF exhibits an inter-chip Hamming distance of 0.4901 and entropy of 0.9997, which proves the randomness of the PUF response. In addition, the proposed PUF exhibits the lowest BER across a wide voltage range (0.9 V-1.3 V) and temperature range (-25 °C - 75 °C) compared with previous PUFs.
Byungkyu Song, Sehee Lim, Seung-Hyuk Kang, Seong-Ook Jung
IEEE Trans. Inf. Forensics Secur.4
2021 Adaptive Sensing Voltage Modulation Technique in Cross-Point OTS-PRAM
abstract
Phase-change random access memory with an ovonic threshold switch (OTS-PRAM) has become increasingly popular as an alternative to resolve the problem caused by the small capacity of dynamic random access memory and the high latency of NAND flash memory in computing systems. However, an OTS has a temperature-dependent OFF-current ( Ioff) and threshold voltage ( VTH). This causes Ioffof a cell ( Ioff_CELL) and VTHof a cell ( VTH_CELL) to become temperature-dependent, inducing a sensing error during read operations. In this article, an adaptive sensing voltage modulation (ASVM) technique is proposed that adaptively controls the bitline and wordline voltage depending on the change in temperature to compensate for the temperature-dependent variation in voltage drop caused by Ioff_CELLand VTH_CELL. The HSPICE simulation results, with an industry-compatible 250-nm complementary metal-oxide-semiconductor process for the 20-nm PRAM technology, show that the OTS-PRAM with the proposed ASVM can achieve a bit error rate below 0.1 ppm within the operating temperature range of 0 °C-85 °C.
Kwang Woo Lee, Hyun Kook Park, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.3
2020 A Read Voltage Modulation Technique for Leakage Current Compensation in Cross-Point OTS-PRAM
abstract
In this paper, a read voltage modulation technique (RVM) is proposed to compensate for leakage current in a cross-point phase change random access memory with an ovonic threshold switch (OTS-PRAM). The leakage current, the sum of off-state current (IOFF) of OTS selectors, causes the voltage drop and increases the variation of sensing voltage (VSENSE) which is the electric potential difference between a selected bit line (BL) and a word line (WL). Eventually, the voltage drop reduces the sensing margin (SM). To compensate for the BL voltage drop, the proposed RVM reduces the VSENSEvariation by applying an adaptive voltage to the selected WL. Thus, a sufficient SM is guaranteed. HSPICE simulation results with industry-compatible 65-nm model parameters show that the cross-point OTS-PRAM with the proposed RVM achieved a remarkable improvement in SM (from 105 mV to 395 mV) in high BL leakage current condition (51.3 uA).
Kwang Woo Lee, Hyun Kook Park, Seong-Ook Jung
ISCAS3
2020 Highly Independent MTJ-Based PUF System Using Diode-Connected Transistor and Two-Step Postprocessing for Improved Response Stability
abstract
In physically unclonable functions (PUFs), generating random cryptographs is required to secure private information. Various memory-based PUFs (MemPUFs), where cryptographs are generated independently from each PUF cell to increase the unpredictability of the cryptographs, have been proposed. Among them, the spin-transfer torque magnetic random-access memory MemPUF generates constant responses under temperature and voltage variations by exploiting a magnetic tunnel junction (MTJ) as the variation source. However, its response stability is diminished by the different characteristics of the two access transistors used in a PUF cell. To solve this problem, a novel PUF array that employs a diode-connected transistor and a shared access transistor, is proposed. In addition, a two-step postprocessing is adopted: 1) a write-back technique that amplifies the initial mismatch of MTJ resistances, and 2) a cell-classification technique that detects unstable PUF cells and discards their responses. The Monte Carlo HSPICE simulation results using industry-compatible 65-nm technology show that the proposed PUF system achieves the highest independence (autocorrelation factor of 0.0306) and the lowest maximum bit error rate (BER) under temperature and supply-voltage variations (<; 0.01% and 0.04% in the ranges of -25 to 75 °C and 0.8-1.2 V, respectively) compared with conventional PUF systems that exploit independent variation sources.
Sehee Lim, Byungkyu Song, Seong-Ook Jung
IEEE Trans. Inf. Forensics Secur.3
2020 A Novel Matchline Scheduling Method for Low-Power and Reliable Search Operation in Cross-Point-Array Nonvolatile Ternary CAM
abstract
Cross-point-array nonvolatile ternary content-addressable memory (CPA nvTCAM) has recently emerged as an alternative to static random-access-memory-type TCAM, based on increased demands for high-capacity and low-power attributes. The CPA structure has various structural weaknesses such as the searchline (SL) combining with the dischargeline and the minimum line pitch of the matchline (ML). This study analyzes these weaknesses in detail for the first time and resolves the issues caused by these weaknesses using the proposed novel ML shield scheduling method with a matching probability-based flexible searching time technique (MLSS + MPFST). The proposed MLSS + MPFST resolves various issues and achieves greater than sixfold smaller cell size (8F2) than the non-CPA nvTCAM with the smallest cell size. To verify the proposed schemes, the Monte Carlo HSPICE simulations were performed using a 22-nm industry-compatible bulk FinFET model parameter in 20-nm resistive memory technology at a circuit level, and the gem5 simulations were performed at a system level. The simulation results indicated that the CPA nvTCAM with the proposed MLSS + MPFST achieved a comparable search operation time of 1 ns and a slight power consumption overhead of 10% but an acceptable system performance overhead of less than 0.6% by resolving the various issues, compared with the non-CPA nvTCAM having the best performance.
Hyun Kook Park, Hong Keun Ahn, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.3
2020 pMOS Pass Gate Local Bitline SRAM Architecture With Virtual $V_{\mathrm{SS}}$ for Near-Threshold Operation
abstract
In this brief, a pMOS pass gate (PPG) local bitline static random access memory (LB SRAM) architecture is proposed to reduce the read delay and resolve the half-select issue with a small area overhead. Virtual VSS write assist is included in the architecture to improve write ability. In 22-nm fin-shaped FET (FinFET) technology, the proposed PPG LB architecture achieves an improved read delay and reduced total operation energy by 44% and 65%, respectively, at 0.4 V, compared to the full-swing LB (FSLB) SRAM architecture.
Tae Woo Oh, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.3
2019 A Decoder for Short BCH Codes With High Decoding Efficiency and Low Power for Emerging Memories
abstract
In this paper, a double-error-correcting and triple-error-detecting (DEC-TED) Bose-Chaudhuri-Hocquenghem (BCH) code decoder with high decoding efficiency and low power for error correction in emerging memories is presented. To increase the decoding efficiency, we propose an adaptive error correction technique for the DEC-TED BCH code that detects the number of errors in a codeword immediately after syndrome generation and applies a different error correction algorithm depending on the error conditions. With the adaptive error correction technique, the average decoding latency and power consumption are significantly reduced owing to the increased decoding efficiency. To further reduce the power consumption, an invalid-transition-inhibition technique is proposed to remove the invalid transitions caused by glitches of syndrome vectors in the error-finding block. Synthesis results with an industry-compatible 65-nm technology library show that the proposed decoders for the (79, 64, 6) BCH code take only 37%-48% average decoding latency and achieve more than 70% power reduction compared to the conventional fully parallel decoder under the 10-4-10-2raw bit-error rate.
Sara Choi, Hong Keun Ahn, Byungkyu Song, Jung Pill Kim, Seung-Hyuk Kang, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.6
2019 Sensing Margin Enhancement Technique Utilizing Boosted Reference Voltage for Low-Voltage and High-Density DRAM
abstract
In the case of dynamic random access memory (DRAM) using a voltage latched sense amplifier, various offset voltage cancellation techniques have been studied to secure the sensing margin. However, as coupling noise and process variations increase with technology scaling, it is impossible to obtain a sufficient sensing margin only by using an offset voltage cancellation technique. In addition, offset voltage cancellation techniques have several problems, such as area overhead and sensing speed slowdown at a low supply voltage. In this paper, we propose a sense amplifier that maximizes the sensing margin by boosting the reference bitline voltage during the charge sharing operation and adopting the offset voltage cancellation technique. The sensing voltage difference of the proposed sense amplifier increases by 50% or more than that of the conventional offset cancellation (OC) sense amplifier at the low supply voltage of 0.9 V, which can improve not only the sensing speed by 2 ns but also the sensing yield by 11.8%. In addition, the proposed sense amplifier achieves a stable sensing yield with a larger cell array height and thus can compensate the area overhead of 44% caused by the OC technique by decreasing the overall cell array height.
Suk Min Kim, Byungkyu Song, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.3
2019 Variation-Tolerant WL Driving Scheme for High-Capacity NAND Flash Memory
abstract
Research on a word-line (WL) driving scheme is essential because the effect of WL parasitic resistance and capacitance (RC) is more severe for high-capacity NAND flash memories. The WL under-driving scheme (WLUDS) mitigates the effect of parasitic RC by reducing the coupling capacitance between WLs. However, WLUDS increases the cell threshold voltage (Vth) distribution because of parasitic RC variation, which causes an overshoot of the programming voltage (VPGM). In this study, we propose the variation-tolerant WL under-driving scheme (VTWLUDS) to reduce the effect of parasitic RC variation and VPGMovershoot through the use of a three-phase VPGMcontrol. We also introduce the fast-verify WL driving scheme (FVWLDS) to reduce the effect of parasitic RC variation in the verify operation. We verified VTWLUDS and FVWLDS by performing an HSPICE simulation with Samsung's transistor model for a NAND peripheral circuit. The simulation results showed that VTWLUDS achieved a sufficient Vthshift during the programming operation regardless of the WL parasitic RC variation. By using VTWLUDS and FVWLDS, we achieved 1304 μs of total programming time (TPROG) for a 512-Gb planar-type NAND flash memory.
Junyoung Ko, Younghwi Yang, Cheon An Lee, Young-Sun Min, Jin-Young Chun, Moosung Kim, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.8
2019 Offset-Canceling Single-Ended Sensing Scheme With One-Bit-Line Precharge Architecture for Resistive Nonvolatile Memory in 65-nm CMOS
abstract
In the design of nonvolatile memory (NVM), the sensing scheme (SS) has become a read-energy bottleneck because the required read-cell current is too large to satisfy a target read yield. This problem is further aggravated by technology scaling because increased process variation and reduced supply voltage (VDD) require more current to satisfy the target read yield. This paper proposes an offset-canceling single-ended SS (OCSE-SS) with one-bit-line precharge architecture (1BLPA) that is intended for use in ultralow power NVM applications. The test chip is fabricated using 65-nm process technology, and the measurement results show that the read energy per bit of the OCSE-SS is 1/3 compared to that of the conventional SS (Conv-SS). The read energy reduction comes from the singleended sensing, offset cancellation, and 1BLPA features. Moreover, when a resistance difference between the data and reference cells is as small as 0.5 kQ, the OCSE-SS reads successfully with a VDD of 1.0 V and a sensing time (tSEN) of 17 ns due to the offset cancellation characteristic, whereas the Conv-SS fails regardless of VDD and tSEN values.
Taehui Na, Byungkyu Song, Sara Choi, Jung Pill Kim, Seung-Hyuk Kang, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.6
2018 SRAM Cell with Data-Aware Power-Gating Write-Asist for Near-Threshold Operation
abstract
This paper proposes a FinFET-based SRAM cell with data-aware power-gating write-assist to achieve both high read stability and write ability by using read-decoupled access transistors and power-gating PMOSs, respectively, for near-threshold operation. By adaptively cutting off the power-gating PMOS depending on the written data, the write disturbance from power supply can be eliminated, which facilitates more reliable write operation without any additional write assist circuit. Bit-interleaving scheme can be implemented in the proposed SRAM for soft error immunity while ensuring sufficient hold stability in half-selected cells during write operation. The proposed SRAM achieves read stability yield of 8.4σ and write ability yield of 6.1σ and consumes 0.47 pJ energy per operation at supply voltage of 0.4 V, a near-threshold voltage, in a 22-nm FinFET technology.
Tae Woo Oh, Seong-Ook Jung
ISCAS2
2018 Sense-Amplifier-Based Flip-Flop With Transition Completion Detection for Low-Voltage Operation
abstract
A novel high-speed and highly reliable sense-amplifier-based flip-flop with transition completion detection (SAFF-TCD) is proposed for low supply voltage (VDD) operation. The SAFF-TCD adopts the internally generated detection signal to indicate the completion of sense-amplifier stage transition. The detection signal gates the pull-down path of the sense-amplifier stage and the slave latch, thus overcoming the operational yield degradation, current contention, and glitches of previous SAFFs. The operational yield, speed, hold time, energy consumption, and area of the proposed and previous FFs are quantitatively compared for a wide range of VDDwith 22-nm FinFET technology. It is shown that the minimum VDDof the SAFF-TCD is 573 mV lower than that of previous SAFFs, which means the SAFF-TCD can operate even when VDDis in the near-threshold or subthreshold region. At 0.3-0.4 V, the SAFF-TCD operates twice as fast as the master-slave-based FF (MSFF) with a practical hold time. Even with these benefits, the energy consumption overhead is limited to less than 20% compared with that of MSFF, and the area is similar to that of previous SAFFs.
Hanwool Jeong, Tae Woo Oh, Seung Chul Song, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.4
2018 All-Digital Process-Variation-Calibrated Timing Generator for ATE With 1.95-ps Resolution and Maximum 1.2-GHz Test Rate
Dong-Hoon Jung, Kyungho Ryu, Jung-Hyun Park, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.4
2017 Power-Gated 9T SRAM Cell for Low-Energy Operation
abstract
This brief proposes a novel power-gated 9T (PG9T) static random access memory (SRAM) cell that uses a read-decoupled access buffer and power-gating transistors to execute reliable read and write operations. The proposed 9T SRAM cell uses bit interleaving to achieve soft error immunity and utilizes a column-based virtual VSS signal to eliminate unnecessary bitline discharges in the unselected columns, thereby reducing the energy consumption. In a 22-nm FinFET technology, the proposed PG9T SRAM cell has a minimum operating voltage of 0.32 V while achieving the 6σ read stability yield. Compared with the previously proposed 9T SRAM cell, the proposed cell consumes 45% and 17% less energy per read and write operation, respectively, at the minimum operating voltage, and has a 12% smaller bit cell area.
Tae Woo Oh, Hanwool Jeong, Kyoman Kang, Younghwi Yang, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.6
2016 Area-optimal sensing circuit designs in deep submicrometer STT-RAM
abstract
As the technology node scales down, a sufficient read current that is capable of achieving a target read yield cannot be used because of the read disturbance problem in spin-transfer-torque random access memory (STT-RAM). As an alternative method, increasing the sensing circuit (SC) area is generally considered because it can reduce the threshold voltage (Vth) variations. However, the increased SC area can adversely reduce the read yield due to the increased load capacitance. The effects of the increased area on read yield can be different according to the SCs because of their own characteristics. In this work, the trends of read yield according to the area are analyzed for two representative SCs, and the areas of two SCs are optimally designed to have high read yield.
Sara Choi, Taehui Na, Seong-Ook Jung, Jung Pill Kim, Seung-Hyuk Kang
ISCAS3
2016 WL under-driving scheme with decremental step voltage and incremental step time for high-capacity NAND flash memory
abstract
In this work, we compared the WL driving schemes in 512 Gb planar NAND with 32 KB page size. In the conventional WL driving scheme, the rising time of the selected WL voltage is very large because of the large coupling capacitance between the selected and unselected WLs. The WL under-driving scheme (WLUDS) reduces the effect of coupling capacitance by using the 2-phase control of unselected WL voltage. However, when WLUDS is used, the relationship between the rising time and overshoot of the selected WL voltage should be considered in order to achieve the small rising time Therefore, we proposes a novel implementation method for WLUDS that controls the under-driving voltage and under-driving timing by using the decremental step voltage and incremental step time (DSVIST) to enhance the rising time considering the overshoot constraint. The HSPICE simulation using the 0.25-μm PTM model with the parasitic RC in 32 KB page size shows that the rising time in the proposed WLUDS with DSVIST is improved to 988 μs compared to 1206 μs in the conventional WL driving scheme.
Junyoung Ko, Younghwi Yang, Seong-Ook Jung, Cheon An Lee, Young-Sun Min, Jin-Young Chun, Moosung Kim
ISCAS3
2016 All-Digital ON-Chip Process Sensor Using Ratioed Inverter-Based Ring Oscillator
abstract
In this paper, an all-digital ON-chip process sensor using a ratioed inverter-based ring oscillator is proposed. Two types of the ratioed inverter-based ring oscillators, nMOS and pMOS types, are proposed to sense process variation. The nMOS (pMOS)-type ring oscillator is designed to improve its sensitivity to the process variation in the nMOS (pMOS) transistors using the ratioed inverter that consists of only nMOS (pMOS) transistors. A compact process sensor can be realized using only these two types of ring oscillators. For a suitable ON-chip implementation, the output of the proposed process sensor is provided with a digital code. The proposed process sensor is fabricated using a 0.13-μm CMOS technology. Measurement results from 30 fabricated chips show that all chips have the same process corner. To verify whether the proposed sensor can properly sense all the process corners, the threshold voltage of the fabricated chips is shifted by body biasing. The verification results show that the measured code error compared with the postlayout simulation is less than 2.92%.
Young-Jae An, Dong-Hoon Jung, Kyungho Ryu, Hyuck-Sang Yim, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.5
2016 Corner-Aware Dynamic Gate Voltage Scheme to Achieve High Read Yield in STT-RAM
abstract
As the technology node scales down, the spin-transfer-torque random access memory (STT-RAM) has been considered as a promising memory solution owing to its scalability. However, the increased process variation and the reduced supply voltage lead to degradation in the sensing yield (SY) as well as an increase in the read disturbance probability. Temperature variation further aggravates this phenomenon. Thus, achieving a target SY with a lower sensing current in all process, voltage, and temperature (PVT) corners has become an important issue in a deep-submicrometer technology node. In this paper, we propose a corner-aware dynamic gate voltage scheme to achieve constant-current sensing, regardless of the PVT variations. By adopting this scheme, the state-of-the-art sensing circuits (SCs) can significantly reduce the sensing current, while achieving the target read yield. The Monte Carlo HSPICE simulation results using industry-compatible 45-nm model parameters show that the offset-canceling dual-stage SC that uses the proposed scheme satisfies a target SY of six-sigma (96.34% for 32 Mb) with two times lower sensing current and two times lower read energy compared with that using a fixed gate voltage.
Sara Choi, Taehui Na, Jung Pill Kim, Seung-Hyuk Kang, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.6
2016 All-Digital 90° Phase-Shift DLL With Dithering Jitter Suppression Scheme
abstract
This paper proposes a 90° phase-shift delay-locked loop (DLL) used in dynamic RAM for data sampling clock generation. The proposed DLL alleviates process variation issues, which are mainly caused by the mismatch between the delay line segments in the previous 90° phase-shift DLLs, and reduces area by adopting a multiplying DLL-based structure. In addition, a novel jitter suppression scheme is also proposed to suppress control code dithering. A stochastic analysis is performed to evaluate the effectiveness of the proposed dithering jitter suppression. The proposed DLL is fabricated using a 45-nm CMOS process on an active area of 69.9 μm × 49.3 μm and utilizes a 1.1 V supply voltage. The proposed DLL has an operating frequency ranging from 500 to 800 MHz and consumes 1.32 mW at 800 MHz. The measured rms and peak-to-peak output jitters are improved by 5.42% to 18.75% and 5.52% to 18.31%, respectively, in the entire operating frequency range.
Dong-Hoon Jung, Kyungho Ryu, Jung-Hyun Park, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.4
2016 Full-Swing Local Bitline SRAM Architecture Based on the 22-nm FinFET Technology for Low-Voltage Operation
abstract
The previously proposed average-8T static random access memory (SRAM) has a competitive area and does not require a write-back scheme. In the case of an average-8T SRAM architecture, a full-swing local bitline (BL) that is connected to the gate of the read buffer can be achieved with a boosted wordline (WL) voltage. However, in the case of an average-8T SRAM based on an advanced technology, such as a 22-nm FinFET technology, where the variation in threshold voltage is large, the boosted WL voltage cannot be used, because it degrades the read stability of the SRAM. Thus, a full-swing local BL cannot be achieved, and the gate of the read buffer cannot be driven by the full supply voltage (VDD), resulting in a considerably large read delay. To overcome the above disadvantage, in this paper, a differential SRAM architecture with a full-swing local BL is proposed. In the proposed SRAM architecture, full swing of the local BL is ensured by the use of cross-coupled pMOSs, and the gate of the read buffer is driven by a full VDD, without the need for the boosted WL voltage. Various configurations of the proposed SRAM architecture, which stores multiple bits, are analyzed in terms of the minimum operating voltage and area per bit. The proposed SRAM that stores four bits in one block can achieve a minimum voltage of 0.42 V and a read delay that is 62.6 times lesser than that of the average-8T SRAM based on the 22-nm FinFET technology.
Kyoman Kang, Hanwool Jeong, Younghwi Yang, Ki-Ryong Kim, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.6
2016 Multiple-Cell Reference Scheme for Narrow Reference Resistance Distribution in Deep Submicrometer STT-RAM
abstract
Spin-transfer-torque random access memory (STT-RAM) has attracted much research interest because of its characteristics of nonvolatility (i.e., zero standby power) and small cell size (i.e., high density and high performance). As the technology node is scaled down, however, the sensing margin of the STT-RAM is degraded because of the increased process variation and reduced supply voltage. To improve the sensing margin, this brief focuses on a reference scheme design capable of reducing the reference resistance distribution. A multiple-cell reference (MCR) scheme is proposed that achieves the narrow reference resistance distribution. Moreover, the MCR scheme does not exhibit parasitic mismatch, regularity problem, read disturbance, and write current degradation, and it also has small area overhead.
Taehui Na, Jung Pill Kim, Seung-Hyuk Kang, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.4
2016 An Offset-Tolerant Dual-Reference-Voltage Sensing Scheme for Deep Submicrometer STT-RAM
abstract
Due to the increased process variation and reduced supply voltage in deep submicrometer technology nodes, an offset-tolerant sensing scheme has become essential. However, most offset-tolerant sensing schemes suffer from inherent performance degradation owing to multiple-stage sensing. In this paper, a dual Vrefsensing scheme (DVSS) that selectively uses an optimal Vrefbetween Vref+and Vref-is proposed. This scheme is tolerant to process variations, and can be used as a spin-transfer-torque random access memory. Because of no additional sensing stage, the offset-tolerant sensing is achieved without sacrificing the performance. The optimal Vrefis selected after fabrication, and the calibrated switch control bit, which contains Vrefselection information, is stored permanently in an on-chip nonvolatile latch. Monte Carlo HSPICE simulation results, using an industry-compatible 45-nm model parameters, show that the proposed DVSS achieves a read yield of 98.24% for 32 Mb (6.1 sigma) with 2× faster sensing speed and 1.5× lower read energy per bit compared with the state-of-the-art offset-tolerant sensing scheme.
Taehui Na, Byungkyu Song, Jung Pill Kim, Seung-Hyuk Kang, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.6
2016 High-Speed, Low-Power, and Highly Reliable Frequency Multiplier for DLL-Based Clock Generator
abstract
A high-speed, low-power, and highly reliable frequency multiplier is proposed for a delay-locked loop-based clock generator to generate a multiplied clock with a high frequency and wide frequency range. The proposed edge combiner achieves a high-speed and highly reliable operation using a hierarchical structure and an overlap canceller. In addition, by applying the logical effort to the pulse generator and multiplication-ratio control logic design, the proposed frequency multiplier minimizes the delay difference between positive- and negative-edge generation paths, which causes a deterministic jitter. Finally, a numerical analysis is performed to analyze and compare the performance of the proposed frequency multiplier with that of previous frequency multipliers. The proposed frequency multiplier is fabricated using a 0.13-μm CMOS process technology, and has the multiplication ratios of 1, 2, 4, 8, and 16, and an output range of 100 MHz-3.3 GHz. The frequency multiplier achieves a power consumption to a frequency ratio of 2.9 μW/MHz.
Kyungho Ryu, Jiwan Jung, Dong-Hoon Jung, Jin Hyuk Kim, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.5
2015 Reference-circuit analysis for high-bandwidth spin transfer torque random access memory
abstract
A global reference-circuit (RC), which means one RC is shared with many sensing circuits (SC), is being considered for high-bandwidth STT-RAMs because of the low power consumption and small area characteristic. However, using the global RC for high-bandwidth STT-RAMs causes a droop effect and coupling noise effect, leading to the significant performance degradation. Thus, the validity of using the global RC should be identified. In this paper, the local RC and various global RCs are introduced, and compared in aspects of area, sensing time, and power consumption. By classification of the merits and demerits of various RCs, we present the following requirements of proper RC for high-bandwidth STT-RAMs: 1) small area, 2) no performance degradation, 3) low power consumption, and 4) process variation tolerant reference signal generation.
Byungkyu Song, Taehui Na, Seong-Ook Jung, Jung Pill Kim, Seung-Hyuk Kang
ISLPED3
2015 An Energy-Efficient All-Digital Time-Domain-Based CMOS Temperature Sensor for SoC Thermal Management
abstract
We propose an all-digital on-chip time-domain temperature sensor for system-on-a-chip (SoC) thermal management. For on-chip purposes, the proposed temperature sensor achieves energy- and area-efficient and fast thermal monitoring by adopting a digitally controlled oscillator (DCO) with the frequency divider and XNOR gate to generate temperature-dependent pulse. The frequency divider with the fine delay unit allows DCO of the proposed structure to have a smaller number of delay cells than a conventional open-loop delay line while maintaining resolution. The use of DCO with frequency divider, which consists of three flip-flops, reduced the required delay line length by 16 times. XNOR gate facilitates the fast thermal monitoring by simply providing the temperature-proportional pulse without additional processing. Temperature measurement results are provided with a digital code generated by a simple counter-based time-to-digital conversion. The proposed temperature sensor is fabricated using 0.13-μm CMOS technology and achieves a low-energy consumption of 2.3 nJ at a conversion rate of 293 kHz with a resolution of 0.72 °C and an area of 0.036 mm2. The proposed sensor also obtains a measurement error of -2.4 °C to 2.16 °C from nine test chips over a temperature range of 20 °C-120 °C, which is suitable for SoC thermal management.
Young-Jae An, Dong-Hoon Jung, Kyungho Ryu, Seung-Han Woo, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.5
2015 Trip-Point Bit-Line Precharge Sensing Scheme for Single-Ended SRAM
abstract
A trip-point bit-line precharge (TBP) sensing scheme is proposed for high-speed single-ended static random-access memory (SRAM). This TBP scheme mitigates the issues of limited performance, power, sensing margin, and area found in the previous single-ended SRAM sensing schemes by biasing the bit-line to a slightly larger value than the trip point of the sense amplifier. Simulation results show that the TBP sensing scheme can reduce the sensing time by 58.5% and 10% compared with the domino and ac-coupled sensing schemes, respectively. Further, compared with the ac-coupled sensing scheme, the proposed scheme offers 10% lower sensing power, 36% lesser area, and a 60 mV lower value of the minimum supply voltage for the target sensing yield.
Hanwool Jeong, Taewon Kim, Taejoong Song, Gyu-Hong Kim, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.5
2015 Architecture-Aware Analytical Yield Model for Read Access in Static Random Access Memory
abstract
We prove analytically that the yield of static random access memory (SRAM) is intrinsically a function of its architecture owing to the correlation among cell failures. In addition, architecture-aware analytical yield models are proposed for read access. The yield results using the proposed models show that the most dominant factor determining yield is the variation in the voltage difference between bitlines due to the cell leakage current variation according to the SRAM architecture. The models also show the possibility that the most dominant factor determining the yield can change with the relative ratios among the amounts of changes in the correlation, recovery sample space, distributions of the sense amplifier enable time, voltage difference between bitlines, as well as sense amplifier offset voltage, memory capacity, and redundancy scheme. The proposed yield models show that combined row and column redundancy ensures the highest yield, whereas column redundancy is the most efficient.
Heechai Kang, Hanwool Jeong, Younghwi Yang, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.5
2015 Level-Converting Retention Flip-Flop for Reducing Standby Power in ZigBee SoCs
abstract
In this paper, we propose a level-converting retention flip-flop (RFF) for ZigBee systems-on-chips (SoCs). The proposed RFF allows the voltage regulator that generates the core supply voltage (VDD,core) to be turned off in the standby mode, and it thus reduces the standby power of the ZigBee SoCs. The logic states are retained in a slave latch composed of thick-oxide transistors using an I/O supply voltage (VDD,IO) that is always turned on. Level-up conversion from VDD,core to VDD,IO is achieved by an embedded nMOS pass-transistor level-conversion scheme that uses a low-only signal-transmitting technique. By embedding a retention latch and level-up converter into the data-to-output path of the proposed RFF, the RFF resolves the problems of the static RAM-based RFF, such as large dc current and low readability caused by threshold drop. The proposed RFF does not also require additional control signals for power mode transitioning. Using 0.13-μm process technology, we implemented an RFF with VDD,core and VDD,IO of 1.2 and 2.5 V, respectively. The maximum operating frequency is 300 MHz. The active energy of the RFF is 191.70 fJ, and its standby power is 350.25 pW.
Jung-Hyun Park, Heechai Kang, Dong-Hoon Jung, Kyungho Ryu, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.5
2015 Single-Ended 9T SRAM Cell for Near-Threshold Voltage Operation With Enhanced Read Performance in 22-nm FinFET Technology
abstract
Although near-threshold (Vth) operation is an attractive method for energy and performance-constrained applications, it suffers from problems in terms of circuit stability, particularly, for static random access memory (SRAM) cells. This brief proposes a near-Vth 9T SRAM cell implemented in a 22-nm FinFET technology. The read buffer of the proposed cell ensures read stability by decoupling the stored node from the read bit-line and improves read performance using a one-transistor read path. Energy and standby power are reduced by eliminating the sub-Vth leakage current in the read buffer. For accurate sensing yield estimation, a new yield-estimation method is also proposed, which considers the dynamic trip voltage. The proposed SRAM cell can achieve a minimum operating voltage of 0.3 V.
Younghwi Yang, Seung Chul Song, Geoffrey Yeap, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.6
2014 High-performance low-power magnetic tunnel junction based non-volatile flip-flop
abstract
In this paper, a novel magnetic tunnel junction (MTJ) based non-volatile flip-flop (NVFF) is proposed. The separated latch and sensing circuit structure maximizes the performance of latch operation, minimizes power consumption, and improves MTJ lifetime. Furthermore, the merged sensing and write circuit structure reduces area overhead. HSPICE simulation results using a 45-nm technology model show that the proposed NVFF achieves three times smaller power delay product with a 2% smaller layout area than the conventional NVFF.
Taehui Na, Kyungho Ryu, Seong-Ook Jung, Jung Pill Kim, Seung-Hyuk Kang
ISCAS4
2014 One-Sided Static Noise Margin and Gaussian-Tail-Fitting Method for SRAM
abstract
In this paper, we propose a method to estimate the read failure rate of a static random access memory (SRAM) cell. Conventional read stability metrics cannot accurately estimate the read failure probability as technology scales down, because some metrics cannot characterize read stability and others can no longer be approximated to a known distribution. We first introduce a one-sided static noise margin (OSNM), whose lower tail region follows a Gaussian distribution, and also propose a Gaussian-tail-fitting method that properly models the distribution of the OSNM at the tail region. It is demonstrated that the OSNM can accurately estimate the SRAM cell yield using the proposed Gaussian-tail-fitting method.
Hanwool Jeong, Younghwi Yang, Junha Lee, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.5
2014 STT-MRAM Sensing Circuit With Self-Body Biasing in Deep Submicron Technologies
abstract
Conventional spin transfer torque MRAM sensing circuits suffer from a small sensing margin and a large sensing margin variation in deep submicron technologies. The small sensing margin issue becomes worse in the low-leakage process technology due to the higher threshold voltage. In this brief, the self-body biasing (self-BB) scheme is proposed to resolve the small sensing margin issue. In the self-BB scheme, the threshold voltage of load pMOS is adaptively controlled by body bias. Although leakage current 'lows through the body due to the positive junction bias voltage, it is well suppressed to less than 1% (0.3 μA) of the sensing current and 'lows only during the sensing operation. To reduce large sensing margin variation, the source degeneration scheme with the longer channel length is used for the load pMOS. The HSPICE simulation results obtained using low-leakage 45-nm model parameters show that the proposed sensing circuit achieves a probability of the read access pass yield (PRAPY Memory) of 100%, whereas the sensing circuit without BB scheme has an PRAPY Memory of 5.8% for a 32-Mb memory with a sensing time of 2 ns.
Kyungho Ryu, Jung Pill Kim, Seung-Hyuk Kang, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.5
2014 An Offset-Canceling Triple-Stage Sensing Circuit for Deep Submicrometer STT-RAM
abstract
Spin-transfer torque random access memory (STT-RAM) is considered to be a leading candidate for next-generation memory. As technology scales, however, the sensing margin of STT-RAM is significantly degraded because of increased process variation. Furthermore, the sensing current should be <;20 μA to protect the read disturbance in the beyond 45-nm technology, leading to a further decrease in the sensing margin. To achieve a target yield of six sigma in the beyond 45-nm technology with a sensing current of <;20 μA, an offset-canceling triple-stage (OCTS) sensing circuit is proposed in this brief. The OCTS sensing circuit can overcome the sensing margin and read disturbance problems by sacrificing the sensing time. Monte Carlo HSPICE simulation results using a 45-nm technology model show that the OCTS sensing circuit achieves a target yield of six sigma (96.74% for 32 Mb) with a sensing current of 20 μA and a sensing time of 6.4 ns.
Taehui Na, Jung Pill Kim, Seung-Hyuk Kang, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.5
2014 Comparative Study of Various Latch-Type Sense Amplifiers
abstract
When the input voltage difference of a sense amplifier (SA) exceeds the offset voltage (VOS), the SA correctly detects it and outputs a large signal. However, when the input voltage is in a certain region, the SA can fail to sense the input voltage difference even if it is sufficiently large. This input voltage region is defined as the sensing dead zone of the SA. Because sensing dead zones differ depending on SAs and the input voltages to the SA differ depending on the memory devices, analyzing the sensing dead zone is very important. In this brief, we analyze the sensing dead zones of the most popular latch-type SAs: voltage- and current-latched SAs. Furthermore, a suitable latch-type SA scheme is suggested for various SA input voltages in terms of sensing delay, power consumption, and PDP, using a 65-nm predictive technology model at a VDD of 1.1 V.
Taehui Na, Seung-Han Woo, Hanwool Jeong, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.5
2013 A comparative study of STT-MTJ based non-volatile flip-flops
abstract
In this paper, we categorize STT-MTJ based non-volatile flip-flops (NV-FF) into two basic structures: merged latch and sensing circuit (MLS) structure and separated latch and sensing circuit (SLS) structure. We also analyze the two structures with various types of sensing and write circuits. HSPICE simulation results using the industry-compatible 45-nm model parameter shows the SLS structure has better performance according to D-Q delay, PDP, and sensing current than the MLS structure because the SLS structure can optimize the FF operation and the sensing operation independently. Among various types of sensing circuit, the cross coupled inverter based sensing circuit including two MTJs and the single ended sensing circuit including two MTJs show better performances on low sensing current and high yield.
Taehui Na, Kyungho Ryu, Seung-Hyuk Kang, Seong-Ook Jung
ISCAS5
2013 ADDLL for Clock-Deskew Buffer in High-Performance SoCs
abstract
In this brief, we propose an all-digital delay locked loop (ADDLL) for a clock-deskew buffer. A low static phase offset at a high operating frequency is achieved by adopting a high-resolution window phase detector (PD) and a tristate-inverter-based ladder type coarse delay line (CDL). The proposed PD generates a high-resolution detection window that is adaptive to the process-voltage-temperature variation and reduces the static phase offset to nearly half of the fine delay line (FDL) resolution using a dual-output FDL. A proposed CDL is adopted in order to attain a small coarse delay step using tristate-inverters. The proposed ADDLL is designed using 0.13- μm process technology with a supply voltage of 1.2 V. The operating frequency range is 700 MHz to 2.0 GHz. The maximum static phase offset is less than 14.75 ps at all conditions and the power consumption is 4.0 mW at 2.0 GHz.
Jung-Hyun Park, Dong-Hoon Jung, Kyungho Ryu, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.4
2012 A Novel Sensing Circuit for Deep Submicron Spin Transfer Torque MRAM (STT-MRAM)
abstract
STT-MRAM has emerged as a compelling candidate for universal memory, but demands an advanced sensing circuit to achieve the proper sensing margin. In addition, STT-MRAM requires low-current sensing to avoid read disturbance. We report a novel sensing circuit that utilizes a source degeneration scheme and a balanced reference scheme. Monte Carlo HSPICE simulation results using 65 nm technology model parameters show that the proposed sensing circuit achieves an read access yield of 96.3% with a sensing current of 43.1 uA at a supply voltage of 1.1 V for 32 M bit, whereas the conventional sensing circuit achieves an read access yield of only 0% (81.5%) with a sensing current of 48.0 uA (64.2 uA) at a supply voltage of 1.1 V (1.6 V).
Kyungho Ryu, Seung-Hyuk Kang, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.4
2012 A Magnetic Tunnel Junction Based Zero Standby Leakage Current Retention Flip-Flop
abstract
Recently, a magnetic tunnel junction (MTJ), which is a strong candidate as a next-generation memory element, has been used not only as a memory cell but also in spintronics logic because of its excellent properties of nonvolatility, no silicon area occupation, and CMOS process compatibility. One of the representative research areas for the spintronics logic is the zero standby leakage retention flip-flop. Conventional zero standby leakage retention flip-flops have several problems, including difficulty in design optimization among the C-Q delay, sensing current, and process variation tolerance, and the insufficient write current. In this paper, a new MTJ based retention flip-flop is presented to solve these problems. The proposed retention flip-flop is designed using industry-compatible 45-nm process technology model. The proposed retention flip-flop achieves a 41.58% reduced C-Q delay and a 67.53% lowered sensing current with a 1.06% increased area compared to the previous retention flip-flop.
Kyungho Ryu, Jiwan Jung, Jung Pill Kim, Seung-Hyuk Kang, Seong-Ook Jung
IEEE Trans. Very Large Scale Integr. Syst.6
2007 Slope interconnect effort: gate-interconnect interdependentdelay model for CMOS logic gates
abstract
We present a circuit delay framework in a closed form that accounts for the dynamic behavior of signal slope in subthreshold (VDD < VT) as well as superthreshold (VDD > VT) regions. The proposed model converts a signal slope into its effective fanout for delay estimation. Simulations show that for ISCAS benchmark circuits, our framework exhibits a speedup of three orders of magnitude over HSPICE with 5% error. Measured results in 65nm show that for a wide range of interconnect lengths and geometries, the proposed model predicts the circuit delay with 5.7% error at the supply voltage of VDD = 1.2V , and with 4.5% error at VDD = 0.4V .
Myeong-Eun Hwang, Seong-Ook Jung, Kaushik Roy 0001
ISLPED2
2005 A 32-bit carry lookahead adder using dual-path all-N logic
abstract
We have developed dual path all-N logic (DPANL) and applied it to 32-bit adder design for higher performance. The speed is significantly enhanced due to reduced capacitance at each evaluation node of dynamic circuits. The power saving is achieved due to reduced adder cell size and minimal race problem. Post-layout simulation results show that this adder can operate at frequencies up to 1.85 GHz for 0.35-/spl mu/m 1P4M CMOS technology and is 32.4% faster than the adder using all-N transistor (ANT). It also consumes 29.2% less power than the ANT adder. A 0.35-/spl mu/m CMOS chip has been fabricated and tested to verify the functionality and performance of the DPANL adder on silicon.
Ge Yang 0004, Seong-Ook Jung, Kwang-Hyun Baek, Soo Hwan Kim, Suki Kim
IEEE Trans. Very Large Scale Integr. Syst.2
2003 Timing constraints for domino logic gates with timing-dependent keepers
abstract
Low threshold voltage (V/sub t/) can be applied to domino logic to improve the performance in dual threshold voltage technology. Then, the keeper transistor should be up-sized to compensate for reduced noise margin due to the significant subthreshold current of low V/sub t/ transistor. However, a large keeper transistor degrades performance. To resolve the tradeoff between performance and noise margin, the authors propose a new domino logic which incorporates a dual keeper structure and delay logic gates. Detailed timing analysis of the proposed domino logic yields optimal timing conditions wherein a contention-free skew-tolerant window is maximized. A broad range of the skew-tolerant window connotes robustness against noise and design parameter variations, while reduced contention between keeper and evaluation NMOS transistors ensures high-speed switching. The authors show that the dual keeper structure increases noise tolerance and delay logic gates fortify signal skew tolerance. Simulation results verify that the proposed domino logic is robust to noise and signal skew while presenting high performance and power efficiency.
Seong-Ook Jung, Ki-Wook Kim
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2003 Minimum delay optimization for domino circuits - a coupling-aware approach
abstract
Minimum delay associated with the hold time requirement is of concern to circuit designers, since race-through hazards are inherent in any multiple clock organization or clock distribution tree irrespective of clock frequency. The monotonic property of domino logic aggravates the min-delay path failure through coupling-induced speedup. To tackle the min-delay problem for domino logic, we propose a min-delay optimization algorithm considering coupling effects. Experimental results indicate that our algorithm yields a significant increase of min-delay without incurring max-delay violation.
Ki-Wook Kim, Seong-Ook Jung, Taewhan Kim 0001
ACM Trans. Design Autom. Electr. Syst.2
2003 Coupling delay optimization by temporal decorrelation using dual threshold voltage technique
abstract
Coupling effect due to line-to-line capacitance is of serious concern in timing analysis of circuits in ultra deep submicrometer CMOS technology. Often coupling delay is heavily dependent on temporal correlation of signal switching in relevant wires. Temporal decorrelation by shifting timing window can alleviate performance degradation induced by tight coupling. This paper presents an algorithm for minimizing circuit delay through timing window modulation in dual V/sub t/ technology. Experimental results on the ISCAS85 benchmark circuits indicate that the critical delay will be reduced significantly when low V/sub t/ is applied properly.
Ki-Wook Kim, Seong-Ook Jung, Taewhan Kim 0001, Prashant Saxena, C. L. Liu 0001, S.-M. S. Kang
IEEE Trans. Very Large Scale Integr. Syst.2
2003 Noise-aware interconnect power optimization in domino logic synthesis
abstract
Realization of high-performance domino logic depends strongly on energy-efficient and noise-tolerant interconnect design in ultradeep submicrometer processes. We characterize the cycle-averaged power model for interconnects accounting for switching statistics and dynamic behaviors. For the sake of signal integrity, cross-coupling effects are also characterized, which reflect logical correlation between adjacent wires. Based on the new models for interconnect power and capacitive crosstalk, we optimize the coupling power consumed by interconnects with crosstalk constraints. Experimental results show that optimized designs save the power consumption about 14% on average.
Ki-Wook Kim, Seong-Ook Jung, Unni Narayanan, C. L. Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2002 Low-swing clock domino logic incorporating dual supply and dual threshold voltages
abstract
High-speed domino logic is now prevailing in performance critical block of a chip. Low Voltage Swing Clock (LVSC) domino logic family is developed for substantial dynamic power saving. To boost up the transition speed in proposed circuitry, a well-established dual threshold voltage technique is exploited. Dual supply voltage technique in the LVSC domino logic is geared to reduce power consumption in clock tree and logic gates effectively. Delay Constrained Power Optimization (DCPO) algorithm allocates low supply voltage to logic gates such that dynamic power consumed by logic gates is minimized. Delay time variations due to gate-to-source voltage change and and input signal arrival time difference are considered for accurate timing analysis in DCPO.
Seong-Ook Jung, Ki-Wook Kim
DAC1
2002 Dual Threshold Voltage Domino Logic Synthesis for High Performance with Noise and Power Constrain
abstract
We introduce a new dual threshold voltage technique for domino logic. Since domino logic is much more sensitive to noise, noise margins have to be taken into account when applying dual threshold voltages to domino logic. To guarantee the signal integrity in domino logic, we carefully consider the effect of transistor sizing and threshold voltage selection. For optimal design, tradeoffs need to be mad? among noise margin, power, and performance. Based on the characteristics of each logic gate, we propose noise and power constrained domino logic synthesis for high performance. ISCAS85 benchmark re sults show that performance can be improved up to 18.62%, with 2% active power increase, while maintaining noise margin.
Seong-Ook Jung, Ki-Wook Kim
DATE1
2002 Noise constrained transistor sizing and power optimization for dual Vst domino logic
abstract
Dynamic logic is susceptible to noise, especially in the ultra-deep submicrometer dual threshold voltage technology. When the dual threshold voltage is applied to the domino logic, noise immunity must be carefully considered since the significant subthreshold current of the low threshold voltage transistor makes the dynamic node much more susceptible to noise. In the first part of this paper, we introduce a new keeper transistor sizing method to determine the optimal keeper transistor size in terms of speed, power, and noise immunity. With the use of data obtained by presimulation, it is unnecessary to simulate all the design corners corresponding to the feasible NMOS evaluation transistor size ranges to find the optimal keeper transistor size. HSPICE simulation results show that the proposed keeper transistor sizing method can be broadly applied to all the domino logic gates. In the second part of this paper, we propose a new dual threshold voltage domino logic synthesis with the keeper transistor sizing to minimize the power consumption while meeting delay and noise constraints. With the optimal keeper transistor size determined by the proposed keeper transistor sizing method, the dual threshold voltage assignment to domino logic can be simplified to the discrete threshold voltage selection. Experimental results for ISCAS85 benchmark circuits show significant savings on leakage power and active power.
Seong-Ook Jung, Ki-Wook Kim
IEEE Trans. Very Large Scale Integr. Syst.1
2001 Coupling Delay Optimization by Temporal Decorrelation using Dual Threshold Voltage Technique
abstract
Coupling effect due to line-to-line capacitance is of serious concern in timing analysis of circuits in ultra deep submicron CMOS technology. Often coupling delay is strongly dependent on temporal correlation of signal switching in relevant wires. Temporal decorrelation by shifting timing window can alleviate performance degradation induced by tight coupling. This paper presents an algorithm for minimizing circuit delay through timing window modulation in dual $V_t$ technology. Experimental results on the ISCAS85 benchmark circuits indicate that the critical delay will be reduced significantly when low $V_t$ is applied properly.
Ki-Wook Kim, Seong-Ook Jung, Prashant Saxena, C. L. Liu 0001
DAC2
2001 Transistor sizing for reliable domino logic design in dual threshold voltage technologies
abstract
Dynamic logic is much susceptible to noise specially in ul tra deep submicron technology The keeper transistor has to be carefully sized to maintain noise margin without much speed penalty In this paper we analyze the keeper tran sistor sizing with respect to the size of NMOS transistors in the evaluation tree Based on the analytical results we propose a keeper transistor sizing method HSPICE simula tion results show that the proposed keeper transistor sizing method can be broadly applied to all domino logic gates
Seong-Ook Jung, Ki-Wook Kim
ACM Great Lakes Symposium on VLSI1
2001 2-level LFSR scheme with asynchronous test pattern transfer for low cost and high efficiency build-in-self-test
Seung-Moon Yoo, Seong-Ook Jung
ACM Great Lakes Symposium on VLSI2
2000 Parallel dynamic logic (PDL) with speed-enhanced skewed static (SSS) logic
abstract
In this paper, we describe parallel dynamic logic (PDL) which exhibits high speed and no charge sharing problem. PDL uses only parallel-connected transistors for logic evaluation and is a good candidate for high-speed low-voltage operation. It has less back-bias effect compared to other logic styles which use stacked transistors. Furthermore, PDL needs no signal ordering nor tapering. PDL with speed-enhanced skewed static logic renders straightforward logic synthesis without area penalty due to logic duplication. Our experimental results on two 32-bit carry look ahead adders using 0.25 /spl mu/m CMOS technology showed that PDL with speed-enhanced skewed static (SSS) logic improves performance over clock-delayed (CD)-domino by 15-27% and power delay by 20-37%.
Chulwoo Kim, Seong-Ook Jung, Kwang-Hyun Baek
ISCAS2
2000 Noise-aware power optimization for on-chip interconnect
abstract
Realization of high-performance domino logic depends strongly on energy-efficient and noise-tolerant interconnect design in ultra deep sub-micron processes. We characterize the cycle-averaged power model for interconnects accounting for switching statistics and dynamic behaviors. For the sake of signal integrity, cross-coupling effects are also characterized which reflect logical correlation between adjacent wires. Based on the new models for interconnect power and capacitive crosstalk, we optimize the coupling power consumed by interconnects with crosstalk constraints. Experimental results show that optimized designs save the power consumption significantly.
Ki-Wook Kim, Seong-Ook Jung, Unni Narayanan, C. L. Liu 0001
ISLPED2