EDBT 2026 Demo / reviewers in the wild / expert
Joo-Hyung Chae
dblp:190/1878
· DBLP profile ↗
14ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0001-6354-5612ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 0.58 Ops/Trans Analog Compute-in-Memory Macro with Input Operand Generator Supporting Wide Boolean Operations for Enhanced Versatility
Eun-Bi Koh, Jeong-Eun Ko, Joo-Hyung Chae |
ISCAS | 3 |
| 2025 | A VM-Terminated PAM-3 Transmitter Using Floating Middle Level With Enhanced Signal Integrity and Energy Efficiency for Low-Power Memory InterfacesabstractThis paper introduces a three-level pulse amplitude modulation single-ended transmitter using VM(half of$\text{V}_{\mathrm {DDQ}}$) termination, designed for next-generation low-power memory interfaces. The proposed signaling strategy, based on the VMtermination, improves signal integrity by reducing signaling current by 41.05% while maintaining the same voltage swing compared to existing VSS-terminated structures for low-power memory interfaces. This reduction in signaling current effectively decreases simultaneous switching output noise. Additionally, this approach alleviates the voltage headroom limitation and enables pre-emphasis adoption, resulting in a larger voltage swing and reduced power consumption compared to de-emphasis. The voltage difference between the drain and source of each pull-up and pull-down driver is balanced, achieving a good ratio of level mismatch (RLM). Moreover, the three-step ZQ calibration and transmission gate-based receiver-side VMtermination help achieve accurate on-resistance and improve output linearity, reducing signal reflection. The prototype chip, fabricated using a 65-nm CMOS process, has an area of 0.0184 mm2. It achieves a data rate of 22.5 Gb/s with an energy efficiency of 0.86 pJ/bit and a measured RLM of 99.2%. Chanheum Han, Ki-Soo Lee, Jun-Cheol Lee, Joo-Hyung Chae |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | A Capacitor-Coupled Offset-Canceled Sense Amplifier for DRAMs With Hidden Offset-Cancellation Time and Cross-Coupled Pre-SensingabstractAs the dynamic random-access memory (DRAM) process is being scaled down, the sensing margin of the bit-line sense amplifier (BLSA) is decreasing. This leads to sensing failure due to the offset and long sensing time of the BLSA. To address this issue, various offset compensation methods and technologies that reduce sensing time have been proposed. However, these technologies still exhibit sensing failure due to long sensing times in low-power and large-capacity memory with low supply voltage and high$C_{BLT}/C_{Cell}$ratio. In addition, previous BLSAs require additional offset compensation time, which further increases the sensing time. A hidden-offset cancellation time and cross-coupled pre-sensing capacitor-coupled offset-canceled sense amplifier (HCP_COSA) are proposed to address these problems. To reduce the sensing time, the separate offset cancellation (OC) cycle is eliminated by performing OC for the cross-coupled inverter and charge-sharing simultaneously. In addition, the cross-coupled inverter operation is used in the pre-sensing phase to further amplify the reference bit line (BLB) voltage with the opposite polarity of the read-out signal, thereby improving both sensing margin and time. The proposed HCP_COSA achieves up to 18.5% faster sensing time at a supply voltage of 0.9 V by increasing the difference between the bit line (BLT) and the BLB by 2.55-3.76 times before starting the main sensing (MS) phase, compared with the previous works. In addition, the proposed architecture increases the number of word lines that can be sensed by up to 33.3% and achieves a sensing yield of 100% even at a low supply voltage of 0.65 V. Ik-Hyeon Jeon, Joo-Hyung Chae |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2025 | Sub 0.1-pJ/bit 14-Gb/s Receiver With Stack-Reduced Slicer Embedding One-Tap DFE for Low-Power Memory InterfacesabstractThis brief presents a single-ended receiver (RX) with a decision feedback equalizer (DFE)-embedded and stack-reduced slicer using a DFE weight selection multiplexer (MUX). The RX employs a quarter-rate clocking architecture to reduce the on-chip clock (CK) frequency and ensure reliable operation under stringent DFE timing constraints. The slicer output is fed back to the DFE weight selection MUX integrated into the second-stage slicer, achieving a short feedback loop latency. In the proposed architecture, the number of stacked transistors in the slicer is reduced to three, thereby reducing the CK-to-Q delay and overall DFE feedback loop latency. This optimized design increases the feedback speed and alleviates DFE timing constraints, ensuring stable operation even at low supply voltages. A prototype RX was fabricated using a 65-nm CMOS process and had an area of 0.004 mm2. The proposed RX achieved a measured bit error rate (BER) below$10^{-12}$at a data rate of 14 Gb/s with an insertion loss of −12 dB and achieved a power efficiency of 0.097 pJ/bit with a supply voltage of 0.75 V. Ki-Soo Lee, Joo-Hyung Chae |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | A Dual-Multiplication-Mode and Reconfigurable Digital Compute-in-Memory Macro Using Precharge-Controlled 4T1C eDRAMabstractRecently, the demand for supporting diverse artificial intelligence (AI) operations within a single edge device has been increasing. However, existing compute-in-memory (CIM) architectures for edge AI are often optimized only for specific lightweight networks, limiting their adaptability. While digital CIM (DCIM) offers high accuracy and exhibits low sensitivity to variations in process, voltage, and temperature, their area and energy inefficiencies remain significant barriers to deployment in edge devices. This study proposes an embedded dynamic random access memory (eDRAM)-based dual-multiplication-mode (DMM) DCIM macro leveraging precharge-controlled 4T1C gain cells. The proposed design enables the AND/XNOR dual-multiplication mode within a 4T1C cell through controlled precharge states. It supports a wide range of neural networks for edge AI, including INT1–8 operations and binary neural networks, such as XNOR-net. An area-efficient adder tree structure reduced the transistor count by 19% compared to conventional adder trees. This structure also accommodates 1–8-bit signed and unsigned inputs and weights without wasting cells, enhancing reconfigurability. Fabricated using a 28-nm CMOS process, the 16-kb DMM-DCIM prototype chip demonstrated area and energy efficiencies of 28.9 TOPS/mm2and 190.3 TOPS/W for 4b–4b multiply-accumulate operations, respectively. Ho-Sung Lee, Ik-Hyeon Jeon, Eun-Bi Koh, Joo-Hyung Chae |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | A Single-Ended PAM-4 Transmitter Using Unstacked Tailless CML Driver and Coefficient-Corrected FFE for Memory InterfacesabstractThis paper presents a single-ended four-level pulse-amplitude modulation (PAM-4) transmitter using an unstacked tailless current-mode logic (CML) driver for memory interfaces. Compared with the voltage-mode (VM) driver commonly used for single-ended memory interfaces, the proposed CML driver has stable termination for impedance matching and a small pre-driver with low dynamic power consumption, which allow the transmitter to achieve a higher data rate and a better total energy efficiency. The unstacked driver structure with an auxiliary leg and a current calibration scheme leads to high PAM-4 linearity by compensating for channel-length modulation that causes current source variation while occupying a small area. The strength of the feed-forward equalization (FFE) distorted by channel-length modulation is also compensated by an additional pulse of the proposed coefficient-corrected equalization. A prototype chip fabricated in a 65-nm CMOS process has an area of 0.0172 mm$^{2}{}$. It achieves a data rate of 34 Gb/s/pin with an energy efficiency of 0.60 pJ/bit and a level separation mismatch ratio (RLM) of 0.987. Yong-Un Jeong, Joo-Hyung Chae |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | Single-Ended Receiver-Side Crosstalk Cancellation With Independent Gain and Timing Control for Minimum Residual FEXTabstractIn multi-lane interfaces using single-ended signaling such as memory interfaces, far-end crosstalk (FEXT) noise of the aggressor signal severely degrades signal integrity of the victim signal. A circuit for crosstalk cancellation (XTC) can reduce FEXT noise. However, the channel spacing makes the flight time difference between the forward and FEXT signals. As a result, the residual FEXT can remain. To further minimize the residual FEXT, this study proposes an XTC method to adjust the amplitude and timing independently. The timing is adjusted according to the passive element’s value of the differentiator, and the amplitude is varied by using the high-frequency boosting circuit. Moreover, a continuous-time linear equalizer (CTLE) and a 1-tap decision feedback equalizer (DFE) are applied to reduce inter-symbol interference. A prototype chip of 2-lane receiver was fabricated in a 55-nm CMOS process to verify this scheme. Using the proposed XTC, CTLE, and 1-tap DFE, timing margins of 0.21 UI and 0.36 UI were achieved in 6-mil and 15-mil channel spacings at 6.4 Gb/s, both at a bit error rate of$10^{-12}$. Yong-Un Jeong, Sungphil Choi, Suhwan Kim 0001, Joo-Hyung Chae |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | A 10 Gb/s/pin Single-Ended Transmitter With Reflection-Aided Duobinary Modulation for Dual-Rank Mobile Memory InterfacesabstractThe dual-rank configuration is one of the parallelization methods to increase the memory capacity for mobile applications. However, the stub mainly composed of the redistribution layer causes resonance reflections and its reflection interval reaches the bit period, which distorts signal and limits the signal bandwidth. A single-ended duobinary transmitter that utilizes the reflection is presented for dual-rank mobile memory interfaces where reflections dominate. The reflection is included in the transfer function for duobinary modulation, which allows the transmitter to have a wide eye opening and high energy efficiency. Two-tap reverse feed-forward equalization and slew-rate control are used to support the duobinary modulation in combination with the reflection. A prototype chip fabricated in a 65 nm CMOS process has an area of 0.0378 mm2and consumes 1.24 pJ/bit. It is verified at data rates of 8, 9 and 10 Gb/s/pin where the flight time of the 9 mm stub is 55.6 ps. Yong-Un Jeong, Sungphil Choi, Joo-Hyung Chae, Jaekwang Yun, Shin-Hyun Jeong, Suhwan Kim 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | A High-Accuracy and Fast-Correction Quadrature Signal Corrector Using an Adaptive Delay Gain Controller for Memory InterfacesabstractIn this paper, a quadrature signal corrector (QSC) with high accuracy and fast correction for memory interfaces is presented. An adaptive delay gain controller in the QSC adjusts each delay gain of four digitally controlled delay lines (DCDLs) separately depending on skew between the quadrature clocks, resulting in short correction time together with low residual skew. To validate the effectiveness of our QSC in memory interfaces, a quarter-rate single-ended 1-tap decision feedback equalizer (DFE) with the QSC was fabricated in a 65nm CMOS process. Using the adaptive delay gain controller, the QSC reduced the skew between the 3 GHz quadrature clocks from a maximum of 21.2 ps to 0.8 ps while correction time was reduced by a factor of 3.9 compared to that without using the adaptive delay gain controller. At 12 Gb/s, the DFE using our QSC achieved a BER of 10-12with an eye width of 140 mUI when the input clock skew is 13.2 ps. Hyunkyu Park 0002, Jae-Whan Lee, Yong-Un Jeong, Shin-Hyun Jeong, Suhwan Kim 0001, Joo-Hyung Chae |
ISCAS | 7 |
| 2019 | A 20Gb/s Dual-Mode PAM4/NRZ Single-Ended Transmitter with RLM CompensationabstractIn this paper, a 20Gb/s dual-mode four-level pulse amplitude modulation (PAM4)/non-return-to-zero (NRZ) single-ended voltage-mode transmitter is proposed. Its output drivers are composed of 60 basic source-series terminated (SST) driver units and 12 additional pull-up (PU) driver units. The additional PU driver units are used to reduce the eye height difference between four amplitude levels of PAM4. Implemented in 65nm CMOS technology, the active area of the transmitter is 0.06mm2including the clock buffer and IQ generator. It draws 61.5mW at 20Gb/s during PAM4 operation and 72mW at 10Gb/s during NRZ operation. Changho Hyun, Hyeongjun Ko, Joo-Hyung Chae, Hyunkyu Park 0002, Suhwan Kim 0001 |
ISCAS | 3 |
| 2019 | A Low-Power and Low-Noise 20: 1 Serializer with Two Calibration Loops in 55-nm CMOSabstractThe increasing data rate of serial links makes it difficult to match timing constraints of serializers in transmitters. Delay compensation clock buffers can alleviate this issue by matching the timing between data and clock. However, these buffers consume significant power and become sources of noise to the transmitter output. The problem is more serious for serializers other than 2n:1, and using only 2n:1 serializer could be a limitation on system design. In this paper, a 20:1 serializer using two calibration feedback loops is presented to solve this issue and reduce power consumption. The two loops detect the phase difference between data and clock, and automatically align the clock phase to the center of the data phase. The loops eliminate the power-consuming clock buffers on the critical clock path and operate at maximum quarter rate, enabling the transmitter to have low power consumption and high performance. A 6.4 Gb/s serializer prototype is fabricated in 55-nm CMOS process with a 1.2 V supply voltage. It achieves 97.5 ps eye width, which is 62.4% of a unit interval (UI) using PRBS-7 data, and its energy efficiency is 1.60 pJ/bit. Yong-Un Jeong, Joo-Hyung Chae, Sungphil Choi, Jaekwang Yun, Shin-Hyun Jeong, Suhwan Kim 0001 |
ISLPED | 2 |
| 2019 | A Quadrature Clock Corrector for DRAM Interfaces, With a Duty-Cycle and Quadrature Phase Detector Based on a Relaxation OscillatorabstractA quadrature clock corrector uses relaxation oscillators to detect duty-cycle and quadrature phase errors by transforming them into pairs of frequencies, which are then digitized and compared. It achieves good detection accuracy and can detect a wide range of duty-cycle and quadrature phase errors. The prototype is implemented in a 55-nm CMOS process with a supply voltage of 1.2 V and occupies an area of 0.003 mm2. The experimental results show that the operation range is from 1 to 3 GHz, the power efficiency is 0.79 mW/GHz, the maximum duty-cycle error is 0.8% at 3 GHz, and the maximum quadrature phase error is 1.1° at 3 GHz. Joo-Hyung Chae, Hyeongjun Ko, Suhwan Kim 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | Energy-Efficient Dynamic Comparator with Active Inductor for Receiver of Memory InterfacesabstractIn this paper, we propose a dynamic comparator that improved the operation performance of receiver (RX) with the effort to reduce power consumption. It is implemented via double-tail StrongARM latch comparator with an active inductor and efforts are made to minimize power consumption for high-speed resulting in better energy efficiency at the targeted high frequency. In this regard, our comparator is suitable for memory application RX to satisfy both low-power and high-speed. It is applied to the single-ended RX designed with a continuous-time linear equalizer, a clock generator and a quarter-rate 2-tap decision-feedback equalizer which is appropriate for the high-frequency memory application. Compared to the conventional one, our design, fabricated in 55nm CMOS process, provides an improvement of 7% in unit interval (UI) margin under the same power consumption and receives up to 10Gb/s PRBS15 data at BER < 10-12 with 0.4 UI margin and energy efficiency of 0.67pJ/bit. Jae-Whan Lee, Joo-Hyung Chae, Hyunkyu Park 0002, Jaekwang Yun, Suhwan Kim 0001 |
ISLPED | 2 |
| 2017 | A 0.13pJ/bit, referenceless transceiver with clock edge modulation for a wired intra-BAN communicationabstractIn this paper, we propose a low power transceiver (TRx) suitable for a wired intra-body area network (BAN) communication. The proposed transceiver is designed with relaxation oscillator which is appropriate for this low frequency (<; 50MHz) application. To lessen the complexity of building this BAN system, we use clock edge modulation (CEM) data as sending or receiving data, and this allows the transceiver to operate without reference clock. The relaxation oscillator in this transceiver is designed to be able to generate CEM data pattern as well as a clock, so this can minimize power consumption in designing additional block related to transmission. Proposed circuit operates up to 36MHz with 1.0V supply voltage. It consumes 1.26uW at an input data rate of 10Mbps and achieves 0.13pJ/bit of energy per bit even though the circuit is implemented in a 0.18μm CMOS technology. Gi-Moon Hong, Mino Kim, Joo-Hyung Chae, Suhwan Kim 0001 |
ISLPED | 4 |