VLDB 2026 Research / reviewers in the wild / expert
Jong-Hyeok Yoon
dblp:149/1461
· DBLP profile ↗
8ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0001-7373-7028ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Hybrid Digital-Analog Compute-in-Memory Using Content-Addressable Memory With Flexible Multi-Bit SlicingabstractCompute-in-memory (CIM) reduces data movement and enhances compute parallelism, making it suitable for AI applications. However, analog CIMs, yet energy-efficient, are vulnerable to PVT variations, while digital CIMs offer robustness but limited efficiency due to their bit-wise computation overhead. To address these challenges, we propose a hybrid CIM architecture that integrates content-addressable memory (CAM) and cluster-based CIM, named CAM-CIM, fabricated in 65nm CMOS technology. The proposed CAM-CIM flexibly slices multi-bit weights, assigning MSBs to CAM and LSBs to CIM, enabling dynamic accuracy-efficiency trade-offs across various bit precisions. A two-stage 8:3 compressor-based adder tree improves CAM efficiency and a reference voltage search algorithm ensures accurate CIM computation with low-bit ADCs. Our CAM-CIM supports 1-8b inputs/weights with reconfigurable compute modes, leveraging ternary-CAM based selective columns and cluster-wise CIM processing to produce multiple trade-off points even in the same bit precision. A prototype chip with a RISC-V controller and custom instructions is demonstrated that shows energy efficiencies of 32.4TOPS/W (8b/8b) and 76.0-354.9TOPS/W (4b/4b) with 0.66% accuracy loss, on average, across a wide range of DNN benchmarks including CNNs and vision transformers on CIFAR and ImageNet datasets. Sangwoo Jung 0001, Dahoon Park, Hyunseob Shin, Jong-Hyeok Yoon, Jaeha Kung 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | CAM-CIM: A Hybrid Compute-in-Memory Using Content-Addressable Memory with Subword Split Mapping for Reduced ADC ResolutionabstractRecently, compute-in-memory (CIM) has become a promising architecture for data-intensive applications such as deep learning. However, analog or digital CIM (ACIM or DCIM) faces some design challenges. ACIMs inherently have non-idealities, which lead to significant accuracy degradation. In addition, a substantial amount of power is consumed by analog-to-digital converters (ADC). On the other hand, DCIMs show an exponential increase in power consumption and computing cycles as the operand bit-width increases, particularly due to an accumulation stage. In this paper, to overcome these challenges, we propose a hybrid DCIM-ACIM architecture that consists of a content addressable memory (CAM) as DCIM and a cluster-based multi-cycle ACIM, called CAM-CIM. As a weight mapping strategy, we present a subword split mapping that assigns some MSBs to DCIM for improved accuracy and the remaining LSBs to ACIM for reduced ADC resolution. The accuracy of using the proposed CAM-CIM array is evaluated on various deep learning benchmarks from CNNs to Swin-Tiny. A 65nm CAM-CIM macro with either 3-bit or 4-bit ADCs shows 10.3 × and 5.4 × improvement in energy efficiency, on average, compared to CAM- and CIM-only architectures, respectively. Compared to recent CIM architectures, CAM-CIM demonstrates 1.4 × higher energy efficiency. Sangwoo Jung 0001, Dahoon Park, Hyunseob Shin, Jong-Hyeok Yoon, Jaeha Kung 0001 |
ISLPED | 7 |
| 2024 | A Dual-Precision and Low-Power CNN Inference Engine Using a Heterogeneous Processing-in-Memory ArchitectureabstractIn this article, we present an energy-scalable CNN model that can adapt to different hardware resource constraints. Specifically, we propose a dual-precision network, named DualNet, that leverages two independent bit-precision paths (INT4 and ternary-binary). DualNet achieves both high accuracy and low complexity by balancing the ratio between two paths. We also present an evolutionary algorithm that allows the automatic search of the optimal ratios. In addition to the novel CNN architecture design, we develop a heterogeneous processing-in-memory (PIM) hardware that integrates SRAM-and eDRAM-based PIMs to efficiently compute two precision paths in parallel. To verify the energy efficiency of DualNet computed on the heterogeneous PIM, we prototyped a test chip in 28nm CMOS technology. To maximize the hardware efficiency, we utilize an improved data mapping scheme achieving the most effective deployment of DualNets on multiple PIM arrays. With the proposed SW-HW co-optimization, we can obtain the most energy-efficient DualNet model operating on the actual PIM hardware. Compared to the other quantized networks with a single bit-precision, DualNet reduces the energy consumption, memory footprint, and latency by 29.0%, 49.5%, 47.3% on average, respectively, for CIFAR-10/100 and ImageNet datasets. Sangwoo Jung 0001, Dahoon Park, Youngjoo Lee 0002, Jong-Hyeok Yoon, Jaeha Kung 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | Improving compute in-memory ECC reliability with successive correctionabstractCompute in-memory (CIM) is an exciting technique that minimizes data transport, maximizes memory throughput, and performs computation on the bitline of memory sub-arrays. This is especially interesting for machine learning applications, where increased memory bandwidth and analog domain computation offer improved area and energy efficiency. Unfortunately, CIM faces new challenges traditional CMOS architectures have avoided. In this work, we explore the impact of device variation (calibrated with measured data on foundry RRAM arrays) and propose a new class of error correcting codes (ECC) for hard and soft errors in CIM. We demonstrate single, double, and triple error correction offering over 16,000× reduction in bit error rate over a design without ECC and over 427× over prior work, while consuming only 29.1% area and 26.3% power overhead. Brian Crafton, Zishen Wan, Samuel Spetalnick, Jong-Hyeok Yoon, Carlos Tokunaga, Vivek De, Arijit Raychowdhury |
DAC | 4 |
| 2022 | Characterization and Mitigation of IR-Drop in RRAM-based Compute In-MemoryabstractCompute in-memory (CIM) is an exciting circuit innovation that promises to increase effective memory bandwidth and perform computation on the bitlines of memory sub-arrays. Utilizing embedded non-volatile memories (eNVM) such as resistive random access memory (RRAM), various forms of neural networks can be implemented. Unfortunately, CIM faces new challenges traditional CMOS architectures have avoided. In this work, we characterize the impact of IR-drop and device variation (calibrated with measured data on foundry RRAM) and evaluate different approaches to write verify. Using various voltages and pulse widths we program cells to offset IR-drop and demonstrate a $136.4 \times $ reduction in BER during CIM. Brian Crafton, Connor Talley, Samuel Spetalnick, Jong-Hyeok Yoon, Arijit Raychowdhury |
ISCAS | 4 |
| 2022 | BitS-Net: Bit-Sparse Deep Neural Network for Energy-Efficient RRAM-Based Compute-In-MemoryabstractThe rising popularity of intelligent mobile devices and the computational cost of deep learning-based models call for efficient and accurate on-device inference schemes. We propose a novel model compression scheme that allows inference to be carried out using bit-level sparsity, which can be efficiently implemented using in-memory computing macros. In this paper, we introduce a method called BitS-Net to leverage the benefits of bit-sparsity (where the number of zeros are more than number of ones in binary representation of weight/activation values) when applied to compute-in-memory (CIM) with resistive RAM (RRAM) to develop energy efficient DNN accelerators operating in the inference mode. We demonstrate that BitS-Net improves the energy efficiency by up to 5x for ResNet models on the ImageNet dataset. Foroozan Karimzadeh, Jong-Hyeok Yoon, Arijit Raychowdhury |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Statistical Optimization of Compute In-Memory Performance Under Device VariationabstractCompute in-memory (CIM) is a promising technique that minimizes data transport, maximizes memory throughput, and performs computation on the bitline of memory sub-arrays. Utilizing embedded non-volatile memories (eNVM) such as resistive random access memory (RRAM), various forms of neural networks can be implemented. Unfortunately, CIM faces new challenges traditional CMOS architectures have avoided. In this work, we explore the impact of device variation (calibrated with measured data on foundry RRAM arrays) and propose a new algorithm based on device variation to increase both performance and accuracy for CIM designs. We demonstrate up to 36% power improvement and 44% performance improvement, while satisfying any error constraint. Brian Crafton, Samuel Spetalnick, Jong-Hyeok Yoon, Arijit Raychowdhury |
ISLPED | 3 |
| 2016 | A 4×10-Gb/s Referenceless-and-Masterless Phase Rotator-Based Parallel Transceiver in 90-nm CMOSabstractA four-parallel 10-Gb/s referenceless-and-masterless phase rotator-based transceiver is presented. Entire lanes operate independently just like the conventional voltage-controlled-oscillator-based parallel referenceless designs while saving power and area. The measured recovered-clock jitter in each lane is 1.24 psrms and the transceiver surpasses the OC-192 jitter-tolerance specification. The power efficiency of the proposed parallel transceiver fabricated in a 90-nm CMOS process is 6.325 mW/(Gb/s). Joon-Yeong Lee, Jaehyeok Yang, Jong-Hyeok Yoon, Soon-Won Kwon, Hyosup Won, Jinho Han, Hyeon-Min Bae |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |