Johnny Rhe

dblp:297/2393 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0003-2603-974XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 5 first-author · 12 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Low-Rank Compression for IMC Arrays
abstract
In this study, we address the challenge of low-rank model compression in the context of in-memory computing (IMC) architectures. Traditional pruning approaches, while effective in model size reduction, necessitate additional peripheral circuitry to manage complex dataflows and mitigate dislocation issues, leading to increased area and energy overheads. To circumvent these drawbacks, we propose leveraging low-rank compression techniques, which, unlike pruning, streamline the dataflow and seamlessly integrate with IMC architectures. However, low-rank compression presents its own set of challenges, namely i) suboptimal IMC array utilization and ii) compromised accuracy. To address these issues, we introduce a novel approach i) employing shift and duplicate kernel (SDK) mapping technique, which exploits idle IMC columns for parallel processing, and ii) group lowrank convolution, which mitigates the information imbalance in the decomposed matrices. Our experimental results demonstrate that our proposed method achieves up to 2.5× speedup or +20.9% accuracy boost over existing pruning techniques.
Kang Eun Jeon, Johnny Rhe, Jong Hwan Ko
DATE2
2025 Row-Column Hybrid Grouping for Fault-Resilient Multi-Bit Weight Representation on IMC Arrays
abstract
This paper addresses two critical challenges in analog In-Memory Computing (IMC) systems that limit their scalability and deployability: the computational unreliability caused by stuck-at faults (SAFs) and the high compilation overhead of existing fault-mitigation algorithms, namely Fault-Free (FF). To overcome these limitations, we first propose a novel multi-bit weight representation technique, termed row-column hybrid grouping, which generalizes conventional column grouping by introducing redundancy across both rows and columns. This structural redundancy enhances fault tolerance and can be effectively combined with existing fault-mitigation solutions. Second, we design a compiler pipeline that reformulates the fault-aware weight decomposition problem as an Integer Linear Programming (ILP) task, enabling fast and scalable compilation through off-the-shelf solvers. Further acceleration is achieved through theoretical insights that identify fault patterns amenable to trivial solutions, significantly reducing computation. Experimental results on convolutional networks and small language models demonstrate the effectiveness of our approach, achieving up to 8%p improvement in accuracy, 150 × faster compilation, and 2 × energy efficiency gain compared to existing baselines.
Kang Eun Jeon, Sangheum Yeon, Jinhee Kim, Hyeonsu Bang, Johnny Rhe, Jong Hwan Ko
ICCAD5
2025 Input/mapping precision controllable digital CIM with adaptive adder tree architecture for flexible DNN inference
Juhong Park, Johnny Rhe, Chanwook Hwang, Jaehyeon So, Jong Hwan Ko
J. Syst. Archit.2
2025 GAROS: Genetic algorithm-aided row-skipping for shift and duplicate kernel mapping in processing-in-memory architectures
Johnny Rhe, Kang Eun Jeon, Jong Hwan Ko
J. Syst. Archit.1
2024 KARS: Kernel-Grouping Aided Row-Skipping for SDK-based Weight Compression in PIM Arrays
abstract
With energy-efficient computation, processing-in-memory (PIM) architectures have been highlighted as one of the most viable candidates to substitute the traditional ones. Recently, shift and duplicate kernel (SDK) mapping method was proposed to enable efficient and fast convolutional neural networks (CNNs) inference in the PIM array. However, since its weight deployment to reuse input data, this method generates idle cells that do not involved in the computation, which leads to an increase of energy consumption. In this paper, we propose a novel weight mapping method called kernel-grouping aided row-skipping (KARS). KARS maximizes utilization by removing idle cells on a PIM array and reduces computing cycles. In comparison to the traditional methods, KARS achieves a speedup by up to 3× at Layer 2 of VGGNet-13 and ResNet-18.
Juhong Park, Johnny Rhe, Jong Hwan Ko
ISCAS2
2024 KERNTROL: Kernel Shape Control Toward Ultimate Memory Utilization for In-Memory Convolutional Weight Mapping
abstract
Processing-in-memory (PIM) architectures have been highlighted as one of the most viable options for faster and more power-efficient computation. Paired with a convolutional weight mapping scheme, PIM arrays can accelerate various deep convolutional neural networks (CNNs) and the applications that adopt them. Recently, shift and duplicate kernel (SDK) convolutional weight mapping scheme was proposed, achieving up to 50% throughput improvement over the prior arts. However, the traditional pattern-based pruning methods, which were adopted for row-skipping and computing cycle reduction, are not optimal for the latest SDK mapping due to the loss of structural regularity caused by the shifted and duplicated kernels. To address this challenge, we propose kernel shape control (KERNTROL), a method where kernel shapes are controlled depending on their mapped columns with the purpose of fostering a structural regularity that is favorable in achieving a high row-skipping ratio and model accuracy. Instead of permanently pruning the weights, KERNTROL with an empty mask (KERNTROL-M) temporarily omits them in the underutilized row using a utilization threshold, thereby preserving important weight elements. However, a significant portion of the memory cells is still underutilized where the threshold is not enforced. To overcome this, we extend KERNTROL-M into KERNTROL with compensatory weights (KERNTROL-C). By populating idle cells with compensatory weights, KERNTROL-C can offset the accuracy drop from weight omission. In comparison to pattern-based pruning approaches, KERNTROL-C achieves simultaneous improvements of up to 36.4% improvement in the compression rate and 5% in model accuracy with up to 100% array utilization.
Johnny Rhe, Kang Eun Jeon, Joo Chan Lee, Seongmoon Jeong, Jong Hwan Ko
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 Kernel Shape Control for Row-Efficient Convolution on Processing-In-Memory Arrays
abstract
Processing-in-memory (PIM) architectures have been highlighted as one of the viable solutions for faster and more power-efficient convolutional neural networks (CNNs) inference. Recently, shift and duplicate kernel (SDK) convolutional weight mapping scheme was proposed, achieving up to 50% through-put improvement over the prior arts. However, the traditional pattern-based pruning methods, which were adopted for row-skipping and computing cycle reduction, are not optimal for the latest SDK mapping due to structural irregularity caused by the shifted and duplicated kernels. To address this issue, we propose a method called kernel shape control (KERNTROL) that aims to promote structural regularity for achieving a high row-skipping ratio and model accuracy. Instead of pruning certain weight elements permanently, KERNTROL controls the kernel shapes through the omission of certain weights based on their mapped columns. In comparison to the latest pattern-based pruning approaches, KERNTROL achieves up to 36.4% improvement in the compression rate, and 38.6% in array utilization with maintaining the original model accuracy.
Johnny Rhe, Kang Eun Jeon, Joo Chan Lee, Seongmoon Jeong, Jong Hwan Ko
ICCAD1
2023 DCR: Decomposition-Aware Column Re-Mapping for Stuck-At-Fault Tolerance in ReRAM Arrays
abstract
The ReRAM-based neuromorphic computing system (NCS) has been widely used as an energy-efficient platform for deep neural network (DNN) acceleration. However, ReRAM commonly suffers from stuck-at-fault (SAF), resulting in permanent device failure. SAF tolerance is an essential task to ensure the reliability of the system by minimizing the DNN inference accuracy degradation. Since hardware-based solutions incur additional overhead and power consumption, it is necessary to seek a solution that can be executed offline to mitigate the impact of SAF. In this work, we propose a decomposition-aware column re-mapping (DCR) for SAF tolerance in analog ReRAM arrays (RAs). Our DCR consists of the column re-mapping technique combined with fault-aware weight decomposition and an advanced sensitivity metric. As a result, it generates a final weight map optimized for the fault map. Our DCR achieves only about 1% loss of inference accuracy on CIFAR-10 and CIFAR-100 for the analog RAs with the SAF rate of 2% and 1%, respectively, without any hardware-based solution or re-training.
Hyeonsu Bang, Kang Eun Jeon, Johnny Rhe, Jong Hwan Ko
ICCD3
2023 Weight-Aware Activation Mapping for Energy-Efficient Convolution on PIM Arrays
abstract
Convolutional weight mapping plays a stapling role in facilitating convolution operations on Processing-in-memory (PIM) architecture which is, at its essence, a matrix-vector multiplication (MVM) accelerator. Despite its importance, convolutional mapping methods are under-studied and existing mapping methods fail to exploit the sparse and redundant characteristics of heavily quantized convolutional weights, leading to low array utilization and ineffectual computations. To address these issues, this paper proposes a novel weight-aware activation mapping method where activations are mapped onto the memory cells instead of the weights. The proposed method significantly reduces the number of computing cycles by skipping zero-valued weights and merging those PIM array rows with the same weight values. Experimental results on ResNet-18 demonstrate that the proposed weight-aware activation mapping can achieve up to 90% energy saving and latency reduction compared to the conventional approaches.
Kang Eun Jeon, Johnny Rhe, Hyeonsu Bang, Jong Hwan Ko
ISLPED2
2023 PAIRS: Pruning-AIded Row-Skipping for SDK-Based Convolutional Weight Mapping in Processing-In-Memory Architectures
abstract
Processing-in-memory (PIM) architecture is becoming a promising candidate for convolutional neural network (CNN) inference. A recent weight mapping method called shift and duplicate kernel (SDK) improves the utilization by the deployment of shifting the same kernels into idle columns. However, this method inevitably generates idle cells with an irregular distribution, which limits reducing the size of the weight matrix. To effectively compress the weight matrix in the PIM array, prior works have introduced a row-wise pruning scheme, one of the structured weight pruning schemes, that aims to skip the operation on a row by zeroing out all weight in the specific row (we call it row-skipping). However, due to the deployment of shifting kernels, SDK mapping complicates zeroing out all the weight in the same row. To address this issue, we propose pruning-aided row-skipping (PAIRS) that effectively reduces the number of rows of convolutional weights that are mapped with SDK mapping. By pairing the SDK mapping-aware pruning pattern design and row-wise pruning, PAIRS achieves a higher row-skipping ratio. In comparison to pruning methods, PAIRS achieves up to$1.95\times$rows skipped and$4\times$higher compression rate with similar or even better inference accuracy.
Johnny Rhe, Kang Eun Jeon, Jong Hwan Ko
ISLPED1
2022 VW-SDK: Efficient Convolutional Weight Mapping Using Variable Windows for Processing-In-Memory Architectures
abstract
With their high energy efficiency, processing-in-memory (PIM) arrays are increasingly used for convolutional neural network (CNN) inference. In PIM-based CNN inference, the computational latency and energy are dependent on how the CNN weights are mapped to the PIM array. A recent study proposed shifted and duplicated kernel (SDK) mapping that reuses the input feature maps with a unit of a parallel window, which is convolved with duplicated kernels to obtain multiple output elements in parallel. However, the existing SDK-based mapping algorithm does not always result in the minimum computing cycles because it only maps a square-shaped parallel window with the entire channels. In this paper, we introduce a novel mapping algorithm called variable-window SDK (VW-SDK), which adaptively determines the shape of the parallel window that leads to the minimum computing cycles for a given convolutional layer and PIM array. By allowing rectangular-shaped windows with partial channels, VW-SDK utilizes the PIM array more efficiently, thereby further reduces the number of computing cycles. The simulation with a$512\times 512$PIM array and Resnet-18 shows that VW-SDK improves the inference speed by$1.69\times$compared to the existing SDK-based algorithm.
Johnny Rhe, Sungmin Moon, Jong Hwan Ko
DATE1
2021 A Charge-Domain Scalable-Weight In-Memory Computing Macro With Dual-SRAM Architecture for Precision-Scalable DNN Accelerators
abstract
This paper presents a charge-domain in-memory computing (IMC) macro for precision-scalable deep neural network accelerators. The proposed Dual-SRAM cell structure with coupling capacitors enables charge-domain multiply and accumulate (MAC) operation with variable-precision signed weights. Unlike prior charge-domain IMC macros that only support binary neural networks or digitally compute weighted sums for MAC operation with multi-bit weights, the proposed macro implements analog weighted sums for energy-efficient bit-scalable MAC operations with a novel series-coupled merging scheme. A test chip with a 16-kb SRAM macro is fabricated in 28-nm FDSOI process, and the measured macro throughput is 125.2-876.5 GOPS for weight bit-precision varying from 2 to 8. The macro also achieves energy efficiency ranging from 18.4 TOPS/W for 8-b weight to 119.2 TOPS/W for 2-b weight.
Eunyoung Lee, Taeyoung Han, Gicheol Shin, Jaerok Kim, Soyoun Jeong, Johnny Rhe, Jaehyun Park 0012, Jong Hwan Ko, Yoonmyung Lee
IEEE Trans. Circuits Syst. I Regul. Pap.8