Jaeyong Jang

dblp:350/3841 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Reliability-Enhanced Offset-Canceling Current-Sampling Sense Amplifier for 2T-2MTJ MRAM PUF
abstract
In this paper, we propose a reliability-enhanced offset-canceling current sampling sense amplifier (REOCCS-SA) with a clamp voltage trimming technique to improve uniformity and uniqueness in spin-transfer-torque magnetic random-access memory (STT-MRAM)-based physically unclonable function (PUF). REOCCS-SA retains the structure of the OCCS-SA while incorporating a preliminary voltage amplification phase to enhance offset tolerance, thereby improving uniformity and uniqueness. Additionally, the clamp voltage trimming technique utilizes a resistor ladder to precisely adjust the clamp voltage, optimizing the reference current to achieve ideal uniformity. HSPICE simulations based on a 28 nm technology model show that the proposed REOCCS-SA reduces the standard deviation of uniformity by 75% and improves inter-Hamming distance (inter-HD) by 50% compared to the OCCS-SA. In addition, compared to other circuits proposed within the past decade, it achieves the lowest figure of merit, demonstrating the highest overall efficiency when operating as an STT-MRAM-based PUF. Furthermore, by applying theV${}_{\mathbf {CMP\_R}}$trimming and the automatic write-back technique, the PUF system using REOCCS-SA exhibits 49.95% uniformity, 50.83% inter-HD, and 0% intra-Hamming distance, confirming its overall robustness and stability.
Mingu Han, Bayartulga Ishdorj, Jaeyong Jang, Taehui Na
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 An eDRAM Digital In-Memory Neural Network Accelerator for High-Throughput and Extended Data Retention Time
abstract
Computing-in-Memory (CIM) optimizes multiply-and-accumulate (MAC) operations for energy-efficient acceleration of neural network models. While SRAM has been a popular choice for CIM designs due to its compatibility with logic processes, its large cell size restricts storage capacity for neural network parameters. Consequently, gain-cell eDRAM, featuring memory cells with only 2–4 transistors, has emerged as an alternative for CIM cells. While digital CIM (DCIM) structure has been actively adopted in SRAM-based CIMs for better accuracy and scalability than analog CIMs (ACIM), previous eDRAM-based CIMs still employed ACIM structure since the eDRAM CIM cells were not able to perform a complete digital logic operation. In this paper, we propose an eDRAM bit cell for more efficient DCIM operations using only 4 transistors. The proposed eDRAM DCIM structure also maintains consistent and accurate output values over time, improving retention times compared to previous eDRAM ACIM designs. We validate our approach by fabricating an eDRAM DCIM macro chip and conducting hardware validation experiments, measuring retention time and neural network accuracy. Experimental results show that the proposed eDRAM DCIM achieves 3× longer retention time than state-of-the-art eDRAM ACIM designs, along with higher throughput without accuracy loss.
Inhwan Lee, Jehun Lee, Jaeyong Jang, Jae-Joon Kim
DATE3
2024 FIGNA: Integer Unit-Based Accelerator Design for FP-INT GEMM Preserving Numerical Accuracy
abstract
The weight-only quantization has emerged as a promising technique for alleviating the computational burden of large language models (LLMs) by employing low-precision integer (INT) weights, while retaining full-precision floating point (FP) activations to ensure inference quality. Despite the memory footprint reduction achieved through decreased bit-precision of weight parameters, the actual computing performance is often not improved significantly due to FP-INT multiply-accumulation (MAC) operations being performed on the floating point unit (FPU) after de quantizing the INT weight values to FP values, owing to the lack of dedicated FP- INT arithmetic units. In this study, we investigate the impact of introducing a dedicated FP-INT unit on overall performance and find that such specialization does not yield substantial improvements. As an alternative approach, we propose FIGNA, an accelerator based on INT units designed specifically for FP- INT MAC operations. A key feature of FIGNA is its ability to achieve the same numerical accuracy as the FPU while relying solely on the integer-unit, a departure from prior methods that relied on integer-units with numerical approximations for FP arithmetic results, albeit claiming similar inference accuracy through dedicated network training. Through comprehensive experiments on FP- INT quantized networks for LLMs, including OPT and BLOOM, we demonstrate the superior performance of FIGNA compared to conventional FPUs in terms of performance per area ($TOPS/mm^{2}$) and energy efficiency (TOPS/W) across various input and weight precision combinations. For instance, in the FP16-INT4 case, FIGNA shows 6.34x higher$TOPS/ mm^{2}$and 2.19x higher TOPS/W compared to the baseline.
Jaeyong Jang, Yulhwa Kim, Juheun Lee, Jae-Joon Kim
HPCA1
2024 Mobileware: Distributed Architecture With Channel Stationary Dataflow for MobileNet Acceleration
abstract
The depthwise separable convolution, a key feature of the MobileNet models, has a different input reuse pattern from the conventional standard convolution, and a smaller number of input/weight pairs are used for a dot product, thereby leading to extremely low MAC utilization. This paper proposes a Mobileware architecture for the high-performance acceleration of the MobileNet workloads. A new channel stationary dataflow architecture distributes the on-chip buffers, and the distributed SRAMs are placed near each PE. By doing so, PEs and SRAMs can communicate with high bandwidth. Our Mobileware architecture shows 1.4-29.5× higher throughput than conventional weight stationary-based hardware architecture, and the proposed design was verified on the Xilinx ZCU102 FPGA evaluation board.
Sungju Ryu, Jaeyong Jang, Youngtaek Oh, Jae-Joon Kim
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Winning Both the Accuracy of Floating Point Activation and the Simplicity of Integer Arithmetic
Yulhwa Kim, Jaeyong Jang, Jehun Lee, Byeongwook Kim, Baeseong Park, Se Jung Kwon, Dongsoo Lee, Jae-Joon Kim
ICLR2