EDBT 2026 Demo / reviewers in the wild / expert
Jaeyong Jang
dblp:350/3841
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reliability-Enhanced Offset-Canceling Current-Sampling Sense Amplifier for 2T-2MTJ MRAM PUFabstractIn this paper, we propose a reliability-enhanced offset-canceling current sampling sense amplifier (REOCCS-SA) with a clamp voltage trimming technique to improve uniformity and uniqueness in spin-transfer-torque magnetic random-access memory (STT-MRAM)-based physically unclonable function (PUF). REOCCS-SA retains the structure of the OCCS-SA while incorporating a preliminary voltage amplification phase to enhance offset tolerance, thereby improving uniformity and uniqueness. Additionally, the clamp voltage trimming technique utilizes a resistor ladder to precisely adjust the clamp voltage, optimizing the reference current to achieve ideal uniformity. HSPICE simulations based on a 28 nm technology model show that the proposed REOCCS-SA reduces the standard deviation of uniformity by 75% and improves inter-Hamming distance (inter-HD) by 50% compared to the OCCS-SA. In addition, compared to other circuits proposed within the past decade, it achieves the lowest figure of merit, demonstrating the highest overall efficiency when operating as an STT-MRAM-based PUF. Furthermore, by applying theV${}_{\mathbf {CMP\_R}}$trimming and the automatic write-back technique, the PUF system using REOCCS-SA exhibits 49.95% uniformity, 50.83% inter-HD, and 0% intra-Hamming distance, confirming its overall robustness and stability. Mingu Han, Bayartulga Ishdorj, Jaeyong Jang, Taehui Na |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | An eDRAM Digital In-Memory Neural Network Accelerator for High-Throughput and Extended Data Retention TimeabstractComputing-in-Memory (CIM) optimizes multiply-and-accumulate (MAC) operations for energy-efficient acceleration of neural network models. While SRAM has been a popular choice for CIM designs due to its compatibility with logic processes, its large cell size restricts storage capacity for neural network parameters. Consequently, gain-cell eDRAM, featuring memory cells with only 2–4 transistors, has emerged as an alternative for CIM cells. While digital CIM (DCIM) structure has been actively adopted in SRAM-based CIMs for better accuracy and scalability than analog CIMs (ACIM), previous eDRAM-based CIMs still employed ACIM structure since the eDRAM CIM cells were not able to perform a complete digital logic operation. In this paper, we propose an eDRAM bit cell for more efficient DCIM operations using only 4 transistors. The proposed eDRAM DCIM structure also maintains consistent and accurate output values over time, improving retention times compared to previous eDRAM ACIM designs. We validate our approach by fabricating an eDRAM DCIM macro chip and conducting hardware validation experiments, measuring retention time and neural network accuracy. Experimental results show that the proposed eDRAM DCIM achieves 3× longer retention time than state-of-the-art eDRAM ACIM designs, along with higher throughput without accuracy loss. Inhwan Lee, Jehun Lee, Jaeyong Jang, Jae-Joon Kim |
DATE | 3 |
| 2024 | FIGNA: Integer Unit-Based Accelerator Design for FP-INT GEMM Preserving Numerical AccuracyabstractThe weight-only quantization has emerged as a promising technique for alleviating the computational burden of large language models (LLMs) by employing low-precision integer (INT) weights, while retaining full-precision floating point (FP) activations to ensure inference quality. Despite the memory footprint reduction achieved through decreased bit-precision of weight parameters, the actual computing performance is often not improved significantly due to FP-INT multiply-accumulation (MAC) operations being performed on the floating point unit (FPU) after de quantizing the INT weight values to FP values, owing to the lack of dedicated FP- INT arithmetic units. In this study, we investigate the impact of introducing a dedicated FP-INT unit on overall performance and find that such specialization does not yield substantial improvements. As an alternative approach, we propose FIGNA, an accelerator based on INT units designed specifically for FP- INT MAC operations. A key feature of FIGNA is its ability to achieve the same numerical accuracy as the FPU while relying solely on the integer-unit, a departure from prior methods that relied on integer-units with numerical approximations for FP arithmetic results, albeit claiming similar inference accuracy through dedicated network training. Through comprehensive experiments on FP- INT quantized networks for LLMs, including OPT and BLOOM, we demonstrate the superior performance of FIGNA compared to conventional FPUs in terms of performance per area ($TOPS/mm^{2}$) and energy efficiency (TOPS/W) across various input and weight precision combinations. For instance, in the FP16-INT4 case, FIGNA shows 6.34x higher$TOPS/ mm^{2}$and 2.19x higher TOPS/W compared to the baseline. Jaeyong Jang, Yulhwa Kim, Juheun Lee, Jae-Joon Kim |
HPCA | 1 |
| 2024 | Mobileware: Distributed Architecture With Channel Stationary Dataflow for MobileNet AccelerationabstractThe depthwise separable convolution, a key feature of the MobileNet models, has a different input reuse pattern from the conventional standard convolution, and a smaller number of input/weight pairs are used for a dot product, thereby leading to extremely low MAC utilization. This paper proposes a Mobileware architecture for the high-performance acceleration of the MobileNet workloads. A new channel stationary dataflow architecture distributes the on-chip buffers, and the distributed SRAMs are placed near each PE. By doing so, PEs and SRAMs can communicate with high bandwidth. Our Mobileware architecture shows 1.4-29.5× higher throughput than conventional weight stationary-based hardware architecture, and the proposed design was verified on the Xilinx ZCU102 FPGA evaluation board. Sungju Ryu, Jaeyong Jang, Youngtaek Oh, Jae-Joon Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Winning Both the Accuracy of Floating Point Activation and the Simplicity of Integer Arithmetic
Yulhwa Kim, Jaeyong Jang, Jehun Lee, Byeongwook Kim, Baeseong Park, Se Jung Kwon, Dongsoo Lee, Jae-Joon Kim |
ICLR | 2 |