EDBT 2026 Demo / reviewers in the wild / expert
Jung Gyu Min
dblp:285/1277
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0002-2041-617XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RISC-V Driven Orchestration of Vector Processing Units and eFlash Compute-in-Memory Arrays for Fast and Accurate Keyword SpottingabstractIn this paper, we propose a computationally efficient keyword spotting (KWS) model, named hybrid reparameterized FSMN (HRepFSMN), by carefully examining the impact of binarization on the accuracy. In particular, we found that binarizing depthwise convolution (DW-Conv) within the previous binarized KWS model, i.e., BiFSMNv2, does not lead to a significant reduction in FLOPs. Therefore, we allow floating-point (FP) operations on less computation-intensive DW-Conv layers while the remaining layers are computed in a binary fashion (hybrid data type). In addition, we remove skip connections, which require data fetching in full precision, by applying a reparameterization technique. More importantly, to efficiently compute the proposed HRepFSMN, we present a RISC-V controlled hardware accelerator that consists of reconfigurable vector processing units for FP operations and eFlash compute-in-memory arrays for binary operations. We extend RISC-V instructions so that the core can efficiently manage both computing fabrics. As a result, our HRepFSMN improves accuracy by 2.57%/4.98% with 24.02×/3.66× speed-up compared to BiFSMNv2/BiFSMNv2_small. By shrinking down our HRepFSMN, we achieve 0.95% higher accuracy with 20.87× speed-up compared to BiFSMNv2_small. Gunil Kang, Dahoon Park, Sangwoo Jung 0001, Jung Gyu Min, Youngjoo Lee 0002, Jaeha Kung 0001 |
ASP-DAC | 6 |
| 2025 | Cost-efficient Processing-in-Memory Architecture with Training-free and Universal Error CompensationabstractDespite the energy efficiency of memory-centric deep neural network (DNN) computations, the nonlinearities inherent in existing processing-in-memory (PIM) architectures cause severe accuracy drops. These imperfections necessitate additional methods to correct inaccurate vector-matrix multiplication (VMM) results. To address this issue without modifying DNN weights, we first propose an input sparsity-based error compensation method. This approach dynamically corrects accumulated errors along the column direction of the non-volatile memory (NVM) array using pre-collected errors and input characteristics. We then present a new PIM architecture along with the proposed compensation scheme by slightly modifying the existing analog-to-digital converter (ADC) or adding a few extra rows to the NVM array. Experimental results show that the proposed work mitigates the nonlinear effects of various emerging memory cells, achieving near-ideal DNN accuracy with negligible hardware overheads. Myeongji Yun, Jung Gyu Min, Sein Oh, Jiwoung Choi, Jang-Sik Lee, Minkyu Je, Youngjoo Lee 0002 |
ISLPED | 2 |
| 2023 | TF-MVP: Novel Sparsity-Aware Transformer Accelerator with Mixed-Length Vector PruningabstractWe present the energy-efficient TF-MVP architecture, a sparsity-aware transformer accelerator, by introducing novel algorithm-hardware co-optimization techniques. From the previous fine-grained pruning map, for the first time, the direction strength is developed to analyze the pruning patterns quantitatively, indicating the major pruning direction and size of each layer. Then, the mixed-length vector pruning (MVP) is proposed to generate the hardware-friendly pruned-transformer model, which is fully supported by our TF-MVP accelerator with the reconfigurable PE structure. Implemented in a 28nm CMOS technology, as a result, TF-MVP achieves 377 GOPs/W for accelerating GPT-2 small model by realizing 4096 multiply-accumulate operators, which is 2.09 times better than the state-of-the-art sparsity-aware transformer accelerator. Eunji Yoo, Gunho Park, Jung Gyu Min, Se Jung Kwon, Baeseong Park, Dongsoo Lee |
DAC | 3 |
| 2023 | Energy-Efficient RISC-V-Based Vector Processor for Cache-Aware Structurally-Pruned TransformersabstractBased on recent RISC-V designs, we present in this paper a low-power vector processor architecture for efficiently deploying vision transformer (ViT) models. To fairly measure the processing efficiency of different processor designs with instruction/data cache memories, we first develop the evaluation framework based on numerous design tools for jointly considering the algorithm, architecture, and circuit performances together, numerically revealing that the previous CSR-based data compression cannot accelerate pruned transformer models at all due to under-utilization of the vector-extended processing units. We then introduce a series of algorithm-hardware co-optimization approaches to greatly minimize cache misses by applying 1) the accuracy-preserved structured ViT pruning, 2) the vertical-CSR (vCSR) data storing format, and 3) vCSR-aware custom memory-accessing instructions. Experimental results show that the proposed optimization schemes eventually improve the processing efficiency of pruned transformers in resource-limited computing platforms, e.g., achieving 11 times lower energy consumption for handling the 0.7-pruned ViT model. Jung Gyu Min, Dongyun Kam, Younghoon Byun, Gunho Park, Youngjoo Lee 0002 |
ISLPED | 1 |
| 2022 | Design and Evaluation Frameworks for Advanced RISC-based Ternary ProcessorabstractIn this paper, we introduce the design and veri-fication frameworks for developing a fully-functional emerging ternary processor. Based on the existing compiling environments for binary processors, for the given ternary instructions, the software-level framework provides an efficient way to convert the given programs to the ternary assembly codes. We also present a hardware-level framework to rapidly evaluate the performance of a ternary processor implemented in arbitrary design technology. As a case study, the fully-functional 9-trit advanced RISC-based ternary (ART-9) core is newly developed by using the proposed frameworks. Utilizing 24 custom ternary instructions, the 5-stage ART-9 prototype architecture is successfully verified by a number of test programs including dhrystone benchmark in a ternary domain, achieving the processing efficiency of 57.8 DMIPS/W and$3.06\times 10^{6}$DMIPS/W in the FPGA-level ternary-logic emulations and the emerging CNTFET ternary gates, respectively. Dongyun Kam, Jung Gyu Min, Jongho Yoon 0001, Sunmean Kim, Seokhyeong Kang, Youngjoo Lee 0002 |
DATE | 2 |