EDBT 2026 Demo / reviewers in the wild / expert
Jungjun Oh
dblp:402/9170
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0006-3998-7385ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GyRot: Leveraging Hidden Synergy Between Rotation and Fine-Grained Group Quantization for Low-Bit LLM InferenceabstractLow-bit quantization is essential for efficient LLM inference, and both rotation and fine-grained group quantization have shown individual promise. However, their combination often leads to accuracy degradation or hardware overhead due to a mismatch between the global nature of rotation and the localized behavior of group scaling. We propose GyRot, a quantization framework and hardware accelerator that bridges this gap through algorithm-hardware co-design. GyRot introduces Coarse Rotation, Fine Grouping (CoRFiG) and Harmonic-Aligned Permutation (HAP) to enable cooperative integration of rotation and group quantization, enhancing quantizability while relaxing scaling factor precision. To further reduce hardware cost, we reformulate asymmetric quantization and introduce a zero-point rounding strategy that enables fully integer dequantization. Implemented on an INT4-based tensor PE architecture, GyRot achieves state-of-the-art 4-bit accuracy across LLaMA-family models, while delivering up to 3.4× speedup and 3.6× energy efficiency over baseline LLM accelerators. These results validate GyRot's practical effectiveness for scalable and energy-efficient LLM deployment. Yuseon Chou, Byeongcheol Kim, Jungjun Oh, Hoi-Jun Yoo |
HPCA | 4 |
| 2026 | SeVeDo: A Heterogeneous Transformer Accelerator for Low-Bit Inference via Hierarchical Group Quantization and SVD-Guided Mixed PrecisionabstractLow-bit quantization is a promising technique for efficient transformer inference by reducing computational and memory overhead. However, aggressive bitwidth reduction remains challenging due to activation outliers, leading to accuracy degradation. Existing methods, such as outlier-handling and group quantization, achieve high accuracy but incur substantial energy consumption. To address this, we propose SeVeDo, an energy-efficient SVD-based heterogeneous accelerator that structurally separates outlier-sensitive components into a high-precision low-rank path, while the remaining computations are executed in a low-bit residual datapath with group quantization. To further enhance efficiency, Hierarchical Group Quantization (HGQ) combines coarse-grained floating-point scaling with fine-grained shifting, effectively reducing dequantization cost. Also, SVD-guided mixed precision (SVD-MP) statically allocates higher bitwidths to precision-sensitive components identified through low-rank decomposition, thereby minimizing floating-point operation cost. Experimental results show that SeVeDo achieves a peak energy efficiency of 13.8TOPS/W, surpassing conventional designs, with 12.7TOPS/W on ViT-Base and 13.4TOPS/W on Llama2-7B benchmarks. Yuseon Choi, Jungjun Oh, Byeongcheol Kim, Hoi-Jun Yoo |
ISCAS | 3 |
| 2026 | An Energy-Efficient High Resolution Vision Transformer Processor Exploiting Token Similarity Beyond Token Merging
Jungjun Oh, Junha Ryu, Byeongcheol Kim, Yuseon Choi, Hoi-Jun Yoo |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | EdgeDiff: Multi-modal Few-step Diffusion Model Accelerator with Mixed-Precision and Reordered Group-Quantization for On-device Generative AI Motivation
Jungjun Oh, Jeonggyu So, Yuseon Choi, Sangyeob Kim, Dongseo Kim, Gwangtae Park, Hoi-Jun Yoo |
HCS | 2 |
| 2025 | A 9.6 TOPS/W Vision Transformer Processor with Hierarchical Token Merging for Similarity-Driven Difference ComputingabstractToken merging is widely used in Vision Transformers(ViT) as an effective method to reduce computation with minimal accuracy drop. However, aggressive token elimination methods like token merging have limitations in applications requiring fine-grained output, such as text-unified object recognition. These methods reach an upper limit on token reduction due to an accuracy loss. This paper proposes exploiting token similarity beyond the merging upper bound to further reduce power consumption. The key features include Hierarchical Token Merging (HTM), Sparsity Separated Accumulation with Sign Magnitude Data Representation(SSA-SM), and Bidirectional Dynamic Allocation (BDA). HTM begins by performing token merging in the first iteration of the transformer block to reduce 34 % of tokens. Then, during the subsequent iterations, similar token difference computing is applied to the remaining tokens and reduces the effective bit by 31 %. SSA-SM then optimizes PE operations to fully exploit this sparsity by reducing the computation logic toggle rate, resulting in a 45 % higher energy efficiency. Lastly, BDA introduces a memory storage technique that ensures sparse and dense inputs are fed separately into the PE, resulting in an additional 14 % power savings. This proposed method improves energy efficiency by 2.39 times and achieves 9.6 TOPS/W for the text-unified object recognition application on the MS-COCO dataset. Jungjun Oh, Byeongcheol Kim, Hoi-Jun Yoo |
ISCAS | 1 |