EDBT 2026 Demo / reviewers in the wild / expert
Seryeong Kim
dblp:352/9921
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0009-9846-894XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 198.7 μJ/token Block Diffusion LLM Processor with Mask Token Similarity-Based Activation Reuse
Yujin Moon, Seryeong Kim, Wonhoon Park, Wooyoung Jo, Yuseon Choi, Sunjoo Whang, Hoi-Jun Yoo |
ISCAS | 3 |
| 2025 | IRIS: A 8.55 mJ/frame Spatial Computing SoC for Real-time Interactable-Rendering and Surface-aware-Modeling with 3D Gaussian Splatting
Seokchan Song, Seryeong Kim, Wonhoon Park, Jongjun Park, Sanghyuk An, Gwangtae Park, Minseo Kim 0001, Hoi-Jun Yoo |
HCS | 2 |
| 2025 | A 51.2 fps Real-Time 3DGS-SLAM Accelerator using Diagonal Feeding with Symmetric Alpha Reuse and Voxel-based 3D Gaussian Cache ManagementabstractThis work presents a high-speed 3D Gaussian Splatting-based SLAM (3DGS-SLAM) accelerator to support dense mapping for mobile devices. 3DGS-SLAM has two main hardware challenges for acceleration: 1) Large α-computation computation. 2) Memory bottleneck caused by irregular memory access and large number of Gaussians. First, diagonal feeding (DF) controller precludes redundant-α computation, and symmetric alpha reuse (SAR) enables reusing computed alpha. This method reduces 35.3% system computation. Second, voxel-based inter-frame caching (VIFC) enables selective inter-frame voxel caching, which reduces 44.0% of external memory access. As a result, the proposed 3DGS-SLAM accelerator achieves 51.2 fps with 0.07µJ/point with support voltage 0.9V, clock frequency 200MHz mapping a high-quality dense-map. Hyungnam Joo, Seryeong Kim, Jongjun Park, Junha Ryu, Hoi-Jun Yoo |
ISCAS | 2 |
| 2024 | A 3.55 mJ/frame Energy-efficient Mixed-Transformer based Semantic Segmentation Accelerator for Mobile DevicesabstractAn energy-efficient semantic segmentation (SS) processor, achieving 3.55 mJ/frame system energy efficiency, is proposed. To address the challenges posed by Mixed Transformer (MiT)-based SS, including high external memory bandwidth requirement and large on-chip memory footprint, we introduce a novel compression method called Chunk-based Bit Plane Compression (CBPC). CBPC leverages the high inter-token locality of feature maps in MiT-based SS, along with the robustness and compression ratio variations based on bit position to achieve a high compression ratio. To support CBPC, we propose an area and power-efficient CBPC encoder/decoder. In addition, a Similar Token Coarse Skipping (STCS) Core is proposed for high throughput. It enables row-wise clock gating and array-wise coarse skipping to reduce redundant computation. By removing redundant computation, the processor achieves higher throughput and lower computation power. The proposed processor reduces 67.6% of EMA power and accomplishes 19.24 TOPS/W core energy efficiency. The proposed processor achieves 44.3% higher system energy efficiency than the previous processors. Jongjun Park, Seryeong Kim, Wonhoon Park, Seokchan Song, Hoi-Jun Yoo |
ISCAS | 2 |
| 2024 | An Energy-Efficient CNN/Transformer Hybrid Neural Semantic Segmentation Processor With Chunk-Based Bit Plane Data Compression and Similarity-Based Token-Level Skipping ExploitationabstractA novel energy-efficient semantic segmentation (SS) processor is proposed for achieving high system energy efficiency on mobile devices. 1) Excessive external memory access and 2) a large amount of redundant computation hinders energy-efficient SS acceleration. Three key features enable real-time energy-efficient CNN/ViT hybrid SS. A new compression method named Chunk-based Bit Plane Compression (CBPC) reduces the memory footprint and energy consumption due to external memory access. CBPC enhances compression ratio by leveraging the high inter-token similarity of feature maps and applying bit plane compression in sign-magnitude data representation, using chunk-wise low-bit plane shared bias. The proposed CBPC encoder/decoder supports CBPC with minimum area overhead. Additionally, the Similar Token Coarse Skipping (STCS) Core enhances the throughput and reduces the computation power by eliminating redundant computations. STCS core employs Row-wise Line Gating for low-power computation and Array-wise Coarse Skipping to minimize redundant computation. As a result, our proposed processor reduces external memory access energy by 67.6% and achieves a core energy efficiency of 19.24 TOPS/W. Our solution achieves 3.55mJ/frame system-level energy efficiency which is 79.7% higher than the previous SOTA SS processor. Jongjun Park, Seryeong Kim, Wonhoon Park, Seokchan Song, Hoi-Jun Yoo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | A Reconfigurable 1T1C eDRAM-based Spiking Neural Network Computing-In-Memory Processor for High System-Level EfficiencyabstractSpiking Neural Network (SNN) Computing-In-Memory (CIM) was proposed for high macro-level energy efficiency. However, system-level energy efficiency is limited by EMA due to a large intermediate activation footprint requirement. To reduce the EMA, a large capacity SNN CIM is needed to load tons of weights in the CIM. This paper proposes a high-density 1T1C eDRAM-based SNN CIM processor for supporting high system-level energy efficiency with two key features: 1) High-density and low-power Reconfigurable Neuro-Cell Array (ReNCA) for memory and SNN peripheral logic using a charge pump and reusing 1T1C cell array, achieving 41% area and 90% power reduction compared to previous work. 2) Reconfigurable CIM architecture with dual-mode ReNCA and Dynamic Adjustable Neuron Link (DAN Link) for layer fusion increases system-level efficiency including intermediate and weight EMA. It achieves$10\times$higher state-of-the-art system-level energy efficiency including EMA. Seryeong Kim, Soyeon Um, Zhiyong Li 0016, Sangyeob Kim, Wooyoung Jo, Hoi-Jun Yoo |
ISCAS | 1 |