VLDB 2026 Research / reviewers in the wild / expert
Sangyeob Kim
dblp:216/3953
· DBLP profile ↗
17ranked-venue papers
6as first author
15since 2021 · last 2025
0000-0002-1783-5296ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 6 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BROCA: A Low-power and Low-latency Conversational Agent RISC-V System-on-Chip for Voice-interactive Mobile Devices
Wooyoung Jo, Seongyon Hong, Beomseok Kwon, Haoyang Sang, Dongseok Im, Sangyeob Kim, Chaeyun Jeong, Yujin Moon, Hoi-Jun Yoo |
HCS | 7 |
| 2025 | A 4.69mW LLM Processor with Binary/Ternary Weights for Billion-Parameter Llama Model
Sangyeob Kim, Jungwan Lee, Byeongju Kim, Hoi-Jun Yoo |
HCS | 1 |
| 2025 | EdgeDiff: Multi-modal Few-step Diffusion Model Accelerator with Mixed-Precision and Reordered Group-Quantization for On-device Generative AI Motivation
Jungjun Oh, Jeonggyu So, Yuseon Choi, Sangyeob Kim, Dongseo Kim, Gwangtae Park, Hoi-Jun Yoo |
HCS | 5 |
| 2025 | A 65.1 TOPS/W Digital CIM Processor for Ultra-Low-Bit Transformers with Multiplexer-based Adder and Scaling Factor-based ReorderingabstractAs large-language-model (LLM) continues to expand in parameter size and improve performance, challenges related to latency and energy efficiency become increasingly significant. While low-bit weight quantization was proposed to alleviate the external memory access bottleneck, it remains insufficient to reduce overall energy consumption. In this paper, we use binarized LLM and design a novel Computing-in-Memory (CIM) processor to reduce both computation and memory energy. We propose three key techniques: 1) Multiplexer-based adder is proposed to replace first stage adders of adder tree with the simple multiplexers and inverter, reducing adder tree energy consumption by 58%, 2) Scaling factor-based reordering is proposed to reduce adder tree bit-width, resulting in an additional 28.6% energy reduction, 3) Cluster-level broadcasting and bit-spatial mapping is proposed to reduce internal memory access and increase bit-scalability. As a result, we can reduce the overall energy consumption by 69.7%, offering a more efficient solution for LLM inference tasks. Nayeong Lee, Sangyeob Kim, Seongyon Hong, Hoi-Jun Yoo |
ISCAS | 2 |
| 2024 | A Low-Power Large-Language-Model Processor with Big-Little Network and Implicit-Weight-Generation for On-Device AI
Sangyeob Kim, Wooyoung Jo, Seongyon Hong, Nayeong Lee, Hoi-Jun Yoo |
HCS | 1 |
| 2024 | NeuGPU: A Neural Graphics Processing Unit for Instant Modeling and Real-Time Rendering on Mobile AR/VR Devicesabstract•Why 3D Modeling using Neural Radiance Field? Junha Ryu, Hankyul Kwon, Wonhoon Park, Zhiyong Li 0016, Beomseok Kwon, Donghyeon Han, Dongseok Im, Sangyeob Kim, Hyungnam Joo, Hoi-Jun Yoo |
HCS | 8 |
| 2024 | Space-Mate: A 303.5mW Real-Time NeRF SLAM Processor with Sparse-Mixture-of-Experts-based AccelerationabstractNeRF-based SLAM for robotic applications face computation barrier Seokchan Song, Haoyang Sang, Dongseok Im, Donghyeon Han, Sangyeob Kim, Hongseok Lee, Hoi-Jun Yoo |
HCS | 5 |
| 2024 | Two-Step Spike Encoding Scheme and Architecture for Highly Sparse Spiking-Neural-NetworkabstractThis paper proposes a two-step spike encoding, which consists of the source encoding and process encoding for energy-efficient spiking-neural-network (SNN) acceleration. The eigen-train generation and its superposition generate spike trains which show high accuracy with low spike ratio. Sparsity boosting (SB) and spike generation skipping (SGS) reduce the number of operations for SNN. Time shrinking multi-level encoding (TS-MLE) compresses the number of spikes in a train along time axis, and spike-level clock skipping (SLCS) decreases the processing time. Eigen-train generation achieves 90.3% accuracy, the same accuracy as CNN, under the condition of 4.18% spike ratio for CIFAR-10 classification. SB reduces spike ratio by 0.49× with only 0.1% accuracy loss, and the SGS reduces the spike ratio by 20.9% with 0.5% accuracy loss. TS- MLE and SLCS increase the throughput of SNN by 2.8× while decreasing the hardware resource for spike generator by 75% compared with previous generators. Sangyeob Kim, Soyeon Um, Hoi-Jun Yoo |
ISCAS | 1 |
| 2023 | A 332 TOPS/W Input/Weight-Parallel Computing-in-Memory Processor with Voltage-Capacitance-Ratio Cell and Time-Based ADCabstractRecent computing-in-memory (CIM) achieves high energy efficiency with charge-domain computation and multi-bit input driving. However, the previous works still require high power consumption and trade computation signal-to-noise ratio (SNR) for energy efficiency. This work proposes an energy-efficient and accurate multi-bit input/weight-parallel CIM processor with four key features: 1) a 10T2C sign-magnitude cell with voltage-capacitance-ratio (VCR) decoding for 5-bit analog inputs with only 2-level supply voltages, 2) a computation word line (CWL) charge reuse method for input driver power reduction, 3) a signal-amplifying noise canceling voltage-to-time converter (SANC-VTC) for SNR improvement, and 4) a distribution-aware time-to-digital converter (DA-TDC) for ADC power reduction. The proposed CIM processor is simulated in 28 nm CMOS technology with 1.25 mm2area. As a result, it achieves 4.44 mW power consumption and 332 TOPS/W energy efficiency with 72.43% benchmark accuracy (@ ImageNet, ResNet50, 5-bit input/5-bit weight). Seongyon Hong, Soyeon Um, Sangyeob Kim, Wooyoung Jo, Hoi-Jun Yoo |
ISCAS | 4 |
| 2023 | A Reconfigurable 1T1C eDRAM-based Spiking Neural Network Computing-In-Memory Processor for High System-Level EfficiencyabstractSpiking Neural Network (SNN) Computing-In-Memory (CIM) was proposed for high macro-level energy efficiency. However, system-level energy efficiency is limited by EMA due to a large intermediate activation footprint requirement. To reduce the EMA, a large capacity SNN CIM is needed to load tons of weights in the CIM. This paper proposes a high-density 1T1C eDRAM-based SNN CIM processor for supporting high system-level energy efficiency with two key features: 1) High-density and low-power Reconfigurable Neuro-Cell Array (ReNCA) for memory and SNN peripheral logic using a charge pump and reusing 1T1C cell array, achieving 41% area and 90% power reduction compared to previous work. 2) Reconfigurable CIM architecture with dual-mode ReNCA and Dynamic Adjustable Neuron Link (DAN Link) for layer fusion increases system-level efficiency including intermediate and weight EMA. It achieves$10\times$higher state-of-the-art system-level energy efficiency including EMA. Seryeong Kim, Soyeon Um, Zhiyong Li 0016, Sangyeob Kim, Wooyoung Jo, Hoi-Jun Yoo |
ISCAS | 6 |
| 2022 | Neuro-CIM: A 310.4 TOPS/W Neuromorphic Computing-in-Memory Processor with Low WL/BL activity and Digital-Analog Mixed-mode Neuron FiringabstractMulti WLs Driving ➔ Low Energy Efficiency by ADC (<100 TOPS/W) Sangyeob Kim, Soyeon Um, Kwantae Kim, Hoi-Jun Yoo |
HCS | 1 |
| 2022 | A Low-Power Graph Convolutional Network Processor With Sparse Grouping for 3D Point Cloud Semantic Segmentation in Mobile DevicesabstractA low-power graph convolutional network (GCN) processor is proposed for accelerating 3D point cloud semantic segmentation (PCSS) in real-time on mobile devices. Three key features enable the low-power GCN-based 3D PCSS. First, the new hardware-friendly GCN algorithm, sparse grouping-based dilated graph convolution (SG-DGC) is proposed. SG-DGC reduces 71.7% of the overall computation and 76.9% of EMA through the sparse grouping of the point cloud. Second, the two-level pipeline (TLP) consisting of the point-level pipeline (PLP) and group-level pipelining (GLP) was proposed to improve low utilization by the imbalanced workload of GCN. The PLP enables point-level module-wise fusion (PMF) which reduces 47.4% of EMA for low power consumption. Also, center point feature reuse (CPFR) reuses computation results of the redundant operation and reduces 11.4% of computation. Finally, the GLP increased the core utilization by 21.1% by balancing the workload of graph generation and graph convolution and enable$1.1\times $higher throughput. The processor is implemented with 65nm CMOS technology, and the 4.0mm23D PCSS processor show 95mW power consumption while operating in real-time of 30.8 fps in the 3D PCSS of the indoor scene with 4k points. Sangyeob Kim, Juhyoung Lee, Hoi-Jun Yoo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | TSUNAMI: Triple Sparsity-Aware Ultra Energy-Efficient Neural Network Training Accelerator With Multi-Modal Iterative PruningabstractThis article proposes the TSUNAMI, which supports an energy-efficient deep-neural-network training. The TSUNAMI supports multi-modal iterative pruning to generate zeros in activation and weight. Tile-based dynamic activation pruning unit and weight memory shared pruning unit eliminate additional memory access. Coarse-zero skipping controller skips multiple unnecessary multiply-and-accumulation (MAC) operations at once, and fine-zero skipping controller skips randomly located unnecessary MAC operations. Weight sparsity balancer solves a utilization degradation caused by weight sparsity imbalance, and the workload of each convolution core is allocated by a random channel allocator. The TSUNAMI achieves an energy efficiency of 3.42 TFLOPS/W at 0.78V and 50MHz with floating-point 8-bit activation and weight. Also, it achieves an energy efficiency of 405.96 TFLOPS/W at 90% sparsity condition. Sangyeob Kim, Juhyoung Lee, Donghyeon Han, Wooyoung Jo, Hoi-Jun Yoo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | An Energy-efficient Floating-Point DNN Processor using Heterogeneous Computing Architecture with Exponent-Computing-in-MemoryabstractAbstract of Proposed FP CIM Processor (1) Heterogeneous FP Computing Arch. : Separate optimization of FP computing: Realize 2 cycles FP MAC w/ CIM (2) Exponent Computing-in-Memory: In-memory AND/NOR + BL charge reusing: Total memory power 46.4% 2) Mantissa Free Exponent Calculation: Removing redundant normalization: Total MAC power 14.4% Juhyoung Lee, Ji-Hoon Kim 0004, Wooyoung Jo, Sangyeob Kim, Donghyeon Han, Jinsu Lee, Hoi-Jun Yoo |
HCS | 4 |
| 2021 | OmniDRL: An Energy-Efficient Mobile Deep Reinforcement Learning Accelerators with Dual-mode Weight Compression and Direct Processing of Compressed DataabstractDeep Reinforcement Learning (DRL)▪ No Pre-labelled Data ➔ Training with Trial-and-errors!– Sequential decision making problems @ Unknown environments– Applications: gaming agent, autonomous systems, agent adaptation Juhyoung Lee, Sangyeob Kim, Ji-Hoon Kim 0004, Wooyoung Jo, Donghyeon Han, Hoi-Jun Yoo |
HCS | 2 |
| 2020 | A 54.7 fps 3D Point Cloud Semantic Segmentation Processor with Sparse Grouping Based Dilated Graph Convolutional Network for Mobile DevicesabstractThe graph convolutional network (GCN) based 3D point cloud semantic segmentation (PCSS) processor for mobile devices is proposed. GCN based 3D PCSS requires a lot of computation, making it unsuitable for real-time operation in mobile devices. For real-time 3D PCSS on mobile devices, this paper proposes two key features: 1) a sparse grouping based dilated graph convolution (SG-DGC) which reduces 71.7% of the overall computation of GCN by simply dividing input point cloud into multiple sparse point cloud. 2) group-level pipelining which improves low pipeline utilization due to the computation imbalance of GCN. Finally, the proposed GCN processor is simulated in 65 nm CMOS technology and occupies 4.0 mm2. The proposed processor consumes 176mW and shows 54.7 frames-per-second (fps) for the 3D point cloud semantic segmentation of indoor scene with 4k points. Sangyeob Kim, Juhyoung Lee, Hoi-Jun Yoo |
ISCAS | 2 |
| 2019 | A 15.2 TOPS/W CNN Accelerator with Similar Feature Skipping for Face Recognition in Mobile DevicesabstractA low-power face recognition processor with similar feature skipping (SFS) and the tile-based clustering algorithm is proposed for high energy efficiency in mobile devices. For higher energy efficiency face recognition (FR) processor, this paper proposes two key features: 1) Tile-based clustering enables to reduce computation overhead of clustering. 2) SFS binary convolution core is proposed to increase energy efficiency, resulting in 15.2 TOPS/W energy efficiency. Implemented with 65 nm CMOS technology, the 6 mm2FR processor achieves 0.26mW power consumption at 1 frames-per-second (fps) always-on face recognition in mobile devices. Sangyeob Kim, Juhyoung Lee, Jinsu Lee, Hoi-Jun Yoo |
ISCAS | 1 |