Seongyon Hong

dblp:343/5202 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0002-8891-735XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SingularBit: Exploiting Synergy of Singular Value Decomposition and Low-Bit Quantization for Weight and KV Compression in LLM Inference
Seongyon Hong, Hyundeok Kong, Jungwan Lee, Hoi-Jun Yoo
ISCA1
2025 BROCA: A Low-power and Low-latency Conversational Agent RISC-V System-on-Chip for Voice-interactive Mobile Devices
Wooyoung Jo, Seongyon Hong, Beomseok Kwon, Haoyang Sang, Dongseok Im, Sangyeob Kim, Chaeyun Jeong, Yujin Moon, Hoi-Jun Yoo
HCS2
2025 A 65.1 TOPS/W Digital CIM Processor for Ultra-Low-Bit Transformers with Multiplexer-based Adder and Scaling Factor-based Reordering
abstract
As large-language-model (LLM) continues to expand in parameter size and improve performance, challenges related to latency and energy efficiency become increasingly significant. While low-bit weight quantization was proposed to alleviate the external memory access bottleneck, it remains insufficient to reduce overall energy consumption. In this paper, we use binarized LLM and design a novel Computing-in-Memory (CIM) processor to reduce both computation and memory energy. We propose three key techniques: 1) Multiplexer-based adder is proposed to replace first stage adders of adder tree with the simple multiplexers and inverter, reducing adder tree energy consumption by 58%, 2) Scaling factor-based reordering is proposed to reduce adder tree bit-width, resulting in an additional 28.6% energy reduction, 3) Cluster-level broadcasting and bit-spatial mapping is proposed to reduce internal memory access and increase bit-scalability. As a result, we can reduce the overall energy consumption by 69.7%, offering a more efficient solution for LLM inference tasks.
Nayeong Lee, Sangyeob Kim, Seongyon Hong, Hoi-Jun Yoo
ISCAS3
2025 A 17.1 TOPS/W FP-INT Transformer Inference Accelerator with Sparsity Boosting and Output Importance-Aware Processing
abstract
This paper presents an energy-efficient FP-INT transformer inference accelerator for diverse applications. The proposed accelerator achieves high energy efficiency with two key features as solutions: 1) Sparsity Boosting Adder Tree (SBAT) to reduce adder tree power by 23.8% by modifying the adder tree structure and booth encoding to maximize sparsity. 2) Output Importance-aware Processing (OIAP) to reduce the Floating-Point Accumulation (FP-ACC) power by 79.6%, dynamically adjusting input block sizes based on the saliency of output channels, thereby reducing the FP-ACC operations by 78.7%. The proposed accelerator is implemented in 28 nm CMOS technology and achieves 17.1 TOPS/W, leveraging the unique FP-INT characteristics on booth encoding and exploiting redundancy based on output saliency in the transformer architecture.
Jeonggyu So, Seongyon Hong, Wooyoung Jo, Hoi-Jun Yoo, Donghyeon Han
ISCAS2
2024 A Low-Power Large-Language-Model Processor with Big-Little Network and Implicit-Weight-Generation for On-Device AI
Sangyeob Kim, Wooyoung Jo, Seongyon Hong, Nayeong Lee, Hoi-Jun Yoo
HCS5
2024 A 28.6 mJ/iter Stable Diffusion Processor for Text-to-Image Generation with Patch Similarity-based Sparsity Augmentation and Text-based Mixed-Precision
abstract
This paper presents an energy-efficient stable diffusion processor for text-to-image generation. While stable diffusion attained attention for high-quality image synthesis results, its inherent characteristics hinder its deployment on mobile platforms. The proposed processor achieves high throughput and energy efficiency with three key features as solutions: 1) Patch similarity-based sparsity augmentation (PSSA) to reduce external memory access (EMA) energy of self-attention score by 60.3 %, leading to 37.8 % total EMA energy reduction. 2) Text-based important pixel spotting (TIPS) to allow 44.8 % of the FFN layer workload to be processed with low-precision activation. 3) Dual-mode bit-slice core (DBSC) architecture to enhance energy efficiency in FFN layers by 43.0 %. The proposed processor is implemented in 28 nm CMOS technology and achieves 3.84 TOPS peak throughput with 225.6 mW average power consumption. In sum, 28.6 mJ/iteration highly energy-efficient text-to-image generation processor can be achieved at MS-COCO dataset.
Wooyoung Jo, Seongyon Hong, Beomseok Kwon, Wonhoon Park, Hoi-Jun Yoo
ISCAS3
2023 A 332 TOPS/W Input/Weight-Parallel Computing-in-Memory Processor with Voltage-Capacitance-Ratio Cell and Time-Based ADC
abstract
Recent computing-in-memory (CIM) achieves high energy efficiency with charge-domain computation and multi-bit input driving. However, the previous works still require high power consumption and trade computation signal-to-noise ratio (SNR) for energy efficiency. This work proposes an energy-efficient and accurate multi-bit input/weight-parallel CIM processor with four key features: 1) a 10T2C sign-magnitude cell with voltage-capacitance-ratio (VCR) decoding for 5-bit analog inputs with only 2-level supply voltages, 2) a computation word line (CWL) charge reuse method for input driver power reduction, 3) a signal-amplifying noise canceling voltage-to-time converter (SANC-VTC) for SNR improvement, and 4) a distribution-aware time-to-digital converter (DA-TDC) for ADC power reduction. The proposed CIM processor is simulated in 28 nm CMOS technology with 1.25 mm2area. As a result, it achieves 4.44 mW power consumption and 332 TOPS/W energy efficiency with 72.43% benchmark accuracy (@ ImageNet, ResNet50, 5-bit input/5-bit weight).
Seongyon Hong, Soyeon Um, Sangyeob Kim, Wooyoung Jo, Hoi-Jun Yoo
ISCAS1