Sangwoo Hwang

dblp:296/1157 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MX-SAFE: Versatile Inference-and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation
abstract
As the demand for deep learning grows, cost reduction through quantization has become essential for both training and inference. In 2022, the Open Compute Project (OCP) consortium standardized narrow precision formats for deep learning, called the microscaling (MX) format. The MX format is a hardware-friendly dynamic quantization scheme that effectively reduces the data size by sharing an 8-bit exponent across multiple operands. The MX format can be categorized into two types with their own strengths: (i) MXINT which focuses on a high precision consisting only of mantissa bits and (ii) MXFP which focuses on a wider dynamic range by allowing local exponent bits. In this work, we present a versatile MXFP format, called MX-SAFE (MXSF in short), that adaptively uses two modes, i.e., a wider mantissa mode (FP8_E2M5) and a subnormal FP mode (FP5_E3M2), to support both training and direct-cast inference. Furthermore, we propose a tile-based block design to increase hardware efficiency by reducing the burden of re-quantization process during the training with the MXSF format. Owing to the use of the proposed MXSF format, 0.05%/11.1% and 3.55%/3.57% improvements in accuracy, on average, for inference/full-training compared to MXFP8_E2M5 and MXFP8_E4M3 are observed, respectively. Moreover, we present a training-inference accelerator that supports the MXSF format and it achieves similar accuracy to the BF16 baseline while using 24.9% less total energy consumption.
Dahoon Park, Jahyun Koo 0002, Sangwoo Hwang, Jaeha Kung 0001
DATE3
2026 GustavSNN: Unleashing the Power of Gustavson's Algorithm on SNN Acceleration with Column-Parallel Tick-Batch Dataflow
abstract
Spiking neural networks (SNNs) require sequential computation over long timesteps, introducing substantial memory and energy overheads due to frequent updates of neuron membrane potentials. Previous SNN accelerators address this by employing tick-batch techniques, which process all timesteps within a layer before moving on to the next. However, existing approaches rely on neuron-centric scheduling, limiting their ability to exploit temporal sparsity. In this work, we propose a novel scheduling approach along with its hardware architecture for Gustavson product (GP)-based SNN acceleration. We introduce a column-parallel tick-batch (CPTB) dataflow that partitions the spike matrix into multiple submatrices and processes each submatrix of a timestep in parallel while maintaining tickbatch semantics. To support this, we present the first GP-based SNN accelerator, named GustavSNN, which avoids accessing the global membrane potential memory by updating neuron states directly in local registers. In addition, we propose a non-zero row vector (NRV) spike format that enables fine-grained skipping of inactive spike rows. As a result, our proposed architecture achieves up to 11.8× higher energy efficiency (GOPS/W) than naïve GP-based accelerator and 1.43× higher energy efficiency compared to state-of-the-art SNN accelerators.
Sangwoo Hwang, Jahyun Koo 0002, Jaeha Kung 0001
HPCA1
2025 Dissecting and Re-Architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMS
abstract
The advancement of large language models has led to models with billions of parameters, significantly increasing memory and compute demands. Serving such models on conventional hardware is challenging due to limited DRAM capacity and high GPU costs. Thus, in this work, we propose offloading the single-batch token generation to a 3D NAND flash processing-in-memory (PIM) device, leveraging its high storage density to overcome the DRAM capacity wall. We explore 3D NAND flash configurations and present a re-architected PIM array with an H-tree network for optimal latency and cell density. Along with the well-chosen PIM array size, we develop operation tiling and mapping methods for LLM layers, achieving a$2.4 \times$speedup over four RTX4090 with vLLM and comparable performance to four A100 with only 4.9% latency overhead. Our detailed area analysis reveals that the proposed 3D NAND flash PIM architecture can be integrated within a$4.98 ~\text{mm}^{2}$die area under the memory array, without extra area overhead.
Yongjoo Jang, Sangwoo Hwang, Sangwoo Jung 0001, Wonbo Shim, Jaeha Kung 0001
ICCD2
2024 SpikedAttention: Training-Free and Fully Spike-Driven Transformer-to-SNN Conversion with Winner-Oriented Spike Shift for Softmax Operation
abstract
Event-driven spiking neural networks(SNNs) are promising neural networks that reduce the energy consumption of continuously growing AI models. Recently, keeping pace with the development of transformers, transformer-based SNNs were presented. Due to the incompatibility of self-attention with spikes, however, existing transformer-based SNNs limit themselves by either restructuring self-attention architecture or conforming to non-spike computations. In this work, we propose a novel transformer-to-SNN conversion method that outputs an end-to-end spike-based transformer, named SpikedAttention. Our method directly converts the well-trained transformer without modifying its attention architecture. For the vision task, the proposed method converts Swin Transformer into an SNN without post-training or conversion-aware training, achieving state-of-the-art SNN accuracy on ImageNet dataset, i.e., 80.0\% with 28.7M parameters. Considering weight accumulation, neuron potential update, and on-chip data movement, SpikedAttention reduces energy consumption by 42\% compared to the baseline ANN, i.e., Swin-T. Furthermore, for the first time, we demonstrate that SpikedAttention successfully converts a BERT model to an SNN with only 0.3\% accuracy loss on average consuming 58\% less energy on GLUE benchmark. Our code is available at Github ( https://github.com/sangwoohwang/SpikedAttention ).
Sangwoo Hwang, Dahoon Park
NeurIPS1
2021 Adaptive Input-to-Neuron Interlink Development in Training of Spike-Based Liquid State Machines
abstract
In this paper, we present a novel approach in developing input-to-neuron interlinks to achieve better accuracy in spike-based liquid state machines. An energy-efficient Spiking Neural Network suffer from lower accuracy in image classification compared to deep learning models. The previous LSM models randomly connect input neurons to excitatory neurons in a liquid. This limits the expressive power of a liquid model as large portion of excitatory neurons become inactive which never fire. To overcome this limitation, we propose an adaptive interlink development method which achieves 3.2% higher classification accuracy than the static LSM model of 3,200 neurons. Also, our hardware implementation on FPGA improves performance by 3.16~4.99 x or 1.47~3.95× over CPU/GPU.
Sangwoo Hwang, Junghyup Lee
ISCAS1