Sungsoo Han

dblp:339/0572 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 BOLD-Q: Blockwise Outlier-aware Logarithmic Dual-Bias Quantization for Hardware-Efficient LLM Inference
abstract
Large language models (LLMs) deliver strong natural language processing performance, but ever-growing parameter counts strain memory and power budgets for on-device deployments. Quantization alleviates these costs; however, the outlier-heavy statistics of LLM activations and weights force calibration-based static schemes to retain high-precision fallbacks for dynamically varying values, yielding heterogeneous execution paths and overheads. Microscaling (MX) applies blockwise dynamic quantization with a per-block shared exponent, achieving a homogeneous execution path. Nevertheless, at 4-bit precision, prior work faces three limitations: (i) fixed bins fail to capture block-specific distributions and outliers; (ii) quantization error due to the limited resolution of shared-exponent scaling; and (iii) a lack of co-design approaches that balance model quality and hardware efficiency. We propose BOLD-Q, an HW/SW co-design quantization framework that combines the logarithmic number system (LNS) with MX. BOLD-Q introduces blockwise Dual-Bias—selected statically for weights via candidate search and dynamically computed for activations—to shift and refine per-block quantization bins, while LNS-based scaling improves distributional fit. On LLaMA-2 7B, BOLD-Q limits perplexity increase to +0.32 (W4/A8) and +0.60 (W4/A4), outperforming same-precision baselines. We further design an LNS-MAC systolic array with a lightweight preprocessing row that derives and broadcasts Dual-Bias, eliminating per-PE bias units; within the array, multiplies become log-domain additions, and rescaling is adder-based. Compared with a baseline, BOLD-Q reduces area by up to 34.0% and energy by 21.4%, enabling a homogeneous, on-device-friendly, low-precision execution path for LLMs. The code is available at https://github.com/IDSL-SeoulTech/BOLD-Q.
Sungsoo Han, Dahun Choi
DATE1
2023 ShakeFlow: Functional Hardware Description with Latency-Insensitive Interface Combinators
abstract
Functional programming’s benefits for hardware description have long been recognized in the literature. In particular, functional hardware description languages provide combinators such as maps and filters to facilitate the compositional description of circuits. However, it is challenging to apply functional programming with combinators to complex circuits with latency-insensitive interfaces such as valid/ready interfaces due to the cyclic nature of their forward and backward ports.
Sungsoo Han, Minseong Jang, Jeehoon Kang
ASPLOS (2)1