EDBT 2026 Demo / reviewers in the wild / expert
Han-Sok Suh
dblp:279/8368
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-4466-4824ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CTDM: Resource-Efficient FPGA-Accelerated Simulation of Large-Scale NPU DesignsabstractThis paper proposes a novel approach to accelerate large Neural Processing Unit (NPU) simulations on FPGA through Chain-based Time-Division Multiplexing (CTDM) and its automatic compiler. CTDM replaces repeated logic patterns with a single logic pattern and register chains, which can take advantage of built-in shift register primitives. It reduces FPGA resource utilization more effectively than conventional multiplexer-based Time-Division Multiplexing (TDM) approaches by minimizing logic overhead and routing congestion. The automated CTDM compiler supports various hardware design languages (HDL) including Verilog, VHDL, high-level synthesis (HLS), and Chisel, as well as a wide range of FPGA devices—from small on-premise boards to server-grade hardware simulators like Synopsys ZeBu. To extend the applicability of CTDM to multi-FPGA systems, we propose a block interleaving technique that hides inter-FPGA link latency and fully utilizes the pipeline in a high-speed serial I/O channel. When applied to NVIDIA Deep Learning Accelerator (NVDLA), CTDM achieved a 66% and 82% reduction in LUT and FF utilization, respectively, and enabled the successful deployment of the largest variant of NVDLA on a single AMD U250 FPGA device. This demonstrated a 3,653× acceleration in NVDLA simulation time over the Synopsys VCS simulator on a CPU. This method has already been implemented for the simulation and verification of our proprietary NPUs. Notably, it enabled the simulation of a 4-die 1024 TFLOPS chiplet using 144 FPGAs on ZeBu 5 server. Hyunje Jo, Han-Sok Suh, Hyun-Seok Heo, Jinseok Kim 0006, Hyunsung Kim 0003, Boeui Hong, Jungju Oh, Sunghyun Park 0006, Jinwook Oh, Sunghwan Jo, Kangwook Lee 0008, Jae-sun Seo |
ICCAD | 2 |
| 2025 | Low-Precision Normalization Algorithm and Accelerator for Neural Network TrainingabstractAs one of the most essential techniques in modern deep learning, normalization layer largely improves the convergence speed and performance of deep neural networks (DNN). However, calculating normalization statistics during training is costly, as the two-pass algorithm requires repeated accumulation on the same data. Although parallelized one-pass algorithms are used for variance calculation, low-precision floating-point arithmetic often leads to catastrophic cancellation and accuracy degradation due to uncentered batch statistics and rounding errors. In this paper, we propose Welford-Pairwise Normalization (WPN), a novel normalization algorithm with accelerator design for low-precision training. WPN resolves numerical instability while performing parallel computation. Implemented on the Xilinx Alveo U280 FPGA, WPN achieves up to 28× throughput improvement compared to standard normalization layer with less than 1% accuracy loss with BEiT3 vision transformer, ResNet, MobileNet, and VGG16 model. Han-Sok Suh, Jian Meng, Jae-sun Seo |
ISCAS | 1 |
| 2023 | FPGA Acceleration of GCN in Light of the Symmetry of Graph Adjacency MatrixabstractGraph Convolutional Neural Networks (GCNs) are widely used to process large-scale graph data. Different from deep neural networks (DNNs), GCNs are sparse, irregular, and unstructured, posing unique challenges to hardware acceleration with regular processing elements (PEs). In particular, the adja-cency matrix of a GCN is extremely sparse, leading to frequent but irregular memory access, low spatial/temporal data locality and poor data reuse. Furthermore, a realistic graph usually consists of unstructured data (e.g., unbalanced distributions), creating significantly different processing times and imbalanced workload for each node in GCN acceleration. To overcome these challenges, we propose an end-to-end hardware-software co-design to accelerate GCNs on resource-constrained FPGAs with the features including: (1) A custom dataflow that leverages symmetry along the diagonal of the adjacency matrix to accelerate feature aggregation for undirected graphs. We utilize either the upper or the lower triangular matrix of the adjacency matrix to perform aggregation in GCN to improve data reuse. (2) Unified compute cores for both aggregation and transform phases, with full support to the symmetry-based dataflow. These cores can be dynamically reconfigured to the systolic mode for transformation or as individual accumulators for aggregation in GCN processing. (3) Preprocessing of the graph in software to rearrange the edges and features to match the custom dataflow. This step improves the regularity in memory access and data reuse in the aggregation phase. Moreover, we quantize the GCN precision from FP32 to INT8 to reduce the memory footprint without losing the inference accuracy. We implement our accelerator design in Intel Stratix10 MX FPGA board with HBM2, and demonstrate$1.3\times-110.5\times$improvement in end-to-end GCN latency as compared to the state-of the-art FPGA implementations, on the graph datasets of Cora, Pubmed, Citeseer and Reddit. Gopikrishnan Raveendran Nair, Han-Sok Suh, Mahantesh Halappanavar, Frank Liu 0001, Jae-sun Seo, Yu Cao 0001 |
DATE | 2 |
| 2023 | Algorithm-hardware Co-optimization for Energy-efficient Drone Detection on Resource-constrained FPGAabstractConvolutional neural network (CNN)-based object detection has achieved very high accuracy; e.g., single-shot multi-box detectors (SSDs) can efficiently detect and localize various objects in an input image. However, they require a high amount of computation and memory storage, which makes it difficult to perform efficient inference on resource-constrained hardware devices such as drones or unmanned aerial vehicles (UAVs). Drone/UAV detection is an important task for applications including surveillance, defense, and multi-drone self-localization and formation control. In this article, we designed and co-optimized an algorithm and hardware for energy-efficient drone detection on resource-constrained FPGA devices. We trained an SSD object detection algorithm with a custom drone dataset. For inference, we employed low-precision quantization and adapted the width of the SSD CNN model. To improve throughput, we use dual-data rate operations for DSPs to effectively double the throughput with limited DSP counts. For different SSD algorithm models, we analyze accuracy or mean average precision (mAP) and evaluate the corresponding FPGA hardware utilization, DRAM communication, and throughput optimization. We evaluated the FPGA hardware for a custom drone dataset, Pascal VOC, and COCO2017. Our proposed design achieves a high mAP of 88.42% on the multi-drone dataset, with a high energy efficiency of 79 GOPS/W and throughput of 158 GOPS using the Xilinx Zynq ZU3EG FPGA device on the Open Vision Computer version 3 (OVC3) platform. Our design achieves 1.1 to 8.7× higher energy efficiency than prior works that used the same Pascal VOC dataset, using the same FPGA device, but at a low-power consumption of 2.54 W. For the COCO dataset, our MobileNet-V1 implementation achieved an mAP of 16.8, and 4.9 FPS/W for energy-efficiency, which is ∼ 1.9× higher than prior FPGA works or other commercial hardware platforms. Han-Sok Suh, Jian Meng, Ty Nguyen, Vijay Kumar 0001, Yu Cao 0001, Jae-sun Seo |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2021 | Algorithm-Hardware Co-Optimization for Energy-Efficient Drone Detection on Resource-Constrained FPGAabstractConvolutional neural network (CNN) based object detection has achieved very high accuracy, e.g. single-shot multi-box detectors (SSD) can efficiently detect and localize various objects in an input image. However, they require a high amount of computation and memory storage, which makes it difficult to perform efficient inference on resource-constrained hardware devices such as drones or unmanned aerial vehicles (UAVs). Drone/UAV detection is an important task for applications including surveillance, defense, and multi-drone self-localization and formation control. In this paper, we designed and co-optimized algorithm and hardware for energy-efficient drone detection on resource-constrained FPGA devices. We trained SSD object detection algorithm with a custom drone dataset. For inference, we employed low-precision quantization and adapted the width of the SSD CNN model. To improve throughput, we use dual-data rate operations for DSPs to effectively double the throughput with limited DSP counts. For different SSD algorithm models, we analyze accuracy or mean average precision (mAP) and evaluate the corresponding FPGA hardware utilization, DRAM communication, throughput optimization. Our proposed design achieves a high mAP of 88.42% on the multi-drone dataset, with a high energy-efficiency of 79 GOPS/W and throughput of 158 GOPS using Xilinx Zynq ZU3EG FPGA device on the Open Vision Computer version 3 (OVC3) platform. Our design achieves 2.7X higher energy efficiency than prior works using the same FPGA device, at a low-power consumption of 1.98 W. Han-Sok Suh, Jian Meng, Ty Nguyen, Shreyas K. Venkataramanaiah, Vijay Kumar 0001, Yu Cao 0001, Jae-sun Seo |
FPT | 1 |
| 2020 | FPGA-based Low-Batch Training Accelerator for Modern CNNs Featuring High Bandwidth MemoryabstractTraining convolutional neural networks (CNNs) requires intensive computations as well as a large amount of storage and memory access. While low bandwidth off-chip memories in prior FPGA works have hindered the system-level performance, modern FPGAs offer high bandwidth memory (HBM2) that unlocks opportunities to improve the throughput/energy of FPGA-based CNN training. This paper presents a FPGA accelerator for CNN training which (1) uses HBM2 for efficient off-chip communication, and (2) supports various training operations (e.g. residual connections, stride-2 convolutions) for modern CNNs. We analyze the impact of HBM2 on CNN training workloads, provide a comprehensive comparison with DDR3, and present the strategies to efficiently use HBM2 features for enhanced CNN training performance. For training ResNet-20/VGG-like CNNs for CIFAR-10 dataset with low batch size of 2, the proposed CNN training accelerator on Intel Stratix-10 MX FPGA demonstrates 1.4/1.7X energy-efficiency improvement compared to Stratix-10 GX FPGA with DDR3 memory, and 4.5/9.7 X energy-efficiency improvement compared to Tesla V100 GPU. Shreyas K. Venkataramanaiah, Han-Sok Suh, Shihui Yin, Eriko Nurvitadhi, Aravind Dasu, Yu Cao 0001, Jae-sun Seo |
ICCAD | 2 |