Seunghyun Moon

dblp:216/3790 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0003-0027-2666ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Bayesian Deep-Learning Processor for Real-Time Bio-Applications With Structured Monte Carlo Dropout for High-Volume Sample Generation
abstract
This work presents a Bayesian neural network (BNN) processor for real-time, edge-based medical applications, designed to generate high-volume inference outputs efficiently, providing rapid and precise uncertainty estimation. To address the computational and memory-intensive demands of BNNs, we propose a structured Monte Carlo (MC) dropout method and an efficient processing technique for depth-wise separable convolution, thereby minimizing both the memory and computational burdens. The proposed processor, implemented in 28-nm LP CMOS, is benchmarked on a 12-lead ECG dataset and demonstrates significant improvements in energy efficiency, achieving a$12.9\times $enhancement over conventional MC dropout methods. Furthermore, the processor incorporates a model-switching technique based on uncertainty estimation, resulting in$3.2\times $lower energy consumption while maintaining robust and accurate ECG classification, even under challenging conditions involving noise and motion artifacts.
Jeong-Min Woo, Seunghyun Moon, Jae-Yoon Sim, Hyunwoo Son
IEEE Trans. Circuits Syst. I Regul. Pap.3
2023 Joint Optimization of Cache Management and Graph Reordering for GCN Acceleration
abstract
Graph Convolutional Networks (GCNs) have demonstrated their efficacy in various real-world applications such as social networks and recommendation systems. Accelerating GCNs presents unique challenges due to their large number of nodes, sparse and heavily skewed connections. The reordering of the adjacency matrix has been the main strategy to effectively reduce the amount of re-access. Existing techniques of the reordering are categorized into i) degree-based sorting to identify high-degree nodes so that their data could be stored in the cache and ii) graph partitioning to maximally reuse the clustered data. However, as connections among the nodes vary significantly, processing various GCNs with a single strategy would cause performance degradation. This paper presents a software/hardware co-optimized platform for processing of general GCNs. We propose a hybrid scheme in the graph reordering that combines a sorting and a clustering in an adaptively optimized two-way partitioning. The two-way partitioning enables an efficient allocation of the on-chip cache memory space to reduce off-chip memory access by 4-to-12 %. The implemented accelerator in 28nm demonstrates full functionalities with improved energy-efficiency by$2.2-\text{to}-3.7\times$compared to the previous GCN accelerators.
Kyeong-Jun Lee, Seunghyun Moon, Jae-Yoon Sim
ISLPED4
2023 Bottleneck-Stationary Compact Model Accelerator With Reduced Requirement on Memory Bandwidth for Edge Applications
abstract
State-of-the-art compact models such as MobileNets and EfficientNets are structured using a linear bottleneck and inverted residuals. Hardware architecture using a single dataflow strategy fails to balance the required memory bandwidth with the given computational resources. This work presents a heterogeneous dual-core accelerator that performs a block-wise pipelined process as a unit using a bottleneck-stationary (BS) dataflow. The BS greatly relieves the requirement on DRAM bandwidth and on-chip SRAM capacity. A look-behind-only attention is also proposed as a co-optimized algorithm. Compared to the state-of-the-art hardware scheme, the proposed accelerator demonstrates a reduction of 1.8-$2.9\times $in latency and 2.2-$3\times $in energy consumption, respectively.For verification, the accelerator with a 16-bit integer precision was implemented using 28nm CMOS process. Measurements show energy efficiencies of 0.5-to-3.75 TOPS/W in a supply voltage range of 0.55-to-1.15V.
Seunghyun Moon, Kyeong-Jun Lee, Jae-Yoon Sim
IEEE Trans. Circuits Syst. I Regul. Pap.2