Bokyoung Seo

dblp:370/8516 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0000-6255-6299ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 A Hardware-Software Co-design of Lightweight Polyp Segmentation Network and Ultra-low-power Processor for Capsule Endoscopy
Ghangmin Yun, Jaekyung Lee, Kyungkeon Chung, Jueun Jung, Bokyoung Seo, Hyejin Lee, Junyoung Park 0002, Kyuho Jason Lee
ISCAS6
2026 HotBa: A Heterogeneous Mamba Accelerator with Δ-Guided Early Rejection for Speculative Decoding
abstract
Mamba has emerged as a promising alternative to Transformers for on-device LLM inference, offering linear complexity and eliminating KV-cache. However, autoregressive decoding reloads full model weights every token, accounting for 93.2% of total inference energy, with no existing accelerator resolving this overhead. We present HotBa, a heterogeneous Mamba accelerator that reduces per-token weight transfer by 82% and redundant computation by 59% via Δ-guided early rejection for wide-tree speculative decoding in Mamba, while a heterogeneous INT8/FP16 core and tree management unit achieve 40.4× area efficiency and 5.18× SSM speedup with 0.4% area overhead. Synthesized in 28nm CMOS, HotBa achieves 75.32 tokens/s with 1.65× speedup and 7.55× energy efficiency over the state-of-the-art Mamba accelerator.
Ghangmin Yun, Jueun Jung, Bokyoung Seo, Chaeyoon Kim, Junghyun Yoo, Kyuho Jason Lee
ISLPED3
2025 MedBiSeNet: Efficient Bilateral Segmentation Network for Real-time Medical Image Processing
abstract
Medical image segmentation is a crucial component in modern computer-aided diagnosis, especially in colonoscopy which has a high failure rate (22-24%) in detecting polyps. Additionally, real-time processing on resource-limited devices is increasingly needed for immediate decision-making and supporting new approaches like capsule endoscopy. However, previous models suffered from either high computational costs or low accuracy due to the difficulty of handling ambiguous boundaries of medical images. This paper presents MedBiSeNet, a high-performance network for real-time medical image segmentation. The MedBiSeNet addresses the challenges with the following features: 1) Boundary-enhanced Bilateral Path to extract accurate boundary information, 2) Noise Refining Feature Fusion Module (NR-FFM) to eliminate unnecessary information and fuse features from bilateral branches. As a result, MedBiSeNet achieves state-of-the-art performance on polyp segmentation datasets (Kvasir-SEG, CVC-ClinicDB, ETIS-LARIBPOLYDB) with dice scores exceeding 0.96 while reducing parameters by 81.9% and computations by 95.4% compared to previous models, achieving 20.7 FPS on NVIDIA's Jetson TX2.
Kyungkeon Chung, Ghangmin Yun, Jaekyung Lee, Jueun Jung, Bokyoung Seo, Hyejin Lee
ISCAS6
2024 LSPU: A 20.7 ms Low-Latency Point Neural Network-Based 3D Perception and Semantic LiDAR SLAM System-on-Chip for Autonomous Driving System
abstract
Intelligent 3D Interaction with Wide & Dynamic Surroundings
Jueun Jung, Seungbin Kim, Bokyoung Seo, Wuyoung Jang, Jeongmin Shin, Donghyeon Han, Kyuho Jason Lee
HCS3
2024 An Energy-Efficient 3D Point Neural Network Accelerator with Fine-grained LiDAR-SoC Pipeline Structure
abstract
3D point neural network (PNN) segmentation using LiDAR data has emerged as a fundamental stage of high-level intelligence algorithms for autonomous applications such as SLAM, path planning, object detection, etc. However, previous processors were not feasible for real-time and low-power 3D PNN systems since they wasted ~100 ms of LiDAR's sensing time and required 107.3 mW of external memory access before PNN processing. Furthermore, their compute-intensive bin partitioning and point sampling methods were not suitable for large-scale outdoor data, causing significant computing power. Therefore, the entire system, from sensing to processing, must be taken into account for 3D PNN processor implementation. This paper proposes L-PNPU, an energy-efficient 3D PNN segmentation processor optimized with the unique mechanical characteristics of LiDAR. It is designed with three key features: 1) Azimuthal bin partitioning to reduce power and latency, 2) Modified PNN algorithm co-optimized with heterogeneous architecture to remove redundant operation and reduce energy, and 3) Fine-grained LiDAR-System-on-Chip (SoC) pipeline structure to enhance the system energy and throughput. At 250 MHz and 1.0V, L-PNPU achieves 1.27M points/s of throughput and 0.51 μJ/point of energy efficiency.
Bokyoung Seo, Jueun Jung, Donghyeon Han, Kyuho Jason Lee
ISLPED1
2024 An Energy-Efficient, Unified CNN Accelerator for Real-Time Multi-Object Semantic Segmentation for Autonomous Vehicle
abstract
An energy-efficient, unified convolutional neural network (CNN) accelerator is proposed with a lightweight RGB-D network to achieve real-time, multi-object semantic segmentation in autonomous electric vehicle system. First, a lightweight Depth-fused Trilateral Network (DTN) is proposed to achieve high accuracy and real-time operation for road and multi-object segmentation at the same time. Optimized with various types of convolution layers and limited hardware resources, the DTN achieves 94.73% accuracy on KITTI Road dataset. Second, the unified CNN processor is designed with dual-mode shift-register-based input reconfiguration units and layer fusion architecture with 2-types of processing elements for depth-wise separable convolution (DSC) to support 5 different types of convolution layers including standard convolution, dilated convolution, transposed convolution, point-wise convolution, and DSC. With flexible architecture, it achieves 17.97$\times$higher throughput with DTN and DSC layer fusion architecture reduces 34.7% of overall external memory access. Implemented with 28nm CMOS technology, the unified CNN processor shows 43.6 mW power consumption and 4.94 TOPS/W energy efficiency. As a result, the proposed system with DTN realizes 40.07 frames-per-second (fps) throughputs in multi-object semantic segmentation application with high resolution driving scenes dataset.
Jueun Jung, Seungbin Kim, Wuyoung Jang, Bokyoung Seo, Kyuho Jason Lee
IEEE Trans. Circuits Syst. I Regul. Pap.4