VLDB 2026 Research / reviewers in the wild / expert
Jueun Jung
dblp:294/2705
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0003-1632-159XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Hardware-Software Co-design of Lightweight Polyp Segmentation Network and Ultra-low-power Processor for Capsule Endoscopy
Ghangmin Yun, Jaekyung Lee, Kyungkeon Chung, Jueun Jung, Bokyoung Seo, Hyejin Lee, Junyoung Park 0002, Kyuho Jason Lee |
ISCAS | 5 |
| 2026 | HotBa: A Heterogeneous Mamba Accelerator with Δ-Guided Early Rejection for Speculative DecodingabstractMamba has emerged as a promising alternative to Transformers for on-device LLM inference, offering linear complexity and eliminating KV-cache. However, autoregressive decoding reloads full model weights every token, accounting for 93.2% of total inference energy, with no existing accelerator resolving this overhead. We present HotBa, a heterogeneous Mamba accelerator that reduces per-token weight transfer by 82% and redundant computation by 59% via Δ-guided early rejection for wide-tree speculative decoding in Mamba, while a heterogeneous INT8/FP16 core and tree management unit achieve 40.4× area efficiency and 5.18× SSM speedup with 0.4% area overhead. Synthesized in 28nm CMOS, HotBa achieves 75.32 tokens/s with 1.65× speedup and 7.55× energy efficiency over the state-of-the-art Mamba accelerator. Ghangmin Yun, Jueun Jung, Bokyoung Seo, Chaeyoon Kim, Junghyun Yoo, Kyuho Jason Lee |
ISLPED | 2 |
| 2025 | BEVSA: A Real-Time Bird's-Eye-View Semantic Segmentation Accelerator for Multi-Camera SystemabstractA bird’s-eye-view (BEV) semantic segmentation accelerator (BEVSA) is proposed for real-time 3D space perception in multi-camera system (MCS). The view transformation from multi-camera-view to BEV obstructs the real-time operation on edge devices through 68.3 ms of time consumption during BEV pooling, which requires sorting and irregular memory access over wide searching space. Moreover, the 69.1% high average input activation sparsity of the segmentation process for the transformed BEV features results in excessive meaningless computations. For the real-time implementation of BEV semantic segmentation on edge platforms, this paper proposes two key features: 1) Block-decomposed hierarchical BEV pooling cluster that partitions the searching space in an MCS-suitable way and supports parallel pooling, achieving $43.2 \times$ speedup for BEV pooling over the edge computing platform; 2) Coarse-to-fine-grained zero skipping convolution cluster, which conducts coarse-grained zero skipping for all-zero channels and tile-wise fine-grained zero skipping, improving the convolution throughput by $1.61 \times$. Implemented with 28 nm technology, and evaluated on two representative MCS datasets, BEVSA finally achieves 23.1 frames-per-second of real-time BEV segmentation throughput with $167.4 \times$ higher energy-per-frame over the edge computing platform. Jueun Jung, Wuyoung Jang, Jihyeon Hwang |
DAC | 2 |
| 2025 | A Real-time Point Cloud Segmentation System with Optimized Ground Estimation Algorithm and Selective Neural NetworkabstractThis paper introduces a fast and accurate point segmentation system for real-time (< 50 ms) 3D-LiDAR semantic segmentation. The real-time application of 3D point-cloud neural networks (PNNs) for semantic segmentation of LiDAR-measured data faces considerable challenges such as memory overhead and processing delay in its implementation on GPUs due to the significant computational requirements. These challenges arise from the large volume of points and their complex spatial relationships, leading to intense computational and memory usage. To facilitate real-time implementation with high accuracy, a selective point segmentation (SPS) system is proposed with 3 key features: 1) Adaptive ground estimation for the surroundings, excluding the ground from PNN inference, thereby reducing latency by 46.0%, and addressing accuracy reductions due to false positives by implementing 2-step bin skipping; 2) Coarse-grained entropy-and-density-based region skipping (RSK) excludes large areas from ground estimation; and 3) Fine-grained bin skipping (BSK) with z-distribution skips non-ground bins within these areas. Together, the system achieves a processing time of 42.24 ms and a 3D semantic segmentation accuracy of 90.69% at the semantic KITTI dataset. Jihyeon Hwang, Jueun Jung, Kyuho Jason Lee |
ISCAS | 3 |
| 2025 | MedBiSeNet: Efficient Bilateral Segmentation Network for Real-time Medical Image ProcessingabstractMedical image segmentation is a crucial component in modern computer-aided diagnosis, especially in colonoscopy which has a high failure rate (22-24%) in detecting polyps. Additionally, real-time processing on resource-limited devices is increasingly needed for immediate decision-making and supporting new approaches like capsule endoscopy. However, previous models suffered from either high computational costs or low accuracy due to the difficulty of handling ambiguous boundaries of medical images. This paper presents MedBiSeNet, a high-performance network for real-time medical image segmentation. The MedBiSeNet addresses the challenges with the following features: 1) Boundary-enhanced Bilateral Path to extract accurate boundary information, 2) Noise Refining Feature Fusion Module (NR-FFM) to eliminate unnecessary information and fuse features from bilateral branches. As a result, MedBiSeNet achieves state-of-the-art performance on polyp segmentation datasets (Kvasir-SEG, CVC-ClinicDB, ETIS-LARIBPOLYDB) with dice scores exceeding 0.96 while reducing parameters by 81.9% and computations by 95.4% compared to previous models, achieving 20.7 FPS on NVIDIA's Jetson TX2. Kyungkeon Chung, Ghangmin Yun, Jaekyung Lee, Jueun Jung, Bokyoung Seo, Hyejin Lee |
ISCAS | 5 |
| 2024 | LSPU: A 20.7 ms Low-Latency Point Neural Network-Based 3D Perception and Semantic LiDAR SLAM System-on-Chip for Autonomous Driving SystemabstractIntelligent 3D Interaction with Wide & Dynamic Surroundings Jueun Jung, Seungbin Kim, Bokyoung Seo, Wuyoung Jang, Jeongmin Shin, Donghyeon Han, Kyuho Jason Lee |
HCS | 1 |
| 2024 | A 422.1 Mpixels/J Tile-based 4K Super Resolution Processor with Variable Bit CompressionabstractA super resolution (SR) accelerator with variable bit compression method is proposed for 4K restoration with > 60 frames-per-second (fps) in mobile devices. Since 4K SR has huge intermediate feature maps, large external memory access (EMA) bandwidth is required that recent mobile processors cannot support. Previous SR processors quantized feature maps and weights for EMA reduction, but it has a trade-off between data compression rate and performance drop. To facilitate > 60 fps 4K SR on mobile processor, this work proposes two key features: 1) Tile-based distribution-aware statistical encoding that results in 61.2% compression rate without information loss; 2) An energy-efficient SR processor which supports variable bit encoding, achieving 59.1% reduction of EMA. Designed with 28 nm CMOS technology, the proposed system can accelerate ×2 scale 4K image restoration at 68.7 fps. It shows 3.69 TOPS of peak performance and 422.1 Mpixels/J of energy efficiency, achieving 1.35× higher energy efficiency than the previous SR processor. Wuyoung Jang, Jinhoon Jo, Jueun Jung, Donghyeon Han, Kyuho Jason Lee |
ISCAS | 4 |
| 2024 | An Energy-Efficient 3D Point Neural Network Accelerator with Fine-grained LiDAR-SoC Pipeline Structureabstract3D point neural network (PNN) segmentation using LiDAR data has emerged as a fundamental stage of high-level intelligence algorithms for autonomous applications such as SLAM, path planning, object detection, etc. However, previous processors were not feasible for real-time and low-power 3D PNN systems since they wasted ~100 ms of LiDAR's sensing time and required 107.3 mW of external memory access before PNN processing. Furthermore, their compute-intensive bin partitioning and point sampling methods were not suitable for large-scale outdoor data, causing significant computing power. Therefore, the entire system, from sensing to processing, must be taken into account for 3D PNN processor implementation. This paper proposes L-PNPU, an energy-efficient 3D PNN segmentation processor optimized with the unique mechanical characteristics of LiDAR. It is designed with three key features: 1) Azimuthal bin partitioning to reduce power and latency, 2) Modified PNN algorithm co-optimized with heterogeneous architecture to remove redundant operation and reduce energy, and 3) Fine-grained LiDAR-System-on-Chip (SoC) pipeline structure to enhance the system energy and throughput. At 250 MHz and 1.0V, L-PNPU achieves 1.27M points/s of throughput and 0.51 μJ/point of energy efficiency. Bokyoung Seo, Jueun Jung, Donghyeon Han, Kyuho Jason Lee |
ISLPED | 2 |
| 2024 | An Energy-Efficient, Unified CNN Accelerator for Real-Time Multi-Object Semantic Segmentation for Autonomous VehicleabstractAn energy-efficient, unified convolutional neural network (CNN) accelerator is proposed with a lightweight RGB-D network to achieve real-time, multi-object semantic segmentation in autonomous electric vehicle system. First, a lightweight Depth-fused Trilateral Network (DTN) is proposed to achieve high accuracy and real-time operation for road and multi-object segmentation at the same time. Optimized with various types of convolution layers and limited hardware resources, the DTN achieves 94.73% accuracy on KITTI Road dataset. Second, the unified CNN processor is designed with dual-mode shift-register-based input reconfiguration units and layer fusion architecture with 2-types of processing elements for depth-wise separable convolution (DSC) to support 5 different types of convolution layers including standard convolution, dilated convolution, transposed convolution, point-wise convolution, and DSC. With flexible architecture, it achieves 17.97$\times$higher throughput with DTN and DSC layer fusion architecture reduces 34.7% of overall external memory access. Implemented with 28nm CMOS technology, the unified CNN processor shows 43.6 mW power consumption and 4.94 TOPS/W energy efficiency. As a result, the proposed system with DTN realizes 40.07 frames-per-second (fps) throughputs in multi-object semantic segmentation application with high resolution driving scenes dataset. Jueun Jung, Seungbin Kim, Wuyoung Jang, Bokyoung Seo, Kyuho Jason Lee |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | A Real-Time Sparsity-Aware 3D-CNN Processor for Mobile Hand Gesture RecognitionabstractA sparsity-aware 3D-convolution neural network (3D-CNN) accelerator is proposed for the real-time mobile hand gesture recognition (HGR) system. The complex computation of 3D-convolution with the video data makes it difficult for real-time operation, especially in a resource-constrained mobile platform. To facilitate real-time implementation of HGR, this paper proposes three key features: 1) Spatio-temporal Variation Encoding and Inter-frame Differential Aware Network for highly sparse and lightweight network, reducing 94.03% parameters with only 2.57% accuracy loss on NvGesture dataset; 2) the ROI-only Computation architecture for utilizing activation sparsity to reduce the number of MAC operations and the external memory bandwidth by 84.3% and 72.3%, respectively; 3) Weight Sparsity-aware PE and Sparsity-distribution-aware Workload Allocation speed up the inference by$19.8\times $. As a result, the low-latency 3D-CNN accelerator utilizes both activation and weight sparsity with data mapping to maximize the reusability of 3D-CNN, achieving$31\times $faster inference than the state-of-the-art. The proposed processor is designed in 65 nm CMOS technology. It consumes 35 mW of power and achieves 46.25 TOPS/W of energy efficiency. As a result, the system realized 1.584 ms inference latency for real-time HGR in a mobile platform. Seungbin Kim, Jueun Jung, Kyuho Jason Lee |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | An Energy-Efficient CNN Accelerator for Multi-object Real-Time Semantic Segmentation in Autonomous VehicleabstractAn energy-efficient convolutional neural network (CNN) accelerator is proposed for real-time segmentation in autonomous electric vehicle (AEV) system. The computation of semantic segmentation with high-resolution images makes it difficult for real-time operation in time-critical and resource-constrained AEV. To facilitate real-time implementation in AEV, this paper proposes two key features: 1) A compressed multi-object Depth-fused Trilateral Network (DTN) with dilated convolution and depthwise separable convolution that reduces 90% of the overall computation of baseline [1] and achieves 94.73% accuracy on KITTI Road dataset; 2) An energy-efficient CNN accelerator, which supports 5 types of CONV’s, achieving 1.33× higher throughput than the previous processor [2]. Finally, the proposed processor is designed in 28 nm CMOS technology. It consumes 65.7 mW of power and achieves 2.91 TOPS/W of energy efficiency. As a result, the system realizes 72.2 and 37 frames-per-second of semantic segmentation for road and multi-objects with high resolution. Jueun Jung, Seungbin Kim, Wuyoung Jang, Hoichang Jeong, Kyuho Jason Lee |
ISCAS | 1 |