EDBT 2026 Demo / reviewers in the wild / expert
Wuyoung Jang
dblp:328/8573
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0000-2088-6775ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BEVSA: A Real-Time Bird's-Eye-View Semantic Segmentation Accelerator for Multi-Camera SystemabstractA bird’s-eye-view (BEV) semantic segmentation accelerator (BEVSA) is proposed for real-time 3D space perception in multi-camera system (MCS). The view transformation from multi-camera-view to BEV obstructs the real-time operation on edge devices through 68.3 ms of time consumption during BEV pooling, which requires sorting and irregular memory access over wide searching space. Moreover, the 69.1% high average input activation sparsity of the segmentation process for the transformed BEV features results in excessive meaningless computations. For the real-time implementation of BEV semantic segmentation on edge platforms, this paper proposes two key features: 1) Block-decomposed hierarchical BEV pooling cluster that partitions the searching space in an MCS-suitable way and supports parallel pooling, achieving $43.2 \times$ speedup for BEV pooling over the edge computing platform; 2) Coarse-to-fine-grained zero skipping convolution cluster, which conducts coarse-grained zero skipping for all-zero channels and tile-wise fine-grained zero skipping, improving the convolution throughput by $1.61 \times$. Implemented with 28 nm technology, and evaluated on two representative MCS datasets, BEVSA finally achieves 23.1 frames-per-second of real-time BEV segmentation throughput with $167.4 \times$ higher energy-per-frame over the edge computing platform. Jueun Jung, Wuyoung Jang, Jihyeon Hwang |
DAC | 3 |
| 2024 | LSPU: A 20.7 ms Low-Latency Point Neural Network-Based 3D Perception and Semantic LiDAR SLAM System-on-Chip for Autonomous Driving SystemabstractIntelligent 3D Interaction with Wide & Dynamic Surroundings Jueun Jung, Seungbin Kim, Bokyoung Seo, Wuyoung Jang, Jeongmin Shin, Donghyeon Han, Kyuho Jason Lee |
HCS | 4 |
| 2024 | A 422.1 Mpixels/J Tile-based 4K Super Resolution Processor with Variable Bit CompressionabstractA super resolution (SR) accelerator with variable bit compression method is proposed for 4K restoration with > 60 frames-per-second (fps) in mobile devices. Since 4K SR has huge intermediate feature maps, large external memory access (EMA) bandwidth is required that recent mobile processors cannot support. Previous SR processors quantized feature maps and weights for EMA reduction, but it has a trade-off between data compression rate and performance drop. To facilitate > 60 fps 4K SR on mobile processor, this work proposes two key features: 1) Tile-based distribution-aware statistical encoding that results in 61.2% compression rate without information loss; 2) An energy-efficient SR processor which supports variable bit encoding, achieving 59.1% reduction of EMA. Designed with 28 nm CMOS technology, the proposed system can accelerate ×2 scale 4K image restoration at 68.7 fps. It shows 3.69 TOPS of peak performance and 422.1 Mpixels/J of energy efficiency, achieving 1.35× higher energy efficiency than the previous SR processor. Wuyoung Jang, Jinhoon Jo, Jueun Jung, Donghyeon Han, Kyuho Jason Lee |
ISCAS | 1 |
| 2024 | An Energy-Efficient, Unified CNN Accelerator for Real-Time Multi-Object Semantic Segmentation for Autonomous VehicleabstractAn energy-efficient, unified convolutional neural network (CNN) accelerator is proposed with a lightweight RGB-D network to achieve real-time, multi-object semantic segmentation in autonomous electric vehicle system. First, a lightweight Depth-fused Trilateral Network (DTN) is proposed to achieve high accuracy and real-time operation for road and multi-object segmentation at the same time. Optimized with various types of convolution layers and limited hardware resources, the DTN achieves 94.73% accuracy on KITTI Road dataset. Second, the unified CNN processor is designed with dual-mode shift-register-based input reconfiguration units and layer fusion architecture with 2-types of processing elements for depth-wise separable convolution (DSC) to support 5 different types of convolution layers including standard convolution, dilated convolution, transposed convolution, point-wise convolution, and DSC. With flexible architecture, it achieves 17.97$\times$higher throughput with DTN and DSC layer fusion architecture reduces 34.7% of overall external memory access. Implemented with 28nm CMOS technology, the unified CNN processor shows 43.6 mW power consumption and 4.94 TOPS/W energy efficiency. As a result, the proposed system with DTN realizes 40.07 frames-per-second (fps) throughputs in multi-object semantic segmentation application with high resolution driving scenes dataset. Jueun Jung, Seungbin Kim, Wuyoung Jang, Bokyoung Seo, Kyuho Jason Lee |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | An Energy-Efficient CNN Accelerator for Multi-object Real-Time Semantic Segmentation in Autonomous VehicleabstractAn energy-efficient convolutional neural network (CNN) accelerator is proposed for real-time segmentation in autonomous electric vehicle (AEV) system. The computation of semantic segmentation with high-resolution images makes it difficult for real-time operation in time-critical and resource-constrained AEV. To facilitate real-time implementation in AEV, this paper proposes two key features: 1) A compressed multi-object Depth-fused Trilateral Network (DTN) with dilated convolution and depthwise separable convolution that reduces 90% of the overall computation of baseline [1] and achieves 94.73% accuracy on KITTI Road dataset; 2) An energy-efficient CNN accelerator, which supports 5 types of CONV’s, achieving 1.33× higher throughput than the previous processor [2]. Finally, the proposed processor is designed in 28 nm CMOS technology. It consumes 65.7 mW of power and achieves 2.91 TOPS/W of energy efficiency. As a result, the system realizes 72.2 and 37 frames-per-second of semantic segmentation for road and multi-objects with high resolution. Jueun Jung, Seungbin Kim, Wuyoung Jang, Hoichang Jeong, Kyuho Jason Lee |
ISCAS | 3 |