Dongseok Im

dblp:244/7623 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-5856-8921ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 5 first-author · 11 since 2021
YearPublicationVenuePosition
2025 BROCA: A Low-power and Low-latency Conversational Agent RISC-V System-on-Chip for Voice-interactive Mobile Devices
Wooyoung Jo, Seongyon Hong, Beomseok Kwon, Haoyang Sang, Dongseok Im, Sangyeob Kim, Chaeyun Jeong, Yujin Moon, Hoi-Jun Yoo
HCS6
2024 NeuGPU: A Neural Graphics Processing Unit for Instant Modeling and Real-Time Rendering on Mobile AR/VR Devices
abstract
•Why 3D Modeling using Neural Radiance Field?
Junha Ryu, Hankyul Kwon, Wonhoon Park, Zhiyong Li 0016, Beomseok Kwon, Donghyeon Han, Dongseok Im, Sangyeob Kim, Hyungnam Joo, Hoi-Jun Yoo
HCS7
2024 Space-Mate: A 303.5mW Real-Time NeRF SLAM Processor with Sparse-Mixture-of-Experts-based Acceleration
abstract
NeRF-based SLAM for robotic applications face computation barrier
Seokchan Song, Haoyang Sang, Dongseok Im, Donghyeon Han, Sangyeob Kim, Hongseok Lee, Hoi-Jun Yoo
HCS3
2024 LUTein: Dense-Sparse Bit-Slice Architecture With Radix-4 LUT-Based Slice-Tensor Processing Units
abstract
Bit-slice architectures have been developed to support various bit-precision and data sparsity of deep neural networks (DNNs). However, because of low-bit precision and a wide range of sparsity of bit-slice computations, bit-slice architectures are challenging in multiplier-, processing element (PE)-, core-, and software (SW)-level designs. First, data sparsity causes a power trade-off between the Radix numbers of a multiplier, and the previous multipliers cannot take advantage of all sparsity ranges. Second, a bit-slice PE which integrated massive lowbit multiplier-and-accumulate (MAC) units brings about large data transactions compared to a fixed bit-width PE. Third, bitslice core architectures only focus on either dense or sparse data computations, limiting the overall performance of bit-slice computations with a wide sparsity range. Lastly, low-bit bit-slice computations cause massive repetitive instruction fetches across hardware units. To solve the challenges, LUTein is proposed. It exploits the new lookup table (LUT)-based computing method to support the Radix-4 Modified Booth algorithm, achieving low power consumption in all sparsity ranges. Moreover, the slice-tensor PE efficiently processes slice-tensor data by sharing hardware units across the Radix-4 LUT-based MAC units. In addition, the LUTein architecture adopts a systolic datapath with a multi-port buffer to exploit both inter-PE data reuse and slice-level sparsity. Lastly, LUTein's instruction set architecture (ISA) and the hierarchical instruction decoder are introduced to alleviate repetitive instruction fetches. As a result, LUTein outperforms the state-of-the-art bit-slice architecture, Sibia, over 1.34× higher energy-efficiency and 1.78× higher area-efficiency.
Dongseok Im, Hoi-Jun Yoo
HPCA1
2024 CamPU: A Multi-Camera Processing Unit for Deep Learning-based 3D Spatial Computing Systems
abstract
A 3D spatial computing system that understands a surrounding environment and interacts with real-world objects has emerged with the development of deep learning technologies. A multi-camera system captures a surrounding view of a scene using multiple cameras, and a deep neural network (DNN) system extracts semantic features from multi-camera images and provides useful information to users. However, processing a multi-camera system requires massive memory accesses as the number of cameras increases while processing a DNN system can improve throughput by exploiting batch processing. This performance gap limits the overall performance of 3D spatial computing systems. To solve this problem, a multi-camera processing unit (CamPU) is proposed. CamPU exploits the inter- and intra-data reuse methods on multi-camera images, minimizing memory accesses for image projection. Moreover, the out-of-order image projection unit with cache memory is designed to increase multi-image projection throughput by avoiding redundant cache accesses and hiding the latency of high-level memory accesses. Lastly, the overlap-aware blending unit speeds up image blending by efficiently handling overlapping regions between adjacent images. The CamPU architecture is evaluated through RTL-level simulation, and the CamPU-integrated DNN platform provides a comprehensive analysis of end-to-end multi-camera deep learning-based 3D spatial systems. Finally, CamPU speedups the overall system performance 2.9 x faster than an NVIDIA RTX2080Ti GPU platform.
Dongseok Im, Hoi-Jun Yoo
MICRO1
2023 Sibia: Signed Bit-slice Architecture for Dense DNN Acceleration with Slice-level Sparsity Exploitation
abstract
Deep neural networks (DNNs) have achieved high performance in many AI fields such as 1-D language, 2-D image, and 3-D point cloud processing applications. Since recent DNN tasks require dense matrix operations with various bit-precision and non-ReLU activation functions, mobile neural processing units (NPUs) suffer from the acceleration of diverse DNN tasks within their limited hardware resources and power budget. Although bit-slice architectures benefit from slice-level computation and slice-level sparsity exploitation, the conventional bit-slice representation is inefficient in bit-slice architectures resulting in poor dense DNN execution. This paper proposes the efficient signed bit-slice architecture, Sibia, with the signed bit-slice representation (SBR) for efficient dense DNN acceleration. The SBR adds a sign bit to each bit-slice and changes signed 11112bit-slice to 00002by borrowing a value of 1 from its lower order of the bit-slice. This scheme generates large numbers of zero bit-slices in dense DNNs even not relying on accuracy-sensitive pruning methods or retraining processes. Moreover, the SBR balances positive and negative values of 2’s complement data, allowing accurate bit-slice-based output speculation that pre-computes high orders of bit-slices. Sibia integrates the signed multiplier-and-accumulate (MAC) units for efficient signed bit-slice computations, and the flexible zero skipping processing element (PE) supports the zero input bit-slice skipping and output skipping for high throughput and energy-efficiency. Additionally, the dynamic sparsity monitoring unit monitors sparsity ratio between input and weight data and determines the more sparse one for zero bit-slice skipping. The heterogeneous network-on-chip (NoC) benefits from data reusability during bit-slice computation, reducing transmission bandwidth. Finally, Sibia outperforms the previous bit-slice architecture, Bit-fusion, over 3.65× higher area-efficiency, 3.88× higher energy-efficiency, and 5.35× higher throughput.
Dongseok Im, Gwangtae Park, Zhiyong Li 0016, Junha Ryu, Hoi-Jun Yoo
HPCA1
2022 HNPU-V2: A 46.6 FPS DNN Training Processor for Real-World Environmental Adaptation based Robust Object Detection on Mobile Devices
abstract
■ Smarter DNNs: # of Parameter ▲
Donghyeon Han, Dongseok Im, Gwangtae Park, Seokchan Song, Juhyoung Lee, Hoi-Jun Yoo
HCS2
2022 DSPU: A 281.6mW Real-Time Deep Learning-Based Dense RGB-D Data Acquisition with Sensor Fusion and 3D Perception System-on-Chip
abstract
3D Data in Mobile Platforms
Dongseok Im, Gwangtae Park, Zhiyong Li 0016, Junha Ryu, Donghyeon Han, Jinsu Lee, Wonhoon Park, Hankyul Kwon, Hoi-Jun Yoo
HCS1
2022 An Efficient High-quality FHD Super-resolution Mobile Accelerator SoC with Hybrid-precision and Energy-efficient Cache
abstract
Super-Resolution on Mobile Platform
Zhiyong Li 0016, Dongseok Im, Donghyeon Han, Hoi-Jun Yoo
HCS3
2021 PNNPU: A Fast and Efficient 3D Point Cloud-based Neural Network Processor with Block-based Point Processing for Regular DRAM Access
abstract
PNN* for Intelligent 3D Vision • Intelligent 3D Vision on Mobile Devices • Accurate & Robust Perception with 3D Structural Information • Mobile 3D Sensor Already Commercialized
Juhyoung Lee, Dongseok Im, Hoi-Jun Yoo
HCS3
2021 A 3.6 TOPS/W Hybrid FP-FXP Deep Learning Processor with Outlier Compensation for Image-to-Image Application
abstract
A Hybrid floating-point (FP) and fixed-point (FXP) deep learning processor with an outlier-aware channel splitting algorithm is proposed for image-to-image applications on mobile devices. Since the high quality of the reconstructed image through deep learning based image-to-image application requires high bit-precision (> FP16), the mobile processor suffers from the high computation power and large external memory access (EMA). In this work, the proposed algorithm reduces 16-bit FP data to 8-bit FXP data, and only few outliers (2. The hierarchical processor successfully demonstrates the x 4 scale Full-HD super-resolution generation achieving 76 frames-per-second (fps) with 133.3 mW power-consumption at 0.9 V supply and 3.6 TOPS/W of energy-efficiency which is × 3.27 higher than the previous 16-bit FXP processor.
Zhiyong Li 0016, Dongseok Im, Jinsu Lee, Hoi-Jun Yoo
ISCAS2
2019 DT-CNN: Dilated and Transposed Convolution Neural Network Accelerator for Real-Time Image Segmentation on Mobile Devices
abstract
A convolution neural network (CNN) accelerator is proposed for real-time image segmentation on mobile devices. The proposed CNN processor cuts down the redundant zero computations in dilated and transposed convolution for higher throughput. As a result, the overall computations of the image segmentation are reduced by 86.6% and the proposed CNN processor boosts up the throughput 6.7×. Moreover, the proposed processor utilizes RoI (Region of Interest) based image segmentation algorithm to reduce the overall computational requirement significantly. Although RoI based image segmentation degrades the image segmentation accuracy, the proposed dilation rate adjustment compensates for the accuracy degradation and maintains the accuracy of the full-size image segmentation. Finally, the proposed CNN processor is simulated in 65 nm CMOS technology, and it occupies 6.8 mm2. The proposed processor consumes 196 mW and shows 211 frames-per-second (fps) at the image segmentation for human body parts.
Dongseok Im, Donghyeon Han, Sungpill Choi, Hoi-Jun Yoo
ISCAS1