VLDB 2026 Research / reviewers in the wild / expert
Dongseok Im
dblp:244/7623
· DBLP profile ↗
12ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-5856-8921ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 5 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BROCA: A Low-power and Low-latency Conversational Agent RISC-V System-on-Chip for Voice-interactive Mobile Devices
Wooyoung Jo, Seongyon Hong, Beomseok Kwon, Haoyang Sang, Dongseok Im, Sangyeob Kim, Chaeyun Jeong, Yujin Moon, Hoi-Jun Yoo |
HCS | 6 |
| 2024 | NeuGPU: A Neural Graphics Processing Unit for Instant Modeling and Real-Time Rendering on Mobile AR/VR Devicesabstract•Why 3D Modeling using Neural Radiance Field? Junha Ryu, Hankyul Kwon, Wonhoon Park, Zhiyong Li 0016, Beomseok Kwon, Donghyeon Han, Dongseok Im, Sangyeob Kim, Hyungnam Joo, Hoi-Jun Yoo |
HCS | 7 |
| 2024 | Space-Mate: A 303.5mW Real-Time NeRF SLAM Processor with Sparse-Mixture-of-Experts-based AccelerationabstractNeRF-based SLAM for robotic applications face computation barrier Seokchan Song, Haoyang Sang, Dongseok Im, Donghyeon Han, Sangyeob Kim, Hongseok Lee, Hoi-Jun Yoo |
HCS | 3 |
| 2024 | LUTein: Dense-Sparse Bit-Slice Architecture With Radix-4 LUT-Based Slice-Tensor Processing UnitsabstractBit-slice architectures have been developed to support various bit-precision and data sparsity of deep neural networks (DNNs). However, because of low-bit precision and a wide range of sparsity of bit-slice computations, bit-slice architectures are challenging in multiplier-, processing element (PE)-, core-, and software (SW)-level designs. First, data sparsity causes a power trade-off between the Radix numbers of a multiplier, and the previous multipliers cannot take advantage of all sparsity ranges. Second, a bit-slice PE which integrated massive lowbit multiplier-and-accumulate (MAC) units brings about large data transactions compared to a fixed bit-width PE. Third, bitslice core architectures only focus on either dense or sparse data computations, limiting the overall performance of bit-slice computations with a wide sparsity range. Lastly, low-bit bit-slice computations cause massive repetitive instruction fetches across hardware units. To solve the challenges, LUTein is proposed. It exploits the new lookup table (LUT)-based computing method to support the Radix-4 Modified Booth algorithm, achieving low power consumption in all sparsity ranges. Moreover, the slice-tensor PE efficiently processes slice-tensor data by sharing hardware units across the Radix-4 LUT-based MAC units. In addition, the LUTein architecture adopts a systolic datapath with a multi-port buffer to exploit both inter-PE data reuse and slice-level sparsity. Lastly, LUTein's instruction set architecture (ISA) and the hierarchical instruction decoder are introduced to alleviate repetitive instruction fetches. As a result, LUTein outperforms the state-of-the-art bit-slice architecture, Sibia, over 1.34× higher energy-efficiency and 1.78× higher area-efficiency. Dongseok Im, Hoi-Jun Yoo |
HPCA | 1 |
| 2024 | CamPU: A Multi-Camera Processing Unit for Deep Learning-based 3D Spatial Computing SystemsabstractA 3D spatial computing system that understands a surrounding environment and interacts with real-world objects has emerged with the development of deep learning technologies. A multi-camera system captures a surrounding view of a scene using multiple cameras, and a deep neural network (DNN) system extracts semantic features from multi-camera images and provides useful information to users. However, processing a multi-camera system requires massive memory accesses as the number of cameras increases while processing a DNN system can improve throughput by exploiting batch processing. This performance gap limits the overall performance of 3D spatial computing systems. To solve this problem, a multi-camera processing unit (CamPU) is proposed. CamPU exploits the inter- and intra-data reuse methods on multi-camera images, minimizing memory accesses for image projection. Moreover, the out-of-order image projection unit with cache memory is designed to increase multi-image projection throughput by avoiding redundant cache accesses and hiding the latency of high-level memory accesses. Lastly, the overlap-aware blending unit speeds up image blending by efficiently handling overlapping regions between adjacent images. The CamPU architecture is evaluated through RTL-level simulation, and the CamPU-integrated DNN platform provides a comprehensive analysis of end-to-end multi-camera deep learning-based 3D spatial systems. Finally, CamPU speedups the overall system performance 2.9 x faster than an NVIDIA RTX2080Ti GPU platform. Dongseok Im, Hoi-Jun Yoo |
MICRO | 1 |
| 2023 | Sibia: Signed Bit-slice Architecture for Dense DNN Acceleration with Slice-level Sparsity ExploitationabstractDeep neural networks (DNNs) have achieved high performance in many AI fields such as 1-D language, 2-D image, and 3-D point cloud processing applications. Since recent DNN tasks require dense matrix operations with various bit-precision and non-ReLU activation functions, mobile neural processing units (NPUs) suffer from the acceleration of diverse DNN tasks within their limited hardware resources and power budget. Although bit-slice architectures benefit from slice-level computation and slice-level sparsity exploitation, the conventional bit-slice representation is inefficient in bit-slice architectures resulting in poor dense DNN execution. This paper proposes the efficient signed bit-slice architecture, Sibia, with the signed bit-slice representation (SBR) for efficient dense DNN acceleration. The SBR adds a sign bit to each bit-slice and changes signed 11112bit-slice to 00002by borrowing a value of 1 from its lower order of the bit-slice. This scheme generates large numbers of zero bit-slices in dense DNNs even not relying on accuracy-sensitive pruning methods or retraining processes. Moreover, the SBR balances positive and negative values of 2’s complement data, allowing accurate bit-slice-based output speculation that pre-computes high orders of bit-slices. Sibia integrates the signed multiplier-and-accumulate (MAC) units for efficient signed bit-slice computations, and the flexible zero skipping processing element (PE) supports the zero input bit-slice skipping and output skipping for high throughput and energy-efficiency. Additionally, the dynamic sparsity monitoring unit monitors sparsity ratio between input and weight data and determines the more sparse one for zero bit-slice skipping. The heterogeneous network-on-chip (NoC) benefits from data reusability during bit-slice computation, reducing transmission bandwidth. Finally, Sibia outperforms the previous bit-slice architecture, Bit-fusion, over 3.65× higher area-efficiency, 3.88× higher energy-efficiency, and 5.35× higher throughput. Dongseok Im, Gwangtae Park, Zhiyong Li 0016, Junha Ryu, Hoi-Jun Yoo |
HPCA | 1 |
| 2022 | HNPU-V2: A 46.6 FPS DNN Training Processor for Real-World Environmental Adaptation based Robust Object Detection on Mobile Devicesabstract■ Smarter DNNs: # of Parameter ▲ Donghyeon Han, Dongseok Im, Gwangtae Park, Seokchan Song, Juhyoung Lee, Hoi-Jun Yoo |
HCS | 2 |
| 2022 | DSPU: A 281.6mW Real-Time Deep Learning-Based Dense RGB-D Data Acquisition with Sensor Fusion and 3D Perception System-on-Chipabstract3D Data in Mobile Platforms Dongseok Im, Gwangtae Park, Zhiyong Li 0016, Junha Ryu, Donghyeon Han, Jinsu Lee, Wonhoon Park, Hankyul Kwon, Hoi-Jun Yoo |
HCS | 1 |
| 2022 | An Efficient High-quality FHD Super-resolution Mobile Accelerator SoC with Hybrid-precision and Energy-efficient CacheabstractSuper-Resolution on Mobile Platform Zhiyong Li 0016, Dongseok Im, Donghyeon Han, Hoi-Jun Yoo |
HCS | 3 |
| 2021 | PNNPU: A Fast and Efficient 3D Point Cloud-based Neural Network Processor with Block-based Point Processing for Regular DRAM AccessabstractPNN* for Intelligent 3D Vision • Intelligent 3D Vision on Mobile Devices • Accurate & Robust Perception with 3D Structural Information • Mobile 3D Sensor Already Commercialized Juhyoung Lee, Dongseok Im, Hoi-Jun Yoo |
HCS | 3 |
| 2021 | A 3.6 TOPS/W Hybrid FP-FXP Deep Learning Processor with Outlier Compensation for Image-to-Image ApplicationabstractA Hybrid floating-point (FP) and fixed-point (FXP) deep learning processor with an outlier-aware channel splitting algorithm is proposed for image-to-image applications on mobile devices. Since the high quality of the reconstructed image through deep learning based image-to-image application requires high bit-precision (> FP16), the mobile processor suffers from the high computation power and large external memory access (EMA). In this work, the proposed algorithm reduces 16-bit FP data to 8-bit FXP data, and only few outliers (2. The hierarchical processor successfully demonstrates the x 4 scale Full-HD super-resolution generation achieving 76 frames-per-second (fps) with 133.3 mW power-consumption at 0.9 V supply and 3.6 TOPS/W of energy-efficiency which is × 3.27 higher than the previous 16-bit FXP processor. Zhiyong Li 0016, Dongseok Im, Jinsu Lee, Hoi-Jun Yoo |
ISCAS | 2 |
| 2019 | DT-CNN: Dilated and Transposed Convolution Neural Network Accelerator for Real-Time Image Segmentation on Mobile DevicesabstractA convolution neural network (CNN) accelerator is proposed for real-time image segmentation on mobile devices. The proposed CNN processor cuts down the redundant zero computations in dilated and transposed convolution for higher throughput. As a result, the overall computations of the image segmentation are reduced by 86.6% and the proposed CNN processor boosts up the throughput 6.7×. Moreover, the proposed processor utilizes RoI (Region of Interest) based image segmentation algorithm to reduce the overall computational requirement significantly. Although RoI based image segmentation degrades the image segmentation accuracy, the proposed dilation rate adjustment compensates for the accuracy degradation and maintains the accuracy of the full-size image segmentation. Finally, the proposed CNN processor is simulated in 65 nm CMOS technology, and it occupies 6.8 mm2. The proposed processor consumes 196 mW and shows 211 frames-per-second (fps) at the image segmentation for human body parts. Dongseok Im, Donghyeon Han, Sungpill Choi, Hoi-Jun Yoo |
ISCAS | 1 |