EDBT 2026 Demo / reviewers in the wild / expert
Juhyoung Lee
dblp:16/1557
· DBLP profile ↗
10ranked-venue papers
4as first author
6since 2021 · last 2022
0000-0002-2100-1024ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | HNPU-V2: A 46.6 FPS DNN Training Processor for Real-World Environmental Adaptation based Robust Object Detection on Mobile Devicesabstract■ Smarter DNNs: # of Parameter ▲ Donghyeon Han, Dongseok Im, Gwangtae Park, Seokchan Song, Juhyoung Lee, Hoi-Jun Yoo |
HCS | 6 |
| 2022 | A Low-Power Graph Convolutional Network Processor With Sparse Grouping for 3D Point Cloud Semantic Segmentation in Mobile DevicesabstractA low-power graph convolutional network (GCN) processor is proposed for accelerating 3D point cloud semantic segmentation (PCSS) in real-time on mobile devices. Three key features enable the low-power GCN-based 3D PCSS. First, the new hardware-friendly GCN algorithm, sparse grouping-based dilated graph convolution (SG-DGC) is proposed. SG-DGC reduces 71.7% of the overall computation and 76.9% of EMA through the sparse grouping of the point cloud. Second, the two-level pipeline (TLP) consisting of the point-level pipeline (PLP) and group-level pipelining (GLP) was proposed to improve low utilization by the imbalanced workload of GCN. The PLP enables point-level module-wise fusion (PMF) which reduces 47.4% of EMA for low power consumption. Also, center point feature reuse (CPFR) reuses computation results of the redundant operation and reduces 11.4% of computation. Finally, the GLP increased the core utilization by 21.1% by balancing the workload of graph generation and graph convolution and enable$1.1\times $higher throughput. The processor is implemented with 65nm CMOS technology, and the 4.0mm23D PCSS processor show 95mW power consumption while operating in real-time of 30.8 fps in the 3D PCSS of the indoor scene with 4k points. Sangyeob Kim, Juhyoung Lee, Hoi-Jun Yoo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | TSUNAMI: Triple Sparsity-Aware Ultra Energy-Efficient Neural Network Training Accelerator With Multi-Modal Iterative PruningabstractThis article proposes the TSUNAMI, which supports an energy-efficient deep-neural-network training. The TSUNAMI supports multi-modal iterative pruning to generate zeros in activation and weight. Tile-based dynamic activation pruning unit and weight memory shared pruning unit eliminate additional memory access. Coarse-zero skipping controller skips multiple unnecessary multiply-and-accumulation (MAC) operations at once, and fine-zero skipping controller skips randomly located unnecessary MAC operations. Weight sparsity balancer solves a utilization degradation caused by weight sparsity imbalance, and the workload of each convolution core is allocated by a random channel allocator. The TSUNAMI achieves an energy efficiency of 3.42 TFLOPS/W at 0.78V and 50MHz with floating-point 8-bit activation and weight. Also, it achieves an energy efficiency of 405.96 TFLOPS/W at 90% sparsity condition. Sangyeob Kim, Juhyoung Lee, Donghyeon Han, Wooyoung Jo, Hoi-Jun Yoo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | PNNPU: A Fast and Efficient 3D Point Cloud-based Neural Network Processor with Block-based Point Processing for Regular DRAM AccessabstractPNN* for Intelligent 3D Vision • Intelligent 3D Vision on Mobile Devices • Accurate & Robust Perception with 3D Structural Information • Mobile 3D Sensor Already Commercialized Juhyoung Lee, Dongseok Im, Hoi-Jun Yoo |
HCS | 2 |
| 2021 | An Energy-efficient Floating-Point DNN Processor using Heterogeneous Computing Architecture with Exponent-Computing-in-MemoryabstractAbstract of Proposed FP CIM Processor (1) Heterogeneous FP Computing Arch. : Separate optimization of FP computing: Realize 2 cycles FP MAC w/ CIM (2) Exponent Computing-in-Memory: In-memory AND/NOR + BL charge reusing: Total memory power 46.4% 2) Mantissa Free Exponent Calculation: Removing redundant normalization: Total MAC power 14.4% Juhyoung Lee, Ji-Hoon Kim 0004, Wooyoung Jo, Sangyeob Kim, Donghyeon Han, Jinsu Lee, Hoi-Jun Yoo |
HCS | 1 |
| 2021 | OmniDRL: An Energy-Efficient Mobile Deep Reinforcement Learning Accelerators with Dual-mode Weight Compression and Direct Processing of Compressed DataabstractDeep Reinforcement Learning (DRL)▪ No Pre-labelled Data ➔ Training with Trial-and-errors!– Sequential decision making problems @ Unknown environments– Applications: gaming agent, autonomous systems, agent adaptation Juhyoung Lee, Sangyeob Kim, Ji-Hoon Kim 0004, Wooyoung Jo, Donghyeon Han, Hoi-Jun Yoo |
HCS | 1 |
| 2020 | A 54.7 fps 3D Point Cloud Semantic Segmentation Processor with Sparse Grouping Based Dilated Graph Convolutional Network for Mobile DevicesabstractThe graph convolutional network (GCN) based 3D point cloud semantic segmentation (PCSS) processor for mobile devices is proposed. GCN based 3D PCSS requires a lot of computation, making it unsuitable for real-time operation in mobile devices. For real-time 3D PCSS on mobile devices, this paper proposes two key features: 1) a sparse grouping based dilated graph convolution (SG-DGC) which reduces 71.7% of the overall computation of GCN by simply dividing input point cloud into multiple sparse point cloud. 2) group-level pipelining which improves low pipeline utilization due to the computation imbalance of GCN. Finally, the proposed GCN processor is simulated in 65 nm CMOS technology and occupies 4.0 mm2. The proposed processor consumes 176mW and shows 54.7 frames-per-second (fps) for the 3D point cloud semantic segmentation of indoor scene with 4k points. Sangyeob Kim, Juhyoung Lee, Hoi-Jun Yoo |
ISCAS | 3 |
| 2019 | A 15.2 TOPS/W CNN Accelerator with Similar Feature Skipping for Face Recognition in Mobile DevicesabstractA low-power face recognition processor with similar feature skipping (SFS) and the tile-based clustering algorithm is proposed for high energy efficiency in mobile devices. For higher energy efficiency face recognition (FR) processor, this paper proposes two key features: 1) Tile-based clustering enables to reduce computation overhead of clustering. 2) SFS binary convolution core is proposed to increase energy efficiency, resulting in 15.2 TOPS/W energy efficiency. Implemented with 65 nm CMOS technology, the 6 mm2FR processor achieves 0.26mW power consumption at 1 frames-per-second (fps) always-on face recognition in mobile devices. Sangyeob Kim, Juhyoung Lee, Jinsu Lee, Hoi-Jun Yoo |
ISCAS | 2 |
| 2018 | A 46.1 fps Global Matching Optical Flow Estimation Processor for Action Recognition in Mobile DevicesabstractA real-time global matching optical flow estimation (OFE) processor is proposed for action recognition in mobile devices. The global OFE requires a large number of external memory accesses (EMAs) and matrix computations, thus it is incompatible on mobile devices with real-time constraints. For real-time OFE on mobile devices, this paper proposes two key features, both of which to reduce the required memory bandwidth and a number of computations: 1) Tile-based hierarchical OFE enables intermediate data to be processed within 328 KB on-chip memory without external memory access. 2) Background skipping eliminates redundant matrix computation for zero optical flow region. Therefore, the proposed features reduce external memory bandwidth and computation by 99.7 % and 50.7 %, respectively. The proposed 4 mm2OFE processor is implemented in 65 nm CMOS technology and it achieves real-time OFE of 46.1 frames-per-second (fps) throughput for an image resolution of QVGA (320 × 240) and the resulting optical flow can be successfully used for action recognition. Juhyoung Lee, Sungpill Choi, Dongjoo Shin, Hoi-Jun Yoo |
ISCAS | 1 |
| 2004 | The Development of Postech Hand 5abstractWe define that the end effector is the device which interacts with the environment or contacts objects to execute tasks. Up to now, many researchers have developed anthropomorphic robotic hands as end effectors. We discuss the problems in the development of a human-scale and motor-driven anthropomorphic robot hand. In this paper, PRH's design concept, kinematic design and developments of the actuator, transmission, and sensing device are presented. By imitating the physiology of human hands, we devised new metacarpalphalangeal joints and interphalangeal joints suitable for human-size motor-driven robot hands. Juhyoung Lee, Youngil Youm, Wan Kyun Chung |
ICRA | 1 |