Jinsu Lee

dblp:176/6044 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
3since 2021 · last 2022
0000-0003-2495-029XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 3 since 2021Computer networks · 1
YearPublicationVenuePosition
2022 DSPU: A 281.6mW Real-Time Deep Learning-Based Dense RGB-D Data Acquisition with Sensor Fusion and 3D Perception System-on-Chip
abstract
3D Data in Mobile Platforms
Dongseok Im, Gwangtae Park, Zhiyong Li 0016, Junha Ryu, Donghyeon Han, Jinsu Lee, Wonhoon Park, Hankyul Kwon, Hoi-Jun Yoo
HCS7
2021 An Energy-efficient Floating-Point DNN Processor using Heterogeneous Computing Architecture with Exponent-Computing-in-Memory
abstract
Abstract of Proposed FP CIM Processor (1) Heterogeneous FP Computing Arch. : Separate optimization of FP computing: Realize 2 cycles FP MAC w/ CIM (2) Exponent Computing-in-Memory: In-memory AND/NOR + BL charge reusing: Total memory power 46.4% 2) Mantissa Free Exponent Calculation: Removing redundant normalization: Total MAC power 14.4%
Juhyoung Lee, Ji-Hoon Kim 0004, Wooyoung Jo, Sangyeob Kim, Donghyeon Han, Jinsu Lee, Hoi-Jun Yoo
HCS7
2021 A 3.6 TOPS/W Hybrid FP-FXP Deep Learning Processor with Outlier Compensation for Image-to-Image Application
abstract
A Hybrid floating-point (FP) and fixed-point (FXP) deep learning processor with an outlier-aware channel splitting algorithm is proposed for image-to-image applications on mobile devices. Since the high quality of the reconstructed image through deep learning based image-to-image application requires high bit-precision (> FP16), the mobile processor suffers from the high computation power and large external memory access (EMA). In this work, the proposed algorithm reduces 16-bit FP data to 8-bit FXP data, and only few outliers (2. The hierarchical processor successfully demonstrates the x 4 scale Full-HD super-resolution generation achieving 76 frames-per-second (fps) with 133.3 mW power-consumption at 0.9 V supply and 3.6 TOPS/W of energy-efficiency which is × 3.27 higher than the previous 16-bit FXP processor.
Zhiyong Li 0016, Dongseok Im, Jinsu Lee, Hoi-Jun Yoo
ISCAS3
2019 A 15.2 TOPS/W CNN Accelerator with Similar Feature Skipping for Face Recognition in Mobile Devices
abstract
A low-power face recognition processor with similar feature skipping (SFS) and the tile-based clustering algorithm is proposed for high energy efficiency in mobile devices. For higher energy efficiency face recognition (FR) processor, this paper proposes two key features: 1) Tile-based clustering enables to reduce computation overhead of clustering. 2) SFS binary convolution core is proposed to increase energy efficiency, resulting in 15.2 TOPS/W energy efficiency. Implemented with 65 nm CMOS technology, the 6 mm2FR processor achieves 0.26mW power consumption at 1 frames-per-second (fps) always-on face recognition in mobile devices.
Sangyeob Kim, Juhyoung Lee, Jinsu Lee, Hoi-Jun Yoo
ISCAS4
2018 A 141.4 mW Low-Power Online Deep Neural Network Training Processor for Real-time Object Tracking in Mobile Devices
abstract
A low-power online deep neural network (DNN) training processor is proposed for a real-time object tracking in mobile devices. For a real-time object tracking, a homogeneous core architecture is proposed to achieve 1.33× higher throughput than previous DNN training processor. To reduce the external memory access (EMA), a binary feedback alignment (BFA) algorithm and an integral run-length compression (iRLC) decoder are proposed. While the BFA reduces the EMA by 11.4% compared to the conventional back-propagation approach, the iRLC decoder achieves 29.7% EMA reduction without throughput degradation. Finally, a dropout controller is proposed and achieves 43.9% power reduction through clock-gating. Implemented with 65 nm CMOS technology, the 4.4 mm2DNN training processor achieves 141.1 mW power consumption at 30.4 frames-per-second (fps) real-time object tracking in mobile devices.
Donghyeon Han, Jinsu Lee, Jinmook Lee, Sungpill Choi, Hoi-Jun Yoo
ISCAS2
2017 A 17.5-fJ/bit Energy-Efficient Analog SRAM for Mixed-Signal Processing
abstract
An energy-efficient analog SRAM (A-SRAM) is proposed to eliminate redundant analog-to-digital (A/D) and digital-to-analog (D/A) conversion in mixed-signal systems, such as neuromorphic chips and neural networks. D/A conversion is integrated into the SRAM readout by charge sharing of the proposed split bitline (BL). Also, A/D conversion is integrated into the SRAM write operation with the successive approximation method in the proposed input-output block. Also, a configurable SRAM bitcell array is proposed to allocate the converted digital data without unfilled bitcells. The multirow access decoder selects multiple bitcells in a single column and configures the bitcell array by controlling the BL switches to split BLs. The proposed A-SRAM is implemented using the 65-nm CMOS technology. It achieves 17.5-fJ/bit energy-efficiency and 21-Gbit/s throughput for the analog readout, which are 64% and 1.3 times better than those of the conventional SRAM followed by a digital-to-analog converter (DAC). Also, the area is reduced by 91% compared with the conventional SRAM with analog-to-digital converter (ADC) and DAC.
Jinsu Lee, Dongjoo Shin, Youchang Kim, Hoi-Jun Yoo
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Practical Residual Error of Interference Cancellation for Spread MSK with a Pseudo-Noise Preamble
abstract
Interference cancellation (IC) has been proposed to improve bandwidth utilization for the majority of modern wireless networks. We claim that the conventional model for IC residual error, which is that the residual error power is simply a fraction of the power of the packet being cancelled, is overly simplified. In this paper, we evaluate the practical residual error of IC with spread MSK modulation in a flat fading environment for a low power wide area (LPWA) wireless sensor network application. In particular, we address the two-packets-overlapping scenario as a function of the signal to noise ratio (SNR) of each packet and the overlapping degree. Using software-defined radios, we demonstrate that the average realistic residual error power based on packet transmission with a single pseudo-noise preamble is not a constant fraction of the power of the packet being canceled. Rather, it depends on the synchronization and channel estimation errors, which in turn, depend on the SINR of the stronger packet preamble.
Qiongjie Lin, Jinsu Lee, Mary Ann Weitnauer
GLOBECOM2
2016 A 17.5 fJ/bit energy-efficient analog SRAM for mixed-signal processing
abstract
An energy-efficient analog SRAM (A-SRAM) is proposed to eliminate redundant analog-to-digital (A/D) and digital-to-analog (D/A) conversions in the mixed-signal processing such as a biomedical and a neural network applications. The D/A and the A/D conversion are integrated into the SRAM readout by the charge sharing of the proposed split bit-line (BL) and the SRAM write by the successive approximation method, respectively. And a data structure is newly proposed to allocate each bit of the input data to the binary-weighted bit-cell array. The proposed A-SRAM is implemented using 65 nm CMOS technology. As a result, it achieves 17.5 fJ/bit read energy-efficiency and 21 Gbit/s read throughput, which are 54% lower and 1.3× higher than the conventional SRAM. Also, the area is reduced by 31% compared to the conventional SRAM with ADC and DAC.
Jinsu Lee, Dongjoo Shin, Youchang Kim, Hoi-Jun Yoo
ISCAS1