Hoi-Jun Yoo

dblp:37/2676 · DBLP profile ↗
← Back
120ranked-venue papers
2as first author
44since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 110 · 2 first-author · 44 since 2021Applied, interdisciplinary, general and emerging computing · 5Graphics, computer vision, multimedia, augmented reality and games · 4Artificial intelligence and machine learning · 2Software engineering, systems software and programming languages · 2 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 GyRot: Leveraging Hidden Synergy Between Rotation and Fine-Grained Group Quantization for Low-Bit LLM Inference
abstract
Low-bit quantization is essential for efficient LLM inference, and both rotation and fine-grained group quantization have shown individual promise. However, their combination often leads to accuracy degradation or hardware overhead due to a mismatch between the global nature of rotation and the localized behavior of group scaling. We propose GyRot, a quantization framework and hardware accelerator that bridges this gap through algorithm-hardware co-design. GyRot introduces Coarse Rotation, Fine Grouping (CoRFiG) and Harmonic-Aligned Permutation (HAP) to enable cooperative integration of rotation and group quantization, enhancing quantizability while relaxing scaling factor precision. To further reduce hardware cost, we reformulate asymmetric quantization and introduce a zero-point rounding strategy that enables fully integer dequantization. Implemented on an INT4-based tensor PE architecture, GyRot achieves state-of-the-art 4-bit accuracy across LLaMA-family models, while delivering up to 3.4× speedup and 3.6× energy efficiency over baseline LLM accelerators. These results validate GyRot's practical effectiveness for scalable and energy-efficient LLM deployment.
Yuseon Chou, Byeongcheol Kim, Jungjun Oh, Hoi-Jun Yoo
HPCA5
2026 SingularBit: Exploiting Synergy of Singular Value Decomposition and Low-Bit Quantization for Weight and KV Compression in LLM Inference
Seongyon Hong, Hyundeok Kong, Jungwan Lee, Hoi-Jun Yoo
ISCA5
2026 SeVeDo: A Heterogeneous Transformer Accelerator for Low-Bit Inference via Hierarchical Group Quantization and SVD-Guided Mixed Precision
abstract
Low-bit quantization is a promising technique for efficient transformer inference by reducing computational and memory overhead. However, aggressive bitwidth reduction remains challenging due to activation outliers, leading to accuracy degradation. Existing methods, such as outlier-handling and group quantization, achieve high accuracy but incur substantial energy consumption. To address this, we propose SeVeDo, an energy-efficient SVD-based heterogeneous accelerator that structurally separates outlier-sensitive components into a high-precision low-rank path, while the remaining computations are executed in a low-bit residual datapath with group quantization. To further enhance efficiency, Hierarchical Group Quantization (HGQ) combines coarse-grained floating-point scaling with fine-grained shifting, effectively reducing dequantization cost. Also, SVD-guided mixed precision (SVD-MP) statically allocates higher bitwidths to precision-sensitive components identified through low-rank decomposition, thereby minimizing floating-point operation cost. Experimental results show that SeVeDo achieves a peak energy efficiency of 13.8TOPS/W, surpassing conventional designs, with 12.7TOPS/W on ViT-Base and 13.4TOPS/W on Llama2-7B benchmarks.
Yuseon Choi, Jungjun Oh, Byeongcheol Kim, Hoi-Jun Yoo
ISCAS5
2026 A 198.7 μJ/token Block Diffusion LLM Processor with Mask Token Similarity-Based Activation Reuse
Yujin Moon, Seryeong Kim, Wonhoon Park, Wooyoung Jo, Yuseon Choi, Sunjoo Whang, Hoi-Jun Yoo
ISCAS8
2026 An Energy-Efficient High Resolution Vision Transformer Processor Exploiting Token Similarity Beyond Token Merging
Jungjun Oh, Junha Ryu, Byeongcheol Kim, Yuseon Choi, Hoi-Jun Yoo
IEEE Trans. Very Large Scale Integr. Syst.7
2025 BROCA: A Low-power and Low-latency Conversational Agent RISC-V System-on-Chip for Voice-interactive Mobile Devices
Wooyoung Jo, Seongyon Hong, Beomseok Kwon, Haoyang Sang, Dongseok Im, Sangyeob Kim, Chaeyun Jeong, Yujin Moon, Hoi-Jun Yoo
HCS11
2025 A 4.69mW LLM Processor with Binary/Ternary Weights for Billion-Parameter Llama Model
Sangyeob Kim, Jungwan Lee, Byeongju Kim, Hoi-Jun Yoo
HCS4
2025 EdgeDiff: Multi-modal Few-step Diffusion Model Accelerator with Mixed-Precision and Reordered Group-Quantization for On-device Generative AI Motivation
Jungjun Oh, Jeonggyu So, Yuseon Choi, Sangyeob Kim, Dongseo Kim, Gwangtae Park, Hoi-Jun Yoo
HCS8
2025 IRIS: A 8.55 mJ/frame Spatial Computing SoC for Real-time Interactable-Rendering and Surface-aware-Modeling with 3D Gaussian Splatting
Seokchan Song, Seryeong Kim, Wonhoon Park, Jongjun Park, Sanghyuk An, Gwangtae Park, Minseo Kim 0001, Hoi-Jun Yoo
HCS8
2025 A 51.2 fps Real-Time 3DGS-SLAM Accelerator using Diagonal Feeding with Symmetric Alpha Reuse and Voxel-based 3D Gaussian Cache Management
abstract
This work presents a high-speed 3D Gaussian Splatting-based SLAM (3DGS-SLAM) accelerator to support dense mapping for mobile devices. 3DGS-SLAM has two main hardware challenges for acceleration: 1) Large α-computation computation. 2) Memory bottleneck caused by irregular memory access and large number of Gaussians. First, diagonal feeding (DF) controller precludes redundant-α computation, and symmetric alpha reuse (SAR) enables reusing computed alpha. This method reduces 35.3% system computation. Second, voxel-based inter-frame caching (VIFC) enables selective inter-frame voxel caching, which reduces 44.0% of external memory access. As a result, the proposed 3DGS-SLAM accelerator achieves 51.2 fps with 0.07µJ/point with support voltage 0.9V, clock frequency 200MHz mapping a high-quality dense-map.
Hyungnam Joo, Seryeong Kim, Jongjun Park, Junha Ryu, Hoi-Jun Yoo
ISCAS5
2025 A 4.21 TFLOPS/W Memory-Efficient LLM Inference Accelerator with Bit-Layered Non-Uniform Quantization
abstract
Non-uniform Quantization (NUQ) is widely used in LLM accelerators due to its high accuracy. However, employing NUQ with models of varying sizes can substantially increase storage requirements on mobile devices. This paper presents a bit-layered NUQ accelerator architecture that supports multiple bit-width configurations while minimizing memory usage. Key features include Reconfigurable Condensed Look-up Accumulator (RCLA), Dual-Sign Path Accumulation (DSPA), and MSB-Sparse Encoding Compression (MSEC). RCLA enables the use of multiple weight precisions within a uniform PE array. In particular, it optimizes PE utilization in high-bit-width NUQ modes, reducing accumulation cycles by 63.2 %. DSPA facilitates energy-efficient computation, resulting in an average power reduction of 40.7 % across various weight modes. MSEC enhances weight compression, reducing the data storage size of each bit plane by up to 47.7 %. The proposed design supports models of different sizes, improves energy efficiency, and reduces memory capacity requirements, rendering it ideal for mobile LLM inference.
Byeongcheol Kim, Sangwoo Ha, Soyeon Um, Kyomin Sohn, Hoi-Jun Yoo
ISCAS6
2025 A 13.8 TOPS/W Polynomial Implicit Neural Representation Accelerator with Tile Similarity Exploitation and LUT-based Matrix Multiplication Reformation
abstract
This paper presents an energy-efficient polynomial implicit neural network (Poly-INR) processor for image generation tasks. Poly-INR can generate high-resolution images with small parameters, but it requires significant computation and has a long inference time, making it unsuitable for mobile device applications. The proposed processor achieves high energy efficiency through the following three key features: 1) Distribution-aware Heterogeneous Tile Processing (DHTP) reduces grid computation by 98.9% with high compression ratio and feature computation by 45.2% with reduced bit precision based on similarity-aware quantization. 2) Reformed Multiply-MAC Core (RMMC) improves core energy efficiency by 3.71× through the reuse of partial products for both grid and feature. 3) Precision-based Tile Reordering Unit (PTRU) reorders the tiles by their precision and similarity for different channels, enhancing throughput by 39.4% with precision-based sorting and an additional 14.0% with similarity consideration. The proposed processor is implemented in 28nm CMOS technology and achieves a peak energy efficiency of 13.8 TOPS/W.
Wonhoon Park, Sanghyuk An, Hoi-Jun Yoo, Donghyeon Han
ISCAS4
2025 A 32.65µm2 Spin/Area Large Scale Ising CIM with Progressive Circular Dataflow and Bi-directional eDRAM Cell Array
abstract
This paper presents a large-scale Ising computing-in-memory (CIM) for a real-world combinatorial optimization problem (COP). While the CIM approach shows promising performance improvement compared to digital-based Ising machines, it can be only used for simple COP due to limited CIM connectivity, bit precision, and graph size. To overcome limitations, the proposed processor supports reconfigurable high-bit Chimera graph topology achieving a small cell area and high energy efficiency through three key features : 1) Progressive circular dataflow reduce area by 58.8% and improved energy efficiency by 44.7% thanks to fully reuse spin and coefficient. 2) Bi-directional eDRAM cell array support bi-directional ising computation for circular dataflow with 3T-2C eDRAM cell. Due to the signed operation with compact coefficient storing, the area was reduced by 36.1%. 3) Reconfigurable spin exchange link and reconfigurable C-2C ladder support various graph sizes and various bit precision of coefficients. In conclusion, the proposed large-scale Ising CIM achieves 32.65µm2spin area and 2.09µW effective spin power which is 3.21× and 5.17× smaller than previous state-of-the art.
Jingu Lee, Sangwoo Ha, Sunjoo Whang, Soyeon Um, Wooyoung Jo, Hoi-Jun Yoo
ISCAS8
2025 A 65.1 TOPS/W Digital CIM Processor for Ultra-Low-Bit Transformers with Multiplexer-based Adder and Scaling Factor-based Reordering
abstract
As large-language-model (LLM) continues to expand in parameter size and improve performance, challenges related to latency and energy efficiency become increasingly significant. While low-bit weight quantization was proposed to alleviate the external memory access bottleneck, it remains insufficient to reduce overall energy consumption. In this paper, we use binarized LLM and design a novel Computing-in-Memory (CIM) processor to reduce both computation and memory energy. We propose three key techniques: 1) Multiplexer-based adder is proposed to replace first stage adders of adder tree with the simple multiplexers and inverter, reducing adder tree energy consumption by 58%, 2) Scaling factor-based reordering is proposed to reduce adder tree bit-width, resulting in an additional 28.6% energy reduction, 3) Cluster-level broadcasting and bit-spatial mapping is proposed to reduce internal memory access and increase bit-scalability. As a result, we can reduce the overall energy consumption by 69.7%, offering a more efficient solution for LLM inference tasks.
Nayeong Lee, Sangyeob Kim, Seongyon Hong, Hoi-Jun Yoo
ISCAS5
2025 A Real-time 4.31 mJ/Frame Neural-3DGS Processor with Voxel Similarity Memory Management and Opacity-based Sparsity Generation
abstract
This work presents an energy-efficient and real-time rendering Neural-3DGS processor for mobile AR/VR devices. While Neural-3DGS shows high-quality and fast rendering, it exhibits low energy efficiency & high latency in mobile implementation. The proposed processor has three key features for the overall processes in Neural-3DGS: 1) Voxel Similarity-aware Memory Management Unit (VSMMU) eliminates redundant operations and achieves 59.5%, 43.3% reduced energy for external memory access and neural-network computation. 2) LUT-based Pre-Sort Unit (LPSU) utilizes pre-computed order to reduce the latency of sorting by 64.3%. 3) Opacity-aware Gaussian Skipping Core (OGSC) exploits sparsity based on opacity and process 63.3% reduced MAC operations. The proposed processor is implemented in 28 nm CMOS technology. It achieves 98.8 FPS for real-time rendering and 4.31 mJ/Frame energy efficiency.
Hongseok Lee, Wonhoon Park, Sanghyuk An, Junha Ryu, Hoi-Jun Yoo
ISCAS5
2025 A 2.67 mJ/frame Video Mamba Accelerator with Importance-aware Redundancy Elimination and SSM Computing Reformulation
abstract
An energy-efficient video understanding processor, SLYTHERIN, is proposed to accelerate the new AI model, mamba, efficiently on edge devices. Mamba, a state-of-the-art model for in-context learning, is designed to replace the transformer, whose computational complexity increases significantly in video applications. However, the acceleration of video mamba on edge devices presents two main challenges: 1) slow inference due to the iterative operation phase of mamba and 2) the large energy consumption caused by external memory access (EMA). To address these challenges, the SLYTHERIN is proposed with 3 key building blocks. 1) A 6-stage pipelined task allocator minimizes computational complexity by dynamically managing redundant computations with importance-aware prediction. 2) Reformed computing SSM engine increases core efficiency by tackling the overheads in mamba’s iterative stages with reordering and distributed L1 cache. 3) Patch data management unit addresses the large EMA with difference-based bit-sliced data compression. It finally achieves 2.67-mJ/frame system energy efficiency with 274 FPS.
Youngjin Moon, Sangwoo Ha, Junha Ryu, Hoi-Jun Yoo, Donghyeon Han
ISCAS5
2025 A 9.6 TOPS/W Vision Transformer Processor with Hierarchical Token Merging for Similarity-Driven Difference Computing
abstract
Token merging is widely used in Vision Transformers(ViT) as an effective method to reduce computation with minimal accuracy drop. However, aggressive token elimination methods like token merging have limitations in applications requiring fine-grained output, such as text-unified object recognition. These methods reach an upper limit on token reduction due to an accuracy loss. This paper proposes exploiting token similarity beyond the merging upper bound to further reduce power consumption. The key features include Hierarchical Token Merging (HTM), Sparsity Separated Accumulation with Sign Magnitude Data Representation(SSA-SM), and Bidirectional Dynamic Allocation (BDA). HTM begins by performing token merging in the first iteration of the transformer block to reduce 34 % of tokens. Then, during the subsequent iterations, similar token difference computing is applied to the remaining tokens and reduces the effective bit by 31 %. SSA-SM then optimizes PE operations to fully exploit this sparsity by reducing the computation logic toggle rate, resulting in a 45 % higher energy efficiency. Lastly, BDA introduces a memory storage technique that ensures sparse and dense inputs are fed separately into the PE, resulting in an additional 14 % power savings. This proposed method improves energy efficiency by 2.39 times and achieves 9.6 TOPS/W for the text-unified object recognition application on the MS-COCO dataset.
Jungjun Oh, Byeongcheol Kim, Hoi-Jun Yoo
ISCAS5
2025 A 17.1 TOPS/W FP-INT Transformer Inference Accelerator with Sparsity Boosting and Output Importance-Aware Processing
abstract
This paper presents an energy-efficient FP-INT transformer inference accelerator for diverse applications. The proposed accelerator achieves high energy efficiency with two key features as solutions: 1) Sparsity Boosting Adder Tree (SBAT) to reduce adder tree power by 23.8% by modifying the adder tree structure and booth encoding to maximize sparsity. 2) Output Importance-aware Processing (OIAP) to reduce the Floating-Point Accumulation (FP-ACC) power by 79.6%, dynamically adjusting input block sizes based on the saliency of output channels, thereby reducing the FP-ACC operations by 78.7%. The proposed accelerator is implemented in 28 nm CMOS technology and achieves 17.1 TOPS/W, leveraging the unique FP-INT characteristics on booth encoding and exploiting redundancy based on output saliency in the transformer architecture.
Jeonggyu So, Seongyon Hong, Wooyoung Jo, Hoi-Jun Yoo, Donghyeon Han
ISCAS6
2025 A 62.8 TOPS/W FP-INT Digital Computing-in Memory Processor with Bit-Reordered Adder Tree and Low Active Hierarchical Accumulator
abstract
This paper introduces an energy-efficient DRAM-based FP-INT Digital Computing-in-Memory (CIM) processor for data-intensive AI workloads, addressing challenges in mixed precision, excess power consumption, and area overhead. The processor features three key innovations: 1) Bit-reordered adder tree (BRAT) that reduces adder bit width through exponent-aware reordering, achieving power and area reductions by 23% and 56%, respectively; 2) Low active hierarchical accumulator (LAHA) that conditionally updates the FP accumulator, minimizing accumulator activity by 57%; and 3) Speculative input block skipping (SIBS) to avoid unnecessary computations, reducing energy consumption by 15%. Implemented in 28nm CMOS technology, the processor achieves 62.8 TOPS/W in FP8 activation and INT4 weight precision for GPT2-Small inference. Our results show an overall energy savings of 31%, making this architecture highly efficient for modern large language models.
Sunjoo Whang, Sangwoo Ha, Soyeon Um, Hoi-Jun Yoo
ISCAS5
2024 A Low-Power Large-Language-Model Processor with Big-Little Network and Implicit-Weight-Generation for On-Device AI
Sangyeob Kim, Wooyoung Jo, Seongyon Hong, Nayeong Lee, Hoi-Jun Yoo
HCS7
2024 NeuGPU: A Neural Graphics Processing Unit for Instant Modeling and Real-Time Rendering on Mobile AR/VR Devices
abstract
•Why 3D Modeling using Neural Radiance Field?
Junha Ryu, Hankyul Kwon, Wonhoon Park, Zhiyong Li 0016, Beomseok Kwon, Donghyeon Han, Dongseok Im, Sangyeob Kim, Hyungnam Joo, Hoi-Jun Yoo
HCS11
2024 Space-Mate: A 303.5mW Real-Time NeRF SLAM Processor with Sparse-Mixture-of-Experts-based Acceleration
abstract
NeRF-based SLAM for robotic applications face computation barrier
Seokchan Song, Haoyang Sang, Dongseok Im, Donghyeon Han, Sangyeob Kim, Hongseok Lee, Hoi-Jun Yoo
HCS7
2024 LUTein: Dense-Sparse Bit-Slice Architecture With Radix-4 LUT-Based Slice-Tensor Processing Units
abstract
Bit-slice architectures have been developed to support various bit-precision and data sparsity of deep neural networks (DNNs). However, because of low-bit precision and a wide range of sparsity of bit-slice computations, bit-slice architectures are challenging in multiplier-, processing element (PE)-, core-, and software (SW)-level designs. First, data sparsity causes a power trade-off between the Radix numbers of a multiplier, and the previous multipliers cannot take advantage of all sparsity ranges. Second, a bit-slice PE which integrated massive lowbit multiplier-and-accumulate (MAC) units brings about large data transactions compared to a fixed bit-width PE. Third, bitslice core architectures only focus on either dense or sparse data computations, limiting the overall performance of bit-slice computations with a wide sparsity range. Lastly, low-bit bit-slice computations cause massive repetitive instruction fetches across hardware units. To solve the challenges, LUTein is proposed. It exploits the new lookup table (LUT)-based computing method to support the Radix-4 Modified Booth algorithm, achieving low power consumption in all sparsity ranges. Moreover, the slice-tensor PE efficiently processes slice-tensor data by sharing hardware units across the Radix-4 LUT-based MAC units. In addition, the LUTein architecture adopts a systolic datapath with a multi-port buffer to exploit both inter-PE data reuse and slice-level sparsity. Lastly, LUTein's instruction set architecture (ISA) and the hierarchical instruction decoder are introduced to alleviate repetitive instruction fetches. As a result, LUTein outperforms the state-of-the-art bit-slice architecture, Sibia, over 1.34× higher energy-efficiency and 1.78× higher area-efficiency.
Dongseok Im, Hoi-Jun Yoo
HPCA2
2024 A 28.6 mJ/iter Stable Diffusion Processor for Text-to-Image Generation with Patch Similarity-based Sparsity Augmentation and Text-based Mixed-Precision
abstract
This paper presents an energy-efficient stable diffusion processor for text-to-image generation. While stable diffusion attained attention for high-quality image synthesis results, its inherent characteristics hinder its deployment on mobile platforms. The proposed processor achieves high throughput and energy efficiency with three key features as solutions: 1) Patch similarity-based sparsity augmentation (PSSA) to reduce external memory access (EMA) energy of self-attention score by 60.3 %, leading to 37.8 % total EMA energy reduction. 2) Text-based important pixel spotting (TIPS) to allow 44.8 % of the FFN layer workload to be processed with low-precision activation. 3) Dual-mode bit-slice core (DBSC) architecture to enhance energy efficiency in FFN layers by 43.0 %. The proposed processor is implemented in 28 nm CMOS technology and achieves 3.84 TOPS peak throughput with 225.6 mW average power consumption. In sum, 28.6 mJ/iteration highly energy-efficient text-to-image generation processor can be achieved at MS-COCO dataset.
Wooyoung Jo, Seongyon Hong, Beomseok Kwon, Wonhoon Park, Hoi-Jun Yoo
ISCAS6
2024 Two-Step Spike Encoding Scheme and Architecture for Highly Sparse Spiking-Neural-Network
abstract
This paper proposes a two-step spike encoding, which consists of the source encoding and process encoding for energy-efficient spiking-neural-network (SNN) acceleration. The eigen-train generation and its superposition generate spike trains which show high accuracy with low spike ratio. Sparsity boosting (SB) and spike generation skipping (SGS) reduce the number of operations for SNN. Time shrinking multi-level encoding (TS-MLE) compresses the number of spikes in a train along time axis, and spike-level clock skipping (SLCS) decreases the processing time. Eigen-train generation achieves 90.3% accuracy, the same accuracy as CNN, under the condition of 4.18% spike ratio for CIFAR-10 classification. SB reduces spike ratio by 0.49× with only 0.1% accuracy loss, and the SGS reduces the spike ratio by 20.9% with 0.5% accuracy loss. TS- MLE and SLCS increase the throughput of SNN by 2.8× while decreasing the hardware resource for spike generator by 75% compared with previous generators.
Sangyeob Kim, Soyeon Um, Hoi-Jun Yoo
ISCAS5
2024 A 3.55 mJ/frame Energy-efficient Mixed-Transformer based Semantic Segmentation Accelerator for Mobile Devices
abstract
An energy-efficient semantic segmentation (SS) processor, achieving 3.55 mJ/frame system energy efficiency, is proposed. To address the challenges posed by Mixed Transformer (MiT)-based SS, including high external memory bandwidth requirement and large on-chip memory footprint, we introduce a novel compression method called Chunk-based Bit Plane Compression (CBPC). CBPC leverages the high inter-token locality of feature maps in MiT-based SS, along with the robustness and compression ratio variations based on bit position to achieve a high compression ratio. To support CBPC, we propose an area and power-efficient CBPC encoder/decoder. In addition, a Similar Token Coarse Skipping (STCS) Core is proposed for high throughput. It enables row-wise clock gating and array-wise coarse skipping to reduce redundant computation. By removing redundant computation, the processor achieves higher throughput and lower computation power. The proposed processor reduces 67.6% of EMA power and accomplishes 19.24 TOPS/W core energy efficiency. The proposed processor achieves 44.3% higher system energy efficiency than the previous processors.
Jongjun Park, Seryeong Kim, Wonhoon Park, Seokchan Song, Hoi-Jun Yoo
ISCAS5
2024 CamPU: A Multi-Camera Processing Unit for Deep Learning-based 3D Spatial Computing Systems
abstract
A 3D spatial computing system that understands a surrounding environment and interacts with real-world objects has emerged with the development of deep learning technologies. A multi-camera system captures a surrounding view of a scene using multiple cameras, and a deep neural network (DNN) system extracts semantic features from multi-camera images and provides useful information to users. However, processing a multi-camera system requires massive memory accesses as the number of cameras increases while processing a DNN system can improve throughput by exploiting batch processing. This performance gap limits the overall performance of 3D spatial computing systems. To solve this problem, a multi-camera processing unit (CamPU) is proposed. CamPU exploits the inter- and intra-data reuse methods on multi-camera images, minimizing memory accesses for image projection. Moreover, the out-of-order image projection unit with cache memory is designed to increase multi-image projection throughput by avoiding redundant cache accesses and hiding the latency of high-level memory accesses. Lastly, the overlap-aware blending unit speeds up image blending by efficiently handling overlapping regions between adjacent images. The CamPU architecture is evaluated through RTL-level simulation, and the CamPU-integrated DNN platform provides a comprehensive analysis of end-to-end multi-camera deep learning-based 3D spatial systems. Finally, CamPU speedups the overall system performance 2.9 x faster than an NVIDIA RTX2080Ti GPU platform.
Dongseok Im, Hoi-Jun Yoo
MICRO2
2024 An Energy-Efficient CNN/Transformer Hybrid Neural Semantic Segmentation Processor With Chunk-Based Bit Plane Data Compression and Similarity-Based Token-Level Skipping Exploitation
abstract
A novel energy-efficient semantic segmentation (SS) processor is proposed for achieving high system energy efficiency on mobile devices. 1) Excessive external memory access and 2) a large amount of redundant computation hinders energy-efficient SS acceleration. Three key features enable real-time energy-efficient CNN/ViT hybrid SS. A new compression method named Chunk-based Bit Plane Compression (CBPC) reduces the memory footprint and energy consumption due to external memory access. CBPC enhances compression ratio by leveraging the high inter-token similarity of feature maps and applying bit plane compression in sign-magnitude data representation, using chunk-wise low-bit plane shared bias. The proposed CBPC encoder/decoder supports CBPC with minimum area overhead. Additionally, the Similar Token Coarse Skipping (STCS) Core enhances the throughput and reduces the computation power by eliminating redundant computations. STCS core employs Row-wise Line Gating for low-power computation and Array-wise Coarse Skipping to minimize redundant computation. As a result, our proposed processor reduces external memory access energy by 67.6% and achieves a core energy efficiency of 19.24 TOPS/W. Our solution achieves 3.55mJ/frame system-level energy efficiency which is 79.7% higher than the previous SOTA SS processor.
Jongjun Park, Seryeong Kim, Wonhoon Park, Seokchan Song, Hoi-Jun Yoo
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 Sibia: Signed Bit-slice Architecture for Dense DNN Acceleration with Slice-level Sparsity Exploitation
abstract
Deep neural networks (DNNs) have achieved high performance in many AI fields such as 1-D language, 2-D image, and 3-D point cloud processing applications. Since recent DNN tasks require dense matrix operations with various bit-precision and non-ReLU activation functions, mobile neural processing units (NPUs) suffer from the acceleration of diverse DNN tasks within their limited hardware resources and power budget. Although bit-slice architectures benefit from slice-level computation and slice-level sparsity exploitation, the conventional bit-slice representation is inefficient in bit-slice architectures resulting in poor dense DNN execution. This paper proposes the efficient signed bit-slice architecture, Sibia, with the signed bit-slice representation (SBR) for efficient dense DNN acceleration. The SBR adds a sign bit to each bit-slice and changes signed 11112bit-slice to 00002by borrowing a value of 1 from its lower order of the bit-slice. This scheme generates large numbers of zero bit-slices in dense DNNs even not relying on accuracy-sensitive pruning methods or retraining processes. Moreover, the SBR balances positive and negative values of 2’s complement data, allowing accurate bit-slice-based output speculation that pre-computes high orders of bit-slices. Sibia integrates the signed multiplier-and-accumulate (MAC) units for efficient signed bit-slice computations, and the flexible zero skipping processing element (PE) supports the zero input bit-slice skipping and output skipping for high throughput and energy-efficiency. Additionally, the dynamic sparsity monitoring unit monitors sparsity ratio between input and weight data and determines the more sparse one for zero bit-slice skipping. The heterogeneous network-on-chip (NoC) benefits from data reusability during bit-slice computation, reducing transmission bandwidth. Finally, Sibia outperforms the previous bit-slice architecture, Bit-fusion, over 3.65× higher area-efficiency, 3.88× higher energy-efficiency, and 5.35× higher throughput.
Dongseok Im, Gwangtae Park, Zhiyong Li 0016, Junha Ryu, Hoi-Jun Yoo
HPCA5
2023 A 332 TOPS/W Input/Weight-Parallel Computing-in-Memory Processor with Voltage-Capacitance-Ratio Cell and Time-Based ADC
abstract
Recent computing-in-memory (CIM) achieves high energy efficiency with charge-domain computation and multi-bit input driving. However, the previous works still require high power consumption and trade computation signal-to-noise ratio (SNR) for energy efficiency. This work proposes an energy-efficient and accurate multi-bit input/weight-parallel CIM processor with four key features: 1) a 10T2C sign-magnitude cell with voltage-capacitance-ratio (VCR) decoding for 5-bit analog inputs with only 2-level supply voltages, 2) a computation word line (CWL) charge reuse method for input driver power reduction, 3) a signal-amplifying noise canceling voltage-to-time converter (SANC-VTC) for SNR improvement, and 4) a distribution-aware time-to-digital converter (DA-TDC) for ADC power reduction. The proposed CIM processor is simulated in 28 nm CMOS technology with 1.25 mm2area. As a result, it achieves 4.44 mW power consumption and 332 TOPS/W energy efficiency with 72.43% benchmark accuracy (@ ImageNet, ResNet50, 5-bit input/5-bit weight).
Seongyon Hong, Soyeon Um, Sangyeob Kim, Wooyoung Jo, Hoi-Jun Yoo
ISCAS6
2023 A Reconfigurable 1T1C eDRAM-based Spiking Neural Network Computing-In-Memory Processor for High System-Level Efficiency
abstract
Spiking Neural Network (SNN) Computing-In-Memory (CIM) was proposed for high macro-level energy efficiency. However, system-level energy efficiency is limited by EMA due to a large intermediate activation footprint requirement. To reduce the EMA, a large capacity SNN CIM is needed to load tons of weights in the CIM. This paper proposes a high-density 1T1C eDRAM-based SNN CIM processor for supporting high system-level energy efficiency with two key features: 1) High-density and low-power Reconfigurable Neuro-Cell Array (ReNCA) for memory and SNN peripheral logic using a charge pump and reusing 1T1C cell array, achieving 41% area and 90% power reduction compared to previous work. 2) Reconfigurable CIM architecture with dual-mode ReNCA and Dynamic Adjustable Neuron Link (DAN Link) for layer fusion increases system-level efficiency including intermediate and weight EMA. It achieves$10\times$higher state-of-the-art system-level energy efficiency including EMA.
Seryeong Kim, Soyeon Um, Zhiyong Li 0016, Sangyeob Kim, Wooyoung Jo, Hoi-Jun Yoo
ISCAS8
2023 A 15.9 mW 96.5 fps Memory-Efficient 3D Reconstruction Processor with Dilation-based TSDF Fusion and Block-Projection Cache System
abstract
A real-time dense 3D reconstruction on lightweight AR headsets is challenging since its memory access surpasses the available memory bandwidth. To solve this problem, the proposed processor integrates two key building blocks - Dilation-based TSDF (D-TSDF) fusion and Block-Projection (BP) engine. D-TSDF projects the depth map in the reverse order of voxel-to-pixel coordinate transformation and dilates it, leading to 96.61% External Memory Access (EMA) reduction with minimum map quality degradation. Second, a specialized BP engine compresses high-resolution occupancy grid by decomposing the 3D bitmap into 2D and 1D vectors, achieving$\times \mathbf{166.09}$reduced memory bandwidth. The proposed processor is implemented in 28nm CMOS technology occupying 1.27 mm2area. As a result, 96.45 fps 3D reconstruction is possible while consuming only 15.94 mW power.
Hankyul Kwon, Gwangtae Park, Junha Ryu, Wooyoung Jo, Hoi-Jun Yoo
ISCAS5
2023 A 5.99 TFLOPS/W Heterogeneous CIM-NPU Architecture for an Energy Efficient Floating-Point DNN Acceleration
abstract
This work presents an energy-efficient digital-based computing-in-memory (CIM) processor to support floating-point (FP) deep neural network (DNN) acceleration. Previous FP-CIM processors have two limitations. Processors with post-alignment shows low throughput due to serial operation, and the other processor with pre-alignment incurs truncation error. To resolve these problems, we focus on the statistics that outlier exists according to shift amount in pre-alignment-based FP operation. As those outlier decreases energy efficiency due to long operation cycles, it needs to be processed separately. The proposed Hetero-FP-CIM integrates both CIM arrays and shared NPU, so they compute both dense inlier and sparse outlier respectively. It also includes efficient weight caching system to avoid entire weight copy in shared NPU. The proposed Hetero-FP-CIM is simulated in 28 nm CMOS technology and occupies 2.7 mm2. As a result, it achieves 5.99 TOPS/W at ImageNet (ResNet50) with bfloat16 representation.
Wonhoon Park, Junha Ryu, Soyeon Um, Wooyoung Jo, Sangyoeb Kim, Hoi-Jun Yoo
ISCAS7
2022 HNPU-V2: A 46.6 FPS DNN Training Processor for Real-World Environmental Adaptation based Robust Object Detection on Mobile Devices
abstract
■ Smarter DNNs: # of Parameter ▲
Donghyeon Han, Dongseok Im, Gwangtae Park, Seokchan Song, Juhyoung Lee, Hoi-Jun Yoo
HCS7
2022 DSPU: A 281.6mW Real-Time Deep Learning-Based Dense RGB-D Data Acquisition with Sensor Fusion and 3D Perception System-on-Chip
abstract
3D Data in Mobile Platforms
Dongseok Im, Gwangtae Park, Zhiyong Li 0016, Junha Ryu, Donghyeon Han, Jinsu Lee, Wonhoon Park, Hankyul Kwon, Hoi-Jun Yoo
HCS10
2022 Neuro-CIM: A 310.4 TOPS/W Neuromorphic Computing-in-Memory Processor with Low WL/BL activity and Digital-Analog Mixed-mode Neuron Firing
abstract
Multi WLs Driving ➔ Low Energy Efficiency by ADC (<100 TOPS/W)
Sangyeob Kim, Soyeon Um, Kwantae Kim, Hoi-Jun Yoo
HCS6
2022 An Efficient High-quality FHD Super-resolution Mobile Accelerator SoC with Hybrid-precision and Energy-efficient Cache
abstract
Super-Resolution on Mobile Platform
Zhiyong Li 0016, Dongseok Im, Donghyeon Han, Hoi-Jun Yoo
HCS5
2022 A 161.6 TOPS/W Mixed-mode Computing-in-Memory Processor for Energy-Efficient Mixed-Precision Deep Neural Networks
abstract
A Mixed-mode Computing-in memory (CIM) processor for the mixed-precision Deep Neural Network (DNN) processing is proposed. Due to the bit-serial processing for the multi-bit data, the previous CIM processors could not exploit the energy-efficient computation of mixed-precision DNNs. This paper proposes an energy-efficient mixed-mode CIM processor with two key features: 1) Mixed-Mode Mixed-precision CIM (M3-CIM) which achieves 55.46% energy efficiency improvement. 2) Digital-CIM for In-memory MAC for the increased throughput of M3-CIM. The proposed CIM processor was simulated in 28nm CMOS technology and occupies 1.96 mm2. It achieves a state-of-the-art energy efficiency of 161.6 TOPS/W with 72.8% accuracy at ImageNet (ResNet50).
Wooyoung Jo, Juhyeong Lee, Soyeon Um, Zhiyong Li 0016, Hoi-Jun Yoo
ISCAS6
2022 A Low-Power Graph Convolutional Network Processor With Sparse Grouping for 3D Point Cloud Semantic Segmentation in Mobile Devices
abstract
A low-power graph convolutional network (GCN) processor is proposed for accelerating 3D point cloud semantic segmentation (PCSS) in real-time on mobile devices. Three key features enable the low-power GCN-based 3D PCSS. First, the new hardware-friendly GCN algorithm, sparse grouping-based dilated graph convolution (SG-DGC) is proposed. SG-DGC reduces 71.7% of the overall computation and 76.9% of EMA through the sparse grouping of the point cloud. Second, the two-level pipeline (TLP) consisting of the point-level pipeline (PLP) and group-level pipelining (GLP) was proposed to improve low utilization by the imbalanced workload of GCN. The PLP enables point-level module-wise fusion (PMF) which reduces 47.4% of EMA for low power consumption. Also, center point feature reuse (CPFR) reuses computation results of the redundant operation and reduces 11.4% of computation. Finally, the GLP increased the core utilization by 21.1% by balancing the workload of graph generation and graph convolution and enable$1.1\times $higher throughput. The processor is implemented with 65nm CMOS technology, and the 4.0mm23D PCSS processor show 95mW power consumption while operating in real-time of 30.8 fps in the 3D PCSS of the indoor scene with 4k points.
Sangyeob Kim, Juhyoung Lee, Hoi-Jun Yoo
IEEE Trans. Circuits Syst. I Regul. Pap.4
2022 TSUNAMI: Triple Sparsity-Aware Ultra Energy-Efficient Neural Network Training Accelerator With Multi-Modal Iterative Pruning
abstract
This article proposes the TSUNAMI, which supports an energy-efficient deep-neural-network training. The TSUNAMI supports multi-modal iterative pruning to generate zeros in activation and weight. Tile-based dynamic activation pruning unit and weight memory shared pruning unit eliminate additional memory access. Coarse-zero skipping controller skips multiple unnecessary multiply-and-accumulation (MAC) operations at once, and fine-zero skipping controller skips randomly located unnecessary MAC operations. Weight sparsity balancer solves a utilization degradation caused by weight sparsity imbalance, and the workload of each convolution core is allocated by a random channel allocator. The TSUNAMI achieves an energy efficiency of 3.42 TFLOPS/W at 0.78V and 50MHz with floating-point 8-bit activation and weight. Also, it achieves an energy efficiency of 405.96 TFLOPS/W at 90% sparsity condition.
Sangyeob Kim, Juhyoung Lee, Donghyeon Han, Wooyoung Jo, Hoi-Jun Yoo
IEEE Trans. Circuits Syst. I Regul. Pap.6
2021 PNNPU: A Fast and Efficient 3D Point Cloud-based Neural Network Processor with Block-based Point Processing for Regular DRAM Access
abstract
PNN* for Intelligent 3D Vision • Intelligent 3D Vision on Mobile Devices • Accurate & Robust Perception with 3D Structural Information • Mobile 3D Sensor Already Commercialized
Juhyoung Lee, Dongseok Im, Hoi-Jun Yoo
HCS4
2021 An Energy-efficient Floating-Point DNN Processor using Heterogeneous Computing Architecture with Exponent-Computing-in-Memory
abstract
Abstract of Proposed FP CIM Processor (1) Heterogeneous FP Computing Arch. : Separate optimization of FP computing: Realize 2 cycles FP MAC w/ CIM (2) Exponent Computing-in-Memory: In-memory AND/NOR + BL charge reusing: Total memory power 46.4% 2) Mantissa Free Exponent Calculation: Removing redundant normalization: Total MAC power 14.4%
Juhyoung Lee, Ji-Hoon Kim 0004, Wooyoung Jo, Sangyeob Kim, Donghyeon Han, Jinsu Lee, Hoi-Jun Yoo
HCS8
2021 OmniDRL: An Energy-Efficient Mobile Deep Reinforcement Learning Accelerators with Dual-mode Weight Compression and Direct Processing of Compressed Data
abstract
Deep Reinforcement Learning (DRL)▪ No Pre-labelled Data ➔ Training with Trial-and-errors!– Sequential decision making problems @ Unknown environments– Applications: gaming agent, autonomous systems, agent adaptation
Juhyoung Lee, Sangyeob Kim, Ji-Hoon Kim 0004, Wooyoung Jo, Donghyeon Han, Hoi-Jun Yoo
HCS7
2021 A 3.6 TOPS/W Hybrid FP-FXP Deep Learning Processor with Outlier Compensation for Image-to-Image Application
abstract
A Hybrid floating-point (FP) and fixed-point (FXP) deep learning processor with an outlier-aware channel splitting algorithm is proposed for image-to-image applications on mobile devices. Since the high quality of the reconstructed image through deep learning based image-to-image application requires high bit-precision (> FP16), the mobile processor suffers from the high computation power and large external memory access (EMA). In this work, the proposed algorithm reduces 16-bit FP data to 8-bit FXP data, and only few outliers (2. The hierarchical processor successfully demonstrates the x 4 scale Full-HD super-resolution generation achieving 76 frames-per-second (fps) with 133.3 mW power-consumption at 0.9 V supply and 3.6 TOPS/W of energy-efficiency which is × 3.27 higher than the previous 16-bit FXP processor.
Zhiyong Li 0016, Dongseok Im, Jinsu Lee, Hoi-Jun Yoo
ISCAS4
2020 Deep Learning Processors for On-Device Intelligence
abstract
Recently, deep learning is influencing not only the technology itself but also our everyday lives. Formerly, most AI functionalities and applications were centralized on datacenters. However, the primary platform for AI has recently shifted to on-devices. With the increasing demand on edge, mobile and IoT AI, conventional hardware solutions face their ordeal because of their low energy efficiency on such power hungry applications. For the past few years, dedicated DNN inference accelerators have been under the spotlight. However, with the rising emphasis on privacy, personalization and local optimization, ability to learn is becoming the second hurdle for "on-device AI." In addition, with the recent developments in hardware research, faster DNN processing speed with low power consumption is achieved, enabling numerous applications on edge and mobile devices, which were formerly not applicable to edge and mobile devices. Applications with humanistic intelligence, which can take users' emotion into account, have been demonstrated, along with GAN and DRL as well as AI models using 3-dimensional data processing for higher accuracy.
Hoi-Jun Yoo
ACM Great Lakes Symposium on VLSI1
2020 A 54.7 fps 3D Point Cloud Semantic Segmentation Processor with Sparse Grouping Based Dilated Graph Convolutional Network for Mobile Devices
abstract
The graph convolutional network (GCN) based 3D point cloud semantic segmentation (PCSS) processor for mobile devices is proposed. GCN based 3D PCSS requires a lot of computation, making it unsuitable for real-time operation in mobile devices. For real-time 3D PCSS on mobile devices, this paper proposes two key features: 1) a sparse grouping based dilated graph convolution (SG-DGC) which reduces 71.7% of the overall computation of GCN by simply dividing input point cloud into multiple sparse point cloud. 2) group-level pipelining which improves low pipeline utilization due to the computation imbalance of GCN. Finally, the proposed GCN processor is simulated in 65 nm CMOS technology and occupies 4.0 mm2. The proposed processor consumes 176mW and shows 54.7 frames-per-second (fps) for the 3D point cloud semantic segmentation of indoor scene with 4k points.
Sangyeob Kim, Juhyoung Lee, Hoi-Jun Yoo
ISCAS4
2020 The Heterogeneous Deep Neural Network Processor With a Non-von Neumann Architecture
abstract
Today's CPUs are general-purpose processors, which have the von Neumann architecture (including the Harvard architectures) to maximize the generality and programmability. On the other hand, application-specific integrated circuits (ASICs) have domain-specific architectures to optimize the cost-effective performance but show very low generality. The combination of generality and ASIC, which usually seemed to have no contact, is expected to be enabled by deep learning (DL). DL, realized with deep neural networks (DNNs), has changed the paradigm of machine learning (ML) and brought significant progress in vision, speech, language processing, and many other applications. DNNs have special features that can be efficiently implemented with dedicated architectures, ASICs. Sharing their special features, DNNs have a wide variety of network architectures, and even the same network architecture can be used for different applications depending on the weight parameters. This paper aims to provide the necessity, validity, and characteristics of the ML-specific integrated circuits (MSICs) that have a different architecture from the von Neumann architecture. MSICs can avoid the overhead from the complex instruction set, instruction decoder, multilevel caches, and branch prediction of the recent von Neumann architecture processors designed for high generality and programmability. We will also discuss the necessity and validity of a heterogeneous architecture in MSIC, starting from the differences between the visual-type information processing and the vector-type information processing, and show the chip implementation results.
Dongjoo Shin, Hoi-Jun Yoo
Proc. IEEE2
2019 93.8% Current Efficiency and 0.672 ns Transient Response Reconfigurable LDO for Wireless Sensor Network Systems
abstract
Current-efficient, fast-transient reconfigurable low-dropout regulator (LDO) is proposed for the wireless sensor network (WSN) system. The proposed LDO is designed and simulated in a 65 nm CMOS process showing the 3 key features: 1) a reconfigurable LDO architecture to achieve both low quiescent current (IQ) and wide bandwidth by adaptively adjusting to different load current conditions, 2) ultra-low-IQ and high PSR regulator for light-load efficient operation utilizing the gain boosting scheme within the flipped voltage follower loop, and 3) fast transient regulator for robust heavy-load operation with level-shifted impedance attenuation buffer (IAB) to reduce the supply ripples generated by wireless transceiver. The proposed LDO shows the state-of-the-art 93.8% current efficiency in the load condition of 1 μA, and achieves 0.672 ns response time even with 10 ns load transition.
Surin Gweon, Jaehyuk Lee, Kwantae Kim, Hoi-Jun Yoo
ISCAS4
2019 DT-CNN: Dilated and Transposed Convolution Neural Network Accelerator for Real-Time Image Segmentation on Mobile Devices
abstract
A convolution neural network (CNN) accelerator is proposed for real-time image segmentation on mobile devices. The proposed CNN processor cuts down the redundant zero computations in dilated and transposed convolution for higher throughput. As a result, the overall computations of the image segmentation are reduced by 86.6% and the proposed CNN processor boosts up the throughput 6.7×. Moreover, the proposed processor utilizes RoI (Region of Interest) based image segmentation algorithm to reduce the overall computational requirement significantly. Although RoI based image segmentation degrades the image segmentation accuracy, the proposed dilation rate adjustment compensates for the accuracy degradation and maintains the accuracy of the full-size image segmentation. Finally, the proposed CNN processor is simulated in 65 nm CMOS technology, and it occupies 6.8 mm2. The proposed processor consumes 196 mW and shows 211 frames-per-second (fps) at the image segmentation for human body parts.
Dongseok Im, Donghyeon Han, Sungpill Choi, Hoi-Jun Yoo
ISCAS5
2019 An Ultra-Low-Power Analog-Digital Hybrid CNN Face Recognition Processor Integrated with a CIS for Always-on Mobile Devices
abstract
An ultra-low-power analog-digital hybrid always-on face recognition (FR) processor integrated with a CMOS image sensor (CIS) is proposed for the wearable mobile devices applications such as user authentication. The proposed processor is the first IC with full process of FR in a single chip. The processor adopts analog-digital hybrid convolution operation for efficient integration of CNN processor with CIS. The analog convolution processor is proposed for the computation of the 1stlayer of CNN and the quantization operation without an ADC that can achieve 15.7% power reduction with 1.3% minimal accuracy loss. In addition, the analog weighted-sum unit with low power (5.18TOPS/W) is proposed with switched-drain regulation (SDR) current mirror which can achieve less than 6% mirroring error. The processor is simulated in 65-nm CMOS technology, 15.84mm2area with 2.5V and 1.2V for analog domain and 0.77-1.1V for digital domain. It consumes 0.6198mW to evaluate one face at 1 fps and achieves 96.18% FR accuracy in LFW dataset.
Ji-Hoon Kim 0004, Kwantae Kim, Hoi-Jun Yoo
ISCAS4
2019 A 15.2 TOPS/W CNN Accelerator with Similar Feature Skipping for Face Recognition in Mobile Devices
abstract
A low-power face recognition processor with similar feature skipping (SFS) and the tile-based clustering algorithm is proposed for high energy efficiency in mobile devices. For higher energy efficiency face recognition (FR) processor, this paper proposes two key features: 1) Tile-based clustering enables to reduce computation overhead of clustering. 2) SFS binary convolution core is proposed to increase energy efficiency, resulting in 15.2 TOPS/W energy efficiency. Implemented with 65 nm CMOS technology, the 6 mm2FR processor achieves 0.26mW power consumption at 1 frames-per-second (fps) always-on face recognition in mobile devices.
Sangyeob Kim, Juhyoung Lee, Jinsu Lee, Hoi-Jun Yoo
ISCAS5
2018 A 141.4 mW Low-Power Online Deep Neural Network Training Processor for Real-time Object Tracking in Mobile Devices
abstract
A low-power online deep neural network (DNN) training processor is proposed for a real-time object tracking in mobile devices. For a real-time object tracking, a homogeneous core architecture is proposed to achieve 1.33× higher throughput than previous DNN training processor. To reduce the external memory access (EMA), a binary feedback alignment (BFA) algorithm and an integral run-length compression (iRLC) decoder are proposed. While the BFA reduces the EMA by 11.4% compared to the conventional back-propagation approach, the iRLC decoder achieves 29.7% EMA reduction without throughput degradation. Finally, a dropout controller is proposed and achieves 43.9% power reduction through clock-gating. Implemented with 65 nm CMOS technology, the 4.4 mm2DNN training processor achieves 141.1 mW power consumption at 30.4 frames-per-second (fps) real-time object tracking in mobile devices.
Donghyeon Han, Jinsu Lee, Jinmook Lee, Sungpill Choi, Hoi-Jun Yoo
ISCAS5
2018 A 46.1 fps Global Matching Optical Flow Estimation Processor for Action Recognition in Mobile Devices
abstract
A real-time global matching optical flow estimation (OFE) processor is proposed for action recognition in mobile devices. The global OFE requires a large number of external memory accesses (EMAs) and matrix computations, thus it is incompatible on mobile devices with real-time constraints. For real-time OFE on mobile devices, this paper proposes two key features, both of which to reduce the required memory bandwidth and a number of computations: 1) Tile-based hierarchical OFE enables intermediate data to be processed within 328 KB on-chip memory without external memory access. 2) Background skipping eliminates redundant matrix computation for zero optical flow region. Therefore, the proposed features reduce external memory bandwidth and computation by 99.7 % and 50.7 %, respectively. The proposed 4 mm2OFE processor is implemented in 65 nm CMOS technology and it achieves real-time OFE of 46.1 frames-per-second (fps) throughput for an image resolution of QVGA (320 × 240) and the resulting optical flow can be successfully used for action recognition.
Juhyoung Lee, Sungpill Choi, Dongjoo Shin, Hoi-Jun Yoo
ISCAS6
2018 A 0.78 mW Low-Power 4.02 High-Compression Ratio Less than 10-6 BER Error-Tolerant Lossless Image Compression Hardware for Wireless Capsule Endoscopy System
abstract
A 0.78mW low power error-tolerant lossless image compressor for a wireless capsule endoscopy system is proposed. In order to achieve high compression ratio of 4.02, this work proposes zero-skipping coding and mode switcher architecture that utilize high image sparsity on endoscopy systems. In addition, a Forward Error Correction (FEC) block is proposed to achieve low Bit-error Rate (BER) under 10−6, the integration error from prediction-based and variable length coding. The proposed error-tolerant lossless image compression hardware is implemented using 1P6M 65nm CMOS process and consumes only 0.78mW power at 40MHz with 2 fps throughput on VGA resolution image, and it is evaluated with real capsule endoscope images.
Kyoung-Rog Lee, Hoi-Jun Yoo
ISCAS3
2017 A Real-Time and Energy-Efficient Embedded System for Intelligent ADAS with RNN-Based Deep Risk Prediction using Stereo Camera
Kyuho Jason Lee, Gyeongmin Choe, Kyeongryeol Bong, In-So Kweon, Hoi-Jun Yoo
ICVS6
2017 A 0.53mW ultra-low-power 3D face frontalization processor for face recognition with human-level accuracy in wearable devices
abstract
An ultra-low-power face frontalization processor (FFP) is proposed for accurate face recognition in wearable devices. 3D face frontalization is essential in face recognition to guarantee human-level accuracy even with rotated or tilted faces. To reduce external memory access (EMA), which causes large power consumption, regression weight quantization with K-means clustering is proposed with the result of 81.25% EMA reduction. In addition, pipelined memory-level zero-skipping regression reduces the EMA by additional 98.43% without latency overhead. Moreover, for low-power consumption of accelerating heterogeneous workload, energy-efficient shared PE array architecture is proposed. While accelerating computation intensive process by allocating large number of PEs for utilizing data-level parallelism, unused PEs are clock-gated for preventing needless power consumption during computationally light process. Proposed workload adaptation with clock-gating showed 37.14% power reduction. The proposed FFP was implemented in 65nm CMOS process, and showed 0.53mW power consumption with 4.73fps throughput, both of which satisfy condition for always-on face recognition in wearable devices.
Jinmook Lee, Kyeongryeol Bong, Hoi-Jun Yoo
ISCAS5
2017 A 17.5-fJ/bit Energy-Efficient Analog SRAM for Mixed-Signal Processing
abstract
An energy-efficient analog SRAM (A-SRAM) is proposed to eliminate redundant analog-to-digital (A/D) and digital-to-analog (D/A) conversion in mixed-signal systems, such as neuromorphic chips and neural networks. D/A conversion is integrated into the SRAM readout by charge sharing of the proposed split bitline (BL). Also, A/D conversion is integrated into the SRAM write operation with the successive approximation method in the proposed input-output block. Also, a configurable SRAM bitcell array is proposed to allocate the converted digital data without unfilled bitcells. The multirow access decoder selects multiple bitcells in a single column and configures the bitcell array by controlling the BL switches to split BLs. The proposed A-SRAM is implemented using the 65-nm CMOS technology. It achieves 17.5-fJ/bit energy-efficiency and 21-Gbit/s throughput for the analog readout, which are 64% and 1.3 times better than those of the conventional SRAM followed by a digital-to-analog converter (DAC). Also, the area is reduced by 91% compared with the conventional SRAM with analog-to-digital converter (ADC) and DAC.
Jinsu Lee, Dongjoo Shin, Youchang Kim, Hoi-Jun Yoo
IEEE Trans. Very Large Scale Integr. Syst.4
2016 An intelligent ADAS processor with real-time semi-global matching and intention prediction for 720p stereo vision
abstract
Presents a collection of slides covering the following topics: Intelligent ADAS Processor; Stereo Vision; and SoC Architecture.
Kyuho Jason Lee, Kyeongryeol Bong, Hoi-Jun Yoo
Hot Chips Symposium4
2016 A fault tolerant cache system of automotive vision processor complying with ISO26262
abstract
With fluctuating voltage, widening operating temperature, and increasing clock frequency, cache systems are becoming increasingly susceptible to transient (soft) errors. Error correction code(ECC) is an attractive approach for transient error detection and correction[1]. However, redundant memory for error correction code has a significant impact on cost and increases the transient error rates. This paper presents a fault tolerant cache system of vision processors for automotive. The cache system has the small redundant memory and decreases the transient error rates. We proposed the mechanism which increases the error recovery rate and is considered with data cache characteristics.
Jinho Han, Young-Su Kwon, Kyeongjin Byun, Hoi-Jun Yoo
ISCAS4
2016 A 43.7 mW 94 fps CMOS image sensor-based stereo matching accelerator with focal-plane rectification and analog census transformation
abstract
The depth information is actively utilized for many applications such as mobile gesture user interface (UI). However, the previous stereo vision systems are unsuitable for the mobile gesture UI due to the long latency and the high-power consumption of external image sensor in embedded environments. In this paper, we propose a CMOS image sensor-based real-time stereo matching accelerator with low power consumption. For real-time operation, the focal-plane rectification is proposed to perform the image readout, the rectification, and the matching cost generation at the same time. Also, a low-power analog census transformation is implemented by simple comparator circuits. The proposed stereo matching CIS, implemented in 65nm CMOS technology, consumes 43.7 mW at 94.1 fps frame rate. It achieves 5.30×103 MDE/J energy efficiency.
Kyeongryeol Bong, Sungpill Choi, Hoi-Jun Yoo
ISCAS4
2016 A 54-μW fast-settling arterial pulse wave sensor for wrist watch type system
abstract
A dedicated ultra-low power arterial pulse wave (APW) sensor for wrist watch type system is implemented in 0.18-μm CMOS technology with 1.8-V supply. A duty cycle controlled (DCC) current source (CS) enables low-power consuming current injection with 98% power reduction. A DC balanced amplifier reduces settling time by 72%, enabling fast APW signal acquisition when motion artifact is occurred. The simulated 2.125-mm2 single chip APW sensor consumes only 54-μW.
Kwantae Kim, Minseo Kim 0001, Kwonjoon Lee, Seung-Tak Ryu, Hoi-Jun Yoo
ISCAS6
2016 30-fps SNR equalized electrical impedance tomography IC with fast-settle filter and adaptive current control for lung monitoring
abstract
Real-time lung electrical impedance tomography (EIT) IC with SNR equalization is implemented in 110-nm CMOS process. The fast-settling high-pass filter (FS-HPF) and fast-settling low-pass filter (FS-LPF) is proposed to reduce the settling-time which takes over 90% of entire EIT operating time. For the FS-LPF, voltage-controlled pseudo-resistor (VCPR) and current DAC-based voltage control circuit are proposed. For accurate image reconstruction, adaptive current control (ACC) scheme is implemented for SNR equalization of sensing electrodes. As a result, 35-μs settling-time is achieved for 100-kHz carrier frequency satisfying 30-fps operation on single receiver channel. The simulation results show that the SNR equalization can reduce the center of mass (COM) error of reconstructed image by 72.8% with ACC.
Jaehyuk Lee, Unsoo Ha, Hoi-Jun Yoo
ISCAS3
2016 A 635 μW non-contact compensation IC for body channel communication
abstract
A low-power non-contact compensation IC for body channel communication (BCC) is proposed. The proposed IC has 3 key building blocks. First, a contact status detection unit (CSDU) based on a capacitance sensor is adopted. Second, non-contact compensation unit (NCU) with inductor and capacitance bank is proposed. Third, a reconfigurable low noise amplifier (LNA) for impedance matching is adopted. The input impedance of the LNA is controlled by a digital controller for high Q compensation. Thanks to the proposed features, 14 dB channel compensation is achieved and the IC consumed only 635 μW in 65 nm CMOS technology with 1.2 V supply. The proposed work is the first BCC IC with non-contact compensation.
Kyoung-Rog Lee, Jaeeun Jang, Hoi-Jun Yoo
ISCAS4
2016 A 48 μW, 8.88 × 10-3 W/W batteryless energy harvesting BCC identification system
abstract
A BCC identification system which is fully compatible with previous radio frequency identification (RFID) systems is proposed in order to reduce power consumption, avoid security breaches, and enhance convenience via an intuitive interface. The BCC identification system is composed of a reader and a tag. The reader sends an RF wave to the tag, receives an identification code, and analyzes the received code. The tag harvests energy from the RF wave transmitted by the reader, and transfers the identification code to the reader. The energy harvester in the BCC identification tag increases power conversion efficiency (PCE) by up to 12% by adaptively changing the number of rectifier stages depending on input power. In addition, a transformer reusing an on-off keying (OOK) BCC transmitter is proposed to inform the reader of completion of energy harvesting of the tag by changing load impedance. As a result, a 48 μW, 8.88×10-3W/W Figure-of-Merit (FoM) BCC identification system is implemented. This system can generate sufficient power in the tag with lower transmitted power from the reader compared to previous RFID systems.
Jihee Lee, Yongsu Lee, Hoi-Jun Yoo
ISCAS4
2016 A 17.5 fJ/bit energy-efficient analog SRAM for mixed-signal processing
abstract
An energy-efficient analog SRAM (A-SRAM) is proposed to eliminate redundant analog-to-digital (A/D) and digital-to-analog (D/A) conversions in the mixed-signal processing such as a biomedical and a neural network applications. The D/A and the A/D conversion are integrated into the SRAM readout by the charge sharing of the proposed split bit-line (BL) and the SRAM write by the successive approximation method, respectively. And a data structure is newly proposed to allocate each bit of the input data to the binary-weighted bit-cell array. The proposed A-SRAM is implemented using 65 nm CMOS technology. As a result, it achieves 17.5 fJ/bit read energy-efficiency and 21 Gbit/s read throughput, which are 54% lower and 1.3× higher than the conventional SRAM. Also, the area is reduced by 31% compared to the conventional SRAM with ADC and DAC.
Jinsu Lee, Dongjoo Shin, Youchang Kim, Hoi-Jun Yoo
ISCAS4
2015 A low-power and real-time augmented reality processor for the next generation smart glasses
Gyeonghoon Kim, Hoi-Jun Yoo
Hot Chips Symposium2
2015 A 124.9fps memory-efficient hand segmentation processor for hand gesture in mobile devices
abstract
Hand gesture recognition is one of emerging Human Computer Interaction (HCI) technologies for the next generation of mobile devices. However, conventional software-oriented approaches spend a considerable time and require a large memory size for hand segmentation, which fails to give real-time interactions between users and mobile devices. Therefore, in this paper, we present a high-throughput and memory-efficient hand segmentation processor. To obtain both of high throughput and high memory-efficiency, we propose a parallelized hand candidate decision and a compressed feedback histogram. As a result, it achieves 124.9 fps with only 26.9 KB on-chip memory, which are 1.39 times faster and 92 time smaller, respectively, compared to the state-of-the-art.
Sungpill Choi, Seongwook Park, Gyeonghoon Kim, Hoi-Jun Yoo
ISCAS4
2015 A 0.54-mW duty controlled RSSI with current reusing technique for human body communication
abstract
A low power adaptive controlled current-reusing received signal strength indicator (RSSI) is proposed for the human body communication (HBC). The proposed RSSI has three low power features. First, the power on controller (PoC) scheme is proposed to achieve the duty control of the RSSI. It significantly reduces the average power consumption of RSSI over 90%. Second, the current stacking scheme is adopted to share both eight rectifiers and eight amplifiers, composing the RSSI. By the current reusing technique, the power consumption of the RSSI is reduced to 45%. In addition, the reconfigurable LNA is used in the front-end of HBC TRX. The RSSI adaptively controls the gain and noise figure of the LNA to optimize the power consumption. The proposed RSSI occupies 0.85mm2in 0.18-μm CMOS technology.
Jaeeun Jang, Yongsu Lee, Hoi-Jun Yoo
ISCAS4
2015 A 24-mW 28-Gb/s wireline receiver with low-frequency equalizing CTLE and 2-tap speculative DFE
abstract
In this paper, a power-efficient equalization techniques is proposed for a high data-rate multi-standard wireline receiver. First, a low-frequency-equalizing continuous-time linear equalizer (LFE-CTLE) compensates for not only the short-term inter-symbol interference (ISI) from high-frequency channel loss but also the long-term ISI from the low-frequency channel loss without additional power consumption compared to the previous CTLE. LFE-CTLE can reduce the required number of taps and power consumption of the following decision feedback equalizer (DFE). A 2-tap speculative DFE adopts 4-phase clocking techniques to reduce the number of summation nodes and latches for low-power consumption. The proposed receiver is designed in 65-nm LP CMOS technology with 1.-2V supply voltage. It can achieve 28-Gb/s data rate with a 24-mW power efficiency.
Minseo Kim 0001, Joonsung Bae, Unsoo Ha, Hoi-Jun Yoo
ISCAS4
2015 A 3.13nJ/sample energy-efficient speech extraction processor for robust speech recognition in mobile head-mounted display systems
abstract
An energy-efficient speech extraction (SE) processor is proposed for the robust speech recognition in the head-mounted display (HMD) systems. Speech extraction is essential for robust speech recognition in noisy environment. For the low-latency speech extraction, FastSE is proposed to overcome 50x larger complex cICA-based selection process which results in <;2ms SE latency. Moreover, a reinforced-FastSE (RFSE) scheme is proposed to achieve 97.2% accuracy with small on-chip memory size of only 33kB for the low-power HMD applications. Also, Reconfigurable matrix operation accelerator (RMAT) is implemented for energy-efficient acceleration of dominant matrix operation on SE. As a result, the proposed SE processor achieves 1.3x lower latency with 4.24x smaller memory compared to the state-of-the-art work, so that speech recognition in noisy environment becomes possible for mobile HMD applications.
Jinmook Lee, Seongwook Park, Injoon Hong, Hoi-Jun Yoo
ISCAS4
2014 An 1.61mW mixed-signal column processor for BRISK feature extraction in CMOS image sensor
abstract
In mobile object recognition (OR) applications, the power consumption of image sensor and data communication between image sensor and digital OR processor becomes crucial as digital OR processor consumes less power in deep sub-micron process. To reduce the amount of data transaction from image sensor to digital OR processor, digital/analog mixed-signal focal-plane processing of Binary Robust Invariant Scalable Keypoints (BRISK) feature extraction in CMOS image sensor (CIS) is proposed. The proposed CIS processor sends BRISK feature vectors instead of the whole image pixel data, resulting in 79% reduction of data communication. In this work, mixed-signal processing of corner detection and successive approximation register (SAR)-based scoring are implemented for BRISK feature point detection. To achieve scale-invariance in object recognition, scale-space is generated and stored in analog line memory. In addition, noise reduction scheme is integrated in column processing chain to remove salt and pepper noise, which degrades recognition accuracy. In a post layout simulation, the proposed system achieves 0.70pW/pixel*frame*feature at 30fps in a 130nm CMOS technology, which is 13.6% lower than the state-of-the-art.
Kyeongryeol Bong, Gyeonghoon Kim, Injoon Hong, Hoi-Jun Yoo
ISCAS4
2014 3.8 mW electrocardiogram (ECG) filtered electrical impedance tomography IC using I/Q homodyne architecture for breast cancer diagnosis
abstract
A low-power electrical impedance tomography (EIT) IC proposed for breast cancer diagnosis is implemented in 180nm CMOS process. For the breast cancer diagnosis, low power and high accuracy is required. The proposed IC reduces the power consumption to 3.8mW using the homodyne conversion mixer to lower the sampling rate of ADC. To gain high accuracy, the adaptive filter using the miller capacitor perfectly filters out the ECG signal. I/Q dual path architecture measures conductivity and permittivity components separately, so distance errors are reduced about 86% in simulation.
Yongsu Lee, Unsoo Ha, Kiseok Song, Hoi-Jun Yoo
ISCAS4
2014 An 1.92mW Feature Reuse Engine based on inter-frame similarity for low-power object recognition in video frames
abstract
A Feature Reuse Engine (FReE) is proposed to achieve low-power object recognition in video frames. Previous object recognition processors perform redundant processing for repeated features of sequential frames in video. Unlike previous works, proposed FReE reuses 58% of features from previous frame with inter-frame similarity. However, false reuse can degrade the object recognition accuracy. By pixel intensity-based Near Pixel Information (NPI), FReE can decide the reusability with 96% of accuracy. Additionally, to minimize latency and power overhead of FReE, a dedicated integral image generator and NPI generator are proposed. As a result, power consumption of object recognition processor is reduced by 31% with the proposed FReE which consumes only 1.92mW in a 130nm CMOS technology.
Dongjoo Shin, Injoon Hong, Hoi-Jun Yoo
ISCAS3
2013 A 0.7pJ/bit 2Gbps self-synchronous serial link receiver using gated-ring oscillator for inductive coupling communication
abstract
A low-energy self-synchronous serial link receiver for inductive coupling communication is implemented in 130nm CMOS process. The gated-ring oscillator (GRO) is proposed to combine three key building blocks (CDR, phase interpolator and PLL) used in conventional receiver resulting in energy consumption reduction. In addition, the start-up time and settling time for clock recovery can be significantly reduced due to the large loop bandwidth characteristics of the proposed GRO. As a result, the proposed serial link receiver achieves 2Gbps data rate and sub-20ns settling time while consuming 0.7pJ/bit from the 1.2V supply.
Unsoo Ha, Hoi-Jun Yoo
ISCAS3
2013 A 34.1fps scale-space processor with two-dimensional cache for real-time object recognition
abstract
A scale-space processor with two-dimensional cache is proposed to achieve real-time object recognition in HD 720p images. Scale-space is the most commonly used concept to achieve scale-invariant property in object recognition, however its high computational cost makes it hard to implement a realtime object recognition processor. We employ hierarchical convolution unit (HCU) which computes multiple pixels in a single cycle with various kernel sizes. In addition, two-dimensional cache (T-Cache) supports accessing vertically consecutive data from any image window size with reduced area. A pre-fetch controller for the proposed cache improves the hit rate by exploiting the sequential access pattern of convolution tasks. As a result, the scale-space processor implemented in a 0.13μm CMOS technology achieves 34.1fps on a HD 720p image while consuming peak power of 84.5mW.
Youchang Kim, Junyoung Park 0002, Hoi-Jun Yoo
ISCAS3
2013 A multi-modal and tunable Radial-Basis-Funtion circuit with supply and temperature compensation
abstract
We propose an analog Radial-Basis-Function (RBF) circuit that generates 4 different types of RBFs, which are spline, Gaussian, multi-quadratic, and log-like spline curves. Moreover, the proposed RBF circuit is designed to have high tunability on centers, heights, and widths. The proposed RBF circuit is also robust to both temperature variation (-37~87□C) and supply voltage variation (1~2V). The sum of area and power consumption of each RBF from 3 different previous works is 13, 622μm2and 121μW, respectively. On the other hand, the proposed circuit occupies only 1,050μm2and consumes 10.5μW which are only 13% and 11.5%, respectively. For its verification, an analog/digital mixed-mode RBF Neural Network (RBFNN) classifier is designed which adopted the proposed RBF circuit.
Kyuho Jason Lee, Junyoung Park 0002, Gyeonghoon Kim, Injoon Hong, Hoi-Jun Yoo
ISCAS5
2013 A 32.8mW 60fps cortical vision processor for spatio-temporal action recognition
abstract
In this paper, we propose a 60fps cortical vision processor modeling a hierarchical object classification model (HMAX) based on spatio-temporality in video stream. It is hard to implement a real-time hardware for HMAX operation due to its high computational cost from 2-D template matching. Three components are proposed with our improved algorithms. Class Refinement Structure (CRS) dramatically reduces a dimension of HMAX descriptors by 97.01% compared to the previous works by exploiting spatio-temporal features in video action recognition. Spatio-Temporal Memory Structure (STMS) adopts spatially adaptive window technique, and it reduces the required on-chip data bandwidth and computations per a template in S2 stage. In addition, a dual image buffer structure also reduces the required off-chip network bandwidth for processing complex hierarchical stages and numerous image spaces. As a result, the 10.8 GOPS cortical vision processor implemented in 0.13μm CMOS process achieves 60frames/sec performances for 256×256 video inputs at 200MHz operating frequency.
Seongwook Park, Junyoung Park 0002, Injoon Hong, Hoi-Jun Yoo
ISCAS4
2013 1.2-mW Online Learning Mixed-Mode Intelligent Inference Engine for Low-Power Real-Time Object Recognition Processor
abstract
Object recognition is computationally intensive and it is challenging to meet 30-f/s real-time processing demands under sub-watt low-power constraints of mobile platforms even for heterogeneous many-core architecture. In this paper, an intelligent inference engine (IIE) is proposed as a hardware controller for a many-core processor to satisfy the requirements of low-power real-time object recognition. The IIE exploits learning and inference capabilities of the neurofuzzy system by adopting the versatile adaptive neurofuzzy inference system (VANFIS) with the proposed hardware-oriented learning algorithm. Using the programmable VANFIS, the IIE can configure its hardware topology adaptively for different target classifications. Its architecture contains analog/digital mixed-mode neurofuzzy circuits for updating online parameters to increase attention efficiency of object recognition process. It is implemented in 0.13-μm CMOS process and achieves 1.2-mW power consumption with 94% average classification accuracy within 1-μs operation delay. The 0.765-mm2IIE achieves 76% attention efficiency and reduces power and processing delay of the 50-mm2image processor by up to 37% and 28%, respectively, when 96% recognition accuracy is achieved.
Jinwook Oh, Seungjin Lee 0001, Hoi-Jun Yoo
IEEE Trans. Very Large Scale Integr. Syst.3
2012 A 39 µW body channel communication wake-up receiver with injection-locking ring-oscillator for wireless body area network
abstract
An ultra-low power wake-up receiver for body channel communication (BCC) is implemented in 0.13 μm CMOS process. The proposed wake-up receiver uses the injection-locking ring-oscillator (ILRO) to replace the RF amplifier with low power consumption. Through the ILRO, the frequency modulated input signal is converted to the full swing rectangular signal which is directly demodulated by the following low power PLL based FSK demodulator. In addition, the relaxed sensitivity and selectivity requirement by the good channel quality of the BCC reduces the power consumption of the receiver. As a result, the proposed wake-up receiver achieves a sensitivity of -55.2 dbm at a data rate of 200 kbps while consuming only 39 μW from the 0.7 V supply.
Joonsung Bae, Hoi-Jun Yoo
ISCAS3
2012 A 2.1µW real-time reconfigurable wearable BAN controller with dual linked list structure
abstract
A real-time reconfigurable network controller is proposed for wearable healthcare applications. Based on the dual linked-list structure, it can add or remove the nodes in the network during operation within 8µs and 500ms, respectively. To achieve low power consumption, three schemes of clock gating, low duty-cycled operation, and fixed-order allocation are proposed. As a result, the proposed network controller consumes only 2.1µW in average with 20MHz clock.
Seulki Lee 0001, Taehwan Roh, Sunjoo Hong, Hoi-Jun Yoo
ISCAS4
2011 A 145µW 8×8 parallel multiplier based on optimized bypassing architecture
abstract
A low-power parallel multiplier based on optimized bypassing architecture (OBA) is proposed. The proposed OBA has two kinds of adder cells to reduce power consumption by 15.7 %. One is the two-dimensional bypassing adder (TDBA) which performs both row and column bypassing scheme simultaneously, and the other is the modified row-bypassing adder (MRBA) for the proposed row-bypassing scheme. In the proposed TDBA and MRBA, the logic evaluation is partially activated by internal tri-state buffers (ITBs) in order to save the switching power dissipation up to 33.7 % and 32.0 %, respectively. Implemented in 0.13 μm CMOS process, the proposed 8x8 parallel multiplier consumes only 145 μW.
Sunjoo Hong, Taehwan Roh, Hoi-Jun Yoo
ISCAS3
2011 A low-energy hybrid radix-4/-8 multiplier for portable multimedia applications
abstract
A hybrid radix-4/-8 multiplier is proposed for portable multimedia applications that demand high speed and low energy operation. Depending on the input pattern, the multiplier operates in the radix-8 mode in 56% of the input cases for low power, but reverts to the radix-4 mode in 44% of the slower input cases for high speed. For this, a mode detection circuit determines the mode signal from the input operand in just 2 gate delays. Based on the mode signal, the radix-4/-8 dual Booth encoder generates encoding signals in a hardware efficient way. Moreover, the carry save adder block is selectively activated to reduce power consumption. Compared to a conventional radix-4 multiplier, the proposed hybrid multiplier architecture consumes 33.5% less power at the expense of just 3.3% additional propagation delay, resulting in 31.3% less energy per operation.
Gyeonghoon Kim, Seungjin Lee 0001, Junyoung Park 0002, Hoi-Jun Yoo
ISCAS4
2011 A 2.4µW 400nC/s constant charge injector for wirelessly-powered electro-acupuncture
abstract
An ultra-low-power constant charge injector (CCI) circuit is presented for wirelessly-powered electro-acupuncture (EA). The CCI adopts adaptive pulse-width and frequency-drift (APF) calibration loop that not only accommodates to body impedance variation (BIV) of 100-200kΩ but also is tolerable to frequency-drift. The proposed 5Hz current starving clock generator combined with sub-Vthreference circuit provides supply voltage dependency of 0.2Hz/V and temperature dependency of 0.018Hz/°C while consuming only 1μW. These variations are finely calibrated by the APF loop to ensure injectable charge intensity as stable as 399.33-400.45nC/s. The low-power low-voltage Gmcircuit in CCI incorporates diode-limited and chopper-modulated inputs to prevent the damage from static electricity of the body while providing linear voltage-to-current conversion. The proposed CCI, simulated in 0.18μm CMOS technology with supply voltage of 1.0V, consumes 2.4μW of power.
Hyungwoo Lee, Kiseok Song, Hoi-Jun Yoo
ISCAS4
2011 24-GOPS 4.5-mm2 Digital Cellular Neural Network for Rapid Visual Attention in an Object-Recognition SoC
abstract
This paper presents the Visual Attention Engine (VAE), which is a digital cellular neural network (CNN) that executes the VA algorithm to speed up object-recognition. The proposed time-multiplexed processing element (TMPE) CNN topology achieves high performance and small area by integrating 4800 (80 × 60) cells and 120 PEs. Pipelined operation of the PEs and single-cycle global shift capability of the cells result in a high PE utilization ratio of 93%. The cells are implemented by 6T static random access memory-based register files and dynamic shift registers to enable a small area of 4.5 mm(2). The bus connections between PEs and cells are optimized to minimize power consumption. The VAE is integrated within an object-recognition system-on-chip (SoC) fabricated in the 0.13- μm complementary metal-oxide-semiconductor process. It achieves 24 GOPS peak performance and 22 GOPS sustained performance at 200 MHz enabling one CNN iteration on an 80 × 60 pixel image to be completed in just 4.3 μs. With VA enabled using the VAE, the workload of the object-recognition SoC is significantly reduced, resulting in 83% higher frame rate while consuming 45% less energy per frame without degradation of recognition accuracy.
Seungjin Lee 0001, Minsu Kim 0004, Kwanho Kim, Joo-Young Kim 0001, Hoi-Jun Yoo
IEEE Trans. Neural Networks5
2010 A 22.4 mW competitive fuzzy edge detection processor for volume rendering
abstract
A low power competitive fuzzy edge detection (C-FED) processor is proposed for gradient calculations in volume rendering. Its linearized fuzzy membership function reduces overall power by 35.1% and the proposed hardware sharing between computation stages reduces power consumption by 18%. Threshold adaptive bit control scheme is proposed to predetermine background pixel with simple operation which results in 13% power reduction. Overall power consumption is reduced by 53.8%. Its power consumption and energy per pixel is 22.4 mW and 0.14nJ/pixel, respectively, at 1.8-V supply. The fabricated processor occupying 450 μm × 450 μm in a 0.18 μm CMOS process achieves 1821.5fps for the input image of 300 × 300 pixels at 200 MHz operating frequency.
Joonsoo Kwon, Minsu Kim 0004, Jinwook Oh, Hoi-Jun Yoo
ISCAS4
2010 Live demonstration: A real-time compensated inductive transceiver for wearable MP3 player system on multi-layered planar fashionable circuit board
abstract
A wearable MP3 player system on a multi-layered common fabric patch is proposed for an unobtrusive usage in daily life. An inductive coupling transceiver is proposed as a wearable wireless connector, and it reduces power consumption below to 185.6μW in total. Also it compensates for the dynamic variation caused by user's activities within 3.96μs in worst case. The complete system is implemented on a 2-layer fabric substrate, and the music playback process is fully demonstrated.
Seulki Lee 0001, Seungwook Paek, Hoi-Jun Yoo
ISCAS3
2010 A real-time compensated inductive transceiver for wearable MP3 player system on multi-layered planar fashionable circuit board
abstract
A wearable MP3 player system on a multi-layered common fabric patch is proposed for an unobtrusive usage in daily life. An inductive coupling transceiver is proposed as a wearable wireless connector in order to eliminate the physical attachment and detachment of the memory card so that enhance the system reliability. It adopts the Pulsed Clock On-Off Keying (PC-OOK) modulation to reduce power consumption below to 185.6μW in total. And Real-time Capacitor Compensation (RCC) scheme compensates for the dynamic variation caused by user's activities within 3.96μs in worst case. The complete system is implemented on a 2-layer fabric substrate, and the music playback process is fully demonstrated.
Seulki Lee 0001, Seungwook Paek, Hoi-Jun Yoo
ISCAS3
2010 A 30fps stereo matching processor based on belief propagation with disparity-parallel PE array architecture
abstract
In this paper, we propose a real-time stereo matching processor based on the belief propagation algorithm. Computationally complex message construction is accelerated by a disparity-parallel PE array architecture, which calculates messages for all disparity levels (1-32) in parallel. A tile-based belief propagation approach reduces the on-chip memory requirements by 95.4% compared to the previous works. In addition, a two-level on-chip buffer and memory access pipelining enable high PE utilization of 89%. As a result, the message construction rate of the PEs is increased by 6.45x compared to previous works. The fabricated processor in a 0.18um CMOS process achieves 30 fps performance for QVGA (320×240) video inputs at 200 MHz operating frequency.
Junyoung Park 0002, Seungjin Lee 0001, Hoi-Jun Yoo
ISCAS3
2010 A lOMb/s 4ns jitter direct conversion low Modulation Index FSK demodulator for low-energy body sensor network
abstract
A high speed and low jitter direct conversion FSK demodulator for low energy body sensor network (BSN) is implemented in 0.18um CMOS technology with 1V supply. Modulation Index (MI) of 1 is used for high data rate, which performs high duty cycling operation for low-energy BSN. To eliminate deterministic jitter at MI of 1, the demodulator employs a delay locked loop (DLL), demodulation logic and a transition predictive detector (TPD), perceiving the exact data transition point. Dual offset-compensated comparators are utilized to reduce the random jitter during transition predictive operation. As a result, the proposed demodulator achieves absolute jitter of 4ns at a 40Mb/s data rate with a 5MHz base band signal.
Taehwan Roh, Joonsung Bae, Hoi-Jun Yoo
ISCAS3
2010 A wirelessly-powered electro-acupuncture based on Adaptive Pulse Width Mono-Phase stimulation
abstract
A wirelessly-powered electro-acupuncture (EA) that dynamically adapts to body-impedance variation (BIV), is proposed. The proposed EA consists of a slender needle, a helical antenna (70 turns, 1mm diameter) using conductive yarn (100pm diameter, 6.6Ω2/m), and a 1.56mm2stimulator chip fabricated in 0.18μm 1P6M CMOS process. A stable supply voltage of 1V is wirelessly generated from 433MHz-ISM band with sensitivity of -16dBm. To deal with BIV in the range of 100KΩ-200kΩ, Adaptive-Pulse-Width (APW) scheme is introduced to maintain constant charge injection of 80nC per stimulation. A pair of EAs forms an EA node, and they operate in Alternate Mono-Phase (AMP) fashion to guarantee the safety by neutralization of the injected charge.
Kiseok Song, Seulki Lee 0001, Hoi-Jun Yoo
ISCAS3
2010 Familiarity based unified visual attention model for fast and robust object recognition
Seungjin Lee 0001, Kwanho Kim, Joo-Young Kim 0001, Minsu Kim 0004, Hoi-Jun Yoo
Pattern Recognit.5
2010 An attention controlled multi-core architecture for energy efficient object recognition
Joo-Young Kim 0001, Sejong Oh, Seungjin Lee 0001, Minsu Kim 0004, Jinwook Oh, Hoi-Jun Yoo
Signal Process. Image Commun.6
2010 Visual Image Processing RAM: Memory Architecture With 2-D Data Location Search and Data Consistency Management for a Multicore Object Recognition Processor
abstract
Abstract-Visual image processing random access memory (VIP-RAM) is proposed for a real-time multicore object recognition processor. It has two key features for the overall processor: 1) single cycle local maximum location search (LMLS) for fast key-point localization in object recognition, and 2) data consistency management (DCM) for producer-consumer data transactions among the processors. To achieve single cycle LMLS operation for a 3 x 3 window, the VIP-RAM adopts a hierarchical three-bank architecture that finds the maximum of each row in each bank first, then finds the final maximum of the window and its address in the top level. To this end, each memory bank embeds specialized logic blocks, such as three successive data read logic and bitwise competition logic comparator. With the single cycle LMLS operation, the key-point localization task is accelerated by 2.6 ? with a 27% reduction of power. For the DCM function, the VIP-RAM includes a valid check unit (VCU) that automatically manages the validity of each 32-bit data. It dynamically updates/checks the validity of the shared data when the producer processor writes the data or the consumer processor reads data. With a customized single-ended memory cell and multibit-line selection logic, the VCU can provide a validity check not only for single data access, but also for multiple data accesses such as burst and LMLS operation. Eliminating data synchronization overhead with the DCM, the VIP-RAM reduces the amount of on-chip data transactions and execution time in producer-consumer data transactions by 22.6% and 15.4%, respectively. The overall object recognition processor that includes eight VIP-RAMs and ten processors is fabricated in 0.18/im complementary metal-oxide-semiconductor technology with the chip size of 7.7 mm ? 5 mm. The VIP-RAM occupies a 1.09 mm ? 0.83 mm die area and dissipates 113.2 mW when it performs the LMLS operation in every cycle at 200 MHz frequency and 1.8-V supply.
Joo-Young Kim 0001, Donghyun Kim 0014, Seungjin Lee 0001, Kwanho Kim, Hoi-Jun Yoo
IEEE Trans. Circuits Syst. Video Technol.5
2010 ECG signal compression and classification algorithm with quad level vector for ECG holter system
abstract
An ECG signal processing method with quad level vector (QLV) is proposed for the ECG holter system. The ECG processing consists of the compression flow and the classification flow, and the QLV is proposed for both flows to achieve better performance with low-computation complexity. The compression algorithm is performed by using ECG skeleton and the Huffman coding. Unit block size optimization, adaptive threshold adjustment, and 4-bit-wise Huffman coding methods are applied to reduce the processing cost while maintaining the signal quality. The heartbeat segmentation and the R-peak detection methods are employed for the classification algorithm. The performance is evaluated by using the Massachusetts Institute of Technology-Boston's Beth Israel Hospital Arrhythmia Database, and the noise robust test is also performed for the reliability of the algorithm. Its average compression ratio is 16.9:1 with 0.641% percentage root mean square difference value and the encoding rate is 6.4 kbps. The accuracy performance of the R-peak detection is 100% without noise and 95.63% at the worst case with -10-dB SNR noise. The overall processing cost is reduced by 45.3% with the proposed compression techniques.
Hyejung Kim, Refet Firat Yazicioglu, Patrick Merken, Chris Van Hoof, Hoi-Jun Yoo
IEEE Trans. Inf. Technol. Biomed.5
2009 An Energy-efficient Dual Sampling SAR ADC with Reduced Capacitive DAC
abstract
This paper presents an energy-efficient SAR ADC which adopts reduced MSB cycling step with dual sampling of the analog signal. By sampling and holding the analog signal asymmetrically at both input sides of comparator, the MSB cycling step can be hidden by hold mode. Benefits from this technique, not only the total capacitance of DAC is reduced by half, but also the average switching energy is reduced by 68% compared with conventional SAR ADC. Moreover, switching energy distribution is more uniform over entire output code compared with previous works.
Binhee Kim, Jerald Yoo, Namjun Cho, Hoi-Jun Yoo
ISCAS5
2009 A 60fps 496mW multi-object recognition processor with workload-aware dynamic power management
abstract
An energy efficient object recognition processor is proposed for real-time visual applications. Its energy efficiency is improved by lowering average power consumption while sustaining high frame rate. To this end, the proposed processor features from all levels of chip design. In architecture level, it performs 3-stage task pipelining for high frame rate operation and workload-aware dynamic power management for low power consumption. In block level, energy efficient special purposed engines are employed while software controlled clock gating is exploited for fine-grained clock control. In circuit level, analog-digital mixed design is used to reduce power with the same performance. As a result, the 49mm2 chip in a 0.13mm technology achieves 60fps object recognition for VGA (640x480) input with 496mW power at the supply of 1.2V. It means only 8.2mJ is dissipated per frame, which is 3.2X more energy efficient than the state of the art.
Joo-Young Kim 0001, Seungjin Lee 0001, Jinwook Oh, Minsu Kim 0004, Hoi-Jun Yoo
ISLPED5
2009 A Configurable Heterogeneous Multicore Architecture With Cellular Neural Network for Real-Time Object Recognition
abstract
As object recognition requires huge computation power to deal with complex image processing tasks, it is very challenging to meet real-time processing demands under low-power constraints for embedded systems. In this paper, a configurable heterogeneous multicore architecture with a dual-mode linear processor array and a cellular neural network on the network-on-chip platform is presented for real-time object recognition. The bio-inspired attention-based object recognition algorithm is devised to reduce computational complexity of the object recognition. The cellular neural network is utilized to accelerate the visual attention algorithm for selecting salient image regions rapidly. The dual-mode parallel processor is configured into single instruction, multiple data (SIMD) or multiple-instruction-multiple-data modes to perform data-intensive image processing operations while exploiting pixel-level and feature-level parallelisms required for the attention-based object recognition. The algorithm's hybrid parallelization strategy on the proposed architecture is adopted to obtain maximum performance improvement. The performance analysis results, using a cycle-accurate architecture simulator, show that the proposed architecture achieves a speedup of 2.8 times for the target algorithm over conventional massively parallel SIMD architecture at low hardware cost overhead. A prototype chip of the proposed architecture, fabricated in 0.13 mum complementary metal-oxide-semiconductor technology, achieves 22 frames/s real-time object recognition with less than 600 mW power consumption.
Kwanho Kim, Seungjin Lee 0001, Joo-Young Kim 0001, Minsu Kim 0004, Hoi-Jun Yoo
IEEE Trans. Circuits Syst. Video Technol.5
2009 A Wearable ECG Acquisition System With Compact Planar-Fashionable Circuit Board-Based Shirt
abstract
A wearable electrocardiogram (ECG) acquisition system implemented with planar-fashionable circuit board (P-FCB)-based shirt is presented. The proposed system removes cumbersome wires from conventional Holter monitor system for convenience. Dry electrodes screen-printed directly on fabric enables long-term monitoring without skin irritation. The ECG monitoring shirt exploits a monitoring chip with a group of electrodes around the body, and both the electrodes and the interconnection are implemented using P-FCB to enhance wearability and to lower production cost. The characteristics of P-FCB electrode are shown, and the prototype hardware is implemented to successfully verify the proposed concept.
Jerald Yoo, Seulki Lee 0001, Hyejung Kim, Hoi-Jun Yoo
IEEE Trans. Inf. Technol. Biomed.5
2009 81.6 GOPS Object Recognition Processor Based on a Memory-Centric NoC
abstract
For mobile intelligent robot applications, an 81.6 GOPS object recognition processor is implemented. Based on an analysis of the target application, the chip architecture and hardware features are decided. The proposed processor aims to support both task-level and data-level parallelism. Ten processing elements are integrated for the task-level parallelism and single instruction multiple data (SIMD) instruction is added to exploit the data-level parallelism. The memory-centric network-on-chip (NoC) is proposed to support efficient pipelined task execution using the ten processing elements. It also provides coherence and consistency schemes tailored for 1-to-N and M-to-1 data transactions in a task-level pipeline. For further performance gain, the visual image processing memory is also implemented. The chip is fabricated in a 0.18-mum CMOS technology and computes the key-point localization stage of the SIFT object recognition twice faster than the 2.3 GHz Core 2 Duo processor.
Donghyun Kim 0014, Kwanho Kim, Joo-Young Kim 0001, Seungjin Lee 0001, Se-Joong Lee, Hoi-Jun Yoo
IEEE Trans. Very Large Scale Integr. Syst.6
2009 A 152-mW Mobile Multimedia SoC With Fully Programmable 3-D Graphics and MPEG4/H.264/JPEG
abstract
This paper presents a low power multimedia system-on-chip (SoC) with full integration with fully programmable 3-D graphics, MPEG4 codec, H.264 decoder, and JPEG codec for mobile devices. The mobile unified shader in 3-D graphics engine provides fully programmable 3-D graphics with 35% area and 28% power reduction. Low-power lighting engine which employs logarithmic number datapath and the specialized lighting instruction enable 9.1 Mvertices/s vertex fill rate, which is 2.5 times improvement compared with previous works including transformations and OpenGL lighting. The SoC consumes less than 152 mW for video applications and less than 195 mW for 3-D graphics applications. The mobile unified shader and merged JPEG/MPEG4 codec reduce the silicon area and the SoC consumes 6.4 mm times 6.4 mm in 0.13 mu m complementary metal-oxide-semiconductor (CMOS) logic process.
Jeong-Ho Woo, Ju-Ho Sohn, Hyejung Kim, Hoi-Jun Yoo
IEEE Trans. Very Large Scale Integr. Syst.4
2008 Vision platform for mobile intelligent robot based on 81.6 GOPS object recognition processor
abstract
To enable power-efficient object recognition of mobile intelligent robots, 81.6GOPS object recognition processor is proposed. Based on analysis of Scale Invariant Feature Transform (SIFT) algorithm, architecture of the proposed processor is designed to support both task and data level parallelism. 10 Processing Elements (PEs) are integrated for task parallelism, and each PE is equipped with SIMD instruction for data parallelism as well. In addition, Visual Image Processing memory replaces complex local maximum pixel search operation with a single read operation for further performance gain. With the proposed processor, we also realized vision platform for real-time SIFT computation of mobile robots. The chip operation is tested up to 200MHz and consumes 540mW in the vision platform at 1.8V supply voltage and 100 MHz operation frequency.
Donghyun Kim 0014, Kwanho Kim, Joo-Young Kim 0001, Seungjin Lee 0001, Hoi-Jun Yoo
DAC5
2008 A 0.6pJ/b 3Gb/s/ch transceiver in 0.18 µm CMOS for 10mm on-chip interconnects
abstract
This paper presents a high speed and low energy transceiver for 10mm long minimum width on-chip global interconnects. To improve the link bandwidth, the transmitter employs a capacitive-resistive pre-emphasis technique and the receiver employs the AC-coupled Resistive Feedback Inverter (RFI) de-emphasis technique. Exploiting two emphasis techniques, the proposed interconnect achieves 1.26GHz bandwidth which is 20 times improved compared to conventional link. As a result, it achieves error-free 3Gb/s data rate and consumes less than 0.6pJ/b during transmission by using low-swing and pulse signaling. The test chip is designed using 1.8V 0.18 μm 6M CMOS technology.
Joonsung Bae, Joo-Young Kim 0001, Hoi-Jun Yoo
ISCAS3
2008 A 6.3nJ/op low energy 160-bit modulo-multiplier for elliptic curve cryptography processor
abstract
A low energy modulo-multiplier is proposed for elliptic curve cryptography (ECC) processor, especially for authentication in mobile device or key encryption in embedded health-care system. The multiplier uses only two 40-bit multipliers to execute 160-bit operation based on the Montgomery modulo-multiplication algorithm. Partial products of multiplication are accumulated with shift registers to get final 160-bit MSB of output value. One modulo-multiplication is executed with 20 clock cycles at 40MHz operating frequency. It consumes 6.3nJ for each modulo-multiplication at 1V supply voltage. It is implemented by using 0.18-μm CMOS process and has 0.7mm × 1.0mm area.
Hyejung Kim, Yongsang Kim, Hoi-Jun Yoo
ISCAS3
2008 A 200Mbps 0.02nJ/b dual-mode inductive coupling transceiver for cm-range interconnection
abstract
A 200 Mbps 0.02 nJ/b dual-mode inductive coupling transceiver is proposed for cm-range inductive coupling interconnection. The parallel capacitor combined with the TX inductor enhances the transmitted signal slew rate so that it increases the transmission distance by twofold. The proposed intersymbol interference (ISI) reduction scheme of the transmitter improves data rate up to 200 Mbps. And the proposed pulse generation scheme allows the transceiver to consume only 0.02 nJ/b energy. The transceiver consumes 0.012 mm2in a TSMC 0.25 um CMOS process.
Seulki Lee 0001, Jerald Yoo, Hoi-Jun Yoo
ISCAS3
2008 Power and Area-Efficient Unified Computation of Vector and Elementary Functions for Handheld 3D Graphics Systems
abstract
A unified computation method of vector and elementary functions is proposed for handheld 3D graphics systems. It unifies vector operations like vector multiply, multiply-and-add, divide, divide-by-square-root, and dot product and elementary functions like trigonometric, inverse trigonometric, hyperbolic, inverse hyperbolic, power (xywith two variables), and logarithm to an arbitrary base into a single four-way arithmetic platform. A number system called the fixed-point hybrid number system (FXP-HNS), which combines the fixed-point number system (FXP) and the logarithmic number system (LNS), is proposed for the power and area-efficient unification. Power and area-efficient logarithmic and antilogarithmic conversion schemes are also proposed for the data conversions between fixed-point and logarithmic numbers in the FXP-HNS and achieve 0.41 percent and 0.08 percent maximum conversion errors, respectively. The unified arithmetic unit based on the proposed schemes is presented with less than 6.3 percent operation error. Its fully pipelined architecture achieves single-cycle throughput with maximum four-cycle latency for all of the supported operations. Comparison results show that the proposed arithmetic unit achieves 30 percent power and 10.9 percent area reductions and runs two times faster than the previous approach.
Byeong-Gyu Nam, Hyejung Kim, Hoi-Jun Yoo
IEEE Trans. Computers3
2007 A Low Power Digital Signal Processor with Adaptive Band Activation for Digital Hearing Aid Chip
abstract
A low power digital signal processor (DSP) for a digital hearing aid chip is presented. The DSP integrates three programmable digital finite impulse response (FIR) filters. Each FIR filter can have one pass-frequency out of seven preset frequencies so that only three FIR filters can have the same flexibility as seven filters. Additionally, a silence mode is defined in which only one filter is activated. A digital voice activity detection circuit is implemented for this purpose. The DSP is implemented as part of a fully integrated digital hearing aid chip. It uses a 0.18 μm CMOS process and occupies an area of 0.5 mm2. Power consumption is 25 μW in normal operating mode and 9 μW in silence mode at 0.9-V supply.
Seungjin Lee 0001, Hoi-Jun Yoo
ISCAS3
2007 A Power Management Unit with Continuous Co-Locking of Clock Frequency and Supply Voltage for Dynamic Voltage and Frequency Scaling
abstract
A power management unit (PMU) architecture is proposed for the domain-specific low power management with dynamic voltage and frequency scaling. The PMU continuously co-locks and dynamically varies the supply voltage and the clock frequency from 89 MHz to 200 MHz and from 1.0V to 1.8V, respectively, in less than 40μs. A 32bit RISC processor is used as power management target device. The PMU, 0.36mm2with 0.18-μm CMOS process, consumes 5mW, and shows -100dBm/Hz phase noise of clock and 160mV load regulation of supply voltage with 100mA load current from the load, RISC processor.
Jeabin Lee, Byeong-Gyu Nam, Seong-Jun Song, Namjun Cho, Hoi-Jun Yoo
ISCAS5
2007 A low power multimedia SoC with fully programmable 3D graphics and MPEG4/H.264/JPEG for mobile devices
abstract
We present a low power multimedia SoC with fully programmable 3D graphics, MPEG4 codec, H.264 decoder and JPEG codec for mobile devices. The unified shader in 3D graphics engine provides fully programmable 3D graphics with 35% area and 28% power reduction. Logarithmic lighting engine and the specialized lighting instruction enable 9.1Mvertices/s vertex throughput. The merged JPEG/MPEG4 codec and the unified shader reduce the silicon area further and the SoC consumes 6.4mm x 6.4mm in 0.13μm CMOS logic process.
Jeong-Ho Woo, Ju-Ho Sohn, Hyejung Kim, Jongcheol Jeong, Euljoo Jeong, Suk Joong Lee, Hoi-Jun Yoo
ISLPED7
2007 Solutions for Real Chip Implementation Issues of NoC and Their Application to Memory-Centric NoC
abstract
This paper describes real chip implementation issues of network-on-chip (NoC) and their solutions along with series of chip design examples. The solutions described in this paper cover both architectural aspects and circuit level techniques for practical chip implementation of NoC. As for architecture level solutions, topology selection, chip-aware protocol design, and on-chip serialization (OCS) for link area reduction are explained. For circuit level techniques, SERDES and synchronizer design, crossbar switch partial activation, and low-voltage link are presented as the foundations for power and area efficient NoC implementation. Regarding presented solutions for NoC implementation, this paper proposes memory centric NoC (MC-NoC) for homogeneous multi processor SoC (MPSoC). Flexibility and feasibility of task mapping on homogeneous SoC is the key feature of the MC-NoC. 8 dual port SRAMs connected to crossbar switches in hierarchical star topology network facilitate data communication between processors, regardless of task mapping into the MC-NoC. Experimental result obtained by mapping edge detection tasks on the MC-NoC in various configurations shows almost constant performance. This result proves the effectiveness of the proposed architecture. The MC-NoC based SoC is also implemented on TSMC 0.18 um process technology
Donghyun Kim 0014, Kwanho Kim, Joo-Young Kim 0001, Seungjin Lee 0001, Hoi-Jun Yoo
NOCS5
2006 A network-on-chip with 3Gbps/wire serialized on-chip interconnect using adaptive control schemes
abstract
An on-chip interconnect is implemented with 3Gbps/wire bandwidth performance with 8:1 serialization scheme. Such high-speed serialization is achieved using a novel serialization scheme, wave-front-train. In order to apply such high-speed link technique to network-on-chip channels, three adaptive control schemes are used: supply voltage dependent reference voltage control, phase compensation scheme with self-calibrating function, and adaptive bandwidth control. The chip is fabricated using 0.18/spl mu/m CMOS technology.
Se-Joong Lee, Kwanho Kim, Hyejung Kim, Namjun Cho, Hoi-Jun Yoo
DATE5
2006 A 372 ps 64-bit adder using fast pull-up logic in 0.18µm CMOS
abstract
This paper presents a 372 ps 64-bit adder using fast pull-up logic (FPL) in 0.18 mum CMOS technology. Fast pull-up logic is devised and applied to decrease pull-up time which is critical in domino-static adder. The implemented adder measures the worst case delay of 372 ps. The adder has a modified tree architecture using load distribution method and has 6 logic stages
Joo-Young Kim 0001, Kangmin Lee, Hoi-Jun Yoo
ISCAS3
2006 A 10µW digital signal processor with adaptive-SNR monitoring for a sub-1V digital hearing aid
abstract
An ultra low-power digital signal processor (DSP) is proposed for the digital hearing aid. The DSP has a SNR monitor to vary its internal clock frequency in accordance with the input signal level. Digital filters use hardwired barrel shifters in place of multipliers, and a parameter ROM provides filter parameters. The clock generator consumes only 1 /spl mu/W at sub-1V. The DSP consumes only 10 /spl mu/W at 0.9-V single supply, and occupies 0.3 mm/sup 2/ with gate count of 10k using 0.18-/spl mu/m CMOS process.
Jerald Yoo, Namjun Cho, Seong-Jun Song, Hoi-Jun Yoo
ISCAS5
2006 Low-power network-on-chip for high-performance SoC design
abstract
An energy-efficient network-on-chip (NoC) is presented for possible application to high-performance system-on-chip (SoC) design. It incorporates heterogeneous intellectual properties (IPs) such as multiple RISCs and SRAMs, a reconfigurable logic array, an off-chip gateway, and a 1.6-GHz phase-locked loop (PLL). Its hierarchically-star-connected on-chip network provides the integrated IPs, which operate at different clock frequencies, with packet-switched serial-communication infrastructure. Various low-power techniques such as low-swing signaling, partially activated crossbar, serial link coding, and clock frequency scaling are devised, and applied to achieve the power-efficient on-chip communications. The 5 /spl times/5 mm/sup 2/ chip containing all the above features is fabricated by 0.18-/spl mu/m CMOS process and successfully measured and demonstrated on a system evaluation board where multimedia applications run. The fabricated chip can deliver 11.2-GB/s aggregated bandwidth at 1.6-GHz signaling frequency. The chip consumes 160 mW and the on-chip network dissipates less than 51 mW.
Kangmin Lee, Se-Joong Lee, Hoi-Jun Yoo
IEEE Trans. Very Large Scale Integr. Syst.3
2004 A low-power graphics LSI integrating 29Mb embedded DRAM for mobile multimedia applications
Ramchan Woo, Sungdae Choi, Ju-Ho Sohn, Seong-Jun Song, Young-Don Bae, Hoi-Jun Yoo
ASP-DAC6
2004 SILENT: serialized low energy transmission coding for on-chip interconnection networks
abstract
On-chip source-synchronous serial communication has many advantages over multi-bit parallel communication in the aspects of skew, crosstalk area cost, wiring difficulty, and clock synchronization. However, the serial wire tends to dissipate more energy than parallel bus due to the bit multiplexing. We propose a coding method to reduce the transmission energy of the serial communication by minimizing the number of transitions on the serial wire. We demonstrate the significant energy saving in a multimedia application, 3D graphics. We also apply the coding technique to a CMOS SoC implementation which integrates various processing units with packet switched on-chip networks.
Kangmin Lee, Se-Joong Lee, Hoi-Jun Yoo
ICCAD3
2002 A practical method to use eDRAM in the shared bus packet switch
abstract
In this paper several methods to use eDRAM (embedded DRAM, on-chip DRAM) In packet switches are analyzed. A practical method using eDRAM as an output queue Is proposed especially in a shared bus packet switch. In the newly proposed output buffer architecture, hierarchical output buffer (HOB), of SRAM plays a role of the small FIFO buffer between a high-speed shared bus and a large eDRAM output buffer. The high density of eDRAM can provide larger capacity than static memories, which results in lower packet loss probability. This paper shows the performance analysis on the proposed HOB switch with the target port speed as 10Gbps for 10 Gigabit Ethernet or OC-192 standards.
Kangmin Lee, Se-Joong Lee, Hoi-Jun Yoo
GLOBECOM3
2001 Single chip 3D rendering engine integrating embedded DRAM frame buffer and Hierarchical Octet Tree (HOT) array processor with bandwidth amplification
abstract
A single chip rendering engine that consist of a DRAM frame buffer, a SRAM serial access memory, pixel/edge processor array and 32b RISC core is proposed for the low power 3D-graphics in portable system. The 56mm2 prototype integrating edge processor, 8 pixel processors, 8 frame buffers and RISC core is fabricated using 0.35um CMOS Embedded Memory Logic (EML) technology.
Yong-Ha Park, Seon-Ho Han, Hoi-Jun Yoo
ASP-DAC3
2000 One chip-low power digital-TCXO with sub-ppm accuracy
abstract
The digital TCXO (DTCXO) has been studied extensively because of its high frequency accuracy and rapid start-up time. The value of compensation capacitance used in the DTCXO is stored in ROM or calculated by computing circuit. In this work, ROM and computing circuit are integrated together to obtain the merit of both schemes; accurate value and high resolution of compensation capacitance, respectively. The DTCXO contains a temperature sensor, A/D converter, controller, EEPROM, capacitor bank, and oscillator. The oscillation frequency can be pulled from /spl plusmn/25 ppm to the required frequency with sub-ppm accuracy. The maximum power consumption of the total chip is 6.6 mW at 3.3 V. The chip, die size of 9mm/sup 2/, is fabricated by a 0.5 /spl mu/m CMOS technology.
Se-Joong Lee, Jinho Han, Seung-Ho Hank, Joe-Ho Lee, Jung-Su Kim, Minkyu Je, Hoi-Jun Yoo
ISCAS7
2000 A 670 ps, 64 bit dynamic low-power adder design
abstract
A 64 bit dynamic low-power adder has been designed and fabricated for 2.5 V 0.25-/spl mu/m 1-poly 5-metal CMOS technology. Fast carry propagation is obtained by fast P generation, parallel quaternary-tree form of group carry (GC) selection and conditional sum selection. The results of proposed adder architecture show that propagation delay, power consumption, and the area are 670 ps, 100 mW, and 0.16 mm/sup 2/, respectively.
Ramchan Woo, Se-Joong Lee, Hoi-Jun Yoo
ISCAS3
1993 A Precision CMOS Voltage Reference with Enhanced Stability for the Application to Advance VLSIs
Hoi-Jun Yoo, Seung-Jun Lee, Jeong-Tae Kwon, Wi-Sik Min, Kye-Hwan Oh
ISCAS1