EDBT 2026 Demo / reviewers in the wild / expert
Hyun Kim 0001
dblp:43/6179-1
· DBLP profile ↗
46ranked-venue papers
4as first author
33since 2021 · last 2026
0000-0002-7962-657XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 2 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 8 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LUT-APP: Dynamic-Precision LUT-based Approximation Unifying Non-Linear Operations in TransformersabstractOn-device transformer inference faces a growing bottleneck in which non-linear functions (e.g., exponential (EXP), reciprocal, reciprocal square root, GeLU, and SiLU) contribute significantly to inference latency as matrix operations become highly optimized. Existing approximation methods either rely on operator-specific datapaths with poor hardware reusability or exhibit a suboptimal accuracy-resource balance with conventional look-up table (LUT)-based piecewise linear approximation (PWL) under stringent edge constraints. This work presents LUT-APP, a unified dynamic-precision LUT-based PWL approximation framework that reconciles accuracy and hardware efficiency across diverse non-linear operators. First, a dynamic fixed-point format (DFF) adaptively allocates bit-width based on input magnitude and parameter scaling to handle the wide dynamic range of EXP. Second, a genetic adaptive differential evolution (GADE) algorithm synthesizes non-uniform PWL segments to minimize approximation error for a given LUT budget. Third, hardware-efficient DFF processing units enable a unified INT8 multiply-add datapath, allowing a single reusable implementation across functions. Experimental results demonstrate that LUT-APP reduces approximation error by up to 6.87× versus state-of-the-art methods while preserving baseline accuracy in large language models and vision transformers without fine-tuning. Hardware synthesis with a 28nm technology shows 4.19× lower area and 3.26× lower power savings than existing LUT-based PWL approaches, validating LUT-APP as a practical, resource-constrained solution for on-device accelerators. We provide the LUT-APP implementation at https://github.com/IDSL-SeoulTech/LUT-APP Seokkyu Yoon, Nam Joon Kim, Hyun Kim 0001 |
DATE | 3 |
| 2026 | SPEED: Structured kernel block pruning with filter groups for efficient and elastic SW-HW co-design in FPGA-based CNN accelerators
Kwanghyun Koo, Sunwoong Kim, Hyun Kim 0001 |
Neurocomputing | 3 |
| 2026 | Accelerating language giants: A survey of optimization strategies for LLM inference on hardware platforms
Young Chan Kim, Seok Kyu Yoon, Soohee Han, Chae Won Park, Jun Oh Park, Jun Ha Ko, Hyun Kim 0001 |
J. Syst. Archit. | 7 |
| 2026 | Towards efficient language giants: A comprehensive survey on structural optimizations and compression techniques for large language models
Gilhyeon Lee, Seonggeun Kim, Kyungmin Goh, Hyun Kim 0001 |
Neural Networks | 5 |
| 2025 | GradQ-ViT: Robust and Efficient Gradient Quantization for Vision TransformersabstractAdvancements in hardware accelerators, such as graphics processing units and neural processing units, have significantly propelled computer vision research. The vision transformer (ViT), leveraging the multi-head self-attention (MHSA) mechanism, has surpassed convolutional neural networks (CNNs) in accuracy but faces challenges in mobile and edge deployment due to its large size and computational demands. In addition, as privacy concerns push for on-device training, research on quantization methods for ViTs, particularly gradient quantization, has gained attention. Unlike CNNs, ViTs face challenges due to outliers and a complex loss landscape. To address this, we propose a gradient quantization framework that stabilizes training by adapting quantization points based on interquartile ranges and constructing an outlier-robust loss function. Additionally, we employ a scaling method to align quantized gradients with original gradients and adaptively assign the learning rate based on quantization error analysis. When quantizing weights, activations, and gradients to INT8, our method improves performance by 0.52% and 0.21% over DeiT-Base and Swin-Base, respectively, and achieves near parity with MobileViT-S with only a 0.09% accuracy drop. Furthermore, a 2.06x speedup was observed when applying our framework to MobileViT in a CUDA 11.8 environment. Dahun Choi, Hyun Kim 0001 |
AAAI | 2 |
| 2025 | LowGradQ: Adaptive Gradient Quantization for Low-Bit CNN Training via Kernel Density Estimation-Guided Thresholding and Hardware-Efficient Stochastic Rounding UnitabstractThis paper proposes a hardware-efficient INT8 training framework with dual-scale adaptive gradient quantization (DAGQ) to cope with the growing need for efficient on-device CNN training. DAGQ captures both small- and large-magnitude gradients, ensuring robust low-bit training with minimal quantization error. Additionally, to reduce the computational and memory demands of stochastic rounding in low-bit training, we introduce a reusable LFSR-based stochastic rounding unit (RLSRU), which efficiently generates and reuses random numbers, minimizing hardware complexity. The proposed framework achieves stable INT8 training across various networks with minimal accuracy loss while being implementable on RTL-based hardware accelerators, making it well-suited for resource-constrained environments. Sangbeom Jeong, Seungil Lee, Hyun Kim 0001 |
DATE | 3 |
| 2025 | TETRIS: On-Device Trainable Energy-Efficient FPGA Accelerator for Trustworthy and Real-Time Instance SegmentationabstractInstance segmentation plays a critical role in high-precision computer vision applications, such as autonomous driving and medical image analysis. As demands for both model accuracy and privacy-preserving solutions continue to rise, on-device training is gaining traction for enabling secure, adaptive learning directly on edge devices. However, the intensive computation and complex dataflows inherent to instance segmentation models pose a major barrier to achieving real-time training in resource-constrained environments. While prior research on on-device training has focused largely on image classification, efficient hardware acceleration for instance segmentation training remains largely unexplored. In this work, we present TETRIS, the first FPGA-based hardware accelerator specifically tailored for training instance segmentation models. TETRIS introduces structural and quantization-aware optimizations, including channel-aware smoothing quantization and distribution-aware tie quantization, to significantly reduce both computational and memory overhead. Moreover, TETRIS adopts a heterogeneous hardware architecture, incorporating a reconfigurable convolution processing unit (RC-PU) that supports variable kernel sizes and a fully pipelined auxiliary processing unit to handle specialized operations efficiently. Experimental evaluation on YOLACT-based instance segmentation tasks demonstrates that TETRIS achieves a peak performance of 805.9 GOPS and an energy efficiency of 69.5 GOPS/W, confirming its ability to support real-time training even in resource-constrained edge environments. Seungil Lee, Juntae Park, Kwanghyun Koo, Gilha Lee, Sangbeom Jeong, Junoh Park, Hyun Kim 0001 |
ICCAD | 7 |
| 2025 | PriME: PIM-Aware Efficient Compression for Memory-Bound Embedding Layers in sLLMsabstractWith the growing demand for on-device AI, increasing efforts have been directed toward deploying lightweight small-scale large language models (sLLMs) on edge and mobile devices to enhance inference performance while minimizing computational cost and latency. As the number of decoder layers in sLLMs decreases, the embedding layer constitutes a substantial portion of the model's overall parameters and memory consumption. Consequently, efficient data compression is crucial; however, existing methods, such as quantization and pruning, drastically degrade accuracy when applied to the error-sensitive embedding layers. Moreover, embedding layer computations exhibit low arithmetic intensity (operations per byte), rendering them memory-bound. This limitation necessitates a shift from conventional von Neumann architectures to processing-in-memory (PIM) architectures. To address these challenges, this paper proposes (1) XOR-based Masking Compression (XMC), a lossless compression algorithm specialized for sLLM embedding layers, and (2) PriME, which integrates XMC with PIM architecture to alleviate memory bottlenecks in embedding layers. XMC enhances zero-bit representation in 16-bit FP data using ADD and XOR masking, achieving an average compression ratio of$1.49 \times$while being implementable with a 3-cycle decompression delay. PriME enables parallel processing of compressed data within PIM, accelerating embedding computations by an average of$4.0 \times$and up to$5.26 \times$compared to GPUs, while simultaneously reducing energy consumption by over 30 %, leading to an average energy efficiency improvement of$6.29 \times$. Designed for broad applicability, PriME is compatible with various sLLMs and holds scalability for extension to multimodal small vision language models, demonstrating its versatility for efficient AI acceleration. Junghyeok Lee, Jihoon Jang 0001, Hyun Kim 0001 |
ICCD | 3 |
| 2025 | LRA-QViT: Integrating Low-Rank Approximation and Quantization for Robust and Efficient Vision TransformersabstractRecently, transformer-based models have demonstrated state-of-the-art performance across various computer vision tasks, including image classification, detection, and segmentation. However, their substantial parameter count poses significant challenges for deployment in resource-constrained environments such as edge or mobile devices. Low-rank approximation (LRA) has emerged as a promising model compression technique, effectively reducing the number of parameters in transformer models by decomposing high-dimensional weight matrices into low-rank representations. Nevertheless, matrix decomposition inherently introduces information loss, often leading to a decline in model accuracy. Furthermore, existing studies on LRA largely overlook the quantization process, which is a critical step in deploying practical vision transformer (ViT) models. To address these challenges, we propose a robust LRA framework that preserves weight information after matrix decomposition and incorporates quantization tailored to LRA characteristics. First, we introduce a reparameterizable branch-based low-rank approximation (RB-LRA) method coupled with weight reconstruction to minimize information loss during matrix decomposition. Subsequently, we enhance model accuracy by integrating RB-LRA with knowledge distillation techniques. Lastly, we present an LRA-aware quantization method designed to mitigate the large outliers generated by LRA, thereby improving the robustness of the quantized model. To validate the effectiveness of our approach, we conducted extensive experiments on the ImageNet dataset using various ViT-based models. Notably, the Swin-B model with RB-LRA achieved a 31.8\% reduction in parameters and a 30.4\% reduction in GFLOPs, with only a 0.03\% drop in accuracy. Furthermore, incorporating the proposed LRA-aware quantization method reduced accuracy loss by an additional 0.83\% compared to naive quantization. Beom Jin Kang, Nam Joon Kim, Hyun Kim 0001 |
ICML | 3 |
| 2025 | XNC: XOR and NOT-Based Lossless Compression for Optimizing Unquantized Embedding Layers in Large Language ModelsabstractAlthough 4-bit quantized small LLMs have been proposed recently, many studies have retained FP16 precision for embedding layers, as they constitute a relatively small proportion of the overall model in existing LLMs and suffer from severe accuracy degradation when quantized. However, in quantized small LLMs, the embedding layer accounts for a substantial proportion of the total model parameters, necessitating its compression. Since embedding layers are sensitive to approximation, lossless compression is more desirable than lossy compression methods such as quantization. While existing lossless compression methods efficiently compress patterns such as zeros, narrow values, or frequently occurring values, embedding layers typically lack these patterns, making effective compression more challenging. In this paper, we propose XOR and NOT-based lossless compression (XNC), which applies XOR operations between adjacent 16-bit blocks and then performs a NOT operation on the result, effectively truncating the upper and lower bits to compress the embedding layer to 9-bit without any loss. The proposed method leverages XOR and NOT operations, enabling easy hardware implementation, with only four cycles required for compression and three cycles for decompression, ensuring efficient data compression without performance degradation. As a result, the proposed compression technique achieves an average compression ratio of 1.34× for the embedding layers of small LLMs without any loss, effectively reducing the model size of 4-bit quantized LLMs by an average of 9.91%. The code is available at https://github.com/IDSL-SeoulTech/XNC. Junghyeok Lee, Jihoon Jang 0001, Hyun Kim 0001 |
ISCAS | 3 |
| 2025 | PIM-BEACON: A Benchmarking and Emulation Framework Supporting Adaptive CONfigurations in DRAM-Based Processing-in-Memory SystemsabstractThe growing demand for data storage and memory bandwidth in large-scale deep neural networks has exacerbated the data movement bottleneck in traditional von Neumann architectures. Processing-in-memory (PIM) technology addresses this challenge by integrating computation within memory, reducing data transfer overhead. Prior research has predominantly relied on in-house PIM simulators for evaluation. However, these simulators often exhibit limited versatility and slow evaluation runtimes, constraining their effectiveness for comprehensive design space exploration. We present PIM-BEACON, a highspeed trace-based PIM emulation platform that ensures high reliability, fidelity, and versatility to address these limitations. PIM-BEACON adopts a modular design by configuring the PIM controller and emulator regions separately and employs SystemVerilog for FPGA-based high-fidelity implementation. To improve evaluation runtime, we introduce an efficient BRAM management mechanism that maximizes FPGA resource utilization. Supporting a wide range of DRAM-based PIM architectures, PIM-BEACON achieves up to 133.61 × faster runtime with only a 1.57 % average performance cycle error rate. Inseong Hwang, Jihoon Jang 0001, Hyun Kim 0001 |
ISPASS | 4 |
| 2025 | HyDRASim: A Versatile and Cycle-Accurate Simulator for Hybrid DRAM PIM-CPU SystemsabstractRecent advancements in deep neural networks (DNNs) have exacerbated the data movement bottleneck inherent in traditional von Neumann architectures. Processing-in-memory (PIM) has emerged as a promising paradigm by integrating computation within memory to alleviate this challenge. However, the lack of versatile and cycle-accurate simulation frameworks significantly limits the evaluation and optimization of diverse PIM designs. Existing in-house PIM simulators tend to be narrowly tailored to specific architectures, lacking generality and extensibility across hybrid computing environments. In this paper, we introduce HyDRASim, a cycle-accurate and extensible PIM-CPU simulator designed to support a wide range of PIM architectural features. HyDRASim integrates ZSim for detailed CPU modeling with DRAMSim3 for accurate memory modeling to enable flexible simulation across CPU-only, PIM-only, and PIM-CPU hybrid configurations. Validation experiments using GEMV workloads of varying sizes demonstrate that HyDRASim achieves cycle-level fidelity, with average latency and power errors of just $5.92 \%$ and $5.61 \%$, respectively, compared to established in-house baselines. Furthermore, system-level performance evaluations with diverse DNN workloads reveal that transformer-based models achieve $2.21 \times$ greater acceleration compared to convolutional neural networks under hybrid configurations modeled with HyDRASim. These results establish HyDRASim as a reliable and powerful tool for accurately modeling, evaluating, and optimizing emerging PIM-CPU hybrid architectures, providing critical insights into the future of memory-centric system design. We open-source HyDRASim at https://github.com/IDSL-SeoulTech/HyDRASim Jihoon Jang 0001, Inseong Hwang, Hyun Kim 0001 |
MASCOTS | 3 |
| 2025 | VFT: A versatile fine-tuning scheme based on feature distribution-aware knowledge distillation for lightweight convolutional neural networksabstractVarious network compression techniques , such as pruning and quantization, are being actively researched in order to lighten convolutional neural networks (CNNs), which have increasingly deep and complex structures accompanied by the achievement of higher accuracy. Since most of these network compression techniques cause a decrease in accuracy, fine-tuning is essential to recover the performance of lightweight models; however, fine-tuning has received limited research attention compared to numerous compression techniques , and thus, performance recovery by fine-tuning has significant room for improvement. In this paper, we analyze the shortcomings of existing fine-tuning methods in terms of loss landscape and introduce a knowledge distillation (KD)-based fine-tuning approach that solves these problems. In particular, to overcome the limitation that KD can be adversely affected by the capacity difference between the teacher and student models or the defined knowledge to be transferred, we propose a feature distribution-aware knowledge distillation (FDKD) method, which defines appropriate supervision in the form of feature distribution to transfer the semantic information from teacher models. Moreover, we also propose a layer-wise FDKD method by exploiting the uniqueness of the lightweight model that the baseline ( i.e. , teacher) and compressed models ( i.e. , student) have the same architecture. Experiments on classification tasks demonstrate the superiority of the proposed method over existing fine-tuning methods, achieving up to 1.99% and 3.83% of accuracy improvement for pruned and quantized models, respectively. The source code for this implementation is available at [ https://github.com/IDSL-SeoulTech/VFT ]. Hyeonseok Hong, Hyun Kim 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | VFF-Net: Evolving forward-forward algorithms into convolutional neural networks for enhanced computational insightsabstractIn recent years, significant efforts have been made to overcome the limitations inherent in the traditional back-propagation (BP) algorithm. These limitations include overfitting, vanishing/exploding gradients, slow convergence, and black-box nature. To address these limitations, alternatives to BP have been explored, the most well-known of which is the forward-forward network (FFN). We propose a visual forward-forward network (VFF-Net) that significantly improves FFNs for deeper networks, focusing on enhancing performance in convolutional neural network (CNN) training. VFF-Net utilizes a label-wise noise labeling method and cosine-similarity-based contrastive loss, which directly uses intermediate features to solve both the input information loss problem and the performance drop problem caused by the goodness function when applied to CNNs. Furthermore, VFF-Net is accompanied by layer grouping, which groups layers with the same output channel for application in well-known existing CNN-based models; this reduces the number of minima that need to be optimized and facilitates the transfer to CNN-based models by demonstrating the effects of ensemble training. VFF-Net improves the test error by up to 8.31% and 3.80% on a model consisting of four convolutional layers compared with the FFN model targeting a conventional CNN on CIFAR-10 and CIFAR-100, respectively. Furthermore, the fully connected layer-based VFF-Net achieved a test error of 1.70% on the MNIST dataset, which is better than that of the existing BP. In conclusion, the proposed VFF-Net significantly reduces the performance gap with BP by improving the FFN and shows the flexibility to be portable to existing CNN-based models. Gilha Lee, Jin Shin, Hyun Kim 0001 |
Neural Networks | 3 |
| 2025 | NPC: A Non-Conflicting Processing-in-Memory Controller in DDR Memory SystemsabstractProcessing-in-Memory (PIM) has emerged as a promising solution to address the memory wall problem. Existing memory interfaces must support new PIM commands to utilize PIM, making the definition of PIM commands according to memory modes a major issue in the development of practical PIM products. For performance and OS-transparency, the memory controller is responsible for changing the memory mode, which requires modifying the controller and resolving conflicts with existing functionalities. Additionally, it must operate to minimize mode transition overhead, which can cause significant performance degradation. In this study, we present NPC, a memory controller designed for mode transition PIM that delivers PIM commands via the DDR interface. NPC issues PIM commands while transparently changing the memory mode with a dedicated scheduling policy that reduces the number of mode transitions with aggregative issuing. Moreover, existing functions, such as refresh, are optimized for PIM operation. We implement NPC in hardware and develop a PIM emulation system to validate it on FPGA platforms. Experimental results reveal that NPC is compatible with existing interfaces and functionality, and the proposed scheduling policy improves performance by 2.2$\boldsymbol{\times}$with balanced fairness, achieving up to 97% of the ideal performance. These findings have the potential to aid the application of PIM in real systems and contribute to the commercialization of mode transition PIM. Seungyong Lee 0003, Chunmyung Park, Woojae Shin, Hyun Kim 0001 |
IEEE Trans. Computers | 7 |
| 2025 | FAB: FPGA-Accelerated Fully-Pipelined Bottleneck Architecture With Batching for High-Performance MobileNetv2 InferenceabstractLightweight neural networks (LWNNs) primarily employ the bottleneck block (BB) introduced in MobileNetv2 or similar architectural structures. However, the channel expansion-reduction process in BB imposes substantial activation memory overhead, a challenge that has not been adequately addressed in prior studies on LWNN accelerators incorporating BB. To overcome this limitation, we propose a fully-pipelined bottleneck architecture (FPB) optimized for the efficient hardware deployment of BB. FPB eliminates the need for intermediate off-chip memory access, effectively addressing deployment challenges associated with BB and enabling an end-to-end accelerator architecture. To enhance hardware efficiency, each FPB core utilizes 2-LUT DSP, Fused-ReLU6, and Q-Residual, optimizing computational performance while minimizing resource consumption. Furthermore, we introduce a batching technique that maximizes the benefits of FPB by ensuring high hardware utilization across FPB cores while enabling the concurrent processing of multiple images. To mitigate the off-chip memory access latency inherently incurred by batching, we propose a stem layer latency hiding technique, which effectively prevents performance degradation. We evaluate the performance of our proposed MobileNetv2 accelerator on the VCU118 board, achieving an energy efficiency of 120.7 GOPS/W at a batch size of 4. This represents an improvement of$1.5\times $to$10.5\times $over prior work. Depending on the batch size configuration, our FAB accelerator achieves a throughput performance ranging from 204.2 GOPS to 772.7 GOPS, demonstrating its high computational efficiency. Young Chan Kim, Nam Joon Kim, Hyun Kim 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | MCM-SR: Multiple Constant Multiplication-Based CNN Streaming Hardware Architecture for Super-ResolutionabstractConvolutional neural network (CNN)-based super-resolution (SR) methods have become prevalent in display devices due to their superior image quality. However, the significant computational demands of CNN-based SR require hardware accelerators for real-time processing. Among the hardware architectures, the streaming architecture can significantly reduce latency and power consumption by minimizing external dynamic random access memory (DRAM) access. Nevertheless, this architecture necessitates a considerable hardware area, as each layer needs a dedicated processing engine. Furthermore, achieving high hardware utilization in this architecture requires substantial design expertise. In this article, we propose methods to reduce the hardware resources of CNN-based SR accelerators by applying the multiple constant multiplication (MCM) algorithm. We propose a loop interchange method for the convolution (CONV) operation to reduce the logic area by 23% and an adaptive loop interchange method for each layer that considers both the static random access memory (SRAM) and logic area simultaneously to reduce the SRAM size by 15%. In addition, we improve the MCM graph exploration speed by$5.4\times $while maintaining the SR quality through beam search when CONV weights are approximated to reduce the hardware resources. Seung-Hwan Bae, Hyun Kim 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | HyQ: Hardware-Friendly Post-Training Quantization for CNN-Transformer Hybrid Networks
Nam Joon Kim, Hyun Kim 0001 |
IJCAI | 3 |
| 2024 | ARC: Adaptive Rounding and Clipping Considering Gradient Distribution for Deep Convolutional Neural Network TrainingabstractIn convolution neural networks (CNNs), quantization is an effective compression method that can conserve hardware resources in convolution operations, which account for the majority of computations, with lower bits. Most quantization studies focused on weight and activation parameters. However, gradient quantization, which is the core of quantization research for CNN training, has a significant impact on network training even with small changes in gradient, making it difficult to achieve excellent accuracy. Although previous works for gradient quantization achieved high accuracy based on stochastic rounding (SR), there are issues of high latency in generating random numbers and difficulty with register transfer level (RTL)-based hardware design. Additionally, the search for a clipping value based on quantization error is effective in the initial training but becomes inadequate after model convergence as the quantization error decreases. In this paper, we address the limitations of SR through an approach based on deterministic rounding, specifically rounding toward zero (RTZ). Additionally, to determine the outliers in a distribution, we search for a suitable clipping value based on the z-score, that would be appropriate even if the network converges. Experimental results show that the proposed method achieves higher accuracy in various vision tasks, such as ResNet, YOLOv5, and YOLACT, and offers robust compatibility. The proposed quantizer was verified through RTL, achieving an accuracy similar to that of SR using resources comparable to nearest rounding. Moreover, when measuring latency on the CPU, the proposed method achieved 43% less latency than SR. Dahun Choi, Hyun Kim 0001 |
ISCAS | 2 |
| 2024 | EDeN: Enabling Low-Power CNN Inference on Edge Devices Using Prefetcher-assisted NVM SystemsabstractThe accuracy of convolutional neural networks (CNNs) has significantly improved over the years. Meanwhile, due to the high portability and usefulness of edge devices, the demand for artificial intelligence (AI) based applications on edge computing devices has been soaring recently. Accordingly, CNN inference has become one of the mainstream AI applications on edge devices. However, the continually increasing leakage power of edge devices drags down the wide deployment of CNN inference applications, as the technology node scales down. Jihoon Jang 0001, Hyokeun Lee, Hyun Kim 0001 |
ISLPED | 3 |
| 2024 | A survey of FPGA and ASIC designs for transformer inference acceleration and optimization
Beom Jin Kang, Haein Lee, Seok Kyu Yoon, Young Chan Kim, Sangbeom Jeong, Seong Jun O, Hyun Kim 0001 |
J. Syst. Archit. | 7 |
| 2024 | Vision transformer models for mobile/edge devices: a survey
Seung Il Lee, Kwanghyun Koo, Gilha Lee, Sangbeom Jeong, Seongjun O, Hyun Kim 0001 |
Multim. Syst. | 7 |
| 2024 | USD: Uncertainty-Based One-Phase Learning to Enhance Pseudo-Label Reliability for Semi-Supervised Object DetectionabstractWith the ease of accessing large unlabeled datasets, studies on semi-supervised learning for object detection (SSOD) have become increasingly popular. Among these SSOD studies, the pseudo-labeling method significantly depends on the accuracy of the pseudo-labels; thus, inaccurate annotations must be filtered to prevent performance degradation. This study classifies annotation errors that occur in pseudo-labeling methods as false negative (FN) and false positive (FP), and solutions to address each type of error are proposed using uncertainty information obtained through Gaussian modeling. Network performance is improved by preventing the background learning of the FN objects based on the uncertainty of the network output. In addition, based on the uncertainty of the annotations, low-reliability annotations are filtered out, and the learning reflectivity of FP objects is determined. Considering the network performance improvement and training complexity, the proposed method employs one-phase learning, including a single pseudo-label update, to achieve maximum performance with the minimum learning process. Moreover, an algorithm is proposed for an optimal update point search to increase the expected performance improvement. Experiments on the Pascal VOC, COCO, and Cityscapes datasets show that the SSD network improves accuracy by 3.3%, 4.7%, and 4.1%, respectively, with negligible computational complexity compared to the baseline. Dayoung Chun, Seungil Lee, Hyun Kim 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Trunk Pruning: Highly Compatible Channel Pruning for Convolutional Neural Networks Without Fine-TuningabstractChannel pruning can efficiently reduce the computation and memory footprint within a reasonable accuracy drop by removing unnecessary channels from convolutional neural networks (CNNs). Among the various channel pruning approaches, sparsity training is the most popular because of its convenient implementation and end-to-end training. It automatically identifies the optimal network structures by applying regularization to parameters. Although this sparsity training has achieved a remarkable performance in terms of the trade-off between accuracy and network size reduction, it needs to be accompanied by a time-consuming fine-tuning process. Moreover, although activation functions with high performance are being continuously developed, the existing sparsity training does not display remarkable scalability for these new activation functions. To address these problems, this study proposes a novel pruning method,trunk pruning, which can produce a compact network by minimizing the accuracy drop during inference even without the fine-tuning process. In the proposed method, one kernel of the next convolutional layer absorbs all the information of the kernels to be pruned, considering the effects of the batch normalization (BN) shift parameters remaining after the sparsity training. Therefore, it is possible to eliminate the fine-tuning process because trunk pruning can effectively reproduce the output of the unpruned network after the sparsity training by removing the pruning loss. Furthermore, because trunk pruning is a technique that can effectively control only the shift parameters of the BN in the CONV layer, it has the significant advantage of being compatible with all BN-based sparsity training schemes and can address various activation functions. Nam Joon Kim, Hyun Kim 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | FieldHAR: A Fully Integrated End-to-End RTL Framework for Human Activity Recognition with Neural Networks from Heterogeneous SensorsabstractIn this work, we propose an open-source scalable end-to-end RTL framework FieldHAR, for complex human activ-ity recognition (HAR) from heterogeneous sensors using artificial neural networks (ANN) optimized for FPGA or ASIC integration. FieldHAR aims to address the lack of apparatus to transform complex HAR methodologies often limited to offline evaluation to efficient runtime edge applications. The framework uses parallel sensor interfaces and integer-based multi-branch convolutional neural networks (CNNs) to support flexible modality extensions with synchronous sampling at the maximum rate of each sensor. To validate the framework, we used a sensor-rich kitchen scenario HAR application which was demonstrated in a previous offline study. Through resource-aware optimizations, with FieldHAR the entire RTL solution was created from data acquisition to ANN inference taking as low as 25% logic elements and 2% memory bits of a low-end Cyclone IV FPGA and less than 1% accuracy loss from the original FP32 precision offline study. The RTL implementation also shows advantages over MCU-based solutions, including superior data acquisition performance and virtually eliminating ANN inference bottleneck. Mengxi Liu 0004, Bo Zhou 0005, Zimin Zhao, Hyeonseok Hong, Hyun Kim 0001, Sungho Suh, Vitor F. Rey, Paul Lukowicz |
ASAP | 5 |
| 2023 | RepSGD: Channel Pruning Using Reparamerization for Accelerating Convolutional Neural NetworksabstractChannel pruning is a popular method for compressing convolutional neural networks (CNNs) while maintaining acceptable accuracy. Most existing channel pruning methods use the approach of zeroing unnecessary filters and then removing them. To address the limitations of existing approaches, methods of creating forcibly filter redundancy and then removing redundant filters have been proposed without heuristic knowledge. However, these methods also use a deformed gradient to make filters identical, and performance degradation is inevitable because the parameters cannot be updated using the original gradients. To solve these problems, this study proposes RepSGD, which can compress CNNs simply and efficiently. RepSGD inserts a new point-wise convolution layer after the existing standard convolution layer. Subsequently, only new point-wise convolution layers are trained to produce filter redundancy (i.e., to make the filters identical), whereas the standard convolution layers are trained using the original gradient. After training, RepSGD merges two consecutive convolution layers into one convolution layer. Subsequently, the redundant filters in the merged convolution layer are pruned. Because RepSGD does not change the original architecture of the CNN, additional inference computation is not required, and it is possible to support training from scratch. In addition, using the original gradient in RepSGD optimizes the objective function of the CNNs better. We show that RepSGD outperforms existing pruning methods in various models and datasets through extensive experiments. Nam Joon Kim, Hyun Kim 0001 |
ISCAS | 2 |
| 2023 | An In-Module Disturbance Barrier for Mitigating Write Disturbance in Phase-Change MemoryabstractWrite disturbance error (WDE) appears as a serious reliability problem preventing phase-change memory (PCM) from general commercialization, and therefore several studies have been proposed to mitigate WDEs. Verify-and-correction (VnC) eliminates WDEs by always verifying the data correctness on neighbors after programming, but incurs significant performance overhead. Encoding-based schemes mitigate WDEs by reducing the number of WDE-vulnerable data patterns; however, mitigation performance notably fluctuates with applications. Moreover, encoding-based schemes still rely on VnC-based schemes. Cache-based schemes lower WDEs by storing data in a write cache, but it requires several megabytes of SRAM to significantly mitigate WDEs. Despite the efforts of previous studies, these methods incur either significant performance or area overhead. Therefore, a new approach, which does not rely on VnC-based schemes or application data patterns, is highly necessary. Furthermore, the new approach should be transparent to processors (i.e., in-module), because the characteristic of WDEs is determined by manufacturers of PCM products. In this paper, we present an in-module disturbance barrier (IMDB) that mitigates WDEs on demand. IMDB includes a two-level hierarchy comprising two SRAM-based tables, whose entries are managed with a dedicated replacement policy that sufficiently utilizes the characteristics of WDEs. The naive implementation of the replacement policy requires hundreds of read ports on SRAM, which is infeasible in real hardware; hence, an approximate comparator is also designed. We also conduct a rigorous exploration of architecture parameters to obtain a cost-effective design. The proposed method significantly reduces WDEs without noticeable speed degradation or additional energy consumption compared to previous methods. Hyokeun Lee, Seungyong Lee 0003, Byeongki Song, Moonsoo Kim, Seokbo Shim, Hyun Kim 0001 |
IEEE Trans. Computers | 7 |
| 2023 | FP-AGL: Filter Pruning With Adaptive Gradient Learning for Accelerating Deep Convolutional Neural NetworksabstractFilter pruning is a technique that reduces computational complexity, inference time, and memory footprint by removing unnecessary filters in convolutional neural networks (CNNs) with an acceptable drop in accuracy, consequently accelerating the network. Unlike traditional filter pruning methods utilizing zeroing-out filters, we propose two techniques to achieve the effect of pruning more filters with less performance degradation, inspired by the existing research on centripetal stochastic gradient descent (C-SGD), wherein the filters are removed only when the ones that need to be pruned have the same value. First, to minimize the negative effect of centripetal vectors that gradually make filters come closer to each other, we redesign the vectors by considering the effect of each vector on the loss-function using the Taylor-based method. Second, we propose an adaptive gradient learning (AGL) technique that updates weights while adaptively changing the gradients. Through AGL, performance degradation can be mitigated because some gradients maintain their original direction, and AGL also minimizes the accuracy loss by perfectly converging the filters, which require pruning, to a single point. Finally, we demonstrate the superiority of the proposed method on various datasets and networks. In particular, on the ILSVRC-2012 dataset, our method removed 52.09% FLOPs with a negligible 0.15% top-1 accuracy drop on ResNet-50. As a result, we achieve the most outstanding performance compared to those reported in previous studies in terms of the trade-off between accuracy and computational complexity. Nam Joon Kim, Hyun Kim 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | GaussianMask: Uncertainty-aware Instance Segmentation based on Gaussian ModelingabstractInstance segmentation, which has been required in various applications in recent years, is aimed at reliable bounding box (bbox) detection (i.e., localization) and stable mask prediction (i.e., segmentation). However, the mask uncertainty problem is still unresolved, which hinders the ability to achieve an accurate instance segmentation. In this paper, we propose GaussianMask, an uncertainty-aware instance segmentation technique based on Gaussian modeling. We can determine the uncertainty of the network through the variance of masks extracted by redesigning the loss function based on Gaussian modeling, and the mask accuracy can be improved by constructing a robust model that adaptively applies this uncertainty to the network. In particular, the additional computations caused during this process are insignificant, and thus negligible for the processing speed of the network. Moreover, GaussianMask has the advantage of being applicable to any network due to its high compatibility. Experimental results show that the mask average precision (AP) of the representative instance segmentation models, YOLACT and Mask R-CNN, increases from 29.7% to 31.8% and from 35.7% to 37.5%, respectively, on the MS-COCO dataset. Seung Il Lee, Hyun Kim 0001 |
ICPR | 2 |
| 2022 | PCMCsim: An Accurate Phase-Change Memory Controller Simulator and its Performance AnalysisabstractWith the growing demand for technology scaling and storage capacity in data centers, phase-change memory (PCM) has garnered attention as a next-generation nonvolatile memory (NVM). However, an accurate simulator that includes the necessary hardware features for PCM is not available, lagging behind current PCM technology. In this study, a functional and cycle-accurate PCM controller simulator, called PCMCsim, is presented to revitalize the related research. The proposed simulator incorporates necessary features for current PCM products and the latest DDR5 specifications. Based on rigorous performance analysis, this study characterizes bottlenecks of the PCM subsystem by sweeping hardware parameters, providing important takeaway messages to designers. Furthermore, the latency is significantly reduced by introducing a dedicated prefetcher into the address translation module. The proposed simulator is validated against a command trace made by a PCM product developer. We release our simulator as open-source software, except for industry-confidential features.11https://github.com/harrylee365/pcmcsim-pub1ic Hyokeun Lee, Hyungsuk Kim, Seokbo Shim, Seungyong Lee 0003, Do-sun Hong, Hyun Kim 0001 |
ISPASS | 7 |
| 2021 | DC-AC: Deep Correlation-Based Adaptive Compression of Feature Map Planes in Convolutional Neural NetworksabstractDeep learning has been successfully deployed to a broad range of applications with its outstanding performance. Supporting an efficient hardware architecture is critical to making effective use of a deep learning approach with proven algorithm performance. One challenge in implementation of deep learning algorithm is to reduce memory bandwidth because a single memory access normally consumes 100* more energy than an arithmetic operation. To reduce the memory bandwidth, deep learning data could be compressed and decompressed before memory write/read operations. Especially, feature maps, which account for a significant portion of the convolutional neural network (CNN), could be compressed further by reducing the correlations between feature map planes. This paper proposes a compression method for feature maps in CNN that adaptively exploits the varying correlation between feature map planes. For every feature map plane, the proposed method searches the most similar plane among nearby planes in the same layer, and compresses the residual of the two planes instead of the plane itself. Experimental results show that the average bit length to store feature maps is reduced by 14.2% compared to the compression without correlation reduction, and the CNN accuracy does not change and additional training is also not required because the proposed method applies lossless compression. Seung-Hwan Bae, Hyun Kim 0001 |
ISCAS | 3 |
| 2021 | Cache Compression with Golomb-Rice Code and Quantization for Convolutional Neural NetworksabstractCache compression schemes reduce the cache miss rate by increasing the effective cache capacity and consequently, reduce memory access and power consumption. Therefore, cache compression is beneficial for applications with heavy memory traffic, including convolutional neural network (CNN). In this paper, a new cache compression of a floating-point number is proposed for CNNs. The exponent is compressed using the Golomb-Rice code, instead of the Huffman code, for an efficient hardware implementation. The compression syntax is carefully designed so that the size of compressed data is not very far from the entropy, which is the theoretical limit, by distinguishing two different types of data used in CNNs. On the other hand, since the mantissa of CNNs data can be hardly compressed by entropy coding, it is simply quantized for data reduction that may not degrade the CNN performance significantly thanks to the error robustness of CNNs. The quantization reduces 23 bits of a mantissa to 4 bits. The experimental results show that the miss rate of a 1 MB compressed cache with the proposed compression method applied is almost similar to that of an uncompressed 2 MB cache without any decrease of the CNN accuracy. Seung-Hwan Bae, Hyun Kim 0001 |
ISCAS | 3 |
| 2021 | Layer-Specific Optimization for Mixed Data Flow With Mixed Precision in FPGA Design for CNN-Based Object DetectorsabstractConvolutional neural networks (CNNs) require both intensive computation and frequent memory access, which lead to a low processing speed and large power dissipation. Although the characteristics of the different layers in a CNN are frequently quite different, previous hardware designs have employed common optimization schemes for them. This paper proposes a layer-specific design that employs different organizations that are optimized for the different layers. The proposed design employs two layer-specific optimizations: layer-specific mixed data flow and layer-specific mixed precision. The mixed data flow aims to minimize the off-chip access while demanding a minimal on-chip memory (BRAM) resource of an FPGA device. The mixed precision quantization is to achieve both a lossless accuracy and an aggressive model compression, thereby further reducing the off-chip access. A Bayesian optimization approach is used to select the best sparsity for each layer, achieving the best trade-off between the accuracy and compression. This mixing scheme allows the entire network model to be stored in BRAMs of the FPGA to aggressively reduce the off-chip access, and thereby achieves a significant performance enhancement. The model size is reduced by 22.66-28.93 times compared to that in a full-precision network with a negligible degradation of accuracy on VOC, COCO, and ImageNet datasets. Furthermore, the combination of mixed dataflow and mixed precision significantly outperforms the previous works in terms of both throughput, off-chip access, and on-chip memory requirement. Duy Thanh Nguyen, Hyun Kim 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | PCM: Precision-Controlled Memory System for Energy Efficient Deep Neural Network TrainingabstractDeep neural network (DNN) training suffers from the significant energy consumption in memory system, and most existing energy reduction techniques for memory system have focused on introducing low precision that is compatible with computing unit (e.g., FP16, FP8). These researches have shown that even in learning the networks with FP16 data precision, it is possible to provide training accuracy as good as FP32, de facto standard of the DNN training. However, our extensive experiments show that we can further reduce the data precision while maintaining the training accuracy of DNNs, which can be obtained by truncating some least significant bits (LSBs) of FP16, named as hard approximation. Nevertheless, the existing hard-ware structures for DNN training cannot efficiently support such low precision. In this work, we propose a novel memory system architecture for GPUs, named as precision-controlled memory system (PCM), which allows for flexible management at the level of hard approximation. PCM provides high DRAM bandwidth by distributing each precision to different channels with as transposed data mapping on DRAM. In addition, PCM supports fine-grained hard approximation in the L1 data cache using software-controlled registers, which can reduce data movement and thereby improve energy saving and system performance. Furthermore, PCM facilitates the reduction of data maintenance energy, which accounts for a considerable portion of memory energy consumption, by controlling refresh period of DRAM. The experimental results show that in training CIFAR-100 dataset on Resnet-20 with precision tuning, PCM achieves energy saving and performance enhancement by 66% and 20%, respectively, without loss of accuracy. Boyeal Kim, Hyun Kim 0001, Duy Thanh Nguyen, Minh-Son Le, Ik Joon Chang, Dohun Kwon, Jin Hyeok Yoo |
DATE | 3 |
| 2020 | An Efficient Sampling Algorithm With a K-NN Expanding Operator for Depth Data Acquisition in a LiDAR SystemabstractThe spatial resolution of a depth-acquisition device, such as a Light Detection and Ranging (LiDAR) sensor, is limited because of the slow acquisition. To accurately reconstruct a depth image from limited spatial resolution, a two-stage sampling process has been widely used. However, two-stage sampling uses an irregular sampling pattern for the sampling operation, which requires complex computation for reconstruction and additional memory space for storage. A mathematical formulation of a LiDAR system demonstrates that two-stage sampling does not satisfy its timing constraint for practical use. To overcome the drawbacks of two-stage sampling, this paper proposes a new sampling method that reduces the computational complexity and memory requirements by generating the optimal representatives of a sampling pattern in down-sample data. A sampling pattern can be derived from a k-NN expanding operation from the downsampled representatives. The proposed algorithm is designed to preserve the object boundary by restricting the expansionoperation only to the object boundary or complex texture. In addition, the proposed algorithm runs in linear-time complexity and reduces the memory requirements using a down-sampling ratio. The experimental results demonstrate that the proposed sampling outperforms grid sampling by at most 7.92 dB. Consequently, the proposed sampling achieves reconstructed quality similar to that of optimal sampling, while substantially reducing the computation time and memory requirements. Xuan Truong Nguyen, Hyun Kim 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | A Low-Cost and High-Throughput FPGA Implementation of the Retinex Algorithm for Real-Time Video EnhancementabstractFor video applications in a special environment such as medical imaging, space exploration, and underwater exploration, the video captured by an image sensor is often deteriorated because of low lighting conditions. Therefore, it is necessary to enhance the part of the image that is too dark to distinguish details while maintaining the remaining part with the same brightness. The retinex algorithm is widely used to restore naturalness of a video, especially exhibiting outstanding performance in the enhancement of a dark area. However, it demands large computational complexity because of its intricate structure, such as the Gaussian filter and exponentiation operations, and consequently, it is difficult to process in real time. This article presents a low-cost and high-throughput design of the retinex video enhancement algorithm. The hardware (HW) design is implemented using a field-programmable gate array (FPGA), and it supports a throughput of 60 frames/s for a$1920 \times 1080$image with negligible latency. The proposed FPGA design minimizes HW resources while maintaining the quality and the performance by using a small line buffer instead of a frame buffer, by applying the concept of approximate computing for the complex Gaussian filter, and by designing a new and nontrivial exponentiation operation. The proposed design makes it possible to significantly reduce HW resources (up to 79.22% of total resources) compared to existing systems and is compatible with commercialized devices through the standard HDMI/DVI video ports. Jin Woo Park, Hyokeun Lee, Boyeal Kim, Dong-Goo Kang, Seung Oh Jin, Hyun Kim 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2019 | Gaussian YOLOv3: An Accurate and Fast Object Detector Using Localization Uncertainty for Autonomous DrivingabstractThe use of object detection algorithms is becoming increasingly important in autonomous vehicles, and object detection at high accuracy and a fast inference speed is essential for safe autonomous driving. A false positive (FP) from a false localization during autonomous driving can lead to fatal accidents and hinder safe and efficient driving. Therefore, a detection algorithm that can cope with mislocalizations is required in autonomous driving applications. This paper proposes a method for improving the detection accuracy while supporting a real-time operation by modeling the bounding box (bbox) of YOLOv3, which is the most representative of one-stage detectors, with a Gaussian parameter and redesigning the loss function. In addition, this paper proposes a method for predicting the localization uncertainty that indicates the reliability of bbox. By using the predicted localization uncertainty during the detection process, the proposed schemes can significantly reduce the FP and increase the true positive (TP), thereby improving the accuracy. Compared to a conventional YOLOv3, the proposed algorithm, Gaussian YOLOv3, improves the mean average precision (mAP) by 3.09 and 3.5 on the KITTI and Berkeley deep drive (BDD) datasets, respectively. Nevertheless, the proposed algorithm is capable of real-time detection at faster than 42 frames per second (fps) and shows a higher accuracy than previous approaches with a similar fps. Therefore, the proposed algorithm is the most suitable for autonomous driving applications. Jiwoong Choi, Dayoung Chun, Hyun Kim 0001 |
ICCV | 3 |
| 2019 | An Effective DRAM Address Remapping for Mitigating Rowhammer ErrorsabstractA rowhammer error represents a loss of data stored in a DRAM cell caused by electromagnetic interference due to repetitive access to the same and/or adjacent rows. Due to the concentrated occurrence of rowhammer errors in specific rows and columns, these errors cannot be corrected by the conventional error correcting code (ECC) commonly used in DRAM devices. Previous techniques avoid these errors by having additional refresh operations that require additional hardware resources and/or power consumption. This paper proposes a different approach to handle rowhammer errors by distributing them across different DRAM rows and columns so that the attack cells are not concentrated on specific rows and columns. To this end, the distribution of rowhammer errors is observed with experiments using several commercial DRAM devices by employing state-of-the-art rowhammer attack techniques. The observation of the rowhammer errors concentrated in specific rows and columns underlies the proposal of an effective DRAM address remapping scheme for re-distribution of rowhammer errors. By using different address mappings to different chips and arrays in a DIMM, the proposed remapping effectively distributes errors over different rows and columns. As a result, the proposed remapping scheme decreases the possibility of multiple errors in a single word, and consequently, reduces uncorrectable errors under single error or single symbol correcting ECC. Experimental results with commercial DIMMs show that the proposed scheme reduces uncorrectable errors by about 95 percent while incurring a small additional hardware cost. Moonsoo Kim, Jungwoo Choi, Hyun Kim 0001 |
IEEE Trans. Computers | 3 |
| 2019 | Integration and Boost of a Read-Modify-Write Module in Phase Change Memory SystemabstractPhase-change memory (PCM) is a non-volatile memory device with favorable characteristics such as persistence, byte-addressability, and lower latency when compared to flash memory. However, it comprises memory cells that have limited lifetime and higher access latency than DRAM. The row buffer size of a PCM is preferred to be larger than 128B to fill the latency gap between two memories and to reduce the metadata overhead incurred by wear leveling. As the cache line size in a general-purpose processor is 64B, a read-modify-write (RMW) module is required to be placed between the processor and the PCM, which in turn induces a performance degradation. To reduce such an overhead and enhance the reliability of a device, this paper presents a new RMW architecture. The proposed model introduces a DRAM cache in the RMW module, which minimizes redundant read operations for write operations by pre-fetching the entire transaction unit instead of merely caching the 64B requested data. Furthermore, a typeless merge operation is performed with the proposed cache by gathering multiple commands accessing consecutive addresses, irrespective of whether they are READ or WRITE. Simulation results indicate that the proposed method enhances the speed by 3.2 times and the reliability by 49 percent as compared to the baseline model. Hyokeun Lee, Moonsoo Kim, Hyunchul Kim, Hyun Kim 0001 |
IEEE Trans. Computers | 4 |
| 2019 | A High-Throughput Hardware Accelerator for Lossless Compression of a DDR4 Command TraceabstractIn a memory system, understanding how the host is stressing the memory is important to improve memory performance. Accordingly, the need for the analysis of memory command trace, which the memory controller sends to the dynamic random access memory, has increased. However, the size of this trace is very large; consequently, a high-throughput hardware (HW) accelerator that can efficiently compress these data in real time is required. This paper proposes a high-throughput HW accelerator for lossless compression of the command trace. The proposed HW is designed in a pipeline structure to process Huffman tree generation, encoding, and stream merge. To avoid the HW cost increase owing to high-throughput processing, a Huffman tree is efficiently implemented by utilizing static random access memory-based queues and bitmaps. In addition, variable length stream merge is performed at a very low cost by reducing the HW wire width using the mathematical properties of Huffman coding and processing the metadata and the Huffman codeword using FIFO separately. Furthermore, to improve the compression efficiency of the DDR4 memory command, the proposed design includes two preprocessing operations, the “don't care bits override” and the “bits arrange,” which utilize the operating characteristics of DDR4 memory. The proposed compression architecture with such preprocessing operations achieves a high throughput of 8 GB/s with a compression ratio of 40.13% on average. Moreover, the total HW resource per throughput of the proposed architecture is superior to the previous implementations. Jiwoong Choi, Boyeal Kim, Hyun Kim 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | A High-Throughput and Power-Efficient FPGA Implementation of YOLO CNN for Object DetectionabstractConvolutional neural networks (CNNs) require numerous computations and external memory accesses. Frequent accesses to off-chip memory cause slow processing and large power dissipation. For real-time object detection with high throughput and power efficiency, this paper presents a Tera-OPS streaming hardware accelerator implementing a you-only-look-once (YOLO) CNN. The parameters of the YOLO CNN are retrained and quantized with the PASCAL VOC data set using binary weight and flexible low-bit activation. The binary weight enables storing the entire network model in block RAMs of a field-programmable gate array (FPGA) to reduce off-chip accesses aggressively and, thereby, achieve significant performance enhancement. In the proposed design, all convolutional layers are fully pipelined for enhanced hardware utilization. The input image is delivered to the accelerator line-by-line. Similarly, the output from the previous layer is transmitted to the next layer line-by-line. The intermediate data are fully reused across layers, thereby eliminating external memory accesses. The decreased dynamic random access memory (DRAM) accesses reduce DRAM power consumption. Furthermore, as the convolutional layers are fully parameterized, it is easy to scale up the network. In this streaming design, each convolution layer is mapped to a dedicated hardware block. Therefore, it outperforms the “one-size-fits-all” designs in both performance and power efficiency. This CNN implemented using VC707 FPGA achieves a throughput of 1.877 tera operations per second (TOPS) at 200 MHz with batch processing while consuming 18.29 W of on-chip power, which shows the best power efficiency compared with the previous research. As for object detection accuracy, it achieves a mean average precision (mAP) of 64.16% for the PASCAL VOC 2007 data set that is only 2.63% lower than the mAP of the same YOLO network with full precision. Duy Thanh Nguyen, Tuan Nghia Nguyen 0002, Hyun Kim 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | An Approximate Memory Architecture for a Reduction of Refresh Power Consumption in Deep Learning ApplicationsabstractA DRAM device requires periodic refresh operations to preserve data integrity, which incurs significant power consumption. This paper proposes a new memory architecture to reduce the power consumption by refresh operations by slowing down the refresh rate. Slow refresh may cause a loss of data stored in a DRAM cell, which affects the correctness of the computation using the lost data. The proposed memory architecture attempts to avoid the problem caused by lost data by taking advantage of the error-tolerant property of deep learning applications that are tolerant to presence of a small amount of errors. For data storage in deep learning applications, the approximate DRAM architecture stores the data in a transposed manner so that data are sorted according to their significance. DRAM organization is modified to support the control of the refresh period according to the significance of stored data. Simulation results with GoogLeNet and VGG-16 show that the power consumption is reduced by 69.68% with a negligible drop of the classification accuracy for both GoogLeNet and VGG-16. Duy Thanh Nguyen, Hyun Kim 0001, Ik Joon Chang |
ISCAS | 2 |
| 2018 | SPIHT Algorithm With Adaptive Selection of Compression Ratio Depending on DWT CoefficientsabstractIn mobile multimedia devices, the frame memory compression (FMC) technique by embedded compression (EC) is becoming an increasingly important video-processing method for reducing the external data bandwidth requirement, which, in turn, results in power savings. Among various EC schemes, the combination of discrete wavelet transform (DWT) and set partitioning in hierarchical trees (SPIHT) is widely used for FMC because it achieves high compression efficiency with low computational complexity. However, there is room for improvement in the conventional DWT and SPIHT algorithm because it compresses all blocks with the same compression ratio without taking into account the correlation between DWT coefficients and the SPIHT algorithm. This study proposes a novel one-dimensional (1-D) DWT and SPIHT algorithm, which enhances the quality of the compressed video by internally applying an adaptive compression ratio for the SPIHT algorithm based on DWT coefficients while keeping the same bit-stream size. The block complexity is predicted from the distribution of DWT coefficients. Then, simple blocks are aggressively compressed with a low compression ratio, while the complex blocks are compressed with a high ratio. Furthermore, to achieve the best video quality, each compression ratio is decided by an optimization technique based on mathematical formulation. Precisely, the logarithm of mean squared error by the SPIHT algorithm is assumed to be linearly correlated with the logarithm of processed DWT coefficients. Experimental results are provided that support the aforementioned model. Compared to the conventional 1-D DWT and SPIHT algorithm, the proposed scheme remarkably improves the video quality by an average of 2.23 dB in peak signal-to-noise ratio when the target compression ratio for the SPIHT algorithm is 5/16. Hyun Kim 0001, Albert No |
IEEE Trans. Multim. | 1 |
| 2016 | A Low-Power Video Recording System With Multiple Operation Modes for H.264 and Light-Weight CompressionabstractAn increasing demand for mobile video recording systems makes it important to reduce power consumption and to increase battery lifetime. The H.264/AVC compression is widely used for many video recording systems because of its high compression efficiency; however, the complex coding structure of H.264/AVC compression requires large power consumption. A light-weight video compression (LWC), based on discrete wavelet transform and set partitioning in hierarchical trees, consumes less power than H.264/AVC compression thanks to its relatively simple coding structure, although its compression efficiency is lower than that of H.264/AVC compression. This paper proposes a low-power video recording system that combines both the H.264/AVC encoder with high compression efficiency and LWC with low power consumption. The LWC is used to compress video data for temporal storage while the H.264/AVC encoder is used for permanent storage of data when some events are detected. For further power reduction, a down-sampling operation is utilized for permanent data storage. For an effective use of the two compressions with the down-sampling operation, an appropriate scheme is selected according to the proportion of long-term to short-term storage and the target bitrate. The proposed system reduces power consumption by up to 72.5% compared to that in a conventional video recording system. Hyun Kim 0001, Chae-Eun Rhee |
IEEE Trans. Multim. | 1 |
| 2015 | An Effective Combination of Power Scaling for H.264/AVC CompressionabstractThis brief proposes a novel method to determine the best combination of operation conditions for multiple power-scaling schemes. The power saving and rate-distortion performances of individual schemes are simulated, and then, the combined effects are modeled to obtain the best operation combination. The optimized combinations are defined as a power-level table. The proposed power-aware design is tested with four popular power-saving schemes and simulations show that a power saving of ~25% is achieved at the sacrifice of .0.172 dB Bjontegaard Delta peak signal-to-noise ratio degradation. Hyun Kim 0001, Chae-Eun Rhee |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | Power-aware design with various low-power algorithms for an H.264/AVC encoderabstractH.264/AVC video compression standard provides high coding efficiency, but requires a considerable amount of complexity and power consumption. This paper presents advanced low-power algorithms for an H.264/AVC encoder and a power-aware design composed of low-power algorithms. Power reduction algorithms with frame memory compression and early skip mode decision are presented, and the search range for motion estimation is reduced for further power reduction. The proposed power aware design controls the power consumption depending on the remaining energy by controlling the operation condition of the proposed low-power algorithms. In order to estimate the power reduction by the proposed algorithms, the power consumed by external memory as well as the bus between an H.264 encoder and an external DRAM is considered. Simulation results show that up to 49.9% of the power consumed by bus and external memory is reduced and that the power consumption from 0% to 41.56% is achieved with a reasonably small degradation of R-D performance. Hyun Kim 0001, Chae-Eun Rhee, Sunwoong Kim |
ISCAS | 1 |