Seung-Hwan Bae

dblp:211/6993 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MCM-SR: Multiple Constant Multiplication-Based CNN Streaming Hardware Architecture for Super-Resolution
abstract
Convolutional neural network (CNN)-based super-resolution (SR) methods have become prevalent in display devices due to their superior image quality. However, the significant computational demands of CNN-based SR require hardware accelerators for real-time processing. Among the hardware architectures, the streaming architecture can significantly reduce latency and power consumption by minimizing external dynamic random access memory (DRAM) access. Nevertheless, this architecture necessitates a considerable hardware area, as each layer needs a dedicated processing engine. Furthermore, achieving high hardware utilization in this architecture requires substantial design expertise. In this article, we propose methods to reduce the hardware resources of CNN-based SR accelerators by applying the multiple constant multiplication (MCM) algorithm. We propose a loop interchange method for the convolution (CONV) operation to reduce the logic area by 23% and an adaptive loop interchange method for each layer that considers both the static random access memory (SRAM) and logic area simultaneously to reduce the SRAM size by 15%. In addition, we improve the MCM graph exploration speed by$5.4\times $while maintaining the SR quality through beam search when CONV weights are approximated to reduce the hardware resources.
Seung-Hwan Bae, Hyun Kim 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2022 Deformable Part Region Learning for Object Detection
abstract
In a convolutional object detector, the detection accuracy can be degraded often due to the low feature discriminability caused by geometric variation or transformation of an object. In this paper, we propose a deformable part region learning in order to allow decomposed part regions to be deformable according to geometric transformation of an object. To this end, we introduce trainable geometric parameters for the location of each part model. Because the ground truth of the part models is not available, we design classification and mask losses for part models, and learn the geometric parameters by minimizing an integral loss including those part losses. As a result, we can train a deformable part region network without extra super-vision and make each part model deformable according to object scale variation. Furthermore, for improving cascade object detection and instance segmentation, we present a Cascade deformable part region architecture which can refine whole and part detections iteratively in the cascade manner. Without bells and whistles, our implementation of a Cascade deformable part region detector achieves better detection and segmentation mAPs on COCO and VOC datasets, compared to the recent cascade and other state-of-the-art detectors.
Seung-Hwan Bae
AAAI1
2022 RT-MOT: Confidence-Aware Real-Time Scheduling Framework for Multi-Object Tracking Tasks
abstract
Different from existing MOT (Multi-Object Tracking) techniques that usually aim at improving tracking accuracy and average FPS, real-time systems such as autonomous vehicles necessitate new requirements of MOT under limited computing resources: (R1) guarantee of timely execution and (R2) high tracking accuracy. In this paper, we propose RT-MOT, a novel system design for multiple MOT tasks, which addresses R1 and R2. Focusing on multiple choices of a workload pair of detection and association, which are two main components of the tracking-by-detection approach for MOT, we tailor a measure of object confidence for RT-MOT and develop how to estimate the measure for the next frame of each MOT task. By utilizing the estimation, we make it possible to predict tracking accuracy variation according to different workload pairs to be applied to the next frame of an MOT task. Next, we develop a novel confidence-aware real-time scheduling framework, which offers an offline timing guarantee for a set of MOT tasks based on non-preemptive fixed-priority scheduling with the smallest workload pair. At run-time, the framework checks the feasibility of a priority-inversion associated with a larger workload pair, which does not compromise the timing guarantee of every task, and then chooses a feasible scenario that yields the largest tracking accuracy improvement based on the proposed prediction. Our experiment results demonstrate that RT-MOT significantly improves overall tracking accuracy by up to 1.5 ×, compared to existing popular tracking-by-detection approaches, while guaranteeing timely execution of all MOT tasks.
Donghwa Kang, Seunghoon Lee 0002, Hoon Sung Chwa, Seung-Hwan Bae, Chang Mook Kang, Jinkyu Lee 0001, Hyeongboo Baek
RTSS4
2021 DC-AC: Deep Correlation-Based Adaptive Compression of Feature Map Planes in Convolutional Neural Networks
abstract
Deep learning has been successfully deployed to a broad range of applications with its outstanding performance. Supporting an efficient hardware architecture is critical to making effective use of a deep learning approach with proven algorithm performance. One challenge in implementation of deep learning algorithm is to reduce memory bandwidth because a single memory access normally consumes 100* more energy than an arithmetic operation. To reduce the memory bandwidth, deep learning data could be compressed and decompressed before memory write/read operations. Especially, feature maps, which account for a significant portion of the convolutional neural network (CNN), could be compressed further by reducing the correlations between feature map planes. This paper proposes a compression method for feature maps in CNN that adaptively exploits the varying correlation between feature map planes. For every feature map plane, the proposed method searches the most similar plane among nearby planes in the same layer, and compresses the residual of the two planes instead of the plane itself. Experimental results show that the average bit length to store feature maps is reduced by 14.2% compared to the compression without correlation reduction, and the CNN accuracy does not change and additional training is also not required because the proposed method applies lossless compression.
Seung-Hwan Bae, Hyun Kim 0001
ISCAS1
2021 Cache Compression with Golomb-Rice Code and Quantization for Convolutional Neural Networks
abstract
Cache compression schemes reduce the cache miss rate by increasing the effective cache capacity and consequently, reduce memory access and power consumption. Therefore, cache compression is beneficial for applications with heavy memory traffic, including convolutional neural network (CNN). In this paper, a new cache compression of a floating-point number is proposed for CNNs. The exponent is compressed using the Golomb-Rice code, instead of the Huffman code, for an efficient hardware implementation. The compression syntax is carefully designed so that the size of compressed data is not very far from the entropy, which is the theoretical limit, by distinguishing two different types of data used in CNNs. On the other hand, since the mantissa of CNNs data can be hardly compressed by entropy coding, it is simply quantized for data reduction that may not degrade the CNN performance significantly thanks to the error robustness of CNNs. The quantization reduces 23 bits of a mantissa to 4 bits. The experimental results show that the miss rate of a 1 MB compressed cache with the proposed compression method applied is almost similar to that of an uncompressed 2 MB cache without any decrease of the CNN accuracy.
Seung-Hwan Bae, Hyun Kim 0001
ISCAS1
2019 Object Detection Based on Region Decomposition and Assembly
abstract
Region-based object detection infers object regions for one or more categories in an image. Due to the recent advances in deep learning and region proposal methods, object detectors based on convolutional neural networks (CNNs) have been flourishing and provided the promising detection results. However, the detection accuracy is degraded often because of the low discriminability of object CNN features caused by occlusions and inaccurate region proposals. In this paper, we therefore propose a region decomposition and assembly detector (R-DAD) for more accurate object detection.In the proposed R-DAD, we first decompose an object region into multiple small regions. To capture an entire appearance and part details of the object jointly, we extract CNN features within the whole object region and decomposed regions. We then learn the semantic relations between the object and its parts by combining the multi-region features stage by stage with region assembly blocks, and use the combined and high-level semantic features for the object classification and localization. In addition, for more accurate region proposals, we propose a multi-scale proposal layer that can generate object proposals of various scales. We integrate the R-DAD into several feature extractors, and prove the distinct performance improvement on PASCAL07/12 and MSCOCO18 compared to the recent convolutional detectors.
Seung-Hwan Bae
AAAI1