EDBT 2026 Demo / reviewers in the wild / expert
Suk-Ju Kang
dblp:99/7096
· DBLP profile ↗
64ranked-venue papers
3as first author
37since 2021 · last 2026
0000-0002-4809-956XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 46 · 2 first-author · 31 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 8 since 2021Systems, architecture and hardware · 7 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Reverse Distillation with Core Exemplar Learning for Unified Multi-Class Anomaly DetectionabstractIn electronics manufacturing, anomaly detection methods face significant challenges due to class distribution imbalance and training instability when handling multiple classes simultaneously under varying imaging conditions. To address these challenges, we propose Reverse Distillation with Core-Exemplar Learning (RDCEL), a unified anomaly detection framework incorporating domain adaptation and novel metric learning strategies. RDCEL integrates unsupervised domain adaptation to align the covariance of the source and target domains and uses soft label-based coreset learning to handle diverse class distributions. It also leverages a coreset repulsion loss to minimize redundancy among coreset representations, fostering a more stable and dispersed embedding space across multiple classes. By aligning spatial statistics across different classes, RDCEL effectively addresses inter-class discrepancies, enabling consistent anomaly scoring under a unified test setting. Extensive experiments show RDCEL significantly outperforms state-of-the-art methods on MVTec AD and VisA datasets, achieving superior accuracy, stable AUROC performance, and faster convergence. Heechul Lim, Min-Soo Kim 0002, Hyun-Boo Lee, Suk-Ju Kang, Kang-Wook Chon, Haeyun Lee |
WACV | 4 |
| 2025 | QSCA: Quantization with Self-Compensating Auxiliary for Monocular Depth EstimationabstractMonocular depth estimation has advanced significantly with foundation models like Depth Anything, leveraging large-scale transformer architectures for the superior generalization.
However, the deployment on resource-constrained devices remains challenging due to the high computation and memory requirement.
Existing quantization methods, such as post-training quantization and quantization-aware training, often face trade-offs between efficiency and accuracy, or require extensive labeled data for retraining.
To address these limitations, we propose Quantization with Self-Compensating Auxiliary for Monocular Depth Estimation (QSCA), a novel framework for 4-bit post-training quantization of Monocular depth estimation models.
Our method integrates a lightweight Self-Compensating Auxiliary (SCA) module into both transformer encoder and decoder blocks, enabling the quantized model to recover from performance degradation without requiring ground truth.
This design enables fast adaptation while preserving structural and spatial consistency in predicted depth maps.
To our knowledge, this is the first framework to successfully apply 4-bit quantization across all layers of large-scale monocular depth estimation models.
Experimental results demonstrate that QSCA significantly improves quantized depth estimation performance. On the NYUv2 dataset, it achieves an 11\% improvement in $\delta_1$ accuracy over existing post-training quantization methods. Jincheol Yang, Matti Zinke, Suk-Ju Kang |
NeurIPS | 4 |
| 2025 | CRAN: Compressed Residual Attention Network for Lightweight Single Image Super-ResolutionabstractWith the growing demand for edge device and mobile environment, the development of lightweight super-resolution (SR) is necessary. However, existing SR approaches have primarily focused on increasing model performance, leading to high computational cost and reliance on high-performance GPUs. To address the issue, we propose a Compressed Residual Attention Network (CRAN), a lightweight SR model designed for high efficiency and performance. CRAN incorporates a Compressed Linear Block (CLB) and a Compressed Residual Attention Block (CRAB), utilizing linear overparameterization and spatial attention mechanism to maximize representational capacity while minimizing computational complexity. Additionally, a feature connection (FC) mechanism is proposed to enhance feature propagation through overparameterized layers without increasing parameter counts. Comprehensive experiments on SR benchmark datasets demonstrate that CRAN achieves superior results compared to existing lightweight SR methods, offering an excellent trade-off between efficiency and performance. Hanni Oh, Yeongje Im, Suk-Ju Kang |
IEEE Signal Process. Lett. | 3 |
| 2025 | Supervised Denoising for Extreme Low-Light Raw VideosabstractDenoising is a critical task in computer vision tasks, especially in challenging environments like extreme low-light conditions. However, the lack of research on denoising raw video in extreme low-light environments is notable, as is the absence of datasets specifically designed for this purpose. The primary challenge in denoising lies in balancing noise removal with detail preservation. Excessive denoising can cause the loss of fine details, while the insufficient denoising fails to adequately suppress noise, leading to degraded performance in downstream tasks such as object detection and recognition [1]. To address these limitations, we present a novel raw video dataset consisting of noisy-clean paired sequences captured under extreme low-light conditions, featuring diverse scenes and extended frames. In addition, we propose an efficient denoising framework tailored for this challenging scenario. Our approach combines shallow denoising, deformable convolution-based temporal alignment, and spatiotemporal attention to reduce noise while preserving texture and temporal consistency effectively. A texture-preserving loss is also proposed to prevent over-smoothing and retain fine details. Our proposed method outperforms state-of-the-art denoising models in terms of PSNR and SSIM on both synthetic and real-world extreme low-light videos, while exhibiting minimal side effects and preserving sharp details, as demonstrated by the quantitative results. Yeongje Im, Jione Pak, Songju Na, Jinhong Park, Jihyung Ryu, Seounghyun Moon, Beomjun Koo, Suk-Ju Kang |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Efficient Monocular Depth-Based Physical Distance Measurement for Low-Depth ScalesabstractIn this study, we introduce the physical distance estimator (PDE), a comprehensive framework designed to precisely measure the physical distance between two points on objects with a low-depth scale. The PDE utilizes monocular depth estimation to reconstruct 2-D depth maps into 3-D point clouds and calculates the physical distance between two points using the 3-D Euclidean distance. To achieve precise measurements for objects with low-depth scales, we configure a novel dataset using an RGB-Depth (RGB-D) camera to capture both RGB images and their depth maps of objects within low-depth scales, simulating various environmental conditions. We obtain a metric depth map by fine-tuning the monocular depth estimation model with the dataset. Furthermore, we applied knowledge distillation and FP16 optimization techniques to reduce the computational cost of PDE while maintaining high accuracy. These optimizations ensure that PDE operates efficiently in resource-constrained industrial environments, enabling real-time performance. Our experimental results show that PDE achieves an mean squared error of 29.60 and an mean absolute error of 3.17, representing improvements of 13.1% and 9.1%, respectively, over the NYUv2 benchmark. These quantitative results highlight the effectiveness and reliability of our method in real-world industrial settings. Jincheol Yang, Matti Zinke, Beoungwoo Kang, Hyung Uk Cho, Hyunyoung Choi, SangGu Lee, Suk-Ju Kang |
IEEE Trans. Ind. Informatics | 7 |
| 2025 | Programmable-Room: Interactive Textured 3D Room Meshes Generation Empowered by Large Language Models
Kyeongbo Kong, Suk-Ju Kang |
IEEE Trans. Multim. | 4 |
| 2025 | Query-Vector-Focused Recurrent Attention for Remaining Useful Life PredictionabstractPrognostics and health management (PHM) plays a crucial role in ensuring the reliability and operational efficiency of industrial systems through continuous monitoring. Among PHM tasks, accurate remaining useful life (RUL) prediction is essential for preventing failures and optimizing maintenance strategies. Recently, attention-based deep learning models have been actively explored for RUL prediction. In attention mechanisms, the query vector enhances the effectiveness of feature selection by determining which parts of the input the model focuses on. However, existing RUL prediction models typically extract query, key, and value features through the same process, leading to a generic attention mechanism that lacks label awareness. As a result, these models may struggle to capture task-specific degradation patterns, which are critical for precise RUL estimation. To address this limitation, we propose a recurrent attention network designed to learn a well-structured query vector. The proposed method explicitly incorporates label characteristics and temporal correlations, improving its ability to focus on task-relevant features. By applying this well-structured query vector within the attention mechanism, our approach effectively enhances feature representation and improves the predictive accuracy of RUL estimation. We evaluate our method on two public benchmark datasets, and experimental results demonstrate that the proposed approach achieves superior performance compared to existing methods. Ye In Park, Suk-Ju Kang |
IEEE Trans. Reliab. | 2 |
| 2024 | Person in Place: Generating Associative Skeleton-Guidance Maps for Human-Object Interaction Image EditingabstractRecently, there were remarkable advances in image editing tasks in various ways. Nevertheless, existing image editing models are not designed for Human-Object Interaction (HOI) image editing. One of these approaches (e.g. ControlNet) employs the skeleton guidance to offer precise representations of human, showing better results in HOI image editing. However, using conventional methods, manually creating HOI skeleton guidance is necessary. This paper proposes the object interactive diffuser with associative attention that considers both the interaction with objects and the joint graph structure, automating the generation of HOI skeleton guidance. Additionally, we propose the HOI loss with novel scaling parameter, demonstrating its effectiveness in generating skeletons that interact better. To evaluate generated object-interactive skeletons, we propose two metrics, top-N accuracy and skeleton probabilistic distance. Our framework integrates object interactive diffuser thatgenerates object-interactive skeletons with previous methods, demonstrating the outstanding results in HOI image editing. Finally, we present potentials of our framework beyond HOI image editing, as applications to human-to-human interaction, skeleton editing, and 3D mesh optimization. The code is available at https://github.com/YangChangHee/CVPR2024_Person-In-Place_RELEASE ChangHee Yang, Chanhee Kang, Kyeongbo Kong, Hanni Oh, Suk-Ju Kang |
CVPR | 5 |
| 2024 | AttentionHand: Text-Driven Controllable Hand Image Generation for 3D Hand Reconstruction in the Wild
Kyeongbo Kong, Suk-Ju Kang |
ECCV (60) | 3 |
| 2024 | Embedding-Free Transformer with Inference Spatial Reduction for Efficient Semantic Segmentation
Hyunwoo Yu, Yubin Cho, Beoungwoo Kang, Seunghun Moon, Kyeongbo Kong, Suk-Ju Kang |
ECCV (42) | 6 |
| 2024 | MetaSeg: MetaFormer-based Global Contexts-aware Network for Efficient Semantic SegmentationabstractBeyond the Transformer, it is important to explore how to exploit the capacity of the MetaFormer, an architecture that is fundamental to the performance improvements of the Transformer. Previous studies have exploited it only for the backbone network. Unlike previous studies, we explore the capacity of the Metaformer architecture more extensively in the semantic segmentation task. We propose a powerful semantic segmentation network, MetaSeg, which leverages the Metaformer architecture from the backbone to the decoder. Our MetaSeg shows that the MetaFormer architecture plays a significant role in capturing the useful contexts for the decoder as well as for the backbone. In addition, recent segmentation methods have shown that using a CNN-based backbone for extracting the spatial information and a decoder for extracting the global information is more effective than using a transformer-based backbone with a CNN-based decoder. This motivates us to adopt the CNN-based backbone using the MetaFormer block and design our MetaFormer-based decoder, which consists of a novel self-attention module to capture the global contexts. To consider both the global contexts extraction and the computational efficiency of the self-attention for semantic segmentation, we propose a Channel Reduction Attention (CRA) module that reduces the channel dimension of the query and key into the one dimension. In this way, our proposed MetaSeg outperforms the previous state-of-the-art methods with more efficient computational costs on popular semantic segmentation and a medical image segmentation benchmark, including ADE20K, Cityscapes, COCO-stuff, and Synapse. Beoungwoo Kang, Seunghun Moon, Yubin Cho, Hyunwoo Yu, Suk-Ju Kang |
WACV | 5 |
| 2024 | Human Motion Aware Text-to-Video Generation with Explicit Camera ControlabstractWith the rise in expectations related to generative models, text-to-video (T2V) models are being actively studied. Existing text-to-video models have limitations such as in generating complex movements replicating human motions. These model often generate unintended human motions, and the scale of the subject is incorrect. To overcome these limitations and generate high-quality videos that depict human motion under plausible viewing angles, we propose a two stage framework in this study. In the first stage a text-driven human motion generation network generates three-dimensional (3D) human motion from input text prompts and then motion-to-skeleton projection module projects generated motions onto a two-dimensional (2D) skeleton. In the second stage, the projected skeletons are used to generate a video in which the movements of a subject are well-represented. We demonstrated that the proposed framework quantitatively and qualitatively outperforms the existing T2V models. Previously reported human motion generation models use texts only or texts and human skeletons. However, our framework only uses texts and outputs a video related to human motion. Moreover, our framework benefits from using skeleton as an additional condition in the text-to-human motion generation networks. To the best of our knowledge, our framework is the first of its kind that uses text-driven human motion generation networks to generate high-quality videos related to human motions. The corresponding codes are available at https://github.com/CSJasper/HMTV. Chanhee Kang, JaeHyuk Park, Daun Jeong, ChangHee Yang, Suk-Ju Kang, Kyeongbo Kong |
WACV | 6 |
| 2024 | Image clustering using generated text centroids
Daehyeon Kong, Kyeongbo Kong, Suk-Ju Kang |
Signal Process. Image Commun. | 3 |
| 2024 | Cross-Aware Early Fusion With Stage-Divided Vision and Language Transformer Encoders for Referring Image SegmentationabstractReferring segmentation aims to segment a target object related to a natural language expression. Key challenges of this task are understanding the meaning of complex and ambiguous language expressions and determining the relevant regions in the image with multiple objects by referring to the expression. Recent models have focused on the early fusion with the language features at the intermediate stage of the vision encoder, but these approaches have a limitation that the language features cannot refer to the visual information. To address this issue, this paper proposes a novel architecture, Cross-aware early fusion with stage-divided Vision and Language Transformer encoders (CrossVLT), which allows both language and vision encoders to perform the early fusion for improving the ability of the cross-modal context modeling. Unlike previous methods, our method enables the vision and language features to refer to each other's information at each stage to mutually enhance the robustness of both encoders. Furthermore, unlike the conventional scheme that relies solely on the high-level features for the cross-modal alignment, we introduce a feature-based alignment scheme that enables the low-level to high-level features of the vision and language encoders to engage in the cross-modal alignment. By aligning the intermediate cross-modal features in all encoder stages, this scheme leads to effective cross-modal fusion. In this way, the proposed approach is simple but effective for referring image segmentation, and it outperforms the previous state-of-the-art methods on three public benchmarks. Yubin Cho, Hyunwoo Yu, Suk-Ju Kang |
IEEE Trans. Multim. | 3 |
| 2024 | Deep Conditional HDRI: Inverse Tone Mapping via Dual Encoder-Decoder Conditioning MethodabstractInverse tone mapping, a technique to restore a high dynamic range (HDR) image from a single low dynamic range (LDR) image, exhibits wide versatility since it may be easily applied to any camera device. Besides, the recent advancement in deep learning has produced great performance improvement in the field of inverse tone mapping. However, it remains a difficult task to accurately restore a wide-range HDR image from a single LDR image. A recent study attempts a spatially adaptive exposure value (EV) condition generated from luminance values to create a pseudo-multi-exposure stack. However, by adopting only luminance values as input, the conditioning method cannot precisely reflect the input image information when generating the EV condition, resulting in the loss of color expression. Moreover, there are some concerns regarding how to apply the EV condition to the image feature. Thus, the key idea of this study is to directly adopt image features in generating EV conditions that are adaptive to both color and brightness. To do this, we design a condition generation network with an encoder-decoder structure and propose a novel multi-exposure stack generation network, which bidirectionally synthesizes the image features and EV-conditioned features. Additionally, to better preserve the feature information in the synthesis of the features, we propose a spatially-adaptive feature transformation block. Our proposed method exhibits outstanding results in restoring the multi-exposure stacks for HDR image synthesis. Furthermore, our method achieves state-of-the-art performance compared to existing methods in multi-exposure stack generation and stack-based HDR restoration. Yoonchan Nam, JoonKyu Kim, Jae-hun Shim, Suk-Ju Kang |
IEEE Trans. Multim. | 4 |
| 2024 | MosaicMVS: Mosaic-Based Omnidirectional Multi-View Stereo for Indoor ScenesabstractWe present MosaicMVS, a novel learning-based depth estimation framework for a mosaic-based omnidirectional multi-view stereo (MVS) camera setup. It uses a regular field of view (FOV) MVS network for an omnidirectional imaging setup with explicit consideration of hypothetical voxel-wise FOV overlaps. The resulting depth predictions are accurate and agree on the omnidirectional multi-view geometry. Unlike existing MVS setups, MosaicMVS camera setup can be easily applied to omnidirectional indoor scenes without having to account for constraints such as intricate epipolar constraints and the distortion of omnidirectional cameras. We validate the effectiveness of our framework on a new challenging indoor dataset in terms of depth estimation, reconstruction, and view synthesis. We also present new evaluation metric to check reconstruction performance using post-processed masks for accurate evaluation without any ground truth depth map or laser-scanned reconstructions. Experimental results show that our framework outperforms the state-of-the-art MVS methods in a large margin in all test scenes. Min-Jung Shin, Woojune Park, Minji Cho, Kyeongbo Kong, Hoseong Son, Joonsoo Kim, Kugjin Yun, Gwangsoon Lee, Suk-Ju Kang |
IEEE Trans. Multim. | 9 |
| 2024 | CMVDE: Consistent Multi-View Video Depth Estimation via Geometric-Temporal Coupling ApproachabstractIn the field of video depth estimation, significant strides have been made with deep learning-based multi-view stereo approaches. However, existing studies struggle to produce consistently accurate depth maps that account for both multi-view geometry and temporal consistency from monocular video contents. To overcome this limitation, we introduce CMVDE, an innovative video depth estimation framework that leverages a multi-view geometric-temporal coupling approach in an end-to-end manner. Our proposed geometric consistency module efficiently generates multi-view geometric features by employing mutual cross-view epipolar attention between adjacent video frames. Additionally, it compresses these features using the novel multi-scale feature compressor, producing an effective input tensor for the subsequent module. Moreover, our framework enhances temporal consistency across consecutive video frames with the temporal consistency module based on convolutional LSTM 1 leveraging previous depth information as geometric guidance. Compared to state-of-the-art models, our approach achieves superior performance in depth quality and consecutive consistency on the ScanNet 2 and 7-Scenes 3 datasets, surpassing previous multi-view video depth estimation methods. Min-Jung Shin, Minji Cho, Joonsoo Kim, Kugjin Yun, Suk-Ju Kang |
IEEE Trans. Multim. | 6 |
| 2023 | FeedFormer: Revisiting Transformer Decoder for Efficient Semantic SegmentationabstractWith the success of Vision Transformer (ViT) in image classification, its variants have yielded great success in many downstream vision tasks. Among those, the semantic segmentation task has also benefited greatly from the advance of ViT variants. However, most studies of the transformer for semantic segmentation only focus on designing efficient transformer encoders, rarely giving attention to designing the decoder. Several studies make attempts in using the transformer decoder as the segmentation decoder with class-wise learnable query. Instead, we aim to directly use the encoder features as the queries. This paper proposes the Feature Enhancing Decoder transFormer (FeedFormer) that enhances structural information using the transformer decoder. Our goal is to decode the high-level encoder features using the lowest-level encoder feature. We do this by formulating high-level features as queries, and the lowest-level feature as the key and value. This enhances the high-level features by collecting the structural information from the lowest-level feature. Additionally, we use a simple reformation trick of pushing the encoder blocks to take the place of the existing self-attention module of the decoder to improve efficiency. We show the superiority of our decoder with various light-weight transformer-based decoders on popular semantic segmentation datasets. Despite the minute computation, our model has achieved state-of-the-art performance in the performance computation trade-off. Our model FeedFormer-B0 surpasses SegFormer-B0 with 1.8% higher mIoU and 7.1% less computation on ADE20K, and 1.7% higher mIoU and 14.4% less computation on Cityscapes, respectively. Code will be released at: https://github.com/jhshim1995/FeedFormer. Jae-hun Shim, Hyunwoo Yu, Kyeongbo Kong, Suk-Ju Kang |
AAAI | 4 |
| 2023 | SEFD: Learning to Distill Complex Pose and OcclusionabstractThis paper addresses the problem of three-dimensional (3D) human mesh estimation in complex poses and occluded situations. Although many improvements have been made in 3D human mesh estimation using the two-dimensional (2D) pose with occlusion between humans, occlusion from complex poses and other objects remains a consistent problem. Therefore, we propose the novel Skinned Multi-Person Linear (SMPL) Edge Feature Distillation (SEFD) that demonstrates robustness to complex poses and occlusions, without increasing the number of parameters compared to the baseline model. The model generates an SMPL overlapping edge similar to the ground truth that contains target person boundary and occlusion information, performing subsequent feature distillation in a simple edge map. We also perform experiments on various benchmarks and exhibit fidelity both qualitatively and quantitatively. Extensive experiments prove that our method outperforms the state-of-the-art method by 2.8% in MPJPE and 1.9% in MPVPE on a benchmark 3DPW dataset in the presence of domain gap. Also, our method is superior in 3DPW-OCC, 3DPW-PC, RH-Dataset, OCHuman, Crowd-Pose, and LSP dataset in which occlusion, complex pose, and domain gap exist. The code and occlusion & complex pose annotation will be available at https://github.com/YangChangHee/ICCV2023_SEFD_RELEASE/. ChangHee Yang, Kyeongbo Kong, Sung-Jun Min, Dongyoon Wee, Ho-Deok Jang, Geonho Cha, Suk-Ju Kang |
ICCV | 7 |
| 2023 | Unifying Domain Adaptation and Energy-Based Techniques for Person SearchabstractPerson search is a challenging task that involves detecting persons and identifying their identities in images. Previous studies are conducted to address the need for large-scale datasets with bounding box and person identity labels via domain adaptation. Recent domain adaptive person search studies rely on softmax-based methods, facing overfitting issues in detection training to source domain. This paper proposes the use of energy-based out-of-distribution detection instead of softmax-based classifier. Our approach separates the distributions of person and background clutter without overfitting issues. The integration of energy-based techniques into the Domain Adaptive Person Search framework improves detection performance, with an average precision increase of 2.11% and 4.25% on CUHKSYSU and PRW datasets. These results highlight the potential of energy-based approaches for domain adaptive person search and pave the way for accurate person search applications in real-world scenarios. Jione Pak, Chang-Ryeol Jeon, Suk-Ju Kang |
VCIP | 3 |
| 2023 | An Unified Framework for Language Guided Image CompletionabstractImage completion is a research field which aims to generate visual contents for unknown regions of an image. Image outpainting and wide-range image blending, which we refer to as extensive painting, are considered challenging because compared to the large unknown regions, relatively less context is provided. Some recent studies have tried to decrease the complexity of extensive painting by generating image hints for the missing regions. In this paper, we introduce a novel modality of hints, the natural language. Moreover, we propose a Captioning-based Extensive Painting (CEP) module, which combines models for two different multi-modal tasks: image captioning and text-guided image completion. In order to generate appropriate captions for masked images, the image captioning model is optimized using self-critical sequence training (SCST) method with random masks. The biggest benefit of our methodology is the accessibility to well-designed image captioning and text-guided image manipulation models such as OFA and GLIDE without the need for additional architectural changes. In evaluation, our model demonstrates remarkable performance even with complicated image datasets both quantitatively and qualitatively. Seong-Hun Jeong, Kyeongbo Kong, Suk-Ju Kang |
WACV | 4 |
| 2023 | Out-of-Focus Image Deblurring for Mobile Display Vision InspectionabstractIn vision inspection tasks, moiré patterns caused by frequency aliasing can severely degrade image quality. To prevent moiré patterns, we used images that were intentionally out-of-focused, and we performed deblurring to restore details during the acquisition of the images. As existing deblurring methods fail to output satisfactory results for low-contrast Mura images, we applied some simple techniques, minimum-maximum normalization, and edge mask fine-tuning to one of the state-of-the-art non-blind deblurring methods by utilizing parametric generalized Gaussian kernels. Structural image details were preserved through edge mask fine-tuning, and image contrast was improved with minimum-maximum normalization. By parameterizing the blur kernel as a generalized Gaussian kernel, we greatly improved the robustness of the blind image deblurring. We evaluated the effects of each module by conducting thorough experiments. The proposed method showed better performance than existing blind deblurring methods for blur-specific no-reference metrics, the image profile, and frequency domain analysis. Sung-Jun Min, Kyeongbo Kong, Suk-Ju Kang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Pseudo-Label-Vector-Guided Parallel Attention Network for Remaining Useful Life PredictionabstractPrognostic health management (PHM) has become important in many industries as a critical technology to increase machine stability and operational efficiency. Recently, various methods using deep learning to estimate the remaining useful life (RUL) as a core task of PHM have been proposed. However, the existing attention methods do not explicitly capture the correlation between temporal and spatial time series, reducing the RUL prediction accuracy. This article proposes a novel RUL prediction algorithm using a spatiotemporal attention mechanism based on the pseudo-label vectors to solve this problem. The proposed attention network uses the pseudo-label vector learned in the intermediate prediction process as a query vector to focus on time sequence data related to the RUL. Therefore, compared with conventional attention models that extract correlations for all the sequences, the proposed model captures features directly related to RUL with less computational cost. Experiments have been performed on two widely used datasets, and the experimental results show that the proposed approach outperforms the state of the art for root-mean-square error, with averages 4.27 and 3039 in the NASA Commercial Modular Aero-Propulsion System Simulation dataset and the IEEE PHM 2012 Prognostic challenge dataset, respectively. In addition, the analysis in the experiment reveals that the proposed model has better interpretability than the existing models by obtaining the correlation between time-series data and the RUL through the attention score in terms of time and features. Ye In Park, Jou Won Song, Suk-Ju Kang |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Human Body-Aware Feature Extractor Using Attachable Feature Corrector for Human Pose EstimationabstractTop-down pose estimation generally employs a person detector and estimates the keypoints of the detected person. This method assumes that only a single person exists within the bounding box cropped by detection. However, this assumption leads to some challenges in practice. First, a loose-fitted bounding box may include certain body parts of a non-target person. Second, spatial interference between several people exists owing to occlusion, so more than a single person can exist in the cropped image. In such scenarios, the pose estimation may falsely predict the keypoints of two or more persons as those of a single person. To tackle these issues, this paper proposes the human body-aware feature extractor based on the global- and local-reasoning features. The global-reasoning feature considers the entire body using transformer's non-local computation property and the local-reasoning feature concentrates on the individual body parts using convolutional neural networks. With those two features, we extract corrected features by filtering unnecessary features and supplementing necessary features using our proposed novel architecture. Hence, the proposed method can focus on the target person's keypoints, thereby mitigating the aforementioned concerns. Our method achieves noticeable improvement when applied to state-of-the-art top-down pose estimation networks. Ginam Kim, Kyeongbo Kong, Jou Won Song, Suk-Ju Kang |
IEEE Trans. Multim. | 5 |
| 2022 | Selective TransHDR: Transformer-Based Selective HDR Imaging Using Ghost Region Mask
Jou Won Song, Ye In Park, Kyeongbo Kong, Jaeho Kwak, Suk-Ju Kang |
ECCV (17) | 5 |
| 2022 | Vision Transformer-Based Retina Vessel Segmentation with Deep Adaptive Gamma CorrectionabstractAccurate segmentation of the retina vessel is essential for the early diagnosis of eye-related diseases. Recently, convolutional neural networks have shown remarkable performance in retina vessel segmentation. However, the complexity of edge structural information and the changeable intensity distribution depending on retina images reduce the performance of the segmentation tasks. This paper proposes two novel deep learning-based modules, channel attention vision transformer (CAViT) and deep adaptive gamma correction (DAGC), to tackle these issues. The CAViT jointly applies the efficient channel attention (ECA) and the vision transformer (ViT), in which the channel attention module considers the interdependency among feature channels and the ViT discriminates meaningful edge structures by considering the global context. The DAGC module provides the optimal gamma correction value for each input image by jointly training a CNN model with the segmentation network so that all the retina images are mapped to a unified intensity distribution. The experimental results show that our proposed method achieves superior performance compared to conventional methods on widely used datasets, DRIVE and CHASE DB1. Hyunwoo Yu, Jae-hun Shim, Jaeho Kwak, Jou Won Song, Suk-Ju Kang |
ICASSP | 5 |
| 2022 | Image-Adaptive Hint Generation via Vision Transformer for OutpaintingabstractImage outpainting has recently received considerable attention because it can be useful in tasks such as image retargeting and panorama image generation. In general, the problem of extending an image beyond its given boundaries is still ill-posed. Conventional methods predominantly attempt image outpainting by using complex network structures. Some recent studies have tried to decrease the problem complexity through the conversion techniques from outpainting to inpainting. Although these methodologies work well in simple cases, their performance reduces considerably for asymmetrical images. This paper proposes a novel hint-based outpainting methodology that can adaptively select the most plausible patches as hints from a given image to reduce the difficulty of outpainting. To estimate high-quality hints, inspired by patch-based image inpainting methods, we utilize Vision Transformer that considers self-attention for each patch. The estimated hints are attached on both boundaries of the input image and the inside missing regions are predicted by using an inpainting network. After finishing the prediction, the output image is obtained by removing the hints. Experiments show that our image-adaptive hint framework, when employed in representative inpainting networks, can consistently improve its performance compared to the other conversion techniques from outpainting to inpainting on SUN and Beach benchmark datasets. Daehyeon Kong, Kyeongbo Kong, Kyunghun Kim, Sung-Jun Min, Suk-Ju Kang |
WACV | 5 |
| 2022 | EAGNet: Elementwise Attentive Gating Network-Based Single Image De-Raining With Rain SimplificationabstractRain streaks are one of the main factors that degrade the performance of computer vision algorithms. Therefore, a preprocessing method is needed to remove rain streaks from rainy images. The main issue of the rain removal task is to prevent over (or under) de-raining. Over de-raining means that the background details are removed along with rain streaks in light rain, and under de-raining means that the rain streaks are not completely removed in heavy rain. These occur as the density of rain and intensity of rain streaks vary. In order to solve this, this paper proposes a two-step rain removal method. The proposed system first estimates the rain streaks image redefined with a simple operation from an input rainy image. The proposed rain streaks image contains rain density and rain streak intensity for the rainy image. By using this, the proposed system can adaptively remove rain streaks from images captured in various rain conditions. In addition, we propose a novel architectural unit, the elementwise attentive gating block, which is an optimized block used to deal with high frequency rain streaks. The proposed block selectively passes the desired components from the input feature maps by applying different weights to each element. It helps to clearly extract the rain streaks, and as a result, there are no traces of rain streaks on the restored image. The proposed method outperforms previous rain removal algorithms for both synthetic and real-world images. Namhyun Ahn, So Yeon Jo, Suk-Ju Kang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Dynamic Hand Gesture Recognition Using Improved Spatio-Temporal Graph Convolutional NetworkabstractHand gesture recognition is essential to human-computer interaction as the most natural way of communicating. Furthermore, with the development of 3D hand pose estimation technology and the performance improvement of low-cost depth cameras, skeleton-based dynamic hand gesture recognition has received much attention. This paper proposes a novel multi-stream improved spatio-temporal graph convolutional network (MS-ISTGCN) for skeleton-based dynamic hand gesture recognition. We adopt an adaptive spatial graph convolution that can learn the relationship between distant hand joints and propose an extended temporal graph convolution with multiple dilation rates that can extract informative temporal features from short to long periods. Furthermore, we add a new attention layer consisting of effective spatio-temporal attention and channel attention between the spatial and temporal graph convolution layers to find and focus on key features. Finally, we propose a multi-stream structure that feeds multiple data modalities (i.e., joints, bones, and motions) as inputs to improve performance using the ensemble technique. Each of the three-stream networks is independently trained and fused to predict the final hand gesture. The performance of the proposed method is verified through extensive experiments with two widely used public dynamic hand gesture datasets: SHREC’17 Track and DHG-14/28. Our proposed method achieves the highest recognition accuracy in various gesture categories for both datasets compared with state-of-the-art methods. Jae-Hun Song, Kyeongbo Kong, Suk-Ju Kang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Deep Arbitrary HDRI: Inverse Tone Mapping With Controllable Exposure ChangesabstractDeep convolutional neural networks (CNNs) have recently made significant advances in the inverse tone mapping technique, which generates a high dynamic range (HDR) image from a single low dynamic range (LDR) image that has lost information in over- and under-exposed regions. The end-to-end inverse tone mapping approach specifies the dynamic range in advance, thereby limiting dynamic range expansion. In contrast, the method of generating multiple exposure LDR images from a single LDR image and subsequently merging them into an HDR image enables flexible dynamic range expansion. However, existing methods for generating multiple exposure LDR images require an additional network for each exposure value to be changed or a process of recursively inferring images that have different exposure values. Therefore, the number of parameters increases significantly due to the use of additional networks, and an error accumulation problem arises due to recursive inference. To solve this problem, we propose a novel network architecture that can control arbitrary exposure values without adding networks or applying recursive inference. The training method of the auxiliary classifier-generative adversarial network structure is employed to generate the image conditioned on the specified exposure. The proposed network uses a newly designed spatially-adaptive normalization to address the limitation of existing methods that cannot sufficiently restore image detail due to the spatially equivariant nature of the convolution. Spatially-adaptive normalization facilitates restoration of the high frequency component by applying different normalization parameters to each element in the feature map according to the characteristics of the input image. Experimental results show that the proposed method outperforms state-of-the-art methods, yielding a 5.48dB higher average peak signal-to-noise ratio, a 0.05 higher average structure similarity index, a 0.28 higher average multi-scale structure similarity index, and a 7.36 higher average HDR-VDP-2 for various datasets. So Yeon Jo, Siyeong Lee, Namhyun Ahn, Suk-Ju Kang |
IEEE Trans. Multim. | 4 |
| 2021 | End-to-End Differentiable Learning to HDR Image Synthesis for Multi-exposure ImagesabstractRecently, high dynamic range (HDR) image reconstruction based on the multiple exposure stack from a given single exposure utilizes a deep learning framework to generate high-quality HDR images. These conventional networks focus on the exposure transfer task to reconstruct the multi-exposure stack. Therefore, they often fail to fuse the multi-exposure stack into a perceptually pleasant HDR image as the inversion artifacts occur. We tackle the problem in stack reconstruction-based methods by proposing a novel framework with a fully differentiable high dynamic range imaging (HDRI) process. By explicitly using the loss, which compares the network's output with the ground truth HDR image, our framework enables a neural network that generates the multiple exposure stack for HDRI to train stably. In other words, our differentiable HDR synthesis layer helps the deep neural network to train to create multi-exposure stacks while reflecting the precise correlations between multi-exposure images in the HDRI process. In addition, our network uses the image decomposition and the recursive process to facilitate the exposure transfer task and to adaptively respond to recursion frequency. The experimental results show that the proposed network outperforms the state-of-the-art quantitative and qualitative results in terms of both the exposure transfer tasks and the whole HDRI process. Jung Hee Kim 0001, Siyeong Lee, Suk-Ju Kang |
AAAI | 3 |
| 2021 | Structured Camera Pose Estimation for Mosaic-Based Omnidirectional ImagingabstractThis paper presents a novel structured camera pose estimation framework for mosaic-based omnidirectional imaging, i.e., producing a wide field of view (FoV) image that covers an entire sphere of the surroundings from a set of regular FoV images. With the effective utilization of geometric priors, the proposed framework exploits an individual image's connected structure while sequentially extracting correspondence between them. In the proposed framework, 2DSfM, a structure from motion method for multi-view images in structured 2D grids, is also proposed. Additionally, we propose a constraint term for rotation vectors in the bundle adjustment process that efficiently incorporates structural priors. We demonstrate our framework on structured omnidirectional image scenes and compare to existing frameworks. The experimental results show that our framework outperforms well-known conventional frameworks regarding both average reprojection error and reconstruction results. Woojune Park, Jung Hee Kim 0001, Suk-Ju Kang, Joonsoo Kim, Kugjin Yun, Won-Sik Cheong |
ISCAS | 3 |
| 2021 | Attention-Based Bidirectional LSTM-CNN Model for Remaining Useful Life EstimationabstractIn many industries, prognostic health management (PHM) technology has become important as a key technology to increase reliability and operational efficiency. Recently, several methods using a deep learning architecture to estimate the remaining useful life (RUL) as a part of the PHM have been presented. However, the limitation of existing methods is that they do not explicitly capture the relationship among different time sequences, which reduces the accuracy of RUL estimation. This paper proposes a novel RUL estimation algorithm using the attention mechanism to solve this problem. The proposed method applies scaled dot product attention to the encoder and the decoder consisting of long short-term memory, convolutional neural network and fully connected layer. The encoder applies self-attention to extract the association between time sequences, and the decoder extracts the association between the target RUL value and the time sequences using the representative vector of the RUL. Therefore, the proposed model has better performance to capture the long-term dependency in the sequence data and outperforms other state-of-the-art models in the experimental results. In addition, the extracted attention map shows that our model has better interpretability for RUL estimation. Jou Won Song, Ye In Park, Jong-Ju Hong, Seonggyun Kim, Suk-Ju Kang |
ISCAS | 5 |
| 2021 | Painting Outside as Inside: Edge Guided Image Outpainting via Bidirectional Rearrangement with Progressive Step LearningabstractImage outpainting is a very intriguing problem as the outside of a given image can be continuously filled by considering as the context of the image. This task has two main challenges. The first is to maintain the spatial consistency in contents of generated regions and the original input. The second is to generate a high-quality large image with a small amount of adjacent information. Conventional image outpainting methods generate inconsistent, blurry, and repeated pixels. To alleviate the difficulty of an outpainting problem, we propose a novel image outpainting method using bidirectional boundary region rearrangement. We rear-range the image to benefit from the image inpainting task by reflecting more directional information. The bidirectional boundary region rearrangement enables the generation of the missing region using bidirectional information similar to that of the image inpainting task, thereby generating the higher quality than the conventional methods using unidirectional information. Moreover, we use the edge map generator that considers images as original input with structural information and hallucinates the edges of unknown regions to generate the image. Our proposed method is compared with other state-of-the-art outpainting and inpainting methods both qualitatively and quantitatively. We further compared and evaluated them using BRISQUE, one of the No-Reference image quality assessment (IQA) metrics, to evaluate the naturalness of the output. The experimental results demonstrate that our method outperforms other methods and generates new images with 360°panoramic characteristics. Kyunghun Kim, Yeohun Yun, Keon-Woo Kang, Kyeongbo Kong, Siyeong Lee, Suk-Ju Kang |
WACV | 6 |
| 2021 | Feedback-based object detection for multi-person pose estimation
Jaeseo Park, Jun Ho Heo, Suk-Ju Kang |
Signal Process. Image Commun. | 3 |
| 2021 | Learning Methodologies to Generate Kernel-Learning-Based Image Downscaler for Arbitrary Scaling FactorsabstractDisplays and content have various resolutions and aspect ratios, requiring an image downscaler to adaptively reduce the image resolution. However, research on downscaling has garnered less attention than upscaling, including super-resolution. In practical display systems, simple interpolation, such as a bicubic filter that cannot preserve image details well, is still widely used for image downscaling rather than frame optimization-based or learning-based methods because of following reasons: frame optimization-based methods can effectively preserve image details after downscaling but are difficult to implement due to hardware costs. Learning-based methods have not been developed because defining a target downscaled image for training is difficult and training all downscaling factors is impossible. We propose a novel kernel-learning-based image downscaler to improve detail-preservation quality while supporting arbitrary downscaling factors using simple linear mapping. For this, a method to produce the ideal target downscaling result considering aliasing artifacts and detail preservation after downscaling is proposed. Then, we propose a training technique using the positional relationship between input and output pixels and a hierarchical region analysis to reproduce target images through simple kernel-based linear mapping. Lastly, a kernel-sharing technique is proposed to generate downscaling results for downscaling factors using a minimum number of trained kernels. In the simulation results, the proposed method demonstrated excellent edge preservation by improving the recall, precision, and F1 score, measuring the edge consistency between input and downscaled images, by up to 0.141, 0.079, 0.053, respectively, compared to benchmark methods. In a paired-comparison-based user study, the proposed method obtained the highest preference among benchmark methods using simple operations. Sung In Cho, Suk-Ju Kang |
IEEE Trans. Image Process. | 2 |
| 2021 | Learning to Generate Multi-Exposure Stacks With Cycle Consistency for High Dynamic Range ImagingabstractInverse tone mapping aims at recovering the lost scene radiances from a single exposure image. With the successful use of deep learning in numerous applications, many inverse tone mapping methods use convolution neural networks in a supervised manner. As these approaches are trained with many pre-fixed high dynamic range (HDR) images, they fail to flexibly expand the dynamic ranges of images. To overcome this limitation, we consider a multiple exposure image synthesis approach for HDR imaging. In particular, we propose a pair of neural networks that represent mappings between images that have exposure levels one unit apart (stop-up/down network). Therefore, it is possible to construct two positive-feedback systems to generate images with greater or lesser exposure. Compared to previous works using the conditional generative adversarial learning framework, the stop-up/down network employs HDR friendly network structures and several techniques to stabilize the training processes. Experiments on HDR datasets demonstrate the advantages of the proposed method compared to conventional methods. Consequently, we apply our approach to restore the full dynamic range of scenes agilely with only two networks and generate photorealistic images in complex lighting situations. Siyeong Lee, So Yeon Jo, Gwon Hwan An, Suk-Ju Kang |
IEEE Trans. Multim. | 4 |
| 2020 | Towards Design Methodology of Efficient Fast Algorithms for Accelerating Generative Adversarial Networks on FPGAsabstractGenerative adversarial networks (GANs) have shown excellent performance in image and speech applications. GANs create impressive data primarily through a new type of operator called deconvolution (DeConv) or transposed convolution (Conv). To implement the DeConv layer in hardware, the state-of-the-art accelerator reduces the high computational complexity via the DeConv-to-Conv conversion and achieves the same results. However, there is a problem that the number of filters increases due to this conversion. Recently, Winograd minimal filtering has been recognized as an effective solution to improve the arithmetic complexity and resource efficiency of the Conv layer. In this paper, we propose an efficient Winograd DeConv accelerator that combines these two orthogonal approaches on FPGAs. Firstly, we introduce a new class of fast algorithm for DeConv layers using Winograd minimal filtering. Since there are regular sparse patterns in Winograd filters, we further amortize the computational complexity by skipping zero weights. Secondly, we propose a new dataflow to prevent resource underutilization by reorganizing the filter layout in the Winograd domain. Finally, we propose an efficient architecture for implementing Winograd DeConv by designing the line buffer and exploring the design space. Experimental results on various GANs show that our accelerator achieves up to 1.78×~8.38× speedup over the state-of-the-art DeConv accelerators. Jung-Woo Chang, Saehyun Ahn, Keon-Woo Kang, Suk-Ju Kang |
ASP-DAC | 4 |
| 2020 | An Efficient Accelerator Design Methodology For Deformable Convolutional NetworksabstractDeformable convolutional networks have demonstrated outstanding performance in object recognition tasks with an effective feature extraction. Unlike standard convolution, the deformable convolution decides the receptive field size using dynamically generated offsets, which leads to an irregular memory access. Especially, the memory access pattern varies both spatially and temporally, making static optimization ineffective. Thus, a naive implementation would lead to an excessive memory footprint. In this paper, we present a novel approach to accelerate deformable convolution on FPGA. First, we propose a novel training method to reduce the size of the receptive field in the deformable convolutional layer without compromising accuracy. By optimizing the receptive field, we can compress the maximum size of the receptive field by 12.6 times. Second, we propose an efficient systolic architecture to maximize its efficiency. We then implement our design on FPGA to support the optimized dataflow. Experimental results show that our accelerator achieves up to 17.25 times speedup over the state-of-the-art accelerator. Saehyun Ahn, Jung-Woo Chang, Suk-Ju Kang |
ICIP | 3 |
| 2020 | Lightweight Deep Neural Network-based Real-Time Pose Estimation on Embedded SystemsabstractThis paper proposes a novel real-time pose estimation system on embedded devices for a driver and a front passenger. The main goal of the proposed system is to operate in real time with limited hardware resources while preserving the high accuracy. The proposed system is divided into an object detection and a pose estimation. In the object detection, we eliminate the redundant and inaccurate bounding boxes by considering the characteristics of the target image domain. In the pose estimation, a single-person pose estimation with a lightweight deep learning model has been proposed and knowledge distillation has been adopted to maximize the performance while maintaining the high speed. In the experimental results, the proposed pose estimation has up to 9 % of the accuracy and the 9 times less computation compared to the previous methods. The operation speed is 195 frame per second on NVIDIA Jetson TX2. Jun Ho Heo, Ginam Kim, Jaeseo Park, Yeonsu Kim, Sung-Sik Cho, Suk-Ju Kang |
IV | 7 |
| 2020 | Tri-level optimization-based image rectification for polydioptric cameras
Siyeong Lee, Gwon Hwan An, Joonsoo Kim, Kugjin Yun, Won-Sik Cheong, Suk-Ju Kang |
Signal Process. Image Commun. | 6 |
| 2020 | Extrapolation-Based Video Retargeting With Backward Warping Using an Image-to-Warping Vector Generation NetworkabstractVideo retargeting is a technique used to transform a given video to a target aspect ratio. Current methods often cause severe visual distortion due to frequent temporal incoherence during the retargeting. In this study, we propose a new extrapolation-based video retargeting method using an image-to-warping vector generation network to maintain temporal coherence and prevent deformation of an input frame by extending the side area of an input frame. Backward warping-based extrapolation is performed using a displacement vector (DV) that is generated by a proposed convolutional neural network (CNN). The DV is defined as the displacement between the current hole to be filled in the extended area and a pixel in the input frame used to fill the hole. We also propose a technique to efficiently train the CNN including a method for ground-truth DV generation. After the extrapolation, we propose a technique for the maintenance of temporal coherence of the extended region and a distortion suppression scheme (DSC) for minimizing visual artifacts. The simulation results demonstrated that the proposed method improved bidirectional similarity (BDS) up to 3.69, which is a measure of the quality of video retargeting, compared with existing video retargeting methods. Sung In Cho, Suk-Ju Kang |
IEEE Signal Process. Lett. | 2 |
| 2020 | An Energy-Efficient FPGA-Based Deconvolutional Neural Networks Accelerator for Single Image Super-ResolutionabstractConvolutional neural networks (CNNs) demonstrate excellent performance in various computer vision applications. In recent years, FPGA-based CNN accelerators have been proposed for optimizing performance and power efficiency. Most accelerators are designed for object detection and recognition algorithms that are performed on low-resolution images. However, real-time image super-resolution (SR) cannot be implemented on a typical accelerator because of the long execution cycles required to generate high-resolution (HR) images, such as those used in ultra-high-definition systems. In this paper, we propose a novel CNN accelerator with efficient parallelization methods for SR applications. First, we propose a new methodology for optimizing the deconvolutional neural networks (DCNNs) used for increasing feature maps. Second, we propose a novel method to optimize CNN dataflow so that the SR algorithm can be driven at low power in display applications. Finally, we quantize and compress a DCNN-based SR algorithm into an optimal model for efficient inference using on-chip memory. We present an energy-efficient architecture for SR and validate our architecture on a mobile panel with quad-high-definition resolution. Our experimental results show that, with the same hardware resources, the proposed DCNN accelerator achieves a throughput up to 108 times greater than that of a conventional DCNN accelerator. In addition, our SR system achieves an energy efficiency of 144.9, 293.0, and 500.2 GOPS/W at SR scale factors of 2, 3, and 4, respectively. Furthermore, we demonstrate that our system can restore HR images to a high quality while greatly reducing the data bit-width and the number of parameters compared with conventional SR algorithms. Jung-Woo Chang, Keon-Woo Kang, Suk-Ju Kang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Object Detection-Based Video Retargeting With Spatial-Temporal ConsistencyabstractThis study proposes a video retargeting method using deep neural network-based object detection. First, the meaningful regions of the input video denoted by bounding boxes of the object detection are extracted. In this case, the area is defined considering the size and number of bounding boxes for objects detected. The bounding boxes of each frame image are considered as regions of interest (RoIs). Second, the Siamese object tracking network is used to address high computational complexity of the object detection network. By dividing the video into scenes, object detection is performed for the first frame image of each scene to obtain the first bounding box. Object tracking is performed for the next sequential frame image until a scene change is detected. Third, the image is resized in the horizontal direction to alter the aspect ratio of the image and obtain the 1D RoIs of the image by projecting bounding boxes in the vertical direction. Then, the proposed method computes the grid map from the 1D RoIs to calculate new coordinates of each column data of the image. Finally, the retargeted video is obtained by rearranging all retargeted frame images. Comparative experiments conducted with various benchmark methods show an average bidirectional similarity score of 1.92, which is higher than other conventional methods. The proposed method was stable and satisfied viewers without causing cognitive discomfort as conventional methods. Seung Joon Lee, Siyeong Lee, Sung In Cho, Suk-Ju Kang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Temporal Incoherence-Free Video Retargeting Using Foreground Aware ExtrapolationabstractVideo retargeting is a method of adjusting the aspect ratio of a given video to the target aspect ratio. However, temporal incoherence of video contents, which can occur frequently by video retargeting, is the most dominant factor that degrades the quality of retargeted videos. Current methods to maintain temporal coherence use the entire frames of the input videos; however, these methods cannot be implemented as on-time systems because of their tremendous computational complexity. As far as we know, there is no existing on-time video retargeting method that can avoid spatial distortion while perfectly maintaining temporal coherence. In this paper, we propose a novel on-time video retargeting method that can perfectly maintain temporal coherence and prevent the spatial distortion by using only two consecutive input frames. In our method, the maximum a posteriori-based foreground aware-block matching is used for the extrapolation that extends the side area of a given video to adjust its aspect ratio to the target. To maintain the temporal coherence of the extended area, the result of block matching for backward warping-based extrapolation of the start frame after the scene change occurs, is reused for the other frames until the next scene change occurs. In addition, we propose a scene scenario-adaptive fallback scheme to prevent severe distortions that can occur with reusing block matching results or extrapolation-based side extension. The simulation results showed that the proposed method greatly improved the bidirectional similarity value, which can measure the quality of video retargeting, by up to 10.26 compared with the existing on-time video retargeting methods. Sung In Cho, Suk-Ju Kang |
IEEE Trans. Image Process. | 2 |
| 2019 | SDCNN: An Efficient Sparse Deconvolutional Neural Network Accelerator on FPGAabstractGenerative adversarial networks (GANs) have shown excellent performance in image generation applications. GAN typically uses a new type of neural network called deconvolutional neural network (DCNN). To implement DCNN in hardware, the state-of-the-art DCNN accelerator optimizes the dataflow using DCNN-to-CNN conversion method. However, this method still requires high computational complexity because the number of feature maps is increased when converted from DCNN to CNN. Recently, pruning has been recognized as an effective solution to reduce the high computational complexity and huge network model size. In this paper, we propose a novel sparse DCNN accelerator (SDCNN) combining these approaches on FPGA. First, we propose a novel dataflow suitable for the sparse DCNN acceleration by loop transformation. Then, we introduce a four stage pipeline for generating the SDCNN model. Finally, we propose an efficient architecture based on SDCNN dataflow. Experimental results on DCGAN show that SDCNN achieves up to 2.63 times speedup over the state-of-the-art DCNN accelerator. Jung-Woo Chang, Keon-Woo Kang, Suk-Ju Kang |
DATE | 3 |
| 2019 | Gradient Prior-Aided CNN Denoiser With Separable Convolution-Based Optimization of Feature DimensionabstractWe propose a novel image denoising method based on a convolutional neural network (CNN), which uses the separable convolution and the gradient prior to reduce the computational complexity while enhancing the denoising performance. The proposed method converts the existing convolution filter in the conventional CNN denoiser to cascaded vertical and horizontal separable convolutions and reduces the number of feature channels between these convolutions by analyzing the distribution of convolution weights. The proposed separable convolution with feature dimension shrinking can greatly reduce the number of multiplications for CNN while minimizing the degradation of denoising quality. In addition, gradients of a given image are used as input for the proposed CNN denoiser by exploiting the relation between an anisotropic diffusion-based denoiser and a residual CNN denoiser to improve the quality of the image denoising. The simulation results showed that the proposed method provided comparable denoising quality while reducing the number of multiplications to 41% compared to the existing state-of-the-art CNN denoiser. Sung In Cho, Suk-Ju Kang |
IEEE Trans. Multim. | 2 |
| 2018 | Optimizing FPGA-based convolutional neural networks accelerator for image super-resolutionabstractConvolutional neural networks (CNN) are widely used in various computer vision applications. Recently, there have been many studies on FPGA-based CNN accelerators to achieve high performance and power efficiency. Most of them have been on CNN-based object detection algorithms, but researches on image super-resolution have been rarely conducted. Fast super-resolution CNN (FSRCNN), well known for CNN-based super-resolution algorithm, are a combination of multiple convolutional layers and a single deconvolutional layer. Since the deconvolutional layer generates high-resolution (HR) output feature maps from low-resolution (LR) input feature maps, its execution cycles are larger than those of the convolutional layer. In this paper, we propose a novel architecture of the FPGA-based CNN accelerator with the efficient parallelization. We develop a method of transforming a deconvolutional layer into a convolutional layer (TDC), a new methodology for the deconvolutional neural networks (DCNN). There is a massive parallelization source in the deconvolutional layer where multiple outputs within the same output feature map are created with the same inputs. When this new parallelization technique is applied to the deconvolutional layer, it generates the LR output feature maps the same as the convolutional layer. Thus, the performance of the accelerator increases without any additional hardware resources because the kernel size required to generate the LR output feature maps is smaller. In addition, if there is a DSP underutilization problem in the deconvolutional layer that some of the processors are in an idle state, the proposed method solves this problem by allowing more output feature maps to be processed in parallel. Experimental results show that the proposed TDC method achieves up to 81 times higher throughput than the state-of-the-art DCNN accelerator with the same hardware resources. We also improve the speed by 7.8 times by having all layers in the hourglass-type FSRCNN to be processed in inter-layer parallelism without additional DSP usage. Jung-Woo Chang, Suk-Ju Kang |
ASP-DAC | 2 |
| 2018 | Deep Recursive HDRI: Inverse Tone Mapping Using Generative Adversarial Networks
Siyeong Lee, Gwon Hwan An, Suk-Ju Kang |
ECCV (2) | 3 |
| 2018 | Driver Identification System Using Convolutional Neural Network with Background Removal-based Infrared Data AugmentationabstractAs the interest of the autonomous driving increases, techniques related to the advanced driver assistance system are evolving together. In this paper, we propose a novel driver identification system using convolutional neural network (CNN) with the background removal-based infrared image data augmentation. It helps to identify who a driver is, and provides the customized driving environment. The process for the proposed identification system is as follows. First, we acquire customized individual infrared images in a driving simulation environment. Second, we augment the large amount of data by using the background removal-based method and several image processing techniques. Third, the augmented data is trained by the low-complexity-based CNN method. Finally, we load all trained weights to the forward network for real-time processing. In the experimental results, the proposed system had the memory resource of 4,795 KB, which are up to 49.0822 times smaller than benchmark algorithms, and the average F1score of 0.9418 for the driver identification accuracy. Sanghyuk Kim, Yunsoo Lee, Namhyun Ahn, Suk-Ju Kang |
Intelligent Vehicles Symposium | 4 |
| 2018 | Geodesic Path-Based Diffusion Acceleration for Image DenoisingabstractWe propose an advanced anisotropic-diffusion (AD)-based approach for an image denoising method, which utilizes a geodesic path to produce single-pass adaptive smoothing by analyzing the diffusion continuity. The proposed method consists of the following four procedures: element-weight determination, geodesic path-based kernel (GPK) generation, single-pass smoothing using the GPK, and post-processing of the GPK smoothing. In the first procedure, weights for neighboring pixels are calculated by diffusivity analysis. In the second procedure, a geodesic path is selected using a geodesic distance that is calculated by a diffusion continuity analysis. In the third procedure, GPK-based smoothing is applied to a given noisy image to extract the noise-free pixel value. Finally, a distant AD that uses the double diffusion length is applied to the resultant image by the GPK filtering to enhance the quality of noise suppression in smooth regions. In addition to the main procedures, schemes for the robust outlier reduction and complexity reduction are introduced. The simulation results showed that the proposed method improved the denoising quality by increasing the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) by up to 4.094 dB and 0.057, respectively, compared to the AD-based benchmark methods. Compared to block-matching and 3-D filtering, the proposed method showed comparable quality of noise reduction with similar PSNR and SSIM values, which it accomplished with much less computation time. Sung In Cho, Suk-Ju Kang |
IEEE Trans. Multim. | 2 |
| 2017 | Census transform-based static caption detection for frame rate up-conversionabstractThis paper presents a new static caption detection method that uses the census transform (CT) and motion vector (MV) for frame rate up-conversion. The proposed method splits a frame into several blocks and detects the static regions using CT and MV. CT is used to consider the spatio-temporal consistency of the texture and MV is used to remove the non-static regions containing moving objects. Next, it corrects the falsely detected regions by performing the outlier removal and inward filling. Finally, it detects the static caption based on the existence of global motion. In the experimental results, the average F1 score of the proposed method was up to 0.704 higher than those of the benchmark methods. Gyu Jin Bae, Young Hwan Kim, Suk-Ju Kang |
ISCAS | 3 |
| 2017 | Sensor-based driver condition recognition using support vector machine for the detection of driver drowsinessabstractDriver's drowsiness is one of the main causes of traffic accidents. It accounts for up to 20% of serious or fatal accidents on the roads. In this paper, we propose a novel sensor-based driver condition recognition to prevent drowsiness-related accidents by using Support Vector Machine (SVM). It helps to determine the driver's condition by using sensors of the wearable device. The process for the proposed warning system is as follows: First, we acquire bio-data from a PhotoPlethysmoGraphy(PPG) sensor in the device to understand the characteristics of the driver's condition. Then, the acquired data is processed through segmentation and averaging to increase classification accuracy. The processed data is used as feature vectors in the SVM for the driver's condition classification. To evaluate performance of proposed method, the following metrics were used: accuracy, error rate, precision and recall. From the calculated results, the proposed method had high accuracy of 96.3 %. Finally, we could create the driver's condition recognition model that can be applied to a system which can alert the driver. Seong-Pil Cheon, Suk-Ju Kang |
Intelligent Vehicles Symposium | 2 |
| 2017 | Image Segmentation Using Linked Mean-Shift Vectors and Global/Local AttributesabstractThis paper proposes novel noniterative mean-shift-based image segmentation that uses global and local attributes. The existing mean-shift-based methods use a fixed range bandwidth, and hence their accuracy is dependent on the range spectrum of an image. To resolve this dependency, this paper proposes to modify the range kernel in the mean-shift process to be anisotropic. The modification is conducted using a global attribute defined as the range covariance matrix of the image. Further, to alleviate oversegmentation, the proposed method merges the segments having similar local attributes more aggressively than other segments. The local attribute for each segment is defined as the sum of the variances of the chromatic components. Finally, to expedite the processing, the proposed method uses a region adjacency graph (RAG) for the merging process, thus differing from the existing linked mean-shift-based methods. In the experiments on the Berkeley segmentation data set, the use of the global and local attributes improved segmentation accuracy; the proposed method outperformed the state-of-the-art linked mean-shift-based method by showing an improvement of 2.15%, 3.16%, 3.32%, and 1.90% in probability rand index, segmentation covering, variation of information, and F-measure, respectively. Further, compared with the benchmark method, which uses the dilating and merging scheme, the proposed method improved the speed of the merging process 42 times by applying the RAG. Hanjoo Cho, Suk-Ju Kang, Young Hwan Kim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Anisotropic diffusion noise filtering using region adaptive smoothing strength
Sanghun Kim, Suk-Ju Kang, Young Hwan Kim |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | Human perception-based image segmentation using optimising of colour quantisationabstractThis study presents an advanced histogram‐based image segmentation method that enhances image segmentation quality, while greatly reducing the computational complexity. Unlike existing histogram‐based methods, the authors optimise the size of bins in the colour histogram by using human perception‐based colour quantisation and the clustering centroids are selected effectively without using a complex process. Additionally, an over‐segmentation removal technique based on connected‐component labelling is employed. This improves the segmentation quality by connectivity analysis. A comparison between the experimental results on the Berkeley Segmentation Dataset by the proposed method and the benchmark methods demonstrated that the proposed method enhanced the segmentation quality by improving the Probabilistic Rand Index and the Segmentation Covering values compared with those of the benchmark methods. The computation time using the proposed method is reduced by up to 91.63% compared with the computation time using benchmark methods. Sung In Cho, Suk-Ju Kang, Young Hwan Kim |
IET Image Process. | 2 |
| 2014 | Dictionary-based anisotropic diffusion for noise reduction
Sung In Cho, Suk-Ju Kang, Hi-Seok Kim, Young Hwan Kim |
Pattern Recognit. Lett. | 2 |
| 2014 | Adaptive Weight Allocation-Based Subpixel Rendering AlgorithmabstractIn this letter, a new approach is presented for adaptive weight allocation-based subpixel rendering in organic light-emitting diode displays. Subpixel rendering is used to enhance the apparent resolution without changing a pixel structure. Existing methods have blurring and color fringing artifact after subpixel rendering. The proposed method, on the other hand, dynamically controls weights of current and neighboring pixels based on color difference. Thus, it preserves the image quality, while increasing the resolution. In experiments, the proposed subpixel rendering improved luminance sharpness by up to 0.043, when compared with the benchmark methods. For chrominance blending, the peak signal-to-noise ratio of the proposed method was up to 6.192 dB higher than those of benchmark methods. Suk-Ju Kang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | Entrance Detection of Buildings Using Multiple Cues
Suk-Ju Kang, Hoang-Hon Trinh, Dae-Nyeon Kim, Kang-Hyun Jo |
ACIIDS (1) | 1 |
| 2010 | Dynamic clipping ratio determination for global backlight dimming in LCDabstractThis paper proposes a dynamic global backlight dimming algorithm based on the analysis of the cumulative distribution function (CDF) of the image histogram. For each intensity level, the convexity of the CDF is calculated using the distance to the uniform histogram curve. Using these values, the proposed algorithm characterizes the shape of the CDF to find the proper clipping ratio. This ratio enables to get the maximum level data beyond which the intensity levels are clipped, as the results of backlight dimming but image distortion. The experimental results show that the proposed algorithm gives a good compromise between backlight power reduction and image quality preservation, quantified by the PSNR metric. Philippe Lavole, Sung-Kyu Lee, Suk-Ju Kang, Young Hwan Kim |
ISCAS | 3 |
| 2010 | Dual Motion Estimation for Frame Rate Up-ConversionabstractIn this letter, we present a new motion estimation algorithm for frame rate up-conversion. The proposed dual motion estimation algorithm enhances the estimation accuracy of motion vectors by using the unidirectional and bidirectional matching ratios of blocks in the previous and current frames. In addition, the proposed motion estimation approach uses motion vector validity to evaluate the accuracy of motion vectors thereby avoiding false motion vectors. In experiments using benchmark image sequences, the proposed motion estimation algorithm improved the average peak signal-to-noise ratio of interpolated frames by up to 2.272 dB, when compared to conventional motion estimation algorithms. For the comparison of the perceptual image quality using the structural similarity, the average value of the proposed dual motion estimation was by up to 0.062 higher than those of the conventional algorithms. Suk-Ju Kang, Sungjoo Yoo, Young Hwan Kim |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Window Extraction Using Geometrical Characteristics of Building Surface
Hoang-Hon Trinh, Dae-Nyeon Kim, Suk-Ju Kang, Kang-Hyun Jo |
ICIC (1) | 3 |
| 2009 | Building-Based Structural Data for Core Functions of Outdoor Scene Analysis
Hoang-Hon Trinh, Dae-Nyeon Kim, Suk-Ju Kang, Kang-Hyun Jo |
ICIC (1) | 3 |
| 2009 | An Efficient Method of Vehicle License Plate Detection Based on HSI Color Model and Histogram
Kaushik Deb, Heechul Lim, Suk-Ju Kang, Kang-Hyun Jo |
IEA/AIE | 3 |