Byung Cheol Song

dblp:66/5310 · DBLP profile ↗
← Back
89ranked-venue papers
21as first author
30since 2021 · last 2026
0000-0001-8742-3433ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 69 · 20 first-author · 16 since 2021Artificial intelligence and machine learning · 27 · 21 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSystems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Variation-aware proxy learning for semantic segmentation
abstract
In semantic segmentation, accurately modeling intra-class variation is essential for capturing fine-grained details and resolving ambiguity near class boundaries. While existing proxy-based embedding methods represent each class with a single prototype, they struggle to reflect diverse intra-class structures, especially in complex scenes. In this paper, we propose a novel representation learning framework called Variation-Aware Proxy Learning, which introduces a representative proxy to encode shared class semantics and multiple variation vectors to capture fine-grained intra-class variations. These components are integrated through a factorized similarity score, enabling more expressive and discriminative embedding structures. To further enhance learning in ambiguous regions, we introduce focal modulation and design a new Compositional Similarity Loss composed of attraction and repulsion terms that adaptively amplify the contribution of hard examples. Our method is model-agnostic and requires no additional inference-time cost. Extensive experiments across multiple segmentation benchmarks—Cityscapes, COCO-Stuff10k, iSAID and ADE20K—and diverse backbones including CNNs and Transformers, demonstrate consistent improvements in mIoU and boundary precision, particularly in challenging regions with high intra-class variability. • Compositional Similarity Loss enables the learning of representative proxies and variation vectors. • Factorized similarity score achieves enhanced class separation with proxy integration. • Focal modulation adaptively emphasizes ambiguous regions with dynamic weighting. • Consistent segmentation gains across diverse real-time and non-real-time models.
Haejun Bae, Byung Cheol Song
Neurocomputing2
2026 Adversary's adversary can be a good friend: Revisiting labels of low-margin examples to reconcile accuracy and robustness
abstract
Adversarial training (AT) is widely recognized as one of the most effective methods for improving the robustness of deep learning models. However, AT suffers from a fundamental trade-off between robustness and generalization, which has motivated various mitigation strategies. Among them, margin-based AT approaches employ loss reweighting, reflecting the idea that more critical examples should contribute larger gradients. Yet, these methods are limited by their exclusive focus on gradient magnitude. In this work, we identify that prior approaches overlook the role of gradient direction, and we provide both theoretical and empirical evidence to support this claim. We argue that both the magnitude and direction of gradients should be considered in adversarial training, and propose a novel label design framework, ADA-Lab ( ADversary’s Adversary for Label adjustment ), which incorporates both aspects to refine supervision for low-margin examples. Specifically, we introduce the concept of the adversary’s adversary to explicitly encode directional information aligned with gradient descent. Our theoretical analysis shows that labels designed using this concept better approximate the true label distribution, especially for low-margin examples (i.e., more important examples). Furthermore, by estimating example importance based on the distance to the decision boundary, our method adaptively controls the degree of label interpolation. Our key novelty lies in introducing direction-aware label refinement based on the adversary’s adversary, a concept that explicitly leverages the gradient descent direction of adversarial inputs to correct label mismatch. This unified design integrates gradient magnitude-based importance weighting and label distribution correction, resulting in improved robustness and generalization, as demonstrated by extensive theoretical and empirical results. • We propose a new perspective on adversarial training by introducing direction-aware label refinement based on the adversary’s adversary concept, which has not been explored in prior margin-based or label smoothing methods. • Our framework unifies the benefits of gradient magnitude-based importance weighting and label distribution correction, offering a principled and scalable approach to improving adversarial robustness. • We provide both theoretical and empirical evidence that reducing margin variance and distribution mismatch leads to a tighter bound on natural and robust risk.
Yoojin Jung, Byung Cheol Song
Neurocomputing3
2026 Mask-based adaptive response distillation for efficient image super-resolution
abstract
Knowledge distillation (KD) is a promising strategy for lightweight image super-resolution (ISR). However, most existing methods rely on vanilla response distillation, which fails to account for the importance of high-frequency structures and treats all pixels equally—often leading to suboptimal learning. To address these issues, we propose a simple yet effective KD framework that is both cost-efficient and architecture-agnostic. Our method introduces: (1) a mask-based adaptive response distillation strategy that emphasizes hard instances and pixels via instance-wise suppression and pixel-wise weighting; and (2) a GT-pretrained student initialization scheme that improves optimization stability and performance. Unlike prior works that add significant training overhead, our framework enhances KD effectiveness without requiring additional modules or labels. Extensive experiments on both CNN and Transformer-based SR models demonstrate that our method consistently outperforms baseline and state-of-the-art KD techniques across multiple datasets and upscaling factors, offering strong generalization and plug-in flexibility.
Suho Son, Jeonghyeok Park, Byung Cheol Song
Neurocomputing3
2026 Structure-aware efficient compression for dental image segmentation using differentiable gates and masked knowledge distillation
Jae-Hwan Han, Tae-Hoon Yong, Soon Hyoung Pyo, Byung Cheol Song
Multim. Tools Appl.5
2026 Controllable Facial Expression Synthesis via implicit keypoints and continuous labels
abstract
Facial Expression Synthesis (FES) is important for applications in virtual reality, human–computer interaction, and digital media. While recent facial animation methods have shown strong visual quality under image-driven settings, continuous-label-driven FES remains challenging because facial expressions must be generated from compact semantic conditions without paired target images, making the problem inherently under-constrained. In this paper, we propose VAPortraits, a framework for controllable facial expression synthesis that manipulates implicit keypoints through an Expression Adjustment Network (EAN). Rather than relying on driving-image supervision, our method learns label-conditioned updates in the implicit keypoint space while preserving the strong generative prior of a fixed keypoint-based synthesis backbone. We consider two continuous-label representations: valence-arousal (VA) and valence-eye-lip (VEL), where the latter replaces arousal with explicit eye and lip openness ratios for more practical control of salient facial motions. To alleviate the trade-off between semantic controllability and identity preservation in this regime, we further introduce a selective regularization loss on statistically identified dead elements. Extensive experiments demonstrate that VAPortraits achieves strong expression accuracy, identity retention, and visual quality, providing an effective solution for realistic and controllable facial expression synthesis under low-dimensional continuous conditions. • Novel framework for facial expression synthesis using continuous labels and keypoints. • Expression Adjustment Network (EAN) enables expression editing with continuous labels. • Selective regularization balances identity preservation and expression flexibility. • VEL model (valence, eye/lip ratios) explicitly controls eye and lip opening. • Framework achieves realistic and diverse expressions with controllable conditions.
Yeong Min Lee, Byung Cheol Song
Pattern Recognit.2
2025 Two is Better than One: Efficient Ensemble Defense for Robust and Compact Models
abstract
Deep learning-based computer vision systems adopt complex and large architectures to improve performance, yet they face challenges in deployment on resource-constrained mobile and edge devices. To address this issue, model compression techniques such as pruning, quantization, and matrix factorization have been proposed; however, these compressed models are often highly vulnerable to adversarial attacks. We introduce the Efficient Ensemble Defense (EED) technique, which diversifies the compression of a single base model based on different pruning importance scores and enhances ensemble diversity to achieve high adversarial robustness and resource efficiency. EED dynamically determines the number of necessary sub-models during the inference stage, minimizing unnecessary computations while maintaining high robustness. On the CIFAR- 10 and SVHN datasets, EED demonstrated state-of-the-art robustness performance compared to existing adversarial pruning techniques, along with an inference speed improvement of up to 1.86 times. This proves that EED is a powerful defense solution in resource-constrained environments.
Yoojin Jung, Byung Cheol Song
CVPR2
2025 MECA: Manipulation With Emotional Intensity-Aware Contrastive Learning and Attention-Based Discriminative Learning
abstract
With recent developments in deep learning, facial expression manipulation (FEM) has become one of the fields receiving great attention. However, many studies focus on learning without considering class distinction in latent space. This paper introduces a representation learning scheme that leverages self-attention and mutual information to effectively account for semantic attributes, such as facial expressions, in the FEM task. Our framework, utilizing attention-based discriminative learning and emotional intensity-aware contrastive learning, is capable of forming a compact embedding space. This compact embedding space can lead to more discerning and richer facial expression synthesis in actual synthesis results. As a result, we have derived facial expression synthesis results that are superior to the previous methods. Also, in terms of the FED metric, which can quantify the degree of facial expression expression in FEM, the proposed method outperforms the other methods. To demonstrate this successful result, we use t-SNE and visualize the actual embedding results for each class. Furthermore, we prove that the latent space formed through the proposed method is also helpful in terms of facial expression recognition.
Byung Cheol Song
IEEE Trans. Affect. Comput.2
2025 DIRE: Enhancing Facial Expression Recognition Through Domain-Invariant Representation Learning for Robust Generalization
abstract
In this paper, we propose DIRE (Domain-Invariant Representation Learning for Expression), a novel approach to enhance the generalizability of facial expression recognition (FER) models in unseen domains. Traditional FER models often struggle with distribution shifts between training and test datasets, leading to significant performance drops. Based on the concept of Single-Source Domain Generalization, we introduce a novel domain augmentation technique that applies pixel-level and feature-level perturbations to domain-variant regions while preserving semantic consistency. Additionally, we incorporate semantic alignment regularization and domain information minimization loss so that domain-invariant features effectively represent facial expressions. Extensive experiments on multiple FER datasets demonstrate that our method significantly improves generalization across diverse target domains, even when trained on a single source domain. The proposed DIRE approach offers a robust solution to real-world FER tasks, where unseen domain generalizability is crucial.
Heeje Kim, Yoojin Jung, Byung Cheol Song
IEEE Trans. Multim.3
2024 All You Need Is Your Voice: Emotional Face Representation with Audio Perspective for Emotional Talking Face Generation
Byung Cheol Song
ECCV (64)2
2024 Beyond superficial emotion recognition: Modality-adaptive emotion recognition system
Dohee Kang, Dae Ha Kim, Taein Kim, Bowon Lee, Deok-Hwan Kim, Byung Cheol Song
Expert Syst. Appl.7
2024 Towards the adversarial robustness of facial expression recognition: Facial attention-aware adversarial training
abstract
Beyond the in-the-lab environment, deep-learning-based facial expression recognition (FER) models that provide reliable performance on wild datasets are gradually becoming applied to the real world. However, the fact that neural networks are inherently vulnerable to digital attacks (e.g., adversarial examples) and their performance is not exposed to external threats reduces the applicability of FER technology. So, we design a so-called test-time attack scenario in which FER models are deceived by superimposing imperceptible perturbation(s) on test images. This scenario, which targets the testing phase in which model weakness is revealed, clearly shows how vulnerable FER models are to external attacks. As a remedy against this attack, we propose a novel method called FAAT, which adversarially trains the model by paying attention to core region(s) of face. FAAT aims to improve model robustness so that the model can be generalized to unseen perturbation(s) while focusing on facial expression-related areas. For example, FAAT’s robustness against PGD attack with a performance improvement of up to 18% is encouraging. Also, various benchmarking results based on our attack scenario analyze the fidelity of prior arts and will promote the development direction of future models.
Daeha Kim, Heeje Kim, Yoojin Jung, Byung Cheol Song
Neurocomputing5
2024 Fast Filter Pruning via Coarse-to-Fine Neural Architecture Search and Contrastive Knowledge Transfer
abstract
Filter pruning is the most representative technique for lightweighting convolutional neural networks (CNNs). In general, filter pruning consists of the pruning and fine-tuning phases, and both still require a considerable computational cost. So, to increase the usability of CNNs, filter pruning itself needs to be lightweighted. For this purpose, we propose a coarse-to-fine neural architecture search (NAS) algorithm and a fine-tuning structure based on contrastive knowledge transfer (CKT). First, candidates of subnetworks are coarsely searched by a filter importance scoring (FIS) technique, and then the best subnetwork is obtained by a fine search based on NAS-based pruning. The proposed pruning algorithm does not require a supernet and adopts a computationally efficient search process, so it can create a pruned network with higher performance at a lower cost than the existing NAS-based search algorithms. Next, a memory bank is configured to store the information of interim subnetworks, i.e., by-products of the above-mentioned subnetwork search phase. Finally, the fine-tuning phase delivers the information of the memory bank through a CKT algorithm. Thanks to the proposed fine-tuning algorithm, the pruned network accomplishes high performance and fast convergence speed because it can take clear guidance from the memory bank. Experiments on various datasets and models prove that the proposed method has a significant speed efficiency with reasonable performance leakage over the state-of-the-art (SOTA) models. For example, the proposed method pruned the ResNet-50 trained on Imagenet-2012 up to 40.01% with no accuracy loss. Also, since the computational cost amounts to only 210 GPU hours, the proposed method is computationally more efficient than SOTA techniques. The source code is publicly available at https://github.com/sseung0703/FFP.
Seunghyun Lee 0001, Byung Cheol Song
IEEE Trans. Neural Networks Learn. Syst.2
2023 Modality-Aware Ood Suppression Using Feature Discrepancy for Multi-Modal Emotion Recognition
abstract
While conventional multi-modal emotion recognition (MER) focuses only on model learning for modality fusion, we are interested in tuning multi-modal data fed to MER model in the testing phase. Tuning the influence of each input may cause MER performance significantly due to the nature of MER datasets consisting of heterogeneous modalities. Thus, we propose a novel approach to detect and suppress a modality that is less useful for emotion prediction based on statistical differences between modality distributions. For the MER dataset annotated with discrete or continuous emotion labels, we experimentally find that the OoD modality adversely affects the prediction. Then, we show that when the proposed suppression method is attached to the backbone techniques in an ad-hoc manner, it can achieve outstanding MER performance improvement of 21%.
Dohee Kang, Somang Kang, Dae Ha Kim, Byung Cheol Song
ICIP4
2023 Foreground-Background Disentanglement based on Image and Feature Co-Learning for 3D-Aware Generative Models
abstract
Recently, studies on generative models using 3D information are active. GIRAFFE, one of the latest 3D-aware generative models, shows better feature disentanglement than existing generative models because it generates an image through volume rendering of independently formed 3D neural feature fields. However, GIRAFFE still suffers from an issue where foreground and background disentanglement is not smooth. In order to accomplish better disentanglement performance than GIRAFFE, we propose co-adversarial learning of the generative model at both image- and feature-levels. As a result of rich simulation experiments, the proposed generative model can produce photo-realistic images with only fewer parameters than existing 3D-aware generative models, along with excellent foreground-background disentanglement performance.
Sanghyuk Lee, Daeha Kim, Byung Cheol Song
VCIP3
2023 Fine Gaze Redirection Learning with Gaze Hardness-aware Transformation
abstract
The gaze redirection is a task to adjust the gaze of a given face or eye image toward the desired direction and aims to learn the gaze direction of a face image through a neural network-based generator. Considering that the prior arts have learned coarse gaze directions, learning fine gaze directions is very challenging. In addition, explicit discriminative learning of high-dimensional gaze features has not been reported yet. This paper presents solutions to overcome the above limitations. First, we propose the feature-level transformation which provides gaze features corresponding to various gaze directions in the latent feature space. Second, we propose a novel loss function for discriminative learning of gaze features. Specifically, features with insignificant or irrelevant effects on gaze (e.g., head pose and appearance) are set as negative pairs, and important gaze features are set as positive pairs, and then pair-wise similarity learning is performed. As a result, the proposed method showed a redirection error of only 2° for the Gaze-Capture dataset. This is a 10% better performance than a state-of-the-art method, i.e., STED. Additionally, the rationale for why latent features of various attributes should be discriminated is presented through activation visualization. Code is available at https://github.com/san9569/Gaze-Redir-Learning
Sangjin Park, Dae Ha Kim, Byung Cheol Song
WACV3
2022 Style-Guided and Disentangled Representation for Robust Image-to-Image Translation
abstract
Recently, various image-to-image translation (I2I) methods have improved mode diversity and visual quality in terms of neural networks or regularization terms. However, conventional I2I methods relies on a static decision boundary and the encoded representations in those methods are entangled with each other, so they often face with ‘mode collapse’ phenomenon. To mitigate mode collapse, 1) we design a so-called style-guided discriminator that guides an input image to the target image style based on the strategy of flexible decision boundary. 2) Also, we make the encoded representations include independent domain attributes. Based on two ideas, this paper proposes Style-Guided and Disentangled Representation for Robust Image-to-Image Translation (SRIT). SRIT showed outstanding FID by 8%, 22.8%, and 10.1% for CelebA-HQ, AFHQ, and Yosemite datasets, respectively. The translated images of SRIT reflect the styles of target domain successfully. This indicates that SRIT shows better mode diversity than previous works.
Jaewoong Choi, Dae Ha Kim, Byung Cheol Song
AAAI3
2022 Emotion-aware Multi-view Contrastive Learning for Facial Emotion Recognition
Dae Ha Kim, Byung Cheol Song
ECCV (13)2
2022 Ensemble Knowledge Guided Sub-network Search and Fine-Tuning for Filter Pruning
Seunghyun Lee 0001, Byung Cheol Song
ECCV (11)2
2022 RPFNET: Complementary Feature Fusion for Hand Gesture Recognition
abstract
Hand gesture recognition (HGR) is one of the most challenging tasks because it is very sensitive to occlusion or background. Various modalities such as RGB, depth, and point cloud as well as their combinations have been proposed to improve the performance of HGR, but the fusion of RGB and point cloud with complementary characteristics has never been attempted. This paper analyzes the synergistic effect of the two complementary modalities, and then proposes a new multi-modal fusion network that quantifies and converges the mutual influence of two modalities. Also, to overcome the inherent limitation that the predicted mutual influence does not match the actual one, we propose the self-labeling-based adaptive guidance. Experimental results show that the proposed method achieved 2.46% higher performance than the SOTA method in the case of the NVGesture dataset.
Do Yeon Kim, Dae Ha Kim, Byung Cheol Song
ICIP3
2022 Image Enhancement for Improved Visibility of Digital Displays Under The Sunlight
abstract
Under the sunlight, images displayed on a digital device are generally perceived to be darker than the original, which leads to a decrease in visibility. So, global luminance compensation or tone mapping adaptive to ambient lighting is required. However, global luminance compensation schemes usually have limitations in chrominance compensation as well as local contrast enhancement. This paper proposes a piece-wise linear curve (PLEC)-based image enhancement to enhance both luminance and chrominance. PLECs are regressed through deep learning. In addition, we present a local contrast enhancement scheme for further visibility improvement. Experimental results show that the proposed method outperforms a prior art in terms of subjective/objective visual quality, with significant run-time reduction.
Heejin Lee, Junmin Lee, Seha Jeong, Seunghyun Lee 0001, Seungwan Yu, Junho Heo, Byung Cheol Song
ICIP7
2022 Optimal Transport-based Identity Matching for Identity-invariant Facial Expression Recognition
abstract
Identity-invariant facial expression recognition (FER) has been one of the challenging computer vision tasks. Since conventional FER schemes do not explicitly address the inter-identity variation of facial expressions, their neural network models still operate depending on facial identity. This paper proposes to quantify the inter-identity variation by utilizing pairs of similar expressions explored through a specific matching process. We formulate the identity matching process as an Optimal Transport (OT) problem. Specifically, to find pairs of similar expressions from different identities, we define the inter-feature similarity as a transportation cost. Then, optimal identity matching to find the optimal flow with minimum transportation cost is performed by Sinkhorn-Knopp iteration. The proposed matching method is not only easy to plug in to other models, but also requires only acceptable computational overhead. Extensive simulations prove that the proposed FER method improves the PCC/CCC performance by up to 10% or more compared to the runner-up on wild datasets. The source code and software demo are available at https://github.com/kdhht2334/ELIM_FER.
Dae Ha Kim, Byung Cheol Song
NeurIPS2
2022 Contextual Gradient Scaling for Few-Shot Learning
abstract
Model-agnostic meta-learning (MAML) is a well-known optimization-based meta-learning algorithm that works well in various computer vision tasks, e.g., few-shot classification. MAML is to learn an initialization so that a model can adapt to a new task in a few steps. However, since the gradient norm of a classifier (head) is much bigger than those of backbone layers, the model focuses on learning the decision boundary of the classifier with similar representations. Furthermore, gradient norms of high-level layers are small than those of the other layers. So, the backbone of MAML usually learns task-generic features, which results in deteriorated adaptation performance in the inner-loop. To resolve or mitigate this problem, we propose contextual gradient scaling (CxGrad), which scales gradient norms of the backbone to facilitate learning task-specific knowledge in the inner-loop. Since the scaling factors are generated from task-conditioned parameters, gradient norms of the backbone can be scaled in a task-wise fashion. Experimental results show that CxGrad effectively encourages the backbone to learn task-specific knowledge in the inner-loop and improves the performance of MAML up to a significant margin in both same- and cross-domain few-shot classification.
Sang Hyuk Lee, Seunghyun Lee 0001, Byung Cheol Song
WACV3
2022 Synthesized rain images for deraining algorithms
abstract
Since most of the rainy scene datasets used for training single image rain removal (SIRR) algorithms are constructed by blending artificial rain streaks with source images, it is difficult for a machine trained with such datasets to understand the patterns of real or realistic rain streaks. So, several studies have been attempted to build a real rainy scene dataset. However, since collecting real rainy scenes itself requires significant costs, the real rainy scene datasets provided by some studies cover only very limited rainy environment(s). This paper presents a new approach to synthesize realistic rainy scenes using GAN, which is a world-first attempt as far as we know. The proposed method builds a representation space to which rain streaks of multiple styles are smoothly mapped by learning the distributions of various rain datasets. The representation space allows control over the generated rain streaks. Also, the proposed method can synthesize multiple rainy scenes per clean (source) scene simultaneously, thereby a synthesized rain image dataset (SyRa) (Dataset can be found here: https://github.com/jaewoong1/SyRa-Synthesized_Rain_dataset) consisting of 11 K clean images and 55 K rainy images was constructed. Finally, this paper provides benchmarking results of several SIRR methods trained with SyRa. This result will be very useful for developing SIRR algorithms that can cope well with the actual rain environment.
Jaewoong Choi, Dae Ha Kim, Sanghyuk Lee, Sang Hyuk Lee, Byung Cheol Song
Neurocomputing5
2022 Balanced knowledge distillation for one-stage object detector
abstract
The latest knowledge distillation (KD) methods have successfully supervised a student model to have a better representation using intermediate layers of a teacher model. However, the previous KD methods did not obtain generalized knowledge for various object scales from a one-stage object detector because the one-stage object detector has a structural property that uses several intermediate layers to extract objects of various scales. In other words, the previous KD methods could not distill and transfer knowledge to intermediate layers of one-stage object detectors in a balanced way. Therefore, we propose a shared knowledge encoder and an averaged prototype transfer to remove or mitigate the distillation and transfer imbalances that adversely affect the KD process. Experimental results show that the proposed KD method outperforms the state-of-the-art methods. For instance, the proposed method provides about 1.3% and 2.2% higher accuracy than the baseline on the PASCAL VOC and MS COCO datasets, respectively.
Sungwook Lee, Seunghyun Lee 0001, Byung Cheol Song
Neurocomputing3
2022 Image based rainfall amount estimation for auto-wiping of vehicles
Jungho Jeon, Dong-Yoon Choi, Jong Min Park, Byung Cheol Song
Neural Comput. Appl.5
2022 Deep Metric Learning With Manifold Class Variability Analysis
abstract
In deep metric learning (DML) techniques, understanding both the local and global characteristics of embedding space is essential. However, conventional DML techniques have two limitations as follows: First, Euclidean distance-based metrics never imply global information such as class variability because they only depend on the physical distance of samples. Second, they assume that the embedding space is simply a vector space which cannot represent complex data features. Therefore, we propose a novel loss function which can fully utilize characteristics of embedding space by using discriminant analysis and nonlinear mapping. With theoretical analysis, the superior performance of the proposed method is verified for the fine-grained retrieval datasets such as Cars196, CUB200-2011, Stanford online products, and In-shop clothes. Source code is available athttps://github.com/kdhht2334/MCVA.
Dae Ha Kim, Byung Cheol Song
IEEE Trans. Multim.2
2022 Knowledge Transfer via Decomposing Essential Information in Convolutional Neural Networks
abstract
Knowledge distillation (KD) from a "teacher" neural network and transfer of the knowledge to a small student network is done to improve the performance of the student network. This method is one of the most popular techniques to lighten convolutional neural networks (CNNs). Many KD algorithms have been proposed recently, but they still cannot properly distill essential knowledge of the teacher network, and the transfer tends to depend on the spatial shape of the teacher's feature map. To solve these problems, we propose a method to transfer knowledge independently of the spatial shape of the teacher's feature map, which is major information obtained by decomposing the feature map through singular value decomposition (SVD). In addition, we present a multitask learning method that enables the student to learn the teacher's knowledge effectively by adaptively adjusting the teacher's constraints to the student's learning speed. Experimental results show that the proposed method performs 2.37% better on the CIFAR100 data set and 2.89% better on the TinyImageNet data set than the state-of-the-art method. The source code is publicly available at https://github.com/sseung0703/KD_methods_with_TF.
Seunghyun Lee 0001, Byung Cheol Song
IEEE Trans. Neural Networks Learn. Syst.2
2021 Contrastive Adversarial Learning for Person Independent Facial Emotion Recognition
abstract
Since most facial emotion recognition (FER) methods significantly rely on supervision information, they have a limit to analyzing emotions independently of persons. On the other hand, adversarial learning is a well-known approach for generalized representation learning because it never requires supervision information. This paper presents a new adversarial learning for FER. In detail, the proposed learning enables the FER network to better understand complex emotional elements inherent in strong emotions by adversarially learning weak emotion samples based on strong emotion samples. As a result, the proposed method can recognize the emotions independently of persons because it understands facial expressions more accurately. In addition, we propose a contrastive loss function for efficient adversarial learning. Finally, the proposed adversarial learning scheme was theoretically verified, and it was experimentally proven to show state of the art (SOTA) performance.
Dae Ha Kim, Byung Cheol Song
AAAI2
2021 Interpretable Embedding Procedure Knowledge Transfer via Stacked Principal Component Analysis and Graph Neural Network
abstract
Knowledge distillation (KD) is one of the most useful techniques for light-weight neural networks. Although neural networks have a clear purpose of embedding datasets into the low-dimensional space, the existing knowledge was quite far from this purpose and provided only limited information. We argue that good knowledge should be able to interpret the embedding procedure. This paper proposes a method of generating interpretable embedding procedure (IEP) knowledge based on principal component analysis, and distilling it based on a message passing neural network. Experimental results show that the student network trained by the proposed KD method improves 2.28% in the CIFAR100 dataset, which is a higher performance than the state-of-the-art (SOTA) method. We also demonstrate that the embedding procedure knowledge is interpretable via visualization of the proposed KD process. The implemented code is available at https://github.com/sseung0703/IEPKT.
Seunghyun Lee 0001, Byung Cheol Song
AAAI2
2021 Virtual sample-based deep metric learning using discriminant analysis
Dae Ha Kim, Byung Cheol Song
Pattern Recognit.2
2020 Deep Learning-Based Pupil Center Detection for Fast and Accurate Eye Tracking System
Kangil Lee 0004, Jungho Jeon, Byung Cheol Song
ECCV (19)3
2020 Channel Pruning Via Gradient Of Mutual Information For Light-Weight Convolutional Neural Networks
abstract
Channel pruning for light-weighting networks is very effective in reducing memory footprint and computational cost. Many channel pruning methods assume that the magnitude of a particular element corresponding to each channel reflects the importance of the channel. Unfortunately, such an assumption does not always hold. To solve this problem, this paper proposes a new method to measure the importance of channels based on gradients of mutual information. The proposed method computes and measures gradients of mutual information during back-propagation by arranging a module capable of estimating mutual information. By using the measured statistics as the importance of the channel, less important channels can be removed. Finally, the fine-tuning enables robust performance restoration of the pruned model. Experimental results show that the proposed method provides better performance with smaller parameter sizes and FLOPs than the conventional schemes.
Min Kyu Lee, Seunghyun Lee 0001, Sang Hyuk Lee, Byung Cheol Song
ICIP4
2020 Slice-Based Super-Resolution Using Light-Weight Network With Relation Loss
abstract
While the performance of convolutional neural networks (CNNs)-based single image super-resolution (SISR) has been greatly improved, the enormous parameter sizes and computational complexity of the underlying CNNs make hardware implementation difficult. Recently, several lightweight SISR methods have been developed, but they still do not consider various structural problems that may occur in hardware implementation. To solve this problem, we propose a slice-based SR using light-weight network (LWN) and a slice-based SR using LWN with relation loss (LWNRL). First, LWN(RL) adopts a slice-based architecture to facilitate system-on-chip (SoC) implementation. Second, LWN(RL) avoids global connection modules that are not suitable for SoC implementation, with minimal performance penalty. Finally, we propose a new loss to improve the performance of LWN without additional cost. Experimental results show that LWNRL achieves significant efficiency of SR model. Especially, the larger the resolution or scale factor, the better the performance of LWNRL than the conventional methods.
Ji-Yun Park, Dong-Yoon Choi, Byung Cheol Song
ICIP3
2020 Real-time purchase behavior recognition system based on deep learning-based object detection and tracking for an unmanned product cabinet
Dae Ha Kim, Seunghyun Lee 0001, Jungho Jeon, Byung Cheol Song
Expert Syst. Appl.4
2020 Eye pupil localization algorithm using convolutional neural networks
Jun Ho Choi, Kangil Lee 0004, Byung Cheol Song
Multim. Tools Appl.3
2020 Semi-supervised learning for facial expression-based emotion recognition in the continuous domain
Dong-Yoon Choi, Byung Cheol Song
Multim. Tools Appl.2
2019 Graph-based Knowledge Distillation by Multi-head Attention Network
Seunghyun Lee 0001, Byung Cheol Song
BMVC2
2019 Visual Scene-aware Hybrid Neural Network Architecture for Video-based Facial Expression Recognition
abstract
With rapid development of deep learning, facial expression recognition (FER) technology has made considerable progress recently. However, since conventional FER techniques are mainly designed and learned for videos which are artificially acquired in a limited environment, they may not operate robustly on videos acquired in a wild environment. To solve this problem, this paper proposes a scene-aware hybrid neural network (NN) having a novel combination of three-dimensional (3D) convolutional NN (CNN), 2D CNN and recurrent NN (RNN). The characteristics of the proposed network are as follows. First, we extract video-based global features and frame-based local features at the same time. In detail, the latent features containing the overall visual scene of a given video are extracted by 3D CNN with auxiliary classifier, and fine-tuned 2D CNN is adopted to extract latent features containing small details from each frame. Second, RNN not only performs temporal domain learning, but also feature-wise fuses two latent features extracted from the networks. For effective fusion, we also present three RNN schemes. Third, the proposed network, in which the above-mentioned methods collaborate, works very robust in a wild environment as well as in a limited environment. Extensive experiments show that the proposed network provides an average accuracy of 49.9% for AFEW dataset, i.e., a representative wild dataset, and an amazing accuracy of 98.2% for another CK+ dataset. We also show that the proposed network outperforms the state-of-the-art network(s).
Min Kyu Lee, Dong-Yoon Choi, Dae Ha Kim, Byung Cheol Song
FG4
2019 Accurate Eye Pupil Localization Using Heterogeneous CNN Models
abstract
Eye pupil localization is one of the indispensable technologies in various computer vision applications such as virtual reality and augmented reality. In general, the algorithm consists of finding the approximate eye region and finding the pupil position by extracting the semantic feature from each eye region. However, the performance is affected not only by illumination and image resolution but also by glasses wear. Therefore, this paper proposes an eye pupil localization algorithm which is robust against the above disturbance conditions and also has high accuracy using heterogeneous CNN models. First, faces in the image and landmarks in the face(s) are detected sequentially, and the eye region is determined based on the landmarks. Especially, if glasses are present, the glasses are removed by GAN to find the correct eye region. Next, the pupil region is segmented using fully convolutional networks. Finally, the position of the segmented pupil is calculated. Experimental results show that the proposed algorithm outperforms the state-of-the-art algorithms for public databases such as BioID and GI4E.
Jun Ho Choi, Kangil Lee 0004, Young Chan Kim, Byung Cheol Song
ICIP4
2019 Facial Expression Recognition via Relation-based Conditional Generative Adversarial Network
abstract
Recognizing emotions by adapting to various human identities is very difficult. In order to solve this problem, this paper proposes a relation-based conditional generative adversarial network (RcGAN), which recognizes facial expressions by using the difference (or relation) between neutral face and expressive face. The proposed method can recognize facial expression or emotion independently of human identity. Experimental results show that the proposed method provides higher accuracies of 97.93% and 82.86% for CK+ and MMI databases, respectively than conventional method.
Byung Cheol Song, Min Kyu Lee, Dong-Yoon Choi
ICMI1
2019 Demosaicking algorithm for white-RGB CFA images
abstract
WRGB colour filter array (CFA) has attracted much attention because it is structurally advantageous to improve image quality in low‐light environment using high sensitivity of W channel. However, the demosaicking techniques for WRGB CFA image sensors developed so far have suffered from blurring at edges and deterioration in image quality due to the lack of correlation between W and RGB channels. In order to overcome the above problems, this paper proposes a correlation error compensation in W channel and a G channel restoration to mitigate blurring via edge adaptive filtering. In addition, the authors propose a brightness enhancement method utilising W channel, while avoiding noise boosting. Experimental results show that the proposed demosaicking algorithm not only shows better subjective visual quality than the existing technique but also has about 15% higher signal‐to‐noise ratio (SNR) than conventional Bayer CFA image.
Jun Ho Choi, Dong-Yoon Choi, Byung Cheol Song
IET Image Process.3
2019 Macro unit-based convolutional neural network for very light-weight deep learning
Dae Ha Kim, Min Kyu Lee, Byung Cheol Song
Image Vis. Comput.4
2018 Self-supervised Knowledge Distillation Using Singular Value Decomposition
Seunghyun Lee 0001, Dae Ha Kim, Byung Cheol Song
ECCV (6)3
2018 Recognizing Fine Facial Micro-Expressions Using Two-Dimensional Landmark Feature
abstract
Emotion recognition based on facial expressions is very important for interaction between human and artificial intelligence (AI) system such as social robots. On the other hand, it is much harder to recognize subtle facial expressions or facial micro-expressions than facial expressions rich in emotional expression in a real environment. In this paper, we propose a two-dimensional (2D) landmark feature for effectively recognizing facial micro-expression. The proposed 2D landmark feature is obtained by converting existing coordinate-based landmark information into 2D image information, and has an advantage of having a unique feature according to emotions regardless of the intensity of facial expression. Thus, we can achieve effective emotion recognition by learning the proposed 2D landmark feature information on a convolutional neural network (CNN) and a long-term term memory (LSTM)-based network. Experimental results show that the proposed method provides more than 77% classification performance for fine facial expression images even when learning with general facial expression images of CK+ dataset.
Dong-Yoon Choi, Dae Ha Kim, Byung Cheol Song
ICIP3
2018 Infrared image super-resolution using auxiliary convolutional neural network and visible image under low-light conditions
Tae Young Han, Dae Ha Kim, Byung Cheol Song
J. Vis. Commun. Image Represent.4
2018 Power-Constrained Image Enhancement Using Multiband Processing for TFT LCD Devices With an Edge LED Backlight Unit
abstract
This paper presents a subband-decomposed image-enhancement algorithm that can preserve the luminance and contrast levels in power-controllable thin-film-transistor liquid-crystal display devices. The proposed algorithm consists of three steps. First, an input image is decomposed to multiscale subbands via subband decomposition. Second, luminance compensation and contrast preservation are achieved simultaneously by adjusting the gain of each subband properly. Third, additional power consumption caused by excessive boosting, while coping with the local dimming through specific clipping, is prevented. The experimental results show that the proposed algorithm provides more natural details and contrast, while preserving more similarities with the original, than the existing methods. Even in terms of the peak signal-to-noise-ratio and the structural similarity index, the proposed algorithm has 27.8% and 3.8% higher values compared to a state-of-the-art method. The real-time operation of the proposed algorithm was verified through intensive optimization and the Compute Unified Device Architecture implementation. In addition, the power-reduction effect of the proposed algorithm on commercial monitors was investigated experimentally.
Dong-Yoon Choi, Byung Cheol Song
IEEE Trans. Circuits Syst. Video Technol.2
2018 Sharpness Enhancement and Super-Resolution of Around-View Monitor Images
abstract
In the wide-angle (WA) images embedded in an around-view monitor system, the subject(s) in the peripheral region is normally small and has little information. Furthermore, since the outer region suffers from the non-uniform blur phenomenon and artifact caused by the inherent optical characteristic of WA lenses, its visual quality tends to deteriorate. According to our experiments, conventional image enhancement techniques rarely improve the degraded visual quality of the outer region of WA images. In order to solve the above-mentioned problem, this paper proposes a joint sharpness enhancement (SE) and super-resolution (SR) algorithm which can improve the sharpness and resolution of WA images together. The proposed SE algorithm improves the sharpness of the deteriorated WA images by exploiting self-similarity. Also, the proposed SR algorithm generates super-resolved images by using high-resolution information which is classified according to the extended local binary pattern-based classifier and learned on a pattern basis. Experimental results show that the proposed scheme effectively improves the sharpness and resolution of the input deteriorated WA images. Even in terms of quantitative metrics such as just noticeable blur, structural similarity, and peak signal-to-noise ratio. Finally, the proposed scheme guarantees real-time processing such that it achieves 720p video at 29 Hz on a low cost GPU platform.
Dong-Yoon Choi, Ji Hoon Choi, Jin Wook Choi, Byung Cheol Song
IEEE Trans. Intell. Transp. Syst.4
2017 CNN-based pre-processing and multi-frame-based view transformation for fisheye camera-based AVM system
abstract
The edges of the wide angle (WA) image generally have poor definition and resolution, which often causes deterioration of the around view monitor (AVM) image quality. This paper proposes a convolutional neural network (CNN)-based preprocessing and a multi-frame-based view transformation to solve this problem, and presents an AVM system based on these methods. First, we analyze the general distortion characteristics of the WA image, and propose a preprocessing using the CNN learning model based on the analysis result. Next, in the view transformation (VT) of the outer edge of the WA image, the inherent problem of low pixel density is solved through motion compensation and hole filling using adjacent frames. Experimental results show that the AVM images by the proposed methods are superior to general AVM images in terms of objective image quality as well as subjective image quality.
Dong-Yoon Choi, Ji Hoon Choi, Jin Wook Choi, Byung Cheol Song
ICIP4
2017 Multi-modal emotion recognition using semi-supervised learning and multiple neural networks in the wild
abstract
Human emotion recognition is a research topic that is receiving continuous attention in computer vision and artificial intelligence domains. This paper proposes a method for classifying human emotions through multiple neural networks based on multi-modal signals which consist of image, landmark, and audio in a wild environment. The proposed method has the following features. First, the learning performance of the image-based network is greatly improved by employing both multi-task learning and semi-supervised learning using the spatio-temporal characteristic of videos. Second, a model for converting 1-dimensional (1D) landmark information of face into two-dimensional (2D) images, is newly proposed, and a CNN-LSTM network based on the model is proposed for better emotion recognition. Third, based on an observation that audio signals are often very effective for specific emotions, we propose an audio deep learning mechanism robust to the specific emotions. Finally, so-called emotion adaptive fusion is applied to enable synergy of multiple networks. In the fifth attempt on the given test set in the EmotiW2017 challenge, the proposed method achieved a classification accuracy of 57.12%.
Dae Ha Kim, Min Kyu Lee, Dong-Yoon Choi, Byung Cheol Song
ICMI4
2017 Fast super-resolution algorithm using rotation-invariant ELBP classifier and hierarchical pattern matching
Dong-Yoon Choi, Byung Cheol Song
J. Vis. Commun. Image Represent.2
2015 Fast super-resolution algorithm using ELBP classifier
abstract
This paper proposes a fast super-resolution (SR) algorithm using content-adaptive two-dimensional (2D) finite impulse response (FIR) filters. The proposed algorithm consists of a learning stage and an inference stage. In the learning stage, we cluster a sufficient number of low-resolution (LR) and high-resolution (HR) patch pairs into a specific number of groups using a specific classifier, and we compute the optimal 2D FIR filter to synthesize a high-quality HR patch from an LR patch per cluster, and store the patch-adaptive 2D FIR filters in a dictionary. In the inference stage, from the dictionary, we find the best matched candidate to each input LR patch in terms of the same classifier as the learning stage, and synthesize the HR patch by using the optimal 2D FIR filter corresponding to the best matched candidate. The experimental results show that the proposed algorithm produces HR images of similar quality to the existing SR methods on a per patch basis, while providing fast running time.
Dong-Yoon Choi, Byung Cheol Song
VCIP2
2015 Multi-frame de-raining algorithm using a motion-compensated non-local mean filter for rainy video sequences
Hak Gu Kim, Seung Ji Seo, Byung Cheol Song
J. Vis. Commun. Image Represent.3
2015 Multi-image high dynamic range algorithm using a hybrid camera
Byungju Lee, Byung Cheol Song
Signal Process. Image Commun.2
2014 Subpixel-based image downsampling algorithm using content-adaptive two-dimensional FIR filters
abstract
This study proposes a subpixel‐based image downsampling algorithm using content‐adaptive two‐dimensional (2D) finite impulse response (FIR) filters. The proposed algorithm consists of a learning stage and an inference stage. In the learning stage, using a sufficient number of low‐resolution (LR) and high‐resolution (HR) patch pairs, the authors compute optimal 2D FIR filters to synthesise LR patches of the highest quality from a specific HR patch and store the patch‐adaptive 2D FIR filters in a dictionary. In the inference stage, they explore candidates that best match to each HR input patch in the dictionary and synthesise LR patches by using their corresponding 2D FIR filters on a subpixel basis. The experimental results show that the proposed algorithm produces higher‐quality LR images on a patch basis than existing methods and entails no blur and aliasing artefacts.
Yeon-Oh Nam, Byung Cheol Song
IET Image Process.2
2014 Fast 3D video stabilization using ROI-based warping
Tae Hwan Lee, Yun-Gu Lee, Byung Cheol Song
J. Vis. Commun. Image Represent.3
2014 Power-Constrained Contrast Enhancement Algorithm Using Multiscale Retinex for OLED Display
abstract
This paper presents a power-constrained contrast enhancement algorithm for organic light-emitting diode display based on multiscale retinex (MSR). In general, MSR, which is the key component of the proposed algorithm, consists of power controllable log operation and subbandwise gain control. First, we decompose an input image to MSRs of different sub-bands, and compute a proper gain for each MSR. Second, we apply a coarse-to-fine power control mechanism, which recomputes the MSRs and gains. This step iterates until the target power saving is accurately accomplished. With video sequences, the contrast levels of adjacent images are determined consistently using temporal coherence in order to avoid flickering artifacts. Finally, we present several optimization skills for real-time processing. Experimental results show that the proposed algorithm provides better visual quality than previous methods, and a consistent power-saving ratio without flickering artifacts, even for video sequences.
Yeon-Oh Nam, Dong-Yoon Choi, Byung Cheol Song
IEEE Trans. Image Process.3
2013 Video deblurring based on bidirectional motion compensation and accurate blur kernel estimation
abstract
This paper presents a video deblurring algorithm utilizing the high resolution information of adjacent unblurred frames. First, two motion-compensated predictors of a blurred frame are derived from its neighboring unblurred frames via bidirectional motion compensation. Then, an accurate blur kernel, which is difficult to directly obtain from the blurred frame itself, is computed between the predictors and the blurred frame. Next, a residual deconvolution is employed to reduce the ringing artifacts inherently caused by conventional deconvolution. The blur kernel estimation and deconvolution processes are iteratively performed for the deblurred frame. Experimental results show that the proposed algorithm provides sharper details and smaller artifacts than the state-of-the-art algorithms.
Dong-Bok Lee, Bo-Young Heo, Byung Cheol Song
ICIP3
2013 A content-adaptive sharpness enhancement algorithm using 2D FIR filters trained by pre-emphasis
Ik Hyun Choi, Yeon-Oh Nam, Byung Cheol Song
J. Vis. Commun. Image Represent.3
2013 Video Deblurring Algorithm Using Accurate Blur Kernel Estimation and Residual Deconvolution Based on a Blurred-Unblurred Frame Pair
abstract
Blurred frames may happen sparsely in a video sequence acquired by consumer devices such as digital camcorders and digital cameras. In order to avoid visually annoying artifacts due to those blurred frames, this paper presents a novel motion deblurring algorithm in which a blurred frame can be reconstructed utilizing the high-resolution information of adjacent unblurred frames. First, a motion-compensated predictor for the blurred frame is derived from its neighboring unblurred frame via specific motion estimation. Then, an accurate blur kernel, which is difficult to directly obtain from the blurred frame itself, is computed using both the predictor and the blurred frame. Next, a residual deconvolution is applied to both of those frames in order to reduce the ringing artifacts inherently caused by conventional deconvolution. The blur kernel estimation and deconvolution processes are iteratively performed for the deblurred frame. Simulation results show that the proposed algorithm provides superior deblurring results over conventional deblurring algorithms while preserving details and reducing ringing artifacts.
Dong-Bok Lee, Shin-Cheol Jeong, Yun-Gu Lee, Byung Cheol Song
IEEE Trans. Image Process.4
2012 Block Adaptive Interpolation Filter Using Trained Dictionary for Sub-Pixel Motion Compensation
abstract
Adaptive interpolation filtering for sub-pel motion compensation is one of key techniques of ITU-T key technology area (KTA) codec. However, the adaptive interpolation filtering has a limitation in coding efficiency because of its frame-based update strategy of filter coefficients. Although switched interpolation filter with offset is presented as a sort of block-adaptive filtering for KTA codec, its coding efficiency is generally lower than that of the best adaptive interpolation filter. In order to overcome such a problem, this paper presents an advanced block-adaptive interpolation filtering using well-trained dictionaries which store optimized filter coefficients. We derive those filter coefficients by using learning-based super-resolution. The proposed block-adaptive interpolation filtering for quarter-pel motion compensation consists of two steps: up-scaling of half-pel accuracy and subsequent up-scaling of quarter-pel accuracy. The dictionary optimized for each step is employed to produce the precise up-scaled pixels. Simulation results show that the proposed algorithm improves higher coding efficiency than the previous adaptive interpolation filters for KTA.
Jaehyun Cho, Shin-Cheol Jeong, Dong-Bok Lee, Byung Cheol Song
IEEE Trans. Circuits Syst. Video Technol.4
2011 Video deblurring algorithm using an adjacent unblurred frame
abstract
Blurred frames may sparsely exist in a video sequence acquired by digital camcorder or digital camera. In order to remove the visually annoying artifact due to those blurred frames, this paper presents a novel motion deblurring algorithm where a blurred frame can be reconstructed utilizing adjacent unblurred frames. Firstly, a motion-compensated predictor of the blurred frame is derived from its neighboring unblurred frame using motion estimation. Then, an accurate blur kernel, which is difficult to obtain from a single blurred frame, is computed using both the predictor and the blurred frame. Next, again using those both frames, a residual deconvolution is proposed to reduce ringing artifacts inherent to conventional deconvolution. Simulation results show that the proposed algorithm provides superior deblurring results over conventional deblurring algorithms while preserving details with reduced ringing artifacts.
Shin-Cheol Jeong, Tae Hwan Lee, Byung Cheol Song, Yun-Gu Lee, Yanglim Choi
VCIP3
2011 Spatio-temporal de-interlacing based on maximum likelihood estimation
abstract
This paper proposes a novel de-interlacing algorithm that can make up motion compensation (MC) errors by using maximum likelihood (ML) estimator. Firstly, a proper registration is performed between current field and its adjacent fields, and the progressive frame corresponding to the current field is found via ML estimator based on the computed registration information. Here, in order to obtain a stable solution, well-known bilateral total variation (BTV)-based regularization is applied. Next, possible feathering artifacts are detected on a block basis effectively. So, edge-directional interpolation is applied to the pixels where feathering artifact may happen, instead of the above-mentioned temporal de-interlacing. Experimental results show that the PSNR of our proposed algorithm is on average 4dB higher than that of previous studies and provides the best visual quality.
Ho-Taek Lee, Tae Hwan Lee, Byung Cheol Song
VCIP3
2011 Video Super-Resolution Algorithm Using Bi-Directional Overlapped Block Motion Compensation and On-the-Fly Dictionary Training
abstract
This paper presents a video super-resolution algorithm to interpolate an arbitrary frame in a low resolution video sequence from sparsely existing high resolution key-frames. First, a hierarchical block-based motion estimation is performed between an input and low resolution key-frames. If the motion-compensated error is small, then an input low resolution patch is temporally super-resolved via bi-directional overlapped block motion compensation. Otherwise, the input patch is spatially super-resolved using the dictionary that has been already learned from the low resolution and its corresponding high resolution key-frame pair. Finally, possible blocking artifacts between temporally super-resolved patches and spatially super-resolved patches are concealed using a specific de-blocking filter. The experimental results show that the proposed algorithm provides significantly better subjective visual quality as well as higher peak-to-peak signal-to-noise ratio than those by previous interpolation algorithms.
Byung Cheol Song, Shin-Cheol Jeong, Yanglim Choi
IEEE Trans. Circuits Syst. Video Technol.1
2009 Low-complexity near-lossless image coder for efficient bus traffic in very large size multimedia SoC
abstract
With dramatic development of image sensor technology, an image resolution of a digital camera has significantly increased recently. However, frequent memory access of such large-size images often brings out huge bus traffic in image processing chips or SoC's embedded in digital cameras. Bus traffic is a critical issue in real-time image processing. This paper presents a computationally efficient near-lossless image coder to mitigate the bus traffic problem. First, the coder can encode a large image data of up to 6K×4K to write to the memory in near-lossless or lossless form with a desired compression ratio. Second, the operating speed of the coder is extremely fast so that it can be easily applied to real-time AV applications with large image resolution. Third, the coder provides much less complexity than typical standard near-lossless coders such as JPEG-LS. So, the proposed coder may cause only a negligible area overhead in SoC implementation. Finally, the coder can adaptively change its compression ratio adaptively with the bus traffic. Simulation results show that the proposed coder provides reasonable coding performance with much less complexity than the existing coders.
Yun-Gu Lee, Byung Cheol Song, Nak Hoon Kim, Woo Hyun Joo
ICIP2
2009 1080P 60HZ intra-frame CODEC based on RGB color space for wireless AV streaming
abstract
To achieve high visual quality of intra-frame coding in order to minimize the visual quality degradation caused by color loss, the authors previously presented an RGB-domain inter-color compensation algorithm using strong correlation between RGB color components. Based on that inter-color compensation algorithm, this paper presents a 1080p 60Hz CODEC system architecture designed to process a bit-rate of up to approximately 100Mbps in real time. Both the encoding and decoding processes are pipelined on a macroblock level. Since syntax processing is a bottleneck to supporting speeds of up to 100Mbps, a high performance context-adaptive variable length coding architecture exploiting the look-ahead technique is included in the proposed design. The final chip implementation can achieve real-time encoding and decoding of 1080p 60Hz videos with reasonable hardware cost and operating clock frequency.
Byung Cheol Song, Yongseok Yi, Yun-Gu Lee, Jun Hyuk Ko
ICIP1
2009 An Intra-Frame Rate Control Algorithm for Ultralow Delay H.264/Advanced Video Coding (AVC)
abstract
This paper presents an intra-fame rate control algorithm for ultralow delay H.264/AVC coding. The main goal of the proposed scheme is to allocate a proper bit budget to each macroblock (MB), based on its complexity (or activity). However, since a coder needs to start encoding a frame before buffering the whole frame for very low delay video streaming applications, the relative complexity of each MB can not be known in advance. So, the proposed algorithm predicts a relative complexity of a current MB from complexities of its spatially/temporally neighboring MBs, and then allocates a proper bit budget to the MB using the predicted complexity. Finally, quantization parameter of each MB is obtained by comparing the generated bits and allocated bit budget, and is refined for improving coding performance, based on information from the previous frame. Even though the required buffer size is only around one-third of average bits per a frame in order to achieve very low coding delay, the proposed algorithm provides reliable coding performance compared to the conventional schemes. Simulation results show that the proposed algorithm can prevent the buffer overflow or underflow and it provides better coding performance.
Yun-Gu Lee, Byung Cheol Song
IEEE Trans. Circuits Syst. Video Technol.2
2008 An intra-frame rate control algorithm for ultra low delay H.264/AVC coding
abstract
In this paper, we present an intra-fame rate control algorithm for ultra low delay H.264/AVC coding. In real time video coding, all the macro-blocks within a current frame may be unavailable before encoding them. Hence the proposed scheme predicts relative complexity of the current block from complexity of available macro-blocks within previous and current frames. Then, the algorithm allocates bits to each macro-block considering the relative complexity between the macro-block and the current frame. Quantization parameter of each macro-block is obtained by comparing the generated and allocated bits and is refined for improving coding performance. The required buffer size is only around one-third of average bits required for encoding a single frame. Simulation results show that the proposed algorithm can prevent the buffer from overflow and underflow while it provides better coding performance.
Yun-Gu Lee, Byung Cheol Song
ICASSP2
2008 A novel CAVLC architecture for H.264 Video encoding at high bit-rate
abstract
In H.264/AVC and the variants, the coding of context-based adaptive variable length codes (CAVLC) is one of the demanding operations, particularly for high bitrates such as 100Mbps. This paper presents a novel architecture that exploits component-level parallelism and pipeline techniques capable of processing high-bitrate video data in a macroblock(MB)-level pipelined CODEC architecture. Additionally, some techniques for efficient CAVLC coding is presented because CAVLC is a dominant part of the syntax coding. The resulting architecture, merged in a MB-level pipelined CODEC system, is capable of coding up to 100Mbps bitstreams in real-time, thus, accommodating the real-time encoding of 1080p 60Hz video.
Yongseok Yi, Byung Cheol Song
ISCAS2
2008 High-Speed CAVLC Encoder for 1080p 60-Hz H.264 Codec
abstract
In H.264/AVC and the variants, the coding of context-based adaptive variable length codes (CAVLC) requires demanding operations, particularly at high bitrates such as 100 Mbps. This letter presents two approaches to accelerate the coding operation substantially. Firstly, in the architectural aspect, we propose component-level parallelism and pipeline techniques capable of processing high-bitrate video data in a macroblock (MB)-level pipelined codec architecture. The second approach focuses on a specific part of the coding process, i.e., the residual block coding, in which the coefficient levels are coded without using look-up tables so we minimize the pertaining logic depth in the critical path, and we achieve higher operating clock frequencies. Additionally, two coefficient levels are processed in parallel by exploiting a look-ahead technique. The resulting architecture, merged in the MB-level pipelined codec system, is capable of coding up to 100 Mbps bitstreams in real-time, thus accommodating the real-time encoding of 1080p@60 Hz video.
Yongseok Yi, Byung Cheol Song
IEEE Signal Process. Lett.2
2008 Block Adaptive Inter-Color Compensation Algorithm for RGB 4: 4: 4 Video Coding
abstract
RGB color space has not been regarded as a proper space from a coding point of a view. However, due to the limited visual quality of YCbCr-domain video coding for high-quality applications, RGB-domain video coding is being newly raised. Recently, several methods have been developed to reduce significant redundancy existing between RGB color planes in an RGB video coder. This paper presents a new algorithm to remove inter-color redundancy, which is based on a linear model having two parameters, i.e., offset and slope information. Those parameters of one color component block are derived from the other color component block. A common coding mode and model parameters optimized for three color components in a macroblock are obtained simultaneously by minimizing a given cost. Then, a residue for each color component block is produced using its predictor based on the linear model, and is encoded. The simulation results show that the proposed algorithm noticeably improves the coding efficiency up to over 20% in comparison with H.264 new amendment.
Byung Cheol Song, Yun-Gu Lee, Nak Hoon Kim
IEEE Trans. Circuits Syst. Video Technol.1
2006 Fast exhaustive multi-resolution search algorithm based on clustering for efficient image retrieval
Byung Cheol Song, Jong Beom Ra
J. Vis. Commun. Image Represent.1
2005 Noise Power Estimation for Effective De-Noising in a Video Encoder
abstract
This paper presents a noise estimation algorithm using multiresolution motion estimation in a video encoder. Firstly, the motion estimator finds minimum block-matching errors at the finest resolution and the middle resolution for each macroblock. Secondly, if the minimum block-matching error at the finest resolution of a certain macroblock is less than a particular threshold, the variance of the inter-macroblock is computed. And then, we employ the square of the minimum block-matching error at the middle resolution as a predictor of the variance of the desired inter-macroblock without noise. The noise variance in the macroblock can be estimated by subtracting the predictor from the variance of the inter-macroblock. Finally, the noise variances estimated only for the well-motion-compensated macroblocks are averaged in each frame. Experimental results show that the proposed noise estimation is very accurate with negligible computational cost.
Byung Cheol Song, Kang Wook Chun
ICASSP (2)1
2005 Transform-domain Wiener filtering for H.264/AVC video encoding and its implementation
abstract
For coding efficiency as well as noise reduction, efficient de-noising needs to be performed prior to video encoding. This paper proposes a transform-domain Wiener filtering scheme in an H.264/AVC video encoder. We show that the generalized Wiener filtering for each integer-transformed block is equivalent to multiplication of multiplication factor (MF) in a block with a proper filter coefficient matrix in a quantization process. Also, we implement efficiently the proposed scheme by employing several predetermined modified MF's for quantization. Experimental results show that the proposed de-noising scheme provides outstanding coding efficiency as well as noise reduction in an H.264/AVC video encoder.
Byung Cheol Song, Nak Hoon Kim, Kang Wook Chun
ICIP (3)1
2005 A rate-constrained fast full-search algorithm based on block sum pyramid
abstract
This paper presents a fast full-search algorithm (FSA) for rate-constrained motion estimation. The proposed algorithm, which is based on the block sum pyramid frame structure, successively eliminates unnecessary search positions according to rate-constrained criterion. This algorithm provides the identical estimation performance to a conventional FSA having rate constraint, while achieving considerable reduction in computation.
Byung Cheol Song, Kang Wook Chun, Jong Beom Ra
IEEE Trans. Image Process.1
2004 Motion-compensated temporal pre-filtering for noise reduction in a video encoder
abstract
For coding efficiency as well as noise reduction, efficient prefiltering needs to be performed prior to video encoding. This paper presents a DCT-domain temporal filtering scheme in an MPEG video encoder. It is proven that the multiplication of every DCT coefficient in an inter block with a proper weight is equivalent to motion compensated temporal filtering in the spatial domain for the inter block. Also, we propose to determine properly the weight by using effective noise estimation scheme. Experimental results show that the proposed prefiltering scheme provides outstanding coding efficiency as well as noise reduction in an MPEG video encoder.
Byung Cheol Song, Kang Wook Chun
ICIP1
2004 Interpolative mode prediction for efficient MPEG-2 video encoding
abstract
This paper presents an algorithm to predict interpolative motion compensation (MC) mode without computing interpolative SADs (Sum of Absolute Difference) for B-picture coding in an MPEG-2 video encoder. Firstly, an initial MC mode is selected among forward frame/field MC and backward frame/field MC for each macroblock (MB) in a Bpicture. This is accomplished by comparing four minimum SADs corresponding to the above-mentioned MC modes, and choosing a minimum SAD among them. Secondly, if the SAD corresponding to the selected mode is less than a specific threshold, we set the selected MC mode to the final MC mode. Otherwise, an MC mode of the current MB is determined between interpolative frame MC and interpolative field MC. If the sum of SADs of forward frame MC mode and backward frame MC mode is less than that of SADs of forward field MC mode and backward field MC mode, the final MC mode is set to an interpolative frame MC. Otherwise, the final MC mode is set to an interpolative field MC mode. Experimental results show that the proposed algorithm is comparable with as brute-force MC mode selection method as in conventional MPEG-2 video encoding. In addition, the proposed algorithm may noticeably reduce the memory bandwidth in implementation of MPEG-2 video encoder.
Byung Cheol Song, Kang Wook Chun
VCIP1
2004 Multi-resolution block matching algorithm and its VLSI architecture for fast motion estimation in an MPEG-2 video encoder
abstract
This paper proposes a high-performance multi-resolution motion estimation algorithm (HMRME) for MPEG-2 video encoding, which satisfies high estimation performance and efficient very large scale integration (VLSI) implementation. HMRME is based on a characteristic that field motion vectors (MVs) are very similar to their corresponding frame MV. Firstly, HMRME performs frame-based motion estimation (ME) as follows: at the coarsest level, two MV candidates are found on the basis of minimum matching error. The two MV candidates from the coarsest level search and the other one based on spatial MV correlation are used as center points for three local searches at the middle level. At the finest level, a frame MV is obtained from a local search around a single candidate from the middle level search. Field MVs are estimated with the single MV candidate from the middle level search of frame ME as initial estimates at the finest level, without any coarser level searches. This paper also describes a VLSI architecture based on HMRME. This architecture is designed to provide a good tradeoff between on-chip memory size and I/O bandwidth with high throughput. We implemented this architecture with about 140 K gates and 2.5 K bytes static random access memory for a large search range of [-192.0, +191.5] by using a synthesizable Verilog HDL.
Byung Cheol Song, Kang Wook Chun
IEEE Trans. Circuits Syst. Video Technol.1
2003 Motion-compensated noise estimation for efficient prefiltering in a video encoder
abstract
For efficient prefiltering prior to video encoding, exact noise strength of an input video sequence should be found, but it is actually a very difficult process. This paper presents an accurate noise estimation scheme using multiresolution motion estimation in a video encoder. Firstly, if the multiresolution motion estimator finds minimum block-matching errors at the finest resolution and the middle resolution, the differences between the two block-matching errors are computed and averaged for all the macroblocks (MBs) in an input frame. Based on a property that the average difference of each frame is linearly proportional to its noisiness, the noise strength of the input frame is precisely estimated. Experimental results show that the proposed noise estimation is very accurate.
Byung Cheol Song, Kang Wook Chun
ICIP (2)1
2003 Virtual frame rate control for efficient MPEG-2 video encoding
Byung Cheol Song, Kang Wook Chun
VCIP1
2002 Efficient video transcoding with scan format conversion
abstract
General-purpose MPEG-2 video transcoders must be able to achieve any conversion between 18 ATSC (Advanced Television System Committee) video formats for DTV (digital television), e.g., scan format, size format, and frame rate format conversion. Especially, scan format conversion is hard to implement because frame rate and size format conversion often happen together. This paper proposes a fast motion estimation (ME) algorithm for MPEG-2 video transcoding supporting scan format conversion. Firstly, we extract and compose a set of candidate motion vectors (MVs) from the input bit-stream to comply with the re-encoding format. Secondly, the best MV is chosen among several candidate MVs by using a weighted median selector. Simulation results show that the proposed ME algorithm reduces significantly transcoding complexity with a minor PSNR degradation.
Byung Cheol Song, Kang Wook Chun
ICIP (1)1
2002 A fast search algorithm for vector quantization using L2-norm pyramid of codewords
abstract
Vector quantization for image compression requires expensive encoding time to find the closest codeword to the input vector. This paper presents a fast algorithm to speed up the closest codeword search process in vector quantization encoding. By using an appropriate topological structure of the codebook, we first derive a condition to eliminate unnecessary matching operations from the search procedure. Then, based on this elimination condition, a fast search algorithm is suggested. Simulation results show that with little preprocessing and memory cost, the proposed search algorithm significantly reduces the encoding complexity while maintaining the same encoding quality as that of the full search algorithm. It is also found that the proposed algorithm outperforms the existing search algorithms.
Byung Cheol Song, Jong Beom Ra
IEEE Trans. Image Process.1
2001 A novel search algorithm based on L2-norm pyramid of codewords for fast vector quantization encoding
abstract
Vector quantization for image compression requires expensive encoding time to find the closest codeword to the input vector. This paper presents a fast algorithm to speed up the closest codeword search process in vector quantization encoding. By using an appropriate topological structure of the codebook, we first derive a condition to eliminate unnecessary matching operations from the search procedure. Then, based on this elimination condition, a fast search algorithm is suggested. Simulation results show that with little preprocessing and memory cost, the proposed search algorithm significantly reduces the encoding complexity while maintaining the same encoding quality as that of the full search algorithm. It is also found that the proposed algorithm outperforms the existing search algorithms.
Byung Cheol Song, Jong Beom Ra
ICIP (2)1
2001 Automatic Shot Change Detection Algorithm Using Multi-stage Clustering for MPEG-Compressed Videos
Byung Cheol Song, Jong Beom Ra
J. Vis. Commun. Image Represent.1
2001 A fast multi-resolution block matching algorithm and its LSI architecture for low bit-rate video coding
abstract
We propose a fast multi-resolution block-matching algorithm (BMA) using multiple motion vector (MV) candidates and spatial correlation in MV fields, called a multi-resolution motion search algorithm (MRMCS). The proposed MRMCS satisfies high estimation performance and efficient LSI implementation. This paper describes the MRMCS with three resolution levels. At the coarsest level, two MV candidates are obtained on the basis of minimum matching error for the next search level. At the middle level, the two candidates selected at the coarsest level and the other one based on spatial MV correlation at the finest level are used as center points for local searches, and a MV candidate is chosen for the next search level. Then, at the finest level, the final MV is obtained from local search around the single candidate obtained at the middle level. This paper also describes an efficient LSI architecture based on the proposed algorithm for low bit-rate video coding. Since this architecture requires a small number of processing elements (PEs) and a small size on-chip memory, MRMCS can be implemented with a much smaller number of gates than other conventional architectures for full-search BMA while keeping a negligible degradation of coding performance. Moreover, the proposed motion estimator can support an advanced prediction mode (8/spl times/8 prediction mode) for H.263 and MPEG-4 video encoding. We implement this architecture with about 25 K gates and 288 bytes of RAM for a search range of [-16.0, +15.5] by using a synthesizable VHDL.
Jae Hun Lee, Kyoung Won Lim, Byung Cheol Song, Jong Beom Ra
IEEE Trans. Circuits Syst. Video Technol.3
2001 A fast multiresolution feature matching algorithm for exhaustive search in large image databases
abstract
Most of the content-based image retrieval systems require a distance computation for each candidate image in the database. As a brute-force approach, the exhaustive search can be employed for this computation. However, this exhaustive search is time-consuming and limits the usefulness of such systems. Thus, there is a growing demand for a fast algorithm which provides the same retrieval results as the exhaustive search. We propose a fast search algorithm based on a multiresolution data structure. The proposed algorithm computes the lower bound of distance at each level and compares it with the latest minimum distance, starting from the low-resolution level. Once it is larger than the latest minimum distance, we can remove the candidates without calculating the full-resolution distance. By doing this, we can dramatically reduce the total computational complexity. It is noticeable that the proposed fast algorithm provides not only the same retrieval results as the exhaustive search, but also a faster searching ability than existing fast algorithms. For additional performance improvement, we can easily combine the proposed algorithm with existing tree-based algorithms. The algorithm can also be used for the fast matching of various features such as luminance histograms, edge histograms, and local binary partition textures.
Byung Cheol Song, Myung Jun Kim, Jong Beom Ra
IEEE Trans. Circuits Syst. Video Technol.1
2000 Fast image retrieval based on K-means clustering and multiresolution data structure for large image databases
Byung Cheol Song, Myung Jun Kim, Jong Beom Ra
VCIP1
2000 A fast multi-resolution block matching algorithm for motion estimation
Byung Cheol Song, Jong Beom Ra
Signal Process. Image Commun.1
1999 A fast motion estimation algorithm based on multi-resolution frame structure
abstract
We present a novel multi-resolution block matching algorithm (BMA) for fast motion estimation. At the coarsest level, a full search BMA (FSBMA) is performed for searching complex or random motion. Concurrently, spatial correlation of motion vector (MV) field is used for searching continuous motion. Here we present an efficient method for searching full resolution MVs without MV decimation even at the coarsest level. After the coarsest level search, two or three initial MV candidates are chosen for the next level. At the further levels, the MV candidates are refined within much smaller search areas. Simulation results show that in comparison with FSBMA, the proposed BMA achieves a speed-up factor over 710 with minor PSNR degradation of 0.2 dB at most, under a normal MPEG-2 coding environment. Furthermore, our scheme is also suitable for hardware implementation due to regular data-flow.
Byung Cheol Song, Jong Beom Ra
ICASSP1
1999 An Efficient Video Down Conversion Algorithm Using Modified IDCT Basis Functions
abstract
A traditional method for down conversion of MPEG compressed video performs low-pass filtering and subsampling after full decompression. However, the method has two shortcomings: large memory requirement and high computational complexity. To overcome the drawbacks, recent research has focussed on down conversion in the DCT domain. Existing DCT domain methods need only a quarter of full frame memory because completely reduced frames are stored in the frame memory. But the loss of field information often causes significant error during motion compensation. To solve this problem, we adopt a hybrid structure where field information is preserved by performing down conversion in the DCT domain only for horizontal direction and down conversion in the spatial domain for vertical direction. Furthermore, we present a novel technique for the horizontal down conversion which adopts a modified 4-IDCT kernel. The modified 4-IDCT kernel outperforms the conventional 4-IDCT kernel in that it performs low-pass filtering and sub-sampling additionally. Experimental results show that the proposed scheme provides high visual quality while maintaining comparable complexity.
Myung Jun Kim, Byung Cheol Song, Sung Kyu Jang, Jong Beom Ra
ICIP (2)2