VLDB 2026 Research / reviewers in the wild / expert
Dae Ha Kim
dblp:207/9644
· DBLP profile ↗
18ranked-venue papers
8as first author
11since 2021 · last 2024
0000-0003-3838-126XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Beyond superficial emotion recognition: Modality-adaptive emotion recognition system
Dohee Kang, Dae Ha Kim, Taein Kim, Bowon Lee, Deok-Hwan Kim, Byung Cheol Song |
Expert Syst. Appl. | 2 |
| 2023 | Modality-Aware Ood Suppression Using Feature Discrepancy for Multi-Modal Emotion RecognitionabstractWhile conventional multi-modal emotion recognition (MER) focuses only on model learning for modality fusion, we are interested in tuning multi-modal data fed to MER model in the testing phase. Tuning the influence of each input may cause MER performance significantly due to the nature of MER datasets consisting of heterogeneous modalities. Thus, we propose a novel approach to detect and suppress a modality that is less useful for emotion prediction based on statistical differences between modality distributions. For the MER dataset annotated with discrete or continuous emotion labels, we experimentally find that the OoD modality adversely affects the prediction. Then, we show that when the proposed suppression method is attached to the backbone techniques in an ad-hoc manner, it can achieve outstanding MER performance improvement of 21%. Dohee Kang, Somang Kang, Dae Ha Kim, Byung Cheol Song |
ICIP | 3 |
| 2023 | Fine Gaze Redirection Learning with Gaze Hardness-aware TransformationabstractThe gaze redirection is a task to adjust the gaze of a given face or eye image toward the desired direction and aims to learn the gaze direction of a face image through a neural network-based generator. Considering that the prior arts have learned coarse gaze directions, learning fine gaze directions is very challenging. In addition, explicit discriminative learning of high-dimensional gaze features has not been reported yet. This paper presents solutions to overcome the above limitations. First, we propose the feature-level transformation which provides gaze features corresponding to various gaze directions in the latent feature space. Second, we propose a novel loss function for discriminative learning of gaze features. Specifically, features with insignificant or irrelevant effects on gaze (e.g., head pose and appearance) are set as negative pairs, and important gaze features are set as positive pairs, and then pair-wise similarity learning is performed. As a result, the proposed method showed a redirection error of only 2° for the Gaze-Capture dataset. This is a 10% better performance than a state-of-the-art method, i.e., STED. Additionally, the rationale for why latent features of various attributes should be discriminated is presented through activation visualization. Code is available at https://github.com/san9569/Gaze-Redir-Learning Sangjin Park, Dae Ha Kim, Byung Cheol Song |
WACV | 2 |
| 2022 | Style-Guided and Disentangled Representation for Robust Image-to-Image TranslationabstractRecently, various image-to-image translation (I2I) methods have improved mode diversity and visual quality in terms of neural networks or regularization terms. However, conventional I2I methods relies on a static decision boundary and the encoded representations in those methods are entangled with each other, so they often face with ‘mode collapse’ phenomenon. To mitigate mode collapse, 1) we design a so-called style-guided discriminator that guides an input image to the target image style based on the strategy of flexible decision boundary. 2) Also, we make the encoded representations include independent domain attributes. Based on two ideas, this paper proposes Style-Guided and Disentangled Representation for Robust Image-to-Image Translation (SRIT). SRIT showed outstanding FID by 8%, 22.8%, and 10.1% for CelebA-HQ, AFHQ, and Yosemite datasets, respectively. The translated images of SRIT reflect the styles of target domain successfully. This indicates that SRIT shows better mode diversity than previous works. Jaewoong Choi, Dae Ha Kim, Byung Cheol Song |
AAAI | 2 |
| 2022 | Emotion-aware Multi-view Contrastive Learning for Facial Emotion Recognition
Dae Ha Kim, Byung Cheol Song |
ECCV (13) | 1 |
| 2022 | RPFNET: Complementary Feature Fusion for Hand Gesture RecognitionabstractHand gesture recognition (HGR) is one of the most challenging tasks because it is very sensitive to occlusion or background. Various modalities such as RGB, depth, and point cloud as well as their combinations have been proposed to improve the performance of HGR, but the fusion of RGB and point cloud with complementary characteristics has never been attempted. This paper analyzes the synergistic effect of the two complementary modalities, and then proposes a new multi-modal fusion network that quantifies and converges the mutual influence of two modalities. Also, to overcome the inherent limitation that the predicted mutual influence does not match the actual one, we propose the self-labeling-based adaptive guidance. Experimental results show that the proposed method achieved 2.46% higher performance than the SOTA method in the case of the NVGesture dataset. Do Yeon Kim, Dae Ha Kim, Byung Cheol Song |
ICIP | 2 |
| 2022 | Optimal Transport-based Identity Matching for Identity-invariant Facial Expression RecognitionabstractIdentity-invariant facial expression recognition (FER) has been one of the challenging computer vision tasks. Since conventional FER schemes do not explicitly address the inter-identity variation of facial expressions, their neural network models still operate depending on facial identity. This paper proposes to quantify the inter-identity variation by utilizing pairs of similar expressions explored through a specific matching process. We formulate the identity matching process as an Optimal Transport (OT) problem. Specifically, to find pairs of similar expressions from different identities, we define the inter-feature similarity as a transportation cost. Then, optimal identity matching to find the optimal flow with minimum transportation cost is performed by Sinkhorn-Knopp iteration. The proposed matching method is not only easy to plug in to other models, but also requires only acceptable computational overhead. Extensive simulations prove that the proposed FER method improves the PCC/CCC performance by up to 10% or more compared to the runner-up on wild datasets. The source code and software demo are available at https://github.com/kdhht2334/ELIM_FER. Dae Ha Kim, Byung Cheol Song |
NeurIPS | 1 |
| 2022 | Synthesized rain images for deraining algorithmsabstractSince most of the rainy scene datasets used for training single image rain removal (SIRR) algorithms are constructed by blending artificial rain streaks with source images, it is difficult for a machine trained with such datasets to understand the patterns of real or realistic rain streaks. So, several studies have been attempted to build a real rainy scene dataset. However, since collecting real rainy scenes itself requires significant costs, the real rainy scene datasets provided by some studies cover only very limited rainy environment(s). This paper presents a new approach to synthesize realistic rainy scenes using GAN, which is a world-first attempt as far as we know. The proposed method builds a representation space to which rain streaks of multiple styles are smoothly mapped by learning the distributions of various rain datasets. The representation space allows control over the generated rain streaks. Also, the proposed method can synthesize multiple rainy scenes per clean (source) scene simultaneously, thereby a synthesized rain image dataset (SyRa) (Dataset can be found here: https://github.com/jaewoong1/SyRa-Synthesized_Rain_dataset) consisting of 11 K clean images and 55 K rainy images was constructed. Finally, this paper provides benchmarking results of several SIRR methods trained with SyRa. This result will be very useful for developing SIRR algorithms that can cope well with the actual rain environment. Jaewoong Choi, Dae Ha Kim, Sanghyuk Lee, Sang Hyuk Lee, Byung Cheol Song |
Neurocomputing | 2 |
| 2022 | Deep Metric Learning With Manifold Class Variability AnalysisabstractIn deep metric learning (DML) techniques, understanding both the local and global characteristics of embedding space is essential. However, conventional DML techniques have two limitations as follows: First, Euclidean distance-based metrics never imply global information such as class variability because they only depend on the physical distance of samples. Second, they assume that the embedding space is simply a vector space which cannot represent complex data features. Therefore, we propose a novel loss function which can fully utilize characteristics of embedding space by using discriminant analysis and nonlinear mapping. With theoretical analysis, the superior performance of the proposed method is verified for the fine-grained retrieval datasets such as Cars196, CUB200-2011, Stanford online products, and In-shop clothes. Source code is available athttps://github.com/kdhht2334/MCVA. Dae Ha Kim, Byung Cheol Song |
IEEE Trans. Multim. | 1 |
| 2021 | Contrastive Adversarial Learning for Person Independent Facial Emotion RecognitionabstractSince most facial emotion recognition (FER) methods significantly rely on supervision information, they have a limit to analyzing emotions independently of persons. On the other hand, adversarial learning is a well-known approach for generalized representation learning because it never requires supervision information. This paper presents a new adversarial learning for FER. In detail, the proposed learning enables the FER network to better understand complex emotional elements inherent in strong emotions by adversarially learning weak emotion samples based on strong emotion samples. As a result, the proposed method can recognize the emotions independently of persons because it understands facial expressions more accurately. In addition, we propose a contrastive loss function for efficient adversarial learning. Finally, the proposed adversarial learning scheme was theoretically verified, and it was experimentally proven to show state of the art (SOTA) performance. Dae Ha Kim, Byung Cheol Song |
AAAI | 1 |
| 2021 | Virtual sample-based deep metric learning using discriminant analysis
Dae Ha Kim, Byung Cheol Song |
Pattern Recognit. | 1 |
| 2020 | Real-time purchase behavior recognition system based on deep learning-based object detection and tracking for an unmanned product cabinet
Dae Ha Kim, Seunghyun Lee 0001, Jungho Jeon, Byung Cheol Song |
Expert Syst. Appl. | 1 |
| 2019 | Visual Scene-aware Hybrid Neural Network Architecture for Video-based Facial Expression RecognitionabstractWith rapid development of deep learning, facial expression recognition (FER) technology has made considerable progress recently. However, since conventional FER techniques are mainly designed and learned for videos which are artificially acquired in a limited environment, they may not operate robustly on videos acquired in a wild environment. To solve this problem, this paper proposes a scene-aware hybrid neural network (NN) having a novel combination of three-dimensional (3D) convolutional NN (CNN), 2D CNN and recurrent NN (RNN). The characteristics of the proposed network are as follows. First, we extract video-based global features and frame-based local features at the same time. In detail, the latent features containing the overall visual scene of a given video are extracted by 3D CNN with auxiliary classifier, and fine-tuned 2D CNN is adopted to extract latent features containing small details from each frame. Second, RNN not only performs temporal domain learning, but also feature-wise fuses two latent features extracted from the networks. For effective fusion, we also present three RNN schemes. Third, the proposed network, in which the above-mentioned methods collaborate, works very robust in a wild environment as well as in a limited environment. Extensive experiments show that the proposed network provides an average accuracy of 49.9% for AFEW dataset, i.e., a representative wild dataset, and an amazing accuracy of 98.2% for another CK+ dataset. We also show that the proposed network outperforms the state-of-the-art network(s). Min Kyu Lee, Dong-Yoon Choi, Dae Ha Kim, Byung Cheol Song |
FG | 3 |
| 2019 | Macro unit-based convolutional neural network for very light-weight deep learning
Dae Ha Kim, Min Kyu Lee, Byung Cheol Song |
Image Vis. Comput. | 1 |
| 2018 | Self-supervised Knowledge Distillation Using Singular Value Decomposition
Seunghyun Lee 0001, Dae Ha Kim, Byung Cheol Song |
ECCV (6) | 2 |
| 2018 | Recognizing Fine Facial Micro-Expressions Using Two-Dimensional Landmark FeatureabstractEmotion recognition based on facial expressions is very important for interaction between human and artificial intelligence (AI) system such as social robots. On the other hand, it is much harder to recognize subtle facial expressions or facial micro-expressions than facial expressions rich in emotional expression in a real environment. In this paper, we propose a two-dimensional (2D) landmark feature for effectively recognizing facial micro-expression. The proposed 2D landmark feature is obtained by converting existing coordinate-based landmark information into 2D image information, and has an advantage of having a unique feature according to emotions regardless of the intensity of facial expression. Thus, we can achieve effective emotion recognition by learning the proposed 2D landmark feature information on a convolutional neural network (CNN) and a long-term term memory (LSTM)-based network. Experimental results show that the proposed method provides more than 77% classification performance for fine facial expression images even when learning with general facial expression images of CK+ dataset. Dong-Yoon Choi, Dae Ha Kim, Byung Cheol Song |
ICIP | 2 |
| 2018 | Infrared image super-resolution using auxiliary convolutional neural network and visible image under low-light conditions
Tae Young Han, Dae Ha Kim, Byung Cheol Song |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Multi-modal emotion recognition using semi-supervised learning and multiple neural networks in the wildabstractHuman emotion recognition is a research topic that is receiving continuous attention in computer vision and artificial intelligence domains. This paper proposes a method for classifying human emotions through multiple neural networks based on multi-modal signals which consist of image, landmark, and audio in a wild environment. The proposed method has the following features. First, the learning performance of the image-based network is greatly improved by employing both multi-task learning and semi-supervised learning using the spatio-temporal characteristic of videos. Second, a model for converting 1-dimensional (1D) landmark information of face into two-dimensional (2D) images, is newly proposed, and a CNN-LSTM network based on the model is proposed for better emotion recognition. Third, based on an observation that audio signals are often very effective for specific emotions, we propose an audio deep learning mechanism robust to the specific emotions. Finally, so-called emotion adaptive fusion is applied to enable synergy of multiple networks. In the fifth attempt on the given test set in the EmotiW2017 challenge, the proposed method achieved a classification accuracy of 57.12%. Dae Ha Kim, Min Kyu Lee, Dong-Yoon Choi, Byung Cheol Song |
ICMI | 1 |