VLDB 2026 Research / reviewers in the wild / expert
Gwantae Kim
dblp:264/5954
· DBLP profile ↗
12ranked-venue papers
6as first author
10since 2021 · last 2024
0000-0002-3239-1865ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards Multi-Domain Face Landmark Detection with Synthetic Data from Diffusion ModelabstractRecently, deep learning-based facial landmark detection for in-the-wild faces has achieved significant improvement. However, there are still challenges in face landmark detection in other domains (e.g. cartoon, caricature, etc). This is due to the scarcity of extensively annotated training data. To tackle this concern, we design a two-stage training approach that effectively leverages limited datasets and the pre-trained diffusion model to obtain aligned pairs of landmarks and face in multiple domains. In the first stage, we train a landmark-conditioned face generation model on a large dataset of real faces. In the second stage, we fine-tune the above model on a small dataset of image-landmark pairs with text prompts for controlling the domain. Our new designs enable our method to generate high-quality synthetic paired datasets from multiple domains while preserving the alignment between landmarks and facial features. Finally, we fine-tuned a pre-trained face landmark detection model on the synthetic dataset to achieve multi-domain face landmark detection. Our qualitative and quantitative results demonstrate that our method outperforms existing methods on multi-domain face landmark detection. Yuanming Li, Gwantae Kim, Jeong-gi Kwak, Bonhwa Ku, Hanseok Ko |
ICASSP | 2 |
| 2024 | Classification and Magnitude Estimation of Global and Local Seismic Events Using Conformer and Low-Rank Adaptation Fine-TuningabstractClassifying seismic events and estimating their magnitude are crucial topics in the study of seismic waves. Due to the disparities between global and local geologic features, models exclusively trained on global data may exhibit suboptimal performance in local contexts. To solve this problem, this letter proposes a method to evaluate the effectiveness of the low-rank adaptation (LoRA) technique in seismic wave research using the convolution-augmented transformer (Conformer). We simplified and modified the Conformer model, reducing the number of parameters by more than 169-fold, and applied the LoRA technique to this model. Experimental results using the Stanford Earthquake Dataset (STEAD) and the Korean Peninsula Earthquake Dataset (KPED) from 2017 to 2018 showed that fine-tuning the model with a significantly reduced number of parameters using the proposed method is suitable for research on seismological applications. Our approach achieved over 99.99% accuracy in seismic event classification for both datasets. Additionally, our model demonstrated a 7% decrease in mean absolute error (MAE) on the STEAD dataset and a 48% decrease on the KPED dataset compared to the state-of-the-art model. Furthermore, the results also indicate that the Conformer is suitable for seismic event classification and magnitude estimation. The model’s performance in the seismic event classification task decreased by 0.1%, despite reducing the number of retrain parameters by 59 times. Additionally, in the magnitude estimation task, there was an 89-fold decrease in the number of retrain parameters, yet the performance decreased by 1%. Yooseok Jin, Gwantae Kim, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | MPE4G : Multimodal Pretrained Encoder for Co-Speech Gesture GenerationabstractWhen virtual agents interact with humans, gestures are crucial to delivering their intentions with speech. Previous multimodal co-speech gesture generation models required encoded features of all modalities to generate gestures. If some input modalities are removed or contain noise, the model may not generate the gestures properly. To acquire robust and generalized encodings, we propose a novel framework with a multimodal pre-trained encoder for co-speech gesture generation. In the proposed method, the multi-head-attention-based encoder is trained with self-supervised learning to contain the information on each modality. Moreover, we collect full-body gestures that consist of 3D joint rotations to improve visualization and apply gestures to the extensible body model. Through the series of experiments and human evaluation, the proposed method renders realistic co-speech gestures not only when all input modalities are given but also when the input modalities are missing or noisy. The project page is available here1 Gwantae Kim, Seonghyeok Noh, Insung Ham, Hanseok Ko |
ICASSP | 1 |
| 2023 | Channel Shuffle Neural Architecture Search for Key Word SpottingabstractThe evolution of Network Architecture (NA) allowed Key-Word Spotting (KWS) to exhibit high performance. Generally, NA for KWS is required to have low parameter and computation complexity maintaining high classification performance. Most of the attempts so far have been based on manual approaches, and often the architectures developed from such efforts dwell in the balance of the performance and the network complexity. Then, several KWS models based on Neural Architecture Search (NAS) technique have been proposed. However, these methods do not consider the number of parameters and FLOPs for NA in the search process and manually adjusted the complexity of NA by reducing the number of cells. It may not produce optimized NA with a balance between network complexity and performance. To develop effective network architecture for KWS, network complexity and performance must be considered. In this letter, we propose Channel Shuffle Neural Architecture Search (CSNAS) with channel weights. CSNAS selects whether each channel of the input feature is reflected in the computation or not in the search process and simultaneously controls the number of parameters, FLOPs, and performance. Experiment results show that CSNAS can generate NA that satisfies complexity and performance conditions, and NAs generated by CSNAS outperform state-of-the-art KWS methods. Bokyeung Lee, Gwantae Kim, Hanseok Ko |
IEEE Signal Process. Lett. | 3 |
| 2022 | 3D Human Motion Generation from the Text Via Gesture Action Classification and the Autoregressive ModelabstractIn this paper, a deep learning-based model for 3D human motion generation from the text is proposed via gesture action classification and an autoregressive model. The model focuses on generating special gestures that express human thinking, such as waving and nodding. To achieve the goal, the proposed method predicts expression from the sentences using a text classification model based on a pretrained language model and generates gestures using the gate recurrent unit-based autoregressive model. Especially, we proposed the loss for the embedding space for restoring raw motions and generating intermediate motions well. Moreover, the novel data augmentation method and stop token are proposed to generate variable length motions. To evaluate the text classification model and 3D human motion generation model, a gesture action classification dataset and action-based gesture dataset are collected. With several experiments, the proposed method successfully generates perceptually natural and realistic 3D human motion from the text. Moreover, we verified the effectiveness of the proposed method using a public-available action recognition dataset to evaluate cross-dataset generalization performance. Gwantae Kim, Youngsuk Ryu, Junyeop Lee, David K. Han, Jeongmin Bae 0006, Hanseok Ko |
ICIP | 1 |
| 2022 | Graph Convolution Networks for Seismic Events Classification Using Raw Waveform Data From Multiple StationsabstractThis letter proposes a multiple station-based seismic event classification model using a deep convolution neural network (CNN) and graph convolution network (GCN). To classify various seismic events, such as natural earthquakes, artificial earthquakes, and noise, the proposed model consists of weight-shared convolution layers, graph convolution layers, and fully connected layers. We employed graph convolution layers in order to aggregate features from multiple stations. Representative experimental results with the Korean peninsula earthquake datasets from 2016 to 2019 showed that the proposed model is superior to the single-station based state-of the-art methods. Moreover, the proposed model significantly reduced false alarms when using continuous waveforms of long duration. The code is available at.1 Gwantae Kim, Bonhwa Ku, Jae-Kwang Ahn, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Prototypical Knowledge Distillation for Noise Robust Keyword SpottingabstractKeyword Spotting (KWS) is an essential component in contemporary audio-based deep learning systems and should be of minimal design when the system is working in streaming and on-device environments. We presented a robust feature extraction with a single-layer dynamic convolution model in our previous work. In this letter, we expand our earlier study into multi-layers of operation and propose a robust Knowledge Distillation (KD) learning method. Based on the distribution between class-centroids and embedding vectors, we compute three distinct distance metrics for the KD training and feature extraction processes. The results indicate that our KD method shows similar KWS performance over state-of-the-art models in terms of KWS but with low computational costs. Furthermore, our proposed method results in a more robust performance in noisy environments than conventional KD methods. Gwantae Kim, Bokyeung Lee, Hanseok Ko |
IEEE Signal Process. Lett. | 2 |
| 2021 | SpecMix : A Mixed Sample Data Augmentation Method for Training with Time-Frequency Domain FeaturesabstractA mixed sample data augmentation strategy is proposed to enhance the performance of models on audio scene classification, sound event classification, and speech enhancement tasks.While there have been several augmentation methods shown to be effective in improving image classification performance, their efficacy toward time-frequency domain features of audio is not assured.We propose a novel audio data augmentation approach named "Specmix" specifically designed for dealing with time-frequency domain features.The augmentation method consists of mixing two different data samples by applying time-frequency masks effective in preserving the spectral correlation of each audio sample.Our experiments on acoustic scene classification, sound event classification, and speech enhancement tasks show that the proposed Specmix improves the performance of various neural network architectures by a maximum of 2.7%. Gwantae Kim, David K. Han, Hanseok Ko |
Interspeech | 1 |
| 2021 | Multifeature Fusion-Based Earthquake Event Classification Using Transfer LearningabstractThis letter proposes a multifeature fusion model using deep convolution neural networks and transfer learning approach for earthquake event classification. There are several feature representations for seismic analysis, such as the time domain, the frequency domain, and the time–frequency domain. To successfully classify various earthquake events, we propose a novel model that combines these features hierarchically. In addition, we apply a transfer learning to mitigate overfitting problem of deep learning model while achieving high classification performance. To evaluate our approach, we conduct experiments with the Korean peninsula earthquake database from 2016 to 2018 and a large earthquake database on the Circum-Pacific belt in 2019. The experimental results show that the proposed method outperforms over the compared state-of-the-art methods. Gwantae Kim, Bonhwa Ku, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2021 | Attention-Based Convolutional Neural Network for Earthquake Event ClassificationabstractThis letter presents a deep convolutional neural network (CNN) with attention module that improves the performance of the classification of various earthquake events. Addressing all possible earthquake events, including not only microearthquakes and artificial-earthquakes but also large-earthquakes, requires both suitable feature expression and a classifier that can effectively discriminate seismic waveforms under adverse conditions. To robustly classify earthquake events, a deep CNN with an attention module was proposed in raw seismic waveforms. Representative experimental results show that the proposed method provides an effective structure for earthquake events classification and, with the Korean peninsula earthquake database from 2016 to 2018, outperforms previous state-of-the-art methods. Bonhwa Ku, Gwantae Kim, Jae-Kwang Ahn, Hanseok Ko |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | Convolutional Recurrent Neural Networks for Earthquake Epicentral Distance Estimation Using Single-Channel Seismic WaveformabstractThis paper proposes a deep learning method for epicentral distance estimation using a single-channel seismic waveform. The model is based on a convolutional recurrent neural network structure to extract spatial and temporal features. Since the proposed model needs only single-channel data, it can also perform the distance estimation even when some channels of the sensor are adversely disabled. To evaluate our approach, we conduct distance estimation experiments with the Korean peninsula earthquake database from 2016 to 2018, which include microearthquakes and distant earthquakes. The epicentral distance estimation by the proposed method show an absolute mean error of 0.50 km with 9.16km standard deviation of error distribution, which shows the best estimation result among the competing model structures. The promising result indicates that the proposed approach can be deployed for epicentral localization task as part of realizing a robust earthquake monitoring system. Gwantae Kim, Bonhwa Ku, Yuanming Li, Jeongki Min, Hanseok Ko |
IGARSS | 1 |
| 2020 | Seismic Signal Synthesis by Generative Adversarial Network with Gated Convolutional Neural Network StructureabstractDetecting earthquake events from seismic time series signal is a challenging task. Recently, detection methods based on machine learning have been developed to improve the accuracy and efficiency. However, accuracy of those methods rely on sufficient amount of high-quality training data. In many situations, the high-quality data is difficulty to obtain. We address and resolve this issue by using a Generative Adversarial Network (GAN) model for seismic signal synthesis. GAN already shows its powerful capability in generating high quality synthetic samples in multiple domains. In this paper, we propose a GAN model with gated CNN which can excellently capture sequential structure of seismic time series. We demonstrate its effectiveness via earthquake classification performance. The results show the synthetic data generated by our model indeed can improve the classification performance over the one trained with only real samples. Yuanming Li, Bonhwa Ku, Gwantae Kim, Jae-Kwang Ahn, Hanseok Ko |
IGARSS | 3 |