Gwantae Kim

dblp:264/5954 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
10since 2021 · last 2024
0000-0002-3239-1865ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Towards Multi-Domain Face Landmark Detection with Synthetic Data from Diffusion Model
abstract
Recently, deep learning-based facial landmark detection for in-the-wild faces has achieved significant improvement. However, there are still challenges in face landmark detection in other domains (e.g. cartoon, caricature, etc). This is due to the scarcity of extensively annotated training data. To tackle this concern, we design a two-stage training approach that effectively leverages limited datasets and the pre-trained diffusion model to obtain aligned pairs of landmarks and face in multiple domains. In the first stage, we train a landmark-conditioned face generation model on a large dataset of real faces. In the second stage, we fine-tune the above model on a small dataset of image-landmark pairs with text prompts for controlling the domain. Our new designs enable our method to generate high-quality synthetic paired datasets from multiple domains while preserving the alignment between landmarks and facial features. Finally, we fine-tuned a pre-trained face landmark detection model on the synthetic dataset to achieve multi-domain face landmark detection. Our qualitative and quantitative results demonstrate that our method outperforms existing methods on multi-domain face landmark detection.
Yuanming Li, Gwantae Kim, Jeong-gi Kwak, Bonhwa Ku, Hanseok Ko
ICASSP2
2024 Classification and Magnitude Estimation of Global and Local Seismic Events Using Conformer and Low-Rank Adaptation Fine-Tuning
abstract
Classifying seismic events and estimating their magnitude are crucial topics in the study of seismic waves. Due to the disparities between global and local geologic features, models exclusively trained on global data may exhibit suboptimal performance in local contexts. To solve this problem, this letter proposes a method to evaluate the effectiveness of the low-rank adaptation (LoRA) technique in seismic wave research using the convolution-augmented transformer (Conformer). We simplified and modified the Conformer model, reducing the number of parameters by more than 169-fold, and applied the LoRA technique to this model. Experimental results using the Stanford Earthquake Dataset (STEAD) and the Korean Peninsula Earthquake Dataset (KPED) from 2017 to 2018 showed that fine-tuning the model with a significantly reduced number of parameters using the proposed method is suitable for research on seismological applications. Our approach achieved over 99.99% accuracy in seismic event classification for both datasets. Additionally, our model demonstrated a 7% decrease in mean absolute error (MAE) on the STEAD dataset and a 48% decrease on the KPED dataset compared to the state-of-the-art model. Furthermore, the results also indicate that the Conformer is suitable for seismic event classification and magnitude estimation. The model’s performance in the seismic event classification task decreased by 0.1%, despite reducing the number of retrain parameters by 59 times. Additionally, in the magnitude estimation task, there was an 89-fold decrease in the number of retrain parameters, yet the performance decreased by 1%.
Yooseok Jin, Gwantae Kim, Hanseok Ko
IEEE Geosci. Remote. Sens. Lett.2
2023 MPE4G : Multimodal Pretrained Encoder for Co-Speech Gesture Generation
abstract
When virtual agents interact with humans, gestures are crucial to delivering their intentions with speech. Previous multimodal co-speech gesture generation models required encoded features of all modalities to generate gestures. If some input modalities are removed or contain noise, the model may not generate the gestures properly. To acquire robust and generalized encodings, we propose a novel framework with a multimodal pre-trained encoder for co-speech gesture generation. In the proposed method, the multi-head-attention-based encoder is trained with self-supervised learning to contain the information on each modality. Moreover, we collect full-body gestures that consist of 3D joint rotations to improve visualization and apply gestures to the extensible body model. Through the series of experiments and human evaluation, the proposed method renders realistic co-speech gestures not only when all input modalities are given but also when the input modalities are missing or noisy. The project page is available here1
Gwantae Kim, Seonghyeok Noh, Insung Ham, Hanseok Ko
ICASSP1
2023 Channel Shuffle Neural Architecture Search for Key Word Spotting
abstract
The evolution of Network Architecture (NA) allowed Key-Word Spotting (KWS) to exhibit high performance. Generally, NA for KWS is required to have low parameter and computation complexity maintaining high classification performance. Most of the attempts so far have been based on manual approaches, and often the architectures developed from such efforts dwell in the balance of the performance and the network complexity. Then, several KWS models based on Neural Architecture Search (NAS) technique have been proposed. However, these methods do not consider the number of parameters and FLOPs for NA in the search process and manually adjusted the complexity of NA by reducing the number of cells. It may not produce optimized NA with a balance between network complexity and performance. To develop effective network architecture for KWS, network complexity and performance must be considered. In this letter, we propose Channel Shuffle Neural Architecture Search (CSNAS) with channel weights. CSNAS selects whether each channel of the input feature is reflected in the computation or not in the search process and simultaneously controls the number of parameters, FLOPs, and performance. Experiment results show that CSNAS can generate NA that satisfies complexity and performance conditions, and NAs generated by CSNAS outperform state-of-the-art KWS methods.
Bokyeung Lee, Gwantae Kim, Hanseok Ko
IEEE Signal Process. Lett.3
2022 3D Human Motion Generation from the Text Via Gesture Action Classification and the Autoregressive Model
abstract
In this paper, a deep learning-based model for 3D human motion generation from the text is proposed via gesture action classification and an autoregressive model. The model focuses on generating special gestures that express human thinking, such as waving and nodding. To achieve the goal, the proposed method predicts expression from the sentences using a text classification model based on a pretrained language model and generates gestures using the gate recurrent unit-based autoregressive model. Especially, we proposed the loss for the embedding space for restoring raw motions and generating intermediate motions well. Moreover, the novel data augmentation method and stop token are proposed to generate variable length motions. To evaluate the text classification model and 3D human motion generation model, a gesture action classification dataset and action-based gesture dataset are collected. With several experiments, the proposed method successfully generates perceptually natural and realistic 3D human motion from the text. Moreover, we verified the effectiveness of the proposed method using a public-available action recognition dataset to evaluate cross-dataset generalization performance.
Gwantae Kim, Youngsuk Ryu, Junyeop Lee, David K. Han, Jeongmin Bae 0006, Hanseok Ko
ICIP1
2022 Graph Convolution Networks for Seismic Events Classification Using Raw Waveform Data From Multiple Stations
abstract
This letter proposes a multiple station-based seismic event classification model using a deep convolution neural network (CNN) and graph convolution network (GCN). To classify various seismic events, such as natural earthquakes, artificial earthquakes, and noise, the proposed model consists of weight-shared convolution layers, graph convolution layers, and fully connected layers. We employed graph convolution layers in order to aggregate features from multiple stations. Representative experimental results with the Korean peninsula earthquake datasets from 2016 to 2019 showed that the proposed model is superior to the single-station based state-of the-art methods. Moreover, the proposed model significantly reduced false alarms when using continuous waveforms of long duration. The code is available at.1
Gwantae Kim, Bonhwa Ku, Jae-Kwang Ahn, Hanseok Ko
IEEE Geosci. Remote. Sens. Lett.1
2022 Prototypical Knowledge Distillation for Noise Robust Keyword Spotting
abstract
Keyword Spotting (KWS) is an essential component in contemporary audio-based deep learning systems and should be of minimal design when the system is working in streaming and on-device environments. We presented a robust feature extraction with a single-layer dynamic convolution model in our previous work. In this letter, we expand our earlier study into multi-layers of operation and propose a robust Knowledge Distillation (KD) learning method. Based on the distribution between class-centroids and embedding vectors, we compute three distinct distance metrics for the KD training and feature extraction processes. The results indicate that our KD method shows similar KWS performance over state-of-the-art models in terms of KWS but with low computational costs. Furthermore, our proposed method results in a more robust performance in noisy environments than conventional KD methods.
Gwantae Kim, Bokyeung Lee, Hanseok Ko
IEEE Signal Process. Lett.2
2021 SpecMix : A Mixed Sample Data Augmentation Method for Training with Time-Frequency Domain Features
abstract
A mixed sample data augmentation strategy is proposed to enhance the performance of models on audio scene classification, sound event classification, and speech enhancement tasks.While there have been several augmentation methods shown to be effective in improving image classification performance, their efficacy toward time-frequency domain features of audio is not assured.We propose a novel audio data augmentation approach named "Specmix" specifically designed for dealing with time-frequency domain features.The augmentation method consists of mixing two different data samples by applying time-frequency masks effective in preserving the spectral correlation of each audio sample.Our experiments on acoustic scene classification, sound event classification, and speech enhancement tasks show that the proposed Specmix improves the performance of various neural network architectures by a maximum of 2.7%.
Gwantae Kim, David K. Han, Hanseok Ko
Interspeech1
2021 Multifeature Fusion-Based Earthquake Event Classification Using Transfer Learning
abstract
This letter proposes a multifeature fusion model using deep convolution neural networks and transfer learning approach for earthquake event classification. There are several feature representations for seismic analysis, such as the time domain, the frequency domain, and the time–frequency domain. To successfully classify various earthquake events, we propose a novel model that combines these features hierarchically. In addition, we apply a transfer learning to mitigate overfitting problem of deep learning model while achieving high classification performance. To evaluate our approach, we conduct experiments with the Korean peninsula earthquake database from 2016 to 2018 and a large earthquake database on the Circum-Pacific belt in 2019. The experimental results show that the proposed method outperforms over the compared state-of-the-art methods.
Gwantae Kim, Bonhwa Ku, Hanseok Ko
IEEE Geosci. Remote. Sens. Lett.1
2021 Attention-Based Convolutional Neural Network for Earthquake Event Classification
abstract
This letter presents a deep convolutional neural network (CNN) with attention module that improves the performance of the classification of various earthquake events. Addressing all possible earthquake events, including not only microearthquakes and artificial-earthquakes but also large-earthquakes, requires both suitable feature expression and a classifier that can effectively discriminate seismic waveforms under adverse conditions. To robustly classify earthquake events, a deep CNN with an attention module was proposed in raw seismic waveforms. Representative experimental results show that the proposed method provides an effective structure for earthquake events classification and, with the Korean peninsula earthquake database from 2016 to 2018, outperforms previous state-of-the-art methods.
Bonhwa Ku, Gwantae Kim, Jae-Kwang Ahn, Hanseok Ko
IEEE Geosci. Remote. Sens. Lett.2
2020 Convolutional Recurrent Neural Networks for Earthquake Epicentral Distance Estimation Using Single-Channel Seismic Waveform
abstract
This paper proposes a deep learning method for epicentral distance estimation using a single-channel seismic waveform. The model is based on a convolutional recurrent neural network structure to extract spatial and temporal features. Since the proposed model needs only single-channel data, it can also perform the distance estimation even when some channels of the sensor are adversely disabled. To evaluate our approach, we conduct distance estimation experiments with the Korean peninsula earthquake database from 2016 to 2018, which include microearthquakes and distant earthquakes. The epicentral distance estimation by the proposed method show an absolute mean error of 0.50 km with 9.16km standard deviation of error distribution, which shows the best estimation result among the competing model structures. The promising result indicates that the proposed approach can be deployed for epicentral localization task as part of realizing a robust earthquake monitoring system.
Gwantae Kim, Bonhwa Ku, Yuanming Li, Jeongki Min, Hanseok Ko
IGARSS1
2020 Seismic Signal Synthesis by Generative Adversarial Network with Gated Convolutional Neural Network Structure
abstract
Detecting earthquake events from seismic time series signal is a challenging task. Recently, detection methods based on machine learning have been developed to improve the accuracy and efficiency. However, accuracy of those methods rely on sufficient amount of high-quality training data. In many situations, the high-quality data is difficulty to obtain. We address and resolve this issue by using a Generative Adversarial Network (GAN) model for seismic signal synthesis. GAN already shows its powerful capability in generating high quality synthetic samples in multiple domains. In this paper, we propose a GAN model with gated CNN which can excellently capture sequential structure of seismic time series. We demonstrate its effectiveness via earthquake classification performance. The results show the synthetic data generated by our model indeed can improve the classification performance over the one trained with only real samples.
Yuanming Li, Bonhwa Ku, Gwantae Kim, Jae-Kwang Ahn, Hanseok Ko
IGARSS3