Woo Hyun Kang

dblp:180/2800 · DBLP profile ↗
← Back
22ranked-venue papers
9as first author
14since 2021 · last 2025
0000-0001-8739-9349ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 11 since 2021Artificial intelligence and machine learning · 15 · 7 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Modality-Agnostic Multimodal Emotion Recognition using a Contrastive Masked Autoencoder
Georgios Chochlakis, Turab Iqbal, Woo Hyun Kang, Zhaocheng Huang
INTERSPEECH3
2024 SWAN: SubWord Alignment Network for HMM-free word timing estimation in end-to-end automatic speech recognition
Woo Hyun Kang, Srikanth Vishnubhotla, Rudolf Braun, Yogesh Virkar, Raghuveer Peri, Kyu J. Han
INTERSPEECH1
2023 Hybrid Neural Network with Cross- and Self-Module Attention Pooling for Text-Independent Speaker Verification
abstract
Extraction of a speaker embedding vector plays an important role in deep learning-based speaker verification. In this contribution, to extract speaker discriminant utterance level embeddings, we propose a hybrid neural network that employs both cross- and self-module attention pooling mechanisms. More specifically, the proposed system incorporates a 2D-Convolution Neural Network (CNN)-based feature extraction module in cascade with a frame-level network, which is composed of a fully Time Delay Neural Network (TDNN) network and a TDNN-Long Short Term Memory (TDNN-LSTM) hybrid network in a parallel manner. The proposed system also employs a multi-level cross- and self-module attention pooling for aggregating the speaker information within an utterance-level context by capturing the complementarity between two parallelly connected modules. In order to evaluate the proposed system, we conduct a set of experiments on the Voxceleb corpus, and the proposed hybrid network is able to outperform the conventional approaches trained on the same dataset.
Jahangir Alam 0001, Woo Hyun Kang, Abderrahim Fathan
ICASSP2
2022 Robust Self-Supervised Speaker Representation Learning Via Instance Mix Regularization
abstract
Over the recent years, various self-supervised contrastive embedding learning methods for deep speaker verification were proposed. The performance of the self-supervised contrastive learning framework highly depends on the data augmentation technique, but due to the sensitive nature of speaker information within the speech signal, most speaker embedding training relies on simple augmentations such as additive noise or simulated reverberation. Thus while the conventional self-supervised speaker embedding systems can yield minimum within-utterance variability, the capability to generalize to out-of-set utterance is limited. In order to alleviate this problem, we propose a novel self-supervised learning framework for speaker verification which combines the angular prototypical loss and the instance mix (i-mix) regularization. The proposed method was evaluated on the VoxCeleb1 dataset and showed noticeable improvement over the standard self-supervised embedding method.
Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan
ICASSP1
2022 Mel-Spectrogram Image-Based End-to-End Audio Deepfake Detection Under Channel-Mismatched Conditions
abstract
This work focuses on the problem of detecting fake audio clips. To improve current audio spoofing detection models, we propose a selection of multiple audio augmentations spe-cially designed to resemble audio spoofing attacks. These augmentations are experimentally found to be very useful and using them achieves a notable performance of 2.8% EER on the ASVspoof 2019 challenge evaluation set. Unlike the widely employed acoustic features, in this paper we explore the use of Mel-spectrogram image features and employ vari-ous audio codecs to achieve robustness to codec and transmission channel variability present in the ASVspoof2021 Evalu-ation set. To better handle spectral information, crucial to de-tect spoofing, we adopt the WaveletCNN and VGG16 archi-tectures which outperform all baselines. Finally, we find that robustness of countermeasure systems degrades dramatically when provided with speech samples degraded through VoIP network transmission or mismatching audio compression.
Abderrahim Fathan, Jahangir Alam 0001, Woo Hyun Kang
ICME3
2022 MIM-DG: Mutual information minimization-based domain generalization for speaker verification
Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan
INTERSPEECH1
2022 Mixup regularization strategies for spoofing countermeasure system
Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan
INTERSPEECH1
2022 End-to-end framework for spoof-aware speaker verification
Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan
INTERSPEECH1
2022 Deep learning-based end-to-end spoken language identification system for domain-mismatched scenario
abstract
Domain mismatch is a critical issue when it comes to spoken language identification. To overcome the domain mismatch problem, we have applied several architectures and deep learning strategies which have shown good results in cross-domain speaker verification tasks to spoken language identification. Our systems were evaluated on the Oriental Language Recognition (OLR) Challenge 2021 Task 1 dataset, which provides a set of cross-domain language identification trials. Among our experimented systems, the best performance was achieved by using the mel frequency cepstral coefficient (MFCC) and pitch features as input and training the ECAPA-TDNN system with a flow-based regularization technique, which resulted in a Cavg of 0.0631 on the OLR 2021 progress set.
Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan
LREC1
2022 Flow-ER: A Flow-Based Embedding Regularization Strategy for Robust Speech Representation Learning
abstract
Over the recent years, various deep learning-based embedding methods were proposed. Although the deep learning-based embedding extraction methods have shown good performance in numerous tasks including speaker verification, language identification and anti-spoofing, their performance is limited when it comes to mismatched conditions due to the variability within them unrelated to the main task. In order to alleviate this problem, we propose a novel training strategy that regularizes the embedding network to have minimum information about the nuisance attributes. To achieve this, our proposed method directly incorporates the information bottleneck scheme into the training process, where the mutual information is estimated using an auxiliary normalizing flow network. The performance of the proposed method is evaluated on different speech processing tasks and found to provide improvement over the standard training strategy in all experimentations.
Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan
SLT1
2021 Giving Space to Your Message: Assistive Word Segmentation for the Electronic Typing of Digital Minorities
abstract
For readability and disambiguation of the written text, appropriate word segmentation is recommended for documentation, and it also holds for the digitized texts. If the language is agglutinative while far from scriptio continua, for instance in the Korean language, the problem becomes more significant. However, some device users these days find it challenging to communicate via key stroking, not only for handicap but also for being unskilled. In this study, we propose a real-time assistive technology that utilizes an automatic word segmentation, designed for digital minorities who are not familiar with electronic typing. We propose a data-driven system trained upon a spoken Korean language corpus with various non-canonical expressions and dialects, guaranteeing the comprehension of contextual information. Through quantitative and qualitative comparison with other text processing toolkits, we show the reliability of the proposed system and its fit with colloquial and non-normalized texts, which fulfills the aim of supportive technology.
Won-Ik Cho, Sung Jun Cheon, Woo Hyun Kang, Ji Won Kim, Nam Soo Kim
Conference on Designing Interactive Systems3
2021 Hybrid Network with Multi-Level Global-Local Statistics Pooling for Robust Text-Independent Speaker Recognition
abstract
In this paper, we propose a new hybrid system for extracting a speaker embedding vector. More specifically, the proposed system employs a multi-level global-local statistics pooling method in order to aggregate the speaker information within short time-span and utterance-level context. In order to evaluate the proposed system, a set of experiments on the NIST SRE 2016, Short-duration speaker verification (SdSV) Challenge 2021, and VoxCeleb datasets were conducted, and the proposed hybrid network was able to outperform the conventional approaches trained on the same dataset. Moreover, our experiments showed that the proposed system is able to achieve stable performance even when using a relatively smaller dataset, which highlights the efficiency of the proposed system in extracting the speaker-dependent information.
Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan
ASRU1
2021 Team02 Text-Independent Speaker Verification System for SdSV Challenge 2021
abstract
In this paper, we provide description of our submitted systems to the Short Duration Speaker Verification (SdSV) Challenge 2021 Task 2. The challenge provides a difficult set of cross-language text-independent speaker verification trials. Our submissions employ ResNet-based embedding networks which are trained using various strategies exploiting both in-domain and out-of-domain datasets. The results show that using the recently proposed joint factor embedding (JFE) scheme can enhance the performance by disentangling the language-dependent information from the speaker embedding. However, upon analyzing the speaker embeddings, it was found that there exists a clear discrepancy between the in-domain and out-of-domain datasets. Therefore, among our submitted systems, the best performance was achieved by pre-training the embedding system using out-of-domain dataset and fine-tuning it with only the in-domain data, which resulted in a MinDCF of 0.142716 on the SdSV2021 evaluation set.
Woo Hyun Kang, Nam Soo Kim
Interspeech1
2021 Gated Recurrent Context: Softmax-Free Attention for Online Encoder-Decoder Speech Recognition
abstract
Recently, attention-based encoder-decoder (AED) models have shown state-of-the-art performance in automatic speech recognition (ASR). As the original AED models with global attentions are not capable of online inference, various online attention schemes have been developed to reduce ASR latency for better user experience. However, a common limitation of the conventional softmax-based online attention approaches is that they introduce an additional hyperparameter related to the length of the attention window, requiring multiple trials of model training for tuning the hyperparameter. In order to deal with this problem, we propose a novel softmax-free attention method and its modified formulation for online attention, which does not need any additional hyperparameter at the training phase. Through a number of ASR experiments, we demonstrate the tradeoff between the latency and performance of the proposed online attention technique can be controlled by merely adjusting a threshold at the test phase. Furthermore, the proposed methods showed competitive performance to the conventional global and online attentions in terms of word-error-rates (WERs).
Hyeon Seung Lee, Woo Hyun Kang, Sung Jun Cheon, Hyeongju Kim, Nam Soo Kim
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Text Matters but Speech Influences: A Computational Analysis of Syntactic Ambiguity Resolution
Jeonghwa Cho, Woo Hyun Kang, Nam Soo Kim
CogSci2
2020 Robust Front-End for Multi-Channel ASR using Flow-Based Density Estimation
Hyeongju Kim, Hyeon Seung Lee, Woo Hyun Kang, Hyung Yong Kim, Nam Soo Kim
IJCAI3
2020 Robust Text-Dependent Speaker Verification via Character-Level Information Preservation for the SdSV Challenge 2020
abstract
This paper describes our submission to Task 1 of the Short-duration Speaker Verification (SdSV) challenge 2020. Task 1 is a text-dependent speaker verification task, where both the speaker and phrase are required to be verified. The submitted systems were composed of TDNN-based and ResNet-based front-end architectures, in which the frame-level features were aggregated with various pooling methods (e.g., statistical, self-attentive, ghostVLAD pooling). Although the conventional pooling methods provide embeddings with a sufficient amount of speaker-dependent information, our experiments show that these embeddings often lack phrase-dependent information. To mitigate this problem, we propose a new pooling and score compensation methods that leverage a CTC-based automatic speech recognition (ASR) model for taking the lexical content into account. Both methods showed improvement over the conventional techniques, and the best performance was achieved by fusing all the experimented systems, which showed 0.0785% MinDCF and 2.23% EER on the challenge's evaluation subset.
Sung Hwan Mun, Woo Hyun Kang, Min Hyun Han, Nam Soo Kim
INTERSPEECH2
2020 SoftFlow: Probabilistic Framework for Normalizing Flow on Manifolds
abstract
Flow-based generative models are composed of invertible transformations between two random variables of the same dimension. Therefore, flow-based models cannot be adequately trained if the dimension of the data distribution does not match that of the underlying target distribution. In this paper, we propose SoftFlow, a probabilistic framework for training normalizing flows on manifolds. To sidestep the dimension mismatch problem, SoftFlow estimates a conditional distribution of the perturbed input data instead of learning the data distribution directly. We experimentally show that SoftFlow can capture the innate structure of the manifold data and generate high-quality samples unlike the conventional flow-based models. Furthermore, we apply the proposed framework to 3D point clouds to alleviate the difficulty of forming thin structures for flow-based models. The proposed model for 3D point clouds, namely SoftPointFlow, can estimate the distribution of various shapes more accurately and achieves state-of-the-art performance in point cloud generation.
Hyeongju Kim, Hyeon Seung Lee, Woo Hyun Kang, Joun Yeop Lee, Nam Soo Kim
NeurIPS3
2019 End-to-End Multi-Channel Speech Enhancement Using Inter-Channel Time-Restricted Attention on Raw Waveform
Hyeon Seung Lee, Hyung Yong Kim, Woo Hyun Kang, Jeunghun Kim, Nam Soo Kim
INTERSPEECH3
2017 Integrated DNN-based model adaptation technique for noise-robust speech recognition
abstract
Since the introduction of deep neural network (DNN)-based acoustic model, robust automatic speech recognition using DNN are being in research. Especially in model adaptation, the techniques utilizing auxiliary context features is known to be a promising technique. Recently, we proposed a technique which is called two-stage noise-aware training (TSNAT). The key idea of TS-NAT is to let the DNN clarify the relationship among noise estimate, noisy features and phonetic target through clean feature representation. However, although TS-NAT enhances the robustness of the DNN, we cannot be certain whether TS-NAT describes the clean feature representation sufficiently. In this paper, we extend TS-NAT using true noise feature and various DNN training techniques. It has been shown that the proposed technique outperforms the conventional DNN-based techniques on Aurora5-task and mismatched noise conditions.
Kang Hyun Lee, Woo Hyun Kang, Tae Gyoon Kang, Nam Soo Kim
ICASSP2
2016 Two-stage noise aware training using asymmetric deep denoising autoencoder
abstract
Ever since the deep neural network (DNN)-based acoustic model appeared, the recognition performance of automatic speech recognition has been greatly improved. Due to this achievement, various researches on DNN-based technique for noise robustness are also in progress. Among these approaches, the noise-aware training (NAT) technique which aims to improve the inherent robustness of DNN using noise estimates has shown remarkable performance. However, despite the great performance, we cannot be certain whether NAT is an optimal method for sufficiently utilizing the inherent robustness of DNN. In this paper, we propose a novel technique which helps the DNN to address the complex connection between the input and target vectors of NAT smoothly. The proposed method outperformed the conventional NAT in Aurora-5 task.
Kang Hyun Lee, Shin Jae Kang, Woo Hyun Kang, Nam Soo Kim
ICASSP3
2016 DNN-Based Feature Enhancement Using Joint Training Framework for Robust Multichannel Speech Recognition
Kang Hyun Lee, Tae Gyoon Kang, Woo Hyun Kang, Nam Soo Kim
INTERSPEECH3