EDBT 2026 Demo / reviewers in the wild / expert
Ju-ho Kim
dblp:239/6451 · also Ju-Ho Kim
· DBLP profile ↗
16ranked-venue papers
5as first author
12since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Diff-SV: A Unified Hierarchical Framework for Noise-Robust Speaker Verification Using Score-Based Diffusion Probabilistic ModelsabstractBackground noise considerably reduces the accuracy and reliability of speaker verification (SV) systems. These challenges can be addressed using a speech enhancement system as a front-end module. Recently, diffusion probabilistic models (DPMs) have exhibited remarkable noise-compensation capabilities in the speech enhancement domain. Building on this success, we propose Diff-SV, a noise-robust SV framework that leverages DPM. Diff-SV unifies a DPM-based speech enhancement system with a speaker embedding extractor, and yields a discriminative and noise-tolerable speaker representation through a hierarchical structure. The proposed model was evaluated under both in-domain and out-of-domain noisy conditions using the VoxCeleb1 test set, an external noise source, and the VOiCES corpus. The obtained experimental results demonstrate that Diff-SV achieves state-of-the-art performance, outperforming recently proposed noise-robust SV systems. Ju-ho Kim, Jungwoo Heo, Hyun-seo Shin, Chan-yeong Lim, Ha-Jin Yu |
ICASSP | 1 |
| 2024 | HM-CONFORMER: A Conformer-Based Audio Deepfake Detection System with Hierarchical Pooling and Multi-Level Classification Token Aggregation MethodsabstractAudio deepfake detection (ADD) is the task of detecting spoofing attacks generated by text-to-speech or voice conversion systems. Spoofing evidence, which helps to distinguish between spoofed and bona-fide utterances, might exist either locally or globally in the input features. To capture these, the Conformer, which consists of Transformers and CNN, possesses a suitable structure. However, since the Conformer was designed for sequence-to-sequence tasks, its direct application to ADD tasks may be sub-optimal. To tackle this limitation, we propose HM-Conformer by adopting two components: (1) Hierarchical pooling method progressively reducing the sequence length to eliminate duplicated information (2) Multi-level classification token aggregation method utilizing classification tokens to gather information from different blocks. Owing to these components, HM-Conformer can efficiently detect spoofing evidence by processing various sequence lengths and aggregating them. In experimental results on the ASVspoof 2021 Deepfake dataset, HM-Conformer achieved a 15.71% EER, showing competitive performance compared to recent systems. Hyun-seo Shin, Jungwoo Heo, Ju-ho Kim, Chan-yeong Lim, Won-Bin Kim, Ha-Jin Yu |
ICASSP | 3 |
| 2024 | Self-supervised speaker verification with relational mask prediction
Ju-ho Kim, Hee-Soo Heo, Bong-Jin Lee, Youngki Kwon, Ha-Jin Yu |
INTERSPEECH | 1 |
| 2024 | MR-RawNet: Speaker verification system with multiple temporal resolutions for variable duration utterances using raw waveforms
Seung-bin Kim, Chan-yeong Lim, Jungwoo Heo, Ju-ho Kim, Hyun-seo Shin, Kyo-Won Koo, Ha-Jin Yu |
INTERSPEECH | 4 |
| 2024 | Improving Noise Robustness in Self-supervised Pre-trained Model for Speaker Verification
Chan-yeong Lim, Hyun-seo Shin, Ju-ho Kim, Jungwoo Heo, Kyo-Won Koo, Seung-bin Kim, Ha-Jin Yu |
INTERSPEECH | 3 |
| 2024 | FA-ExU-Net: The Simultaneous Training of an Embedding Extractor and Enhancement Model for a Speaker Verification System Robust to Short Noisy UtterancesabstractSpeaker verification (SV) technology has the potential to enhance personalization and security in various applications, such as voice assistants, forensics, and access control. However, several challenges hinder the practical application of SV systems, including limitations and distortions in speaker information due to short utterances and noisy environments. Furthermore, these two factors often coexist in real-world situations, resulting in a significant performance degradation of SV systems. Despite the significance of these obstacles, each factor is independently studied, and the co-occurrence of both factors is rarely investigated. Here, we propose a novel SV framework, feature aggregated extended U-Net (FA-ExU-Net), which simultaneously addresses both the challenges by building on the success of prior research on each factor. The FA-ExU-Net incorporates an iterative and hierarchical feature aggregation scheme, a target task-specific feature enhancement module, and a multi-scale feature aggregator for extracting information-rich embeddings. Our proposed system outperforms the recent baseline models based on four evaluation criteria: generalizability, short utterance performance, capacity to handle noisy environments, and robustness to short utterances in noisy environments. We demonstrate the effectiveness of the proposed model through comparison and ablation experiments and intuitive visualizations. The proposed novel approach is expected to contribute to the development of more robust and accurate SV models for practical applications. Our training codes are available athttps://github.com/wngh1187/FA-ExU-Net. Ju-ho Kim, Jungwoo Heo, Hyun-seo Shin, Chan-yeong Lim, Ha-Jin Yu |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | One-Step Knowledge Distillation and Fine-Tuning in Using Large Pre-Trained Self-Supervised Learning Models for Speaker Verification
Jungwoo Heo, Chan-yeong Lim, Ju-ho Kim, Hyun-seo Shin, Ha-Jin Yu |
INTERSPEECH | 3 |
| 2022 | RawNeXt: Speaker Verification System For Variable-Duration Utterances With Deep Layer Aggregation And Extended Dynamic Scaling PoliciesabstractDespite achieving satisfactory performance in speaker verification using deep neural networks, variable-duration utterances remain a challenge that threatens the robustness of systems. To deal with this issue, we propose a speaker verification system called RawNeXt that can handle input raw waveforms of arbitrary length by employing the following two components: (1) A deep layer aggregation strategy enhances speaker information by iteratively and hierarchically aggregating features of various time scales and spectral channels output from blocks. (2) An extended dynamic scaling policy flexibly processes features according to the length of the utterance by selectively merging the activations of different resolution branches in each block. Owing to these two components, our proposed model can extract speaker embeddings rich in time-spectral information and operate dynamically on length variations. Experimental results on the VoxCeleb1 test set consisting of various duration utterances demonstrate that RawNeXt achieves state-of-the-art performance compared to the recently proposed systems. Our code and trained model weights are available at https://github.com/wngh1187/RawNeXt. Ju-ho Kim, Hye-Jin Shim, Jungwoo Heo, Ha-Jin Yu |
ICASSP | 1 |
| 2022 | Attentive Max Feature Map and Joint Training for Acoustic Scene ClassificationabstractVarious attention mechanisms are being widely applied to acoustic scene classification. However, we empirically found that the attention mechanism can excessively discard potentially valuable information, despite improving performance. We propose the attentive max feature map that combines two effective techniques, attention and a max feature map, to further elaborate the attention mechanism and mitigate the above-mentioned phenomenon. We also explore various joint training methods, including multi-task learning, that allocate additional abstract labels for each audio recording. Our proposed system demonstrates competitive performance with much larger state-of-the-art systems for single systems on Subtask A of the DCASE 2020 challenge by applying the two proposed techniques using relatively fewer parameters. Furthermore, adopting the proposed attentive max feature map, our team placed fourth in the recent DCASE 2021 challenge. Hye-Jin Shim, Jee-Weon Jung, Ju-ho Kim, Ha-Jin Yu |
ICASSP | 3 |
| 2022 | Two Methods for Spoofing-Aware Speaker Verification: Multi-Layer Perceptron Score Fusion Model and Integrated Embedding ProjectorabstractThe use of deep neural networks (DNN) has dramatically elevated the performance of automatic speaker verification (ASV) over the last decade. However, ASV systems can be easily neutralized by spoofing attacks. Therefore, the Spoofing-Aware Speaker Verification (SASV) challenge is designed and held to promote development of systems that can perform ASV considering spoofing attacks by integrating ASV and spoofing countermeasure (CM) systems. In this paper, we propose two back-end systems: multi-layer perceptron score fusion model (MSFM) and integrated embedding projector (IEP). The MSFM, score fusion back-end system, derived SASV score utilizing ASV and CM scores and embeddings. On the other hand,IEP combines ASV and CM embeddings into SASV embedding and calculates final SASV score based on the cosine similarity. We effectively integrated ASV and CM systems through proposed MSFM and IEP and achieved the SASV equal error rates 0.56%, 1.32% on the official evaluation trials of the SASV 2022 challenge. Jungwoo Heo, Ju-ho Kim, Hyun-seo Shin |
INTERSPEECH | 2 |
| 2022 | Extended U-Net for Speaker Verification in Noisy EnvironmentsabstractBackground noise is a well-known factor that deteriorates the accuracy and reliability of speaker verification (SV) systems by blurring speech intelligibility.Various studies have used separate pretrained enhancement models as the front-end module of the SV system in noisy environments, and these methods effectively remove noises.However, the denoising process of independent enhancement models not tailored to the SV task can also distort the speaker information included in utterances.We argue that the enhancement network and speaker embedding extractor should be fully jointly trained for SV tasks under noisy conditions to alleviate this issue.Therefore, we proposed a U-Net-based integrated framework that simultaneously optimizes speaker identification and feature enhancement losses.Moreover, we analyzed the structural limitations of using U-Net directly for noise SV tasks and further proposed Extended U-Net to reduce these drawbacks.We evaluated the models on the noise-synthesized VoxCeleb1 test set and VOiCES development set recorded in various noisy scenarios.The experimental results demonstrate that the U-Net-based fully joint training framework is more effective than the baseline, and the extended U-Net exhibited state-of-the-art performance versus the recently proposed compensation systems. Ju-ho Kim, Jungwoo Heo, Hye-Jin Shim, Ha-Jin Yu |
INTERSPEECH | 1 |
| 2021 | DCASENET: An Integrated Pretrained Deep Neural Network for Detecting and Classifying Acoustic Scenes and EventsabstractAlthough acoustic scenes and events include many related tasks, their combined detection and classification have been scarcely investigated. We propose three architectures of deep neural networks that are integrated to simultaneously perform acoustic scene classification, audio tagging, and sound event detection. The first two architectures are inspired by human cognitive processes. The first architecture resembles the short-term perception for scene classification of adults, who can detect various sound events that are then used to identify the acoustic scene. The second architecture resembles the long-term learning of babies, being also the concept under-lying self-supervised learning. Babies first observe the effects of abstract notions such as gravity and then learn specific tasks using such perceptions. The third architecture adds a few layers to the second one that solely perform a single task before its corresponding output layer. The aim is to build an integrated system that can serve as a pretrained model to perform the three abovementioned tasks. Experiments on three datasets demonstrate that the proposed architecture, called DcaseNet, can be either directly used for any of the tasks while providing suitable results or fine-tuned to improve the performance of one task. The code and pretrained DcaseNet weights are available at https://github.com/Jungjee/DcaseNet. Jee-Weon Jung, Hye-Jin Shim, Ju-ho Kim, Ha-Jin Yu |
ICASSP | 3 |
| 2020 | Improved RawNet with Feature Map Scaling for Text-Independent Speaker Verification Using Raw WaveformsabstractRecent advances in deep learning have facilitated the design of speaker verification systems that directly input raw waveforms.For example, RawNet [1] extracts speaker embeddings from raw waveforms, which simplifies the process pipeline and demonstrates competitive performance.In this study, we improve RawNet by scaling feature maps using various methods.The proposed mechanism utilizes a scale vector that adopts a sigmoid non-linear function.It refers to a vector with dimensionality equal to the number of filters in a given feature map.Using a scale vector, we propose to scale the feature map multiplicatively, additively, or both.In addition, we investigate replacing the first convolution layer with the sinc-convolution layer of SincNet.Experiments performed on the VoxCeleb1 evaluation dataset demonstrate the effectiveness of the proposed methods, and the best performing system reduces the equal error rate by half compared to the original RawNet.Expanded evaluation results obtained using the VoxCeleb1-E and VoxCeleb-H protocols marginally outperform existing state-ofthe-art systems. Jee-Weon Jung, Seung-bin Kim, Hye-Jin Shim, Ju-ho Kim, Ha-Jin Yu |
INTERSPEECH | 4 |
| 2020 | Acoustic Scene Classification Using Audio TaggingabstractAcoustic scene classification systems using deep neural networks classify given recordings into pre-defined classes. In this study, we propose a novel scheme for acoustic scene classification which adopts an audio tagging system inspired by the human perception mechanism. When humans identify an acoustic scene, the existence of different sound events provides discriminative information which affects the judgement. The proposed framework mimics this mechanism using various approaches. Firstly, we employ three methods to concatenate tag vectors extracted using an audio tagging system with an intermediate hidden layer of an acoustic scene classification system. We also explore the multi-head attention on the feature map of an acoustic scene classification system using tag vectors. Experiments conducted on the detection and classification of acoustic scenes and events 2019 task 1-a dataset demonstrate the effectiveness of the proposed scheme. Concatenation and multi-head attention show a classification accuracy of 75.66 % and 75.58 %, respectively, compared to 73.63 % accuracy of the baseline. The system with the proposed two approaches combined demonstrates an accuracy of 76.75 %. Jee-Weon Jung, Hye-Jin Shim, Ju-ho Kim, Seung-bin Kim, Ha-Jin Yu |
INTERSPEECH | 3 |
| 2020 | Segment Aggregation for Short Utterances Speaker Verification Using Raw WaveformsabstractMost studies on speaker verification systems focus on longduration utterances, which are composed of sufficient phonetic information.However, the performances of these systems are known to degrade when short-duration utterances are inputted due to the lack of phonetic information as compared to the long utterances.In this paper, we propose a method that compensates for the performance degradation of speaker verification for short utterances, referred to as "segment aggregation".The proposed method adopts an ensemble-based design to improve the stability and accuracy of speaker verification systems.The proposed method segments an input utterance into several short utterances and then aggregates the segment embeddings extracted from the segmented inputs to compose a speaker embedding.Then, this method simultaneously trains the segment embeddings and the aggregated speaker embedding.In addition, we also modified the teacher-student learning method for the proposed method.Experimental results on different input duration using the VoxCeleb1 test set demonstrate that the proposed technique improves speaker verification performance by about 45.37% relatively compared to the baseline system with 1-second test utterance condition. Seung-bin Kim, Jee-Weon Jung, Hye-Jin Shim, Ju-ho Kim, Ha-Jin Yu |
INTERSPEECH | 4 |
| 2019 | RawNet: Advanced End-to-End Deep Neural Network Using Raw Waveforms for Text-Independent Speaker VerificationabstractRecently, direct modeling of raw waveforms using deep neural networks has been widely studied for a number of tasks in audio domains.In speaker verification, however, utilization of raw waveforms is in its preliminary phase, requiring further investigation.In this study, we explore end-to-end deep neural networks that input raw waveforms to improve various aspects: front-end speaker embedding extraction including model architecture, pre-training scheme, additional objective functions, and back-end classification.Adjustment of model architecture using a pre-training scheme can extract speaker embeddings, giving a significant improvement in performance.Additional objective functions simplify the process of extracting speaker embeddings by merging conventional two-phase processes: extracting utterance-level features such as i-vectors or x-vectors and the feature enhancement phase, e.g., linear discriminant analysis.Effective back-end classification models that suit the proposed speaker embedding are also explored.We propose an end-toend system that comprises two deep neural networks, one frontend for utterance-level speaker embedding extraction and the other for back-end classification.Experiments conducted on the VoxCeleb1 dataset demonstrate that the proposed model achieves state-of-the-art performance among systems without data augmentation.The proposed system is also comparable to the state-of-the-art x-vector system that adopts data augmentation. Jee-Weon Jung, Hee-Soo Heo, Ju-ho Kim, Hye-Jin Shim, Ha-Jin Yu |
INTERSPEECH | 3 |