EDBT 2026 Demo / reviewers in the wild / expert
Ha-Jin Yu
dblp:54/5374
· DBLP profile ↗
49ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0003-3657-0665ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 7 first-author · 19 since 2021Artificial intelligence and machine learning · 34 · 7 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NM-FlowGAN: Pixel-wise noise and spatial correlation modeling for sRGB noise without paired images in generation time
Young-Joo Han, Ha-Jin Yu |
Expert Syst. Appl. | 2 |
| 2025 | SV-Mixer: Replacing the Transformer Encoder with Lightweight MLPs for Self-Supervised Model Compresison in Speaker VerificationabstractSelf-supervised learning (SSL) has pushed speaker verification accuracy close to state-of-the-art levels, but the Transformer backbones used in most SSL encoders hinder on-device and real-time deployment. Prior compression work trims layer depth or width yet still inherits the quadratic cost of self-attention. We propose SV-Mixer, the first fully MLPbased student encoder for SSL distillation. SV-Mixer replaces Transformer with three lightweight modules: Multi-Scale Mixing for multi-resolution temporal features, Local-Global Mixing for frame-to-utterance context, and Group Channel Mixing for spectral subspaces. Distilled from WavLM, SV-Mixer outperforms a Transformer student by $14.6 \%$ while cutting parameters and GMACs by over half, and at $\mathbf{7 5 \%}$ compression, it closely matches the teacher’s performance. Our results show that attention-free SSL students can deliver teacher-level accuracy with hardwarefriendly footprints, opening the door to robust on-device speaker verification. Jungwoo Heo, Hyun-seo Shin, Chan-yeong Lim, Kyo-Won Koo, Seung-bin Kim, Jisoo Son, Ha-Jin Yu |
ASRU | 7 |
| 2025 | Token-based Attractors and Cross-attention in Spoof DiarizationabstractSpoof diarization identifies “what spoofed when” in a given speech by temporally locating spoofed regions and determining their manipulation techniques. As a first step toward this task, prior work proposed a two-branch model for localization and spoof type clustering, which laid the foundation for spoof diarization. However, its simple structure limits the ability to capture complex spoofing patterns and lacks explicit reference points for distinguishing between bona fide and various spoofing types. To address these limitations, our approach introduces learnable tokens where each token represents acoustic features of bona fide and spoofed speech. These attractors interact with frame-level embeddings to extract discriminative representations, improving separation between genuine and generated speech. Vast experiments on PartialSpoof dataset consistently demon-strate that our approach outperforms existing methods in bona fide detection and spoofing method clustering. Kyo-Won Koo, Chan-yeong Lim, Jee-Weon Jung, Hye-Jin Shim, Ha-Jin Yu |
ASRU | 5 |
| 2025 | Enhancing Audio Deepfake Detection by Improving Representation Similarity of Bonafide Speech
Seung-bin Kim, Hyun-seo Shin, Jungwoo Heo, Chan-yeong Lim, Kyo-Won Koo, Jisoo Son, Sanghyun Hong 0001, Souhwan Jung, Ha-Jin Yu |
INTERSPEECH | 9 |
| 2025 | SEED: Speaker Embedding Enhancement Diffusion Model
Kihyun Nam, Jungwoo Heo, Jee-Weon Jung, Gangin Park, Chaeyoung Jung, Ha-Jin Yu, Joon Son Chung |
INTERSPEECH | 6 |
| 2024 | Diff-SV: A Unified Hierarchical Framework for Noise-Robust Speaker Verification Using Score-Based Diffusion Probabilistic ModelsabstractBackground noise considerably reduces the accuracy and reliability of speaker verification (SV) systems. These challenges can be addressed using a speech enhancement system as a front-end module. Recently, diffusion probabilistic models (DPMs) have exhibited remarkable noise-compensation capabilities in the speech enhancement domain. Building on this success, we propose Diff-SV, a noise-robust SV framework that leverages DPM. Diff-SV unifies a DPM-based speech enhancement system with a speaker embedding extractor, and yields a discriminative and noise-tolerable speaker representation through a hierarchical structure. The proposed model was evaluated under both in-domain and out-of-domain noisy conditions using the VoxCeleb1 test set, an external noise source, and the VOiCES corpus. The obtained experimental results demonstrate that Diff-SV achieves state-of-the-art performance, outperforming recently proposed noise-robust SV systems. Ju-ho Kim, Jungwoo Heo, Hyun-seo Shin, Chan-yeong Lim, Ha-Jin Yu |
ICASSP | 5 |
| 2024 | HM-CONFORMER: A Conformer-Based Audio Deepfake Detection System with Hierarchical Pooling and Multi-Level Classification Token Aggregation MethodsabstractAudio deepfake detection (ADD) is the task of detecting spoofing attacks generated by text-to-speech or voice conversion systems. Spoofing evidence, which helps to distinguish between spoofed and bona-fide utterances, might exist either locally or globally in the input features. To capture these, the Conformer, which consists of Transformers and CNN, possesses a suitable structure. However, since the Conformer was designed for sequence-to-sequence tasks, its direct application to ADD tasks may be sub-optimal. To tackle this limitation, we propose HM-Conformer by adopting two components: (1) Hierarchical pooling method progressively reducing the sequence length to eliminate duplicated information (2) Multi-level classification token aggregation method utilizing classification tokens to gather information from different blocks. Owing to these components, HM-Conformer can efficiently detect spoofing evidence by processing various sequence lengths and aggregating them. In experimental results on the ASVspoof 2021 Deepfake dataset, HM-Conformer achieved a 15.71% EER, showing competitive performance compared to recent systems. Hyun-seo Shin, Jungwoo Heo, Ju-ho Kim, Chan-yeong Lim, Won-Bin Kim, Ha-Jin Yu |
ICASSP | 6 |
| 2024 | Self-supervised speaker verification with relational mask prediction
Ju-ho Kim, Hee-Soo Heo, Bong-Jin Lee, Youngki Kwon, Ha-Jin Yu |
INTERSPEECH | 6 |
| 2024 | MR-RawNet: Speaker verification system with multiple temporal resolutions for variable duration utterances using raw waveforms
Seung-bin Kim, Chan-yeong Lim, Jungwoo Heo, Ju-ho Kim, Hyun-seo Shin, Kyo-Won Koo, Ha-Jin Yu |
INTERSPEECH | 7 |
| 2024 | Improving Noise Robustness in Self-supervised Pre-trained Model for Speaker Verification
Chan-yeong Lim, Hyun-seo Shin, Ju-ho Kim, Jungwoo Heo, Kyo-Won Koo, Seung-bin Kim, Ha-Jin Yu |
INTERSPEECH | 7 |
| 2024 | FA-ExU-Net: The Simultaneous Training of an Embedding Extractor and Enhancement Model for a Speaker Verification System Robust to Short Noisy UtterancesabstractSpeaker verification (SV) technology has the potential to enhance personalization and security in various applications, such as voice assistants, forensics, and access control. However, several challenges hinder the practical application of SV systems, including limitations and distortions in speaker information due to short utterances and noisy environments. Furthermore, these two factors often coexist in real-world situations, resulting in a significant performance degradation of SV systems. Despite the significance of these obstacles, each factor is independently studied, and the co-occurrence of both factors is rarely investigated. Here, we propose a novel SV framework, feature aggregated extended U-Net (FA-ExU-Net), which simultaneously addresses both the challenges by building on the success of prior research on each factor. The FA-ExU-Net incorporates an iterative and hierarchical feature aggregation scheme, a target task-specific feature enhancement module, and a multi-scale feature aggregator for extracting information-rich embeddings. Our proposed system outperforms the recent baseline models based on four evaluation criteria: generalizability, short utterance performance, capacity to handle noisy environments, and robustness to short utterances in noisy environments. We demonstrate the effectiveness of the proposed model through comparison and ablation experiments and intuitive visualizations. The proposed novel approach is expected to contribute to the development of more robust and accurate SV models for practical applications. Our training codes are available athttps://github.com/wngh1187/FA-ExU-Net. Ju-ho Kim, Jungwoo Heo, Hyun-seo Shin, Chan-yeong Lim, Ha-Jin Yu |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | SS-BSN: Attentive Blind-Spot Network for Self-Supervised Denoising with Nonlocal Self-SimilarityabstractRecently, numerous studies have been conducted on supervised learning-based image denoising methods. However, these methods rely on large-scale noisy-clean image pairs, which are difficult to obtain in practice. Denoising methods with self-supervised training that can be trained with only noisy images have been proposed to address the limitation. These methods are based on the convolutional neural network (CNN) and have shown promising performance. However, CNN-based methods do not consider using nonlocal self-similarities essential in the traditional method, which can cause performance limitations. This paper presents self-similarity attention (SS-Attention), a novel self-attention module that can capture nonlocal self-similarities to solve the problem. We focus on designing a lightweight self-attention module in a pixel-wise manner, which is nearly impossible to implement using the classic self-attention module due to the quadratically increasing complexity with spatial resolution. Furthermore, we integrate SS-Attention into the blind-spot network called self-similarity-based blind-spot network (SS-BSN). We conduct the experiments on real-world image denoising tasks. The proposed method quantitatively and qualitatively outperforms state-of-the-art methods in self-supervised denoising on the Smartphone Image Denoising Dataset (SIDD) and Darmstadt Noise Dataset (DND) benchmark datasets. Young-Joo Han, Ha-Jin Yu |
IJCAI | 2 |
| 2023 | One-Step Knowledge Distillation and Fine-Tuning in Using Large Pre-Trained Self-Supervised Learning Models for Speaker Verification
Jungwoo Heo, Chan-yeong Lim, Ju-ho Kim, Hyun-seo Shin, Ha-Jin Yu |
INTERSPEECH | 5 |
| 2022 | AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention NetworksabstractArtefacts that differentiate spoofed from bona-fide utterances can reside in specific temporal or spectral intervals. Their reliable detection usually depends upon computationally demanding ensemble systems where each subsystem is tuned to some specific artefacts. We seek to develop an efficient, single system that can detect a broad range of different spoofing attacks without score-level ensembles. We propose a novel heterogeneous stacking graph attention layer that models artefacts spanning heterogeneous temporal and spectral intervals with a heterogeneous attention mechanism and a stack node. With a new max graph operation that involves a competitive mechanism and a new readout scheme, our approach, named AASIST, outperforms the current state-of-the-art by 20% relative. Even a lightweight variant, AASIST-L, with only 85k parameters, outperforms all competing systems. Jee-Weon Jung, Hee-Soo Heo, Hemlata Tak, Hye-Jin Shim, Joon Son Chung, Bong-Jin Lee, Ha-Jin Yu, Nicholas W. D. Evans |
ICASSP | 7 |
| 2022 | RawNeXt: Speaker Verification System For Variable-Duration Utterances With Deep Layer Aggregation And Extended Dynamic Scaling PoliciesabstractDespite achieving satisfactory performance in speaker verification using deep neural networks, variable-duration utterances remain a challenge that threatens the robustness of systems. To deal with this issue, we propose a speaker verification system called RawNeXt that can handle input raw waveforms of arbitrary length by employing the following two components: (1) A deep layer aggregation strategy enhances speaker information by iteratively and hierarchically aggregating features of various time scales and spectral channels output from blocks. (2) An extended dynamic scaling policy flexibly processes features according to the length of the utterance by selectively merging the activations of different resolution branches in each block. Owing to these two components, our proposed model can extract speaker embeddings rich in time-spectral information and operate dynamically on length variations. Experimental results on the VoxCeleb1 test set consisting of various duration utterances demonstrate that RawNeXt achieves state-of-the-art performance compared to the recently proposed systems. Our code and trained model weights are available at https://github.com/wngh1187/RawNeXt. Ju-ho Kim, Hye-Jin Shim, Jungwoo Heo, Ha-Jin Yu |
ICASSP | 4 |
| 2022 | Graph Attentive Feature Aggregation for Text-Independent Speaker VerificationabstractThe objective of this paper is to combine multiple frame-level features into a single utterance-level representation considering pair-wise relationships. For this purpose, we propose a novel graph attentive feature aggregation module by interpreting each frame-level feature as a node of a graph. The inter-relationship between all possible pairs of features, typically exploited indirectly, can be directly modeled using a graph. The module comprises a graph attention layer and a graph pooling layer followed by a readout operation. The graph attention layer first models the non-Euclidean data manifold between different nodes. Then, the graph pooling layer discards less informative nodes considering the significance of the nodes. Finally, the readout operation combines the remaining nodes into a single representation. We employ two recent systems, SE-ResNet and RawNet2, with different input features and architectures and demonstrate that the proposed feature aggregation module consistently shows a relative improvement over 10%, compared to the baseline. Hye-Jin Shim, Jungwoo Heo, Jae-Han Park, Ga-Hui Lee, Ha-Jin Yu |
ICASSP | 5 |
| 2022 | Attentive Max Feature Map and Joint Training for Acoustic Scene ClassificationabstractVarious attention mechanisms are being widely applied to acoustic scene classification. However, we empirically found that the attention mechanism can excessively discard potentially valuable information, despite improving performance. We propose the attentive max feature map that combines two effective techniques, attention and a max feature map, to further elaborate the attention mechanism and mitigate the above-mentioned phenomenon. We also explore various joint training methods, including multi-task learning, that allocate additional abstract labels for each audio recording. Our proposed system demonstrates competitive performance with much larger state-of-the-art systems for single systems on Subtask A of the DCASE 2020 challenge by applying the two proposed techniques using relatively fewer parameters. Furthermore, adopting the proposed attentive max feature map, our team placed fourth in the recent DCASE 2021 challenge. Hye-Jin Shim, Jee-Weon Jung, Ju-ho Kim, Ha-Jin Yu |
ICASSP | 4 |
| 2022 | SASV 2022: The First Spoofing-Aware Speaker Verification ChallengeabstractThe first spoofing-aware speaker verification (SASV) challenge aims to integrate research efforts in speaker verification and anti-spoofing.We extend the speaker verification scenario by introducing spoofed trials to the usual set of target and impostor trials.In contrast to the established ASVspoof challenge where the focus is upon separate, independently optimised spoofing detection and speaker verification sub-systems, SASV targets the development of integrated and jointly optimised solutions.Pre-trained spoofing detection and speaker verification models are provided as open source and are used in two baseline SASV solutions.Both models and baselines are freely available to participants and can be used to develop back-end fusion approaches or end-to-end solutions.Using the provided common evaluation protocol, 23 teams submitted SASV solutions.When assessed with target, bona fide non-target and spoofed non-target trials, the top-performing system reduces the equal error rate of a conventional speaker verification system from 23.83% to 0.13%.SASV challenge results are a testament to the reliability of today's state-of-the-art approaches to spoofing detection and speaker verification. Jee-Weon Jung, Hemlata Tak, Hye-Jin Shim, Hee-Soo Heo, Bong-Jin Lee, Soo-Whan Chung, Ha-Jin Yu, Nicholas W. D. Evans, Tomi Kinnunen |
INTERSPEECH | 7 |
| 2022 | Extended U-Net for Speaker Verification in Noisy EnvironmentsabstractBackground noise is a well-known factor that deteriorates the accuracy and reliability of speaker verification (SV) systems by blurring speech intelligibility.Various studies have used separate pretrained enhancement models as the front-end module of the SV system in noisy environments, and these methods effectively remove noises.However, the denoising process of independent enhancement models not tailored to the SV task can also distort the speaker information included in utterances.We argue that the enhancement network and speaker embedding extractor should be fully jointly trained for SV tasks under noisy conditions to alleviate this issue.Therefore, we proposed a U-Net-based integrated framework that simultaneously optimizes speaker identification and feature enhancement losses.Moreover, we analyzed the structural limitations of using U-Net directly for noise SV tasks and further proposed Extended U-Net to reduce these drawbacks.We evaluated the models on the noise-synthesized VoxCeleb1 test set and VOiCES development set recorded in various noisy scenarios.The experimental results demonstrate that the U-Net-based fully joint training framework is more effective than the baseline, and the extended U-Net exhibited state-of-the-art performance versus the recently proposed compensation systems. Ju-ho Kim, Jungwoo Heo, Hye-Jin Shim, Ha-Jin Yu |
INTERSPEECH | 4 |
| 2021 | Graph Attention Networks for Speaker VerificationabstractThis work presents a novel back-end framework for speaker verification using graph attention networks. Segment-wise speaker embeddings extracted from multiple crops within an utterance are interpreted as node representations of a graph. The proposed framework inputs segment-wise speaker embeddings from an enrollment and a test utterance and directly outputs a similarity score. We first construct a graph using segment-wise speaker embeddings and then input these to graph attention networks. After a few graph attention layers with residual connections, each node is projected into a one-dimensional space using affine transform, followed by a readout operation resulting in a scalar similarity score. To enable successful adaptation for speaker verification, we propose techniques such as separating trainable weights for attention map calculations between segment-wise speaker embeddings from different utterances. The effectiveness of the proposed framework is validated using three different speaker embedding extractors trained with different architectures and objective functions. Experimental results demonstrate consistent improvement over various baseline back-end classifiers, with an average equal error rate improvement of 20% over the cosine similarity back-end without test time augmentation. Jee-Weon Jung, Hee-Soo Heo, Ha-Jin Yu, Joon Son Chung |
ICASSP | 3 |
| 2021 | DCASENET: An Integrated Pretrained Deep Neural Network for Detecting and Classifying Acoustic Scenes and EventsabstractAlthough acoustic scenes and events include many related tasks, their combined detection and classification have been scarcely investigated. We propose three architectures of deep neural networks that are integrated to simultaneously perform acoustic scene classification, audio tagging, and sound event detection. The first two architectures are inspired by human cognitive processes. The first architecture resembles the short-term perception for scene classification of adults, who can detect various sound events that are then used to identify the acoustic scene. The second architecture resembles the long-term learning of babies, being also the concept under-lying self-supervised learning. Babies first observe the effects of abstract notions such as gravity and then learn specific tasks using such perceptions. The third architecture adds a few layers to the second one that solely perform a single task before its corresponding output layer. The aim is to build an integrated system that can serve as a pretrained model to perform the three abovementioned tasks. Experiments on three datasets demonstrate that the proposed architecture, called DcaseNet, can be either directly used for any of the tasks while providing suitable results or fine-tuned to improve the performance of one task. The code and pretrained DcaseNet weights are available at https://github.com/Jungjee/DcaseNet. Jee-Weon Jung, Hye-Jin Shim, Ju-ho Kim, Ha-Jin Yu |
ICASSP | 4 |
| 2020 | Multiple Points Input For Convolutional Neural Networks in Replay Attack DetectionabstractThe models based on convolutional neural network (CNN) have shown remarkable performance in spoofing detection for automatic speaker verification. In order to input data into CNN-based models in mini-batch unit, the shape of all data in each mini-batch must be equal. Therefore, the method to make all data have the same length should be preceded because speeches have variable lengths. Segmentation is one of the methods to make the lengths of all data be equal. It divides the data into multiple segments using sliding window. Then, the models take one segment as input. However, it means that the amount of information that can be considered at one time is limited. We proposed the multiple points input method to increase the amount of information that can be considered at one time. The CNNs get input from multiple points in an utterance that are separated far enough to have different characteristics. The experimental results on ASVspoof 2019 physical access scenarios showed that our proposed method reduced the relative equal error rate by about 44% compared to the baseline. Sung-Hyun Yoon, Ha-Jin Yu |
ICASSP | 2 |
| 2020 | Improved RawNet with Feature Map Scaling for Text-Independent Speaker Verification Using Raw WaveformsabstractRecent advances in deep learning have facilitated the design of speaker verification systems that directly input raw waveforms.For example, RawNet [1] extracts speaker embeddings from raw waveforms, which simplifies the process pipeline and demonstrates competitive performance.In this study, we improve RawNet by scaling feature maps using various methods.The proposed mechanism utilizes a scale vector that adopts a sigmoid non-linear function.It refers to a vector with dimensionality equal to the number of filters in a given feature map.Using a scale vector, we propose to scale the feature map multiplicatively, additively, or both.In addition, we investigate replacing the first convolution layer with the sinc-convolution layer of SincNet.Experiments performed on the VoxCeleb1 evaluation dataset demonstrate the effectiveness of the proposed methods, and the best performing system reduces the equal error rate by half compared to the original RawNet.Expanded evaluation results obtained using the VoxCeleb1-E and VoxCeleb-H protocols marginally outperform existing state-ofthe-art systems. Jee-Weon Jung, Seung-bin Kim, Hye-Jin Shim, Ju-ho Kim, Ha-Jin Yu |
INTERSPEECH | 5 |
| 2020 | Acoustic Scene Classification Using Audio TaggingabstractAcoustic scene classification systems using deep neural networks classify given recordings into pre-defined classes. In this study, we propose a novel scheme for acoustic scene classification which adopts an audio tagging system inspired by the human perception mechanism. When humans identify an acoustic scene, the existence of different sound events provides discriminative information which affects the judgement. The proposed framework mimics this mechanism using various approaches. Firstly, we employ three methods to concatenate tag vectors extracted using an audio tagging system with an intermediate hidden layer of an acoustic scene classification system. We also explore the multi-head attention on the feature map of an acoustic scene classification system using tag vectors. Experiments conducted on the detection and classification of acoustic scenes and events 2019 task 1-a dataset demonstrate the effectiveness of the proposed scheme. Concatenation and multi-head attention show a classification accuracy of 75.66 % and 75.58 %, respectively, compared to 73.63 % accuracy of the baseline. The system with the proposed two approaches combined demonstrates an accuracy of 76.75 %. Jee-Weon Jung, Hye-Jin Shim, Ju-ho Kim, Seung-bin Kim, Ha-Jin Yu |
INTERSPEECH | 5 |
| 2020 | Segment Aggregation for Short Utterances Speaker Verification Using Raw WaveformsabstractMost studies on speaker verification systems focus on longduration utterances, which are composed of sufficient phonetic information.However, the performances of these systems are known to degrade when short-duration utterances are inputted due to the lack of phonetic information as compared to the long utterances.In this paper, we propose a method that compensates for the performance degradation of speaker verification for short utterances, referred to as "segment aggregation".The proposed method adopts an ensemble-based design to improve the stability and accuracy of speaker verification systems.The proposed method segments an input utterance into several short utterances and then aggregates the segment embeddings extracted from the segmented inputs to compose a speaker embedding.Then, this method simultaneously trains the segment embeddings and the aggregated speaker embedding.In addition, we also modified the teacher-student learning method for the proposed method.Experimental results on different input duration using the VoxCeleb1 test set demonstrate that the proposed technique improves speaker verification performance by about 45.37% relatively compared to the baseline system with 1-second test utterance condition. Seung-bin Kim, Jee-Weon Jung, Hye-Jin Shim, Ju-ho Kim, Ha-Jin Yu |
INTERSPEECH | 5 |
| 2020 | Self-Supervised Pre-Training with Acoustic Configurations for Replay Spoofing DetectionabstractConstructing a dataset for replay spoofing detection requires a physical process of playing an utterance and re-recording it, presenting a challenge to the collection of large-scale datasets.In this study, we propose a self-supervised framework for pretraining acoustic configurations using datasets published for other tasks, such as speaker verification.Here, acoustic configurations refer to the environmental factors generated during the process of voice recording but not the voice itself, including microphone types, place and ambient noise levels.Specifically, we select pairs of segments from utterances and train deep neural networks to determine whether the acoustic configurations of the two segments are identical.We validate the effectiveness of the proposed method based on the ASVspoof 2019 physical access dataset utilizing two well-performing systems.The experimental results demonstrate that the proposed method outperforms the baseline approach by 30%. Hye-Jin Shim, Hee-Soo Heo, Jee-Weon Jung, Ha-Jin Yu |
INTERSPEECH | 4 |
| 2019 | Short Utterance Compensation in Speaker Verification via Cosine-Based Teacher-Student Learning of Speaker EmbeddingsabstractThe short duration of an input utterance is one of the most critical threats that degrade the performance of speaker verification systems. This study aimed to develop an integrated text-independent speaker verification system that inputs utterances with short duration of 2 seconds or less. We propose an approach using a teacher-student learning framework for this goal, applied to short utterance compensation for the first time in our knowledge. The core concept of the proposed system is to conduct the compensation throughout the network that extracts the speaker embedding, mainly in phonetic-level, rather than compensating via a separate system after extracting the speaker embedding. In the proposed architecture, phonetic-level features where each feature represents a segment of 130 ms are extracted using convolutional layers. A layer of gated recurrent units extracts an utterance-level feature using phonetic-level features. The proposed approach also adopts a new objective function for teacher-student learning that considers both Kullback-Leibler divergence of output layers and cosine distance of speaker embeddings layers. Experiments were conducted using deep neural networks that take raw waveforms as input, and output speaker embeddings on VoxCelebl dataset. The proposed model showed 16.6 % relative improvement compared to a baseline approach. Jee-Weon Jung, Hee-Soo Heo, Hye-Jin Shim, Ha-Jin Yu |
ASRU | 4 |
| 2019 | Acoustic Scene Classification Using Teacher-Student Learning with Soft-LabelsabstractAcoustic scene classification identifies an input segment into one of the pre-defined classes using spectral information. The spectral information of acoustic scenes may not be mutually exclusive due to common acoustic properties across different classes, such as babble noises included in both airports and shopping malls. However, conventional training procedure based on one-hot labels does not consider the similarities between different acoustic scenes. We exploit teacher-student learning with the purpose to derive soft-labels that consider common acoustic properties among different acoustic scenes. In teacher-student learning, the teacher network produces soft-labels, based on which the student network is trained. We investigate various methods to extract soft-labels that better represent similarities across different scenes. Such attempts include extracting soft-labels from multiple audio segments that are defined as an identical acoustic scene. Experimental results demonstrate the potential of our approach, showing a classification accuracy of 77.36 % on the DCASE 2018 task 1 validation set. Hee-Soo Heo, Jee-Weon Jung, Hye-Jin Shim, Ha-Jin Yu |
INTERSPEECH | 4 |
| 2019 | End-to-End Losses Based on Speaker Basis Vectors and All-Speaker Hard Negative Mining for Speaker VerificationabstractIn recent years, speaker verification has primarily performed using deep neural networks that are trained to output embeddings from input features such as spectrograms or Mel-filterbank energies.Studies that design various loss functions, including metric learning have been widely explored.In this study, we propose two end-to-end loss functions for speaker verification using the concept of speaker bases, which are trainable parameters.One loss function is designed to further increase the interspeaker variation, and the other is designed to conduct the identical concept with hard negative mining.Each speaker basis is designed to represent the corresponding speaker in the process of training deep neural networks.In contrast to the conventional loss functions that can consider only a limited number of speakers included in a mini-batch, the proposed loss functions can consider all the speakers in the training set regardless of the mini-batch composition.In particular, the proposed loss functions enable hard negative mining and calculations of betweenspeaker variations with consideration of all speakers.Through experiments on VoxCeleb1 and VoxCeleb2 datasets, we confirmed that the proposed loss functions could supplement conventional softmax and center loss functions. Hee-Soo Heo, Jee-Weon Jung, Il-Ho Yang, Sung-Hyun Yoon, Hye-Jin Shim, Ha-Jin Yu |
INTERSPEECH | 6 |
| 2019 | RawNet: Advanced End-to-End Deep Neural Network Using Raw Waveforms for Text-Independent Speaker VerificationabstractRecently, direct modeling of raw waveforms using deep neural networks has been widely studied for a number of tasks in audio domains.In speaker verification, however, utilization of raw waveforms is in its preliminary phase, requiring further investigation.In this study, we explore end-to-end deep neural networks that input raw waveforms to improve various aspects: front-end speaker embedding extraction including model architecture, pre-training scheme, additional objective functions, and back-end classification.Adjustment of model architecture using a pre-training scheme can extract speaker embeddings, giving a significant improvement in performance.Additional objective functions simplify the process of extracting speaker embeddings by merging conventional two-phase processes: extracting utterance-level features such as i-vectors or x-vectors and the feature enhancement phase, e.g., linear discriminant analysis.Effective back-end classification models that suit the proposed speaker embedding are also explored.We propose an end-toend system that comprises two deep neural networks, one frontend for utterance-level speaker embedding extraction and the other for back-end classification.Experiments conducted on the VoxCeleb1 dataset demonstrate that the proposed model achieves state-of-the-art performance among systems without data augmentation.The proposed system is also comparable to the state-of-the-art x-vector system that adopts data augmentation. Jee-Weon Jung, Hee-Soo Heo, Ju-ho Kim, Hye-Jin Shim, Ha-Jin Yu |
INTERSPEECH | 5 |
| 2019 | Replay Attack Detection with Complementary High-Resolution Information Using End-to-End DNN for the ASVspoof 2019 ChallengeabstractIn this study, we concentrate on replacing the process of extracting hand-crafted acoustic feature with end-to-end DNN using complementary high-resolution spectrograms.As a result of advance in audio devices, typical characteristics of a replayed speech based on conventional knowledge alter or diminish in unknown replay configurations.Thus, it has become increasingly difficult to detect spoofed speech with a conventional knowledge-based approach.To detect unrevealed characteristics that reside in a replayed speech, we directly input spectrograms into an end-to-end DNN without knowledge-based intervention.Explorations dealt in this study that differentiates from existing spectrogram-based systems are twofold: complementary information and high-resolution.Spectrograms with different information are explored, and it is shown that additional information such as the phase information can be complementary.High-resolution spectrograms are employed with the assumption that the difference between a bona-fide and a replayed speech exists in the details.Additionally, to verify whether other features are complementary to spectrograms, we also examine raw waveform and an i-vector based system.Experiments conducted on the ASVspoof 2019 physical access challenge show promising results, where t-DCF and equal error rates are 0.0570 and 2.45 % for the evaluation set, respectively. Jee-Weon Jung, Hye-Jin Shim, Hee-Soo Heo, Ha-Jin Yu |
INTERSPEECH | 4 |
| 2018 | A Complete End-to-End Speaker Verification System Using Deep Neural Networks: From Raw Signals to Verification ResultabstractEnd-to-end systems using deep neural networks have been widely studied in the field of speaker verification. Raw audio signal processing has also been widely studied in the fields of automatic music tagging and speech recognition. However, as far as we know, end-to-end systems using raw audio signals have not been explored in speaker verification. In this paper, a complete end-to-end speaker verification system is proposed, which inputs raw audio signals and outputs the verification results. A pre-processing layer and the embedded speaker feature extraction models were mainly investigated. The proposed pre-emphasis layer was combined with a strided convolution layer for pre-processing at the first two hidden layers. In addition, speaker feature extraction models using convolutionallayer and long short-term memory are proposed to be embedded in the proposed end-to-end system. Jee-Weon Jung, Hee-Soo Heo, Il-Ho Yang, Hye-Jin Shim, Ha-Jin Yu |
ICASSP | 5 |
| 2018 | Avoiding Speaker Overfitting in End-to-End DNNs Using Raw Waveform for Text-Independent Speaker Verification
Jee-Weon Jung, Hee-Soo Heo, Il-Ho Yang, Hye-Jin Shim, Ha-Jin Yu |
INTERSPEECH | 5 |
| 2017 | Applying compensation techniques on i-vectors extracted from short-test utterances for speaker verification using deep neural networkabstractWe propose a method to improve speaker verification performance when a test utterance is very short. In some situations with short test utterances, performance of ivector/probabilistic linear discriminant analysis systems degrades. The proposed method transforms short-utterance feature vectors to adequate vectors using a deep neural network, which compensate for short utterances. To reduce the dimensionality of the search space, we extract several principal components from the residual vectors between every long utterance i-vector in a development set and its truncated short utterance i-vector. Then an input i-vector of the network is transformed by linear combination of these directions. In this case, network outputs correspond to weights for linear combination of principal components. We use public speech databases to evaluate the method. The experimental results on short2-10sec condition (det6, male portion) of the NIST 2008 speaker recognition evaluation corpus show that the proposed method reduces the minimum detection cost relative to the baseline system, which uses linear discriminant analysis transformed i-vectors as features. Il-Ho Yang, Hee-Soo Heo, Sung-Hyun Yoon, Ha-Jin Yu |
ICASSP | 4 |
| 2017 | Joint Training of Expanded End-to-End DNN for Text-Dependent Speaker Verification
Hee-Soo Heo, Jee-Weon Jung, Il-Ho Yang, Sung-Hyun Yoon, Ha-Jin Yu |
INTERSPEECH | 5 |
| 2017 | Histogram equalization using a reduced feature set of background speakers' utterances for speaker recognitionabstractWe propose a method for histogram equalization using supplement sets to improve the performance of speaker recognition when the training and test utterances are very short. The supplement sets are derived using outputs of selection or clustering algorithms from the background speakers’ utterances. The proposed approach is used as a feature normalization method for building histograms when there are insufficient input utterance samples. In addition, the proposed method is used as an i-vector normalization method in an i-vector-based probabilistic linear discriminant analysis (PLDA) system, which is the current state-of-the-art for speaker verification. The ranks of sample values for histogram equalization are estimated in ascending order from both the input utterances and the supplement set. New ranks are obtained by computing the sum of different kinds of ranks. Subsequently, the proposed method determines the cumulative distribution function of the test utterance using the newly defined ranks. The proposed method is compared with conventional feature normalization methods, such as cepstral mean normalization (CMN), cepstral mean and variance normalization (MVN), histogram equalization (HEQ), and the European Telecommunications Standards Institute (ETSI) advanced front-end methods. In addition, performance is compared for a case in which the greedy selection algorithm is used with fuzzy C -means and K -means algorithms. The YOHO and Electronics and Telecommunications Research Institute (ETRI) databases are used in an evaluation in the feature space. The test sets are simulated by the Opus VoIP codec. We also use the 2008 National Institute of Standards and Technology (NIST) speaker recognition evaluation (SRE) corpus for the i-vector system. The results of the experimental evaluation demonstrate that the average system performance is improved when the proposed method is used, compared to the conventional feature normalization methods. Myung-Jae Kim, Il-Ho Yang, Min-Seok Kim, Ha-Jin Yu |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2016 | Advanced b-vector system based deep neural network as classifier for speaker verificationabstractFew studies on speaker verification have directly used a deep neural network (DNN) as a classifier. It is difficult to directly apply a DNN as a discriminative model to speaker-verification tasks because the training data for each speaker are very limited. Therefore, a b-vector has been proposed to solve the problem. However, the DNN with the b-vectors showed lower performance than the conventional i-vector probabilistic linear-discriminant analysis (PLDA) system. In this paper, we propose an improved version of the b-vector DNN system, which incorporates the background speakers' information into the DNN. In this study, each input feature is paired with a representative background speaker's feature vectors, and a b-vector is extracted from each pair; thus, feeding background information into the DNN. We confirmed that the performance improvements of the proposed system compensate for the shortcomings of conventional b-vectors in experiments carried out using the National Institute of Standards and Technology 2008 Speaker-Recognition Evaluation tests. Hee-Soo Heo, Il-Ho Yang, Myung-Jae Kim, Sung-Hyun Yoon, Ha-Jin Yu |
ICASSP | 5 |
| 2010 | Kernel multimodal discriminant analysis for speaker verificationabstractIn this paper, we propose a robust speaker feature extraction method using kernel multimodal Fisher discriminant analysis (kernel MFDA). Kernel MFDA has been designed to have the characteristics both of kernel principal component analysis (kernel PCA) and kernel Fisher discriminant analysis (kernel FDA). Therefore, the feature vectors extracted by kernel MFDA are denoised as well as discriminated. For evaluation, we compare our proposed method with principal component analysis (PCA) and kernel PCA on the speaker verification systems. Min-Seok Kim, Il-Ho Yang, Ha-Jin Yu |
ICASSP | 3 |
| 2009 | Robust Speaker Identification Using Multimodal Discriminant Analysis with KernelsabstractIn this paper, we propose kernel multimodal fisher discriminant analysis (kernel MFDA), a new non-linear feature transformation method, which can be applied to large-scale problems such as speaker recognition tasks. Our proposed method has characteristics of kernel fisher discriminant analysis (kernel FDA) as well as kernel principal component analysis (kernel PCA). The memory requirement of our proposed method is much lower than the other kernel methods. In the experiments, we apply our proposed method to a speaker identification task, and then we compare the accuracy of this method with kernel FDA and kernel PCA in clean and noisy environments. As the results, our proposed method outperforms kernel PCA. Min-Seok Kim, Il-Ho Yang, Ha-Jin Yu |
ICTAI | 3 |
| 2008 | Robust Speaker Identification Using Greedy Kernel PCAabstractWe propose a robust speaker identification system in noisy environments using greedy kernel principal component analysis. We expect that kernel PCA can project important information to some axes and the noise to some other axes in the arbitrary high dimensional space resulting in denoising of the input features. However, it is not easy to use kernel PCA for speaker identification because the storage required for the kernel matrix grows quadratically, and the computational cost grows linearly with the number of training vectors. Therefore, we use greedy kernel PCA which can approximate kernel PCA with small representation error. In the experiments, we compare the accuracy of the greedy kernel PCA with that of the baseline Gaussian mixture models using MFCCs and PCA in noisy environment. As the results, the greedy kernel PCA outperforms conventional methods. Min-Seok Kim, Il-Ho Yang, Ha-Jin Yu |
ICTAI (2) | 3 |
| 2007 | A New Feature Transformation Method Based on Rotation for Speaker IdentificationabstractIn this paper, we propose a new feature transformation method that is optimized for diagonal covariance Gaussian mixture models which is used for a speaker identification system. We first define an object function as the distances between the Gaussian mixture components and rotate each plane in the feature space to maximize the object function. The optimal degrees of the rotations are found using the particle swarm optimization algorithm. We applied the transformation to a speaker identification task in unknown noisy environments. The proposed transformation is compared with conventional principle component analysis and linear discriminant analysis. The results show that the proposed feature transformation method outperformed existing methods in very high noise environment. Min-Seok Kim, Ha-Jin Yu |
ICTAI (1) | 2 |
| 2002 | A training prompts generation algorithm for connected spoken word recognition
Ha-Jin Yu, Jin Suk Kim |
INTERSPEECH | 1 |
| 2000 | Large vocabulary Korean continuous speech recognition using a one-pass algorithm
Ha-Jin Yu, Joon-Mo Hong, Jong-Seok Lee |
INTERSPEECH | 1 |
| 2000 | A neural network for 500 word vocabulary word spotting using non-uniform units
Ha-Jin Yu, Yung-Hwan Oh |
Neural Networks | 1 |
| 1998 | Automatic recognition of Korean broadcast news speechabstractThis paper describes preliminary results of automatic recognition of Korean broadcast-news speech. We have been working on flexible vocabulary isolated-word speech recognition, and the same HMM models are used for broadcast-news continuous speech recognition. The recognizer is trained by using phonetically balanced isolated words speech, rather than the broadcast news speech itself. In this research, we use several different lexica to investigate the recognition performance according to the length of the words. We also propose a long-distance bigram language model, which can be used at the first stage of the search, so that it can reduce the recognition errors caused by earlier pruning of correct hypothesis. 1. Ha-Jin Yu, Jae-Seung Choi, Joon-Mo Hong, Kew-Suh Park, Jong-Seok Lee, Hee-Youn Lee |
ICSLP | 1 |
| 1997 | A neural network for 500 vocabulary word spotting using acoustic sub-word unitsabstractA neural network model based on a non-uniform unit for speaker-independent continuous speech recognition is proposed. The functions of the neural network model include segmenting the input speech into sub-word units, classifying the units and detecting words, and each of them is implemented by a module. The recognition unit we propose can include an arbitrary number of phonemes in a unit, so that it can absorb co-articulation effects which spread for several phonemes. The unit classifier module separates the speech into stationary and transition parts and use different parameters for them. The word detector module can learn all the pronunciation variations in the training data. The system is evaluated on a subset of the TIMIT speech data. Ha-Jin Yu, Yung-Hwan Oh |
ICASSP | 1 |
| 1996 | A neural network using acoustic sub-word units for continuous speech recognition
Ha-Jin Yu, Yung-Hwan Oh |
ICSLP | 1 |
| 1995 | Estimating Fuzzy Phoneme Similarity Relations for Continuous Speech Recognition
Ha-Jin Yu, Sung-Joo Kim, Yung-Hwan Oh |
IEA/AIE | 1 |
| 1995 | A neural network using non-uniform units for continuous speech recognition
Ha-Jin Yu, Yung-Hwan Oh |
EUROSPEECH | 1 |