EDBT 2026 Demo / reviewers in the wild / expert
Hye-Jin Shim
dblp:35/262 · also Hye-jin Shim
· DBLP profile ↗
32ranked-venue papers
5as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 5 first-author · 17 since 2021Artificial intelligence and machine learning · 24 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards explainable spoofed speech attribution and detection: A probabilistic approach for characterizing speech synthesizer componentsabstractWe propose an explainable probabilistic framework for characterizing spoofed speech by decomposing it into probabilistic attribute embeddings. Unlike raw high-dimensional countermeasure embeddings, which lack interpretability, the proposed probabilistic attribute embeddings aim to detect specific speech synthesizer components, represented through high-level attributes and their corresponding values. We use these probabilistic embeddings with four classifier back-ends to address two downstream tasks: spoofing detection and spoofing attack attribution. The former is the well-known bonafide-spoof detection task, whereas the latter seeks to identify the source method (generator) of a spoofed utterance. We additionally use Shapley values, a widely used technique in machine learning, to quantify the relative contribution of each attribute value to the decision-making process in each task. Results on the ASVspoof2019 dataset demonstrate the substantial role of waveform generator, conversion model outputs, and inputs in spoofing detection; and inputs, speaker, and duration modeling in spoofing attack attribution. In the detection task, the probabilistic attribute embeddings achieve 99.7% balanced accuracy and 0.22% equal error rate (EER), closely matching the performance of raw embeddings (99.9% balanced accuracy and 0.22% EER). Similarly, in the attribution task, our embeddings achieve 90.23% balanced accuracy and 2.07% EER, compared to 90.16% and 2.11% with raw embeddings. These results demonstrate that the proposed framework is both inherently explainable by design and capable of achieving performance comparable to raw CM embeddings. Jagabandhu Mishra, Manasi Chhibber, Hye-Jin Shim, Tomi Kinnunen |
Comput. Speech Lang. | 3 |
| 2026 | ASVspoof 5: Design, collection and validation of resources for spoofing, deepfake, and adversarial attack detection using crowdsourced speechabstractASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake attacks as well as the design of detection solutions. We introduce the ASVspoof 5 database which is generated in a crowdsourced fashion from data collected in diverse acoustic conditions (cf. studio-quality data for earlier ASVspoof databases) and from ∼ 2,000 speakers (cf. ∼ 100 earlier). The database contains attacks generated with 32 different algorithms, also crowdsourced, and optimised to varying degrees using new surrogate detection models. Among them are attacks generated with a mix of legacy and contemporary text-to-speech synthesis and voice conversion models, in addition to adversarial attacks which are incorporated for the first time. ASVspoof 5 protocols comprise seven speaker-disjoint partitions. They include two distinct partitions for the training of different sets of attack models, two more for the development and evaluation of surrogate detection models, and then three additional partitions which comprise the ASVspoof 5 training, development and evaluation sets. An auxiliary set of data collected from an additional 30k speakers can also be used to train speaker encoders for the implementation of attack algorithms. Also described herein is an experimental validation of the new ASVspoof 5 database using a set of automatic speaker verification and spoof/deepfake baseline detectors. With the exception of protocols and tools for the generation of spoofed/deepfake speech, the resources described in this paper, already used by participants of the ASVspoof 5 challenge in 2024, are now all freely available to the community. Xin Wang 0037, Héctor Delgado, Hemlata Tak, Jee-Weon Jung, Hye-Jin Shim, Massimiliano Todisco, Ivan Kukanov, Xuechen Liu 0001, Md. Sahidullah, Tomi Kinnunen, Nicholas W. D. Evans, Kong-Aik Lee, Junichi Yamagishi, Myeonghun Jeong, Yongyi Zang, Soumi Maiti, Florian Lux, Nicolas Müller, Wangyou Zhang, Chengzhe Sun 0001, Shuwei Hou, Siwei Lyu, Sébastien Le Maguer, Hanjie Guo, Vishwanath Pratap Singh |
Comput. Speech Lang. | 5 |
| 2025 | Token-based Attractors and Cross-attention in Spoof DiarizationabstractSpoof diarization identifies “what spoofed when” in a given speech by temporally locating spoofed regions and determining their manipulation techniques. As a first step toward this task, prior work proposed a two-branch model for localization and spoof type clustering, which laid the foundation for spoof diarization. However, its simple structure limits the ability to capture complex spoofing patterns and lacks explicit reference points for distinguishing between bona fide and various spoofing types. To address these limitations, our approach introduces learnable tokens where each token represents acoustic features of bona fide and spoofed speech. These attractors interact with frame-level embeddings to extract discriminative representations, improving separation between genuine and generated speech. Vast experiments on PartialSpoof dataset consistently demon-strate that our approach outperforms existing methods in bona fide detection and spoofing method clustering. Kyo-Won Koo, Chan-yeong Lim, Jee-Weon Jung, Hye-Jin Shim, Ha-Jin Yu |
ASRU | 4 |
| 2025 | VERSA-v2: A Modular and Scalable Toolkit for Speech and Audio Evaluation with Expanded Metrics, Visualization, and LLM IntegrationabstractWe present VERSA-v2, a major upgrade of the Versatile Evaluation of Speech and Audio (VERSA) toolkit for standardized and scalable evaluation across speech, audio, and music tasks. It features a modular, object-oriented architecture that simplifies metric integration and now supports over 100 metrics, organized into curated task-specific packs. VERSA-v2 also introduces interactive visualizations, per-metric profiling, and prompt-based evaluation using both text- and audio-based large language models (LLMs). These advancements make VERSA-v2 a robust, extensible, and LLM-enabled platform for comprehensive and interpretable speech and audio evaluation. Jiatong Shi, Bo-Hao Su, Shikhar Bharadwaj, Shih-Heng Wang, Jionghao Hang, Wei Wang 0010, Wenhao Feng, Yuxun Tang, Nezih Topaloglu, Siddhant Arora, Jinchuan Tian, Hye-Jin Shim, Wangyou Zhang, Wen-Chin Huang, Shinji Watanabe 0001 |
ASRU | 15 |
| 2025 | Geolocation-Aware Robust Spoken Language IdentificationabstractWhile Self-supervised Learning (SSL) has significantly improved Spoken Language Identification (LID), existing models often struggle to consistently classify dialects and accents of the same language as a unified class. To address this challenge, we propose geolocation-aware LID, a novel approach that incorporates language-level geolocation information into the SSL-based LID model. Specifically, we introduce geolocation prediction as an auxiliary task and inject the predicted vectors into intermediate representations as conditioning signals. This explicit conditioning encourages the model to learn more unified representations for dialectal and accented variations. Experiments across six multilingual datasets demonstrate that our approach improves robustness to intra-language variations and unseen domains, achieving new state-of-the-art accuracy on FLEURS (97.7%) and 9.7% relative improvement on ML-SUPERB 2.0 dialect set. Qingzheng Wang, Hye-Jin Shim, Jiancheng Sun, Shinji Watanabe 0001 |
ASRU | 2 |
| 2025 | An Explainable Probabilistic Attribute Embedding Approach for Spoofed Speech CharacterizationabstractWe propose a novel approach for spoofed speech characterization through explainable probabilistic attribute embeddings. In contrast to high-dimensional raw embeddings extracted from a spoofing countermeasure (CM) whose dimensions are not easy to interpret, the probabilistic attributes are designed to gauge the presence or absence of sub-components that make up a specific spoofing attack. These attributes are then applied to two downstream tasks: spoofing detection and attack attribution. To enforce interpretability also to the back-end, we adopt a decision tree classifier. Our experiments on the ASVspoof2019 dataset with spoof CM embeddings extracted from three models (AASIST, Rawboost-AASIST, SSL-AASIST) suggest that the performance of the attribute embeddings are on par with the original raw spoof CM embeddings for both tasks. The best performance achieved with the proposed approach for spoofing detection and attack attribution, in terms of accuracy, is 99.7% and 99.2%, respectively, compared to 99.7% and 94.7% using the raw CM embeddings. To analyze the relative contribution of each attribute, we estimate their Shapley values. Attributes related to acoustic feature prediction, waveform generation (vocoder), and speaker modeling are found important for spoofing detection; while duration modeling, vocoder, and input type play a role in spoofing attack attribution. Manasi Chhibber, Jagabandhu Mishra, Hye-Jin Shim, Tomi Kinnunen |
ICASSP | 3 |
| 2025 | The Text-to-speech in the Wild (TITW) Database
Jee-Weon Jung, Wangyou Zhang, Soumi Maiti, Yihan Wu 0008, Xin Wang 0037, Yuta Matsunaga, Seyun Um, Jinchuan Tian, Hye-Jin Shim, Nicholas W. D. Evans, Joon Son Chung, Shinnosuke Takamichi, Shinji Watanabe 0001 |
INTERSPEECH | 10 |
| 2025 | Uni-VERSA: Versatile Speech Assessment with a Unified Network
Jiatong Shi, Hye-Jin Shim, Shinji Watanabe 0001 |
INTERSPEECH | 2 |
| 2025 | ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric EstimationabstractSpeech signal analysis poses significant challenges, particularly in tasks such as speech quality evaluation and profiling, where the goal is to predict multiple perceptual and objective metrics. For instance, metrics like PESQ (Perceptual Evaluation of Speech Quality), STOI (Short-Time Objective Intelligibility), and MOS (Mean Opinion Score) each capture different aspects of speech quality. However, these metrics often have different scales, assumptions, and dependencies, making joint estimation non-trivial. To address these issues, we introduce ARECHO (Autoregressive Evaluation via Chain-based Hypothesis Optimization), a chain-based, versatile evaluation system for speech assessment grounded in autoregressive dependency modeling. ARECHO is distinguished by three key innovations: (1) a comprehensive speech information tokenization pipeline; (2) a dynamic classifier chain that explicitly captures inter-metric dependencies; and (3) a two-step confidence-oriented decoding algorithm that enhances inference reliability. Experiments demonstrate that ARECHO significantly outperforms the baseline framework across diverse evaluation scenarios, including enhanced speech analysis, speech generation evaluation, and noisy speech evaluation. Furthermore, its dynamic dependency modeling improves interpretability by capturing inter-metric relationships. Across tasks, ARECHO offers reference-free evaluation using its dynamic classifier chain to support subset queries (single or multiple metrics) and reduces error propagation via confidence-oriented decoding. Jiatong Shi, Bo-Hao Su, Hye-Jin Shim, Jinchuan Tian, Samuele Cornell, Siddhant Arora, Shinji Watanabe 0001 |
NeurIPS | 4 |
| 2024 | To what extent can ASV systems naturally defend against spoofing attacks?
Jee-Weon Jung, Xin Wang 0037, Nicholas W. D. Evans, Shinji Watanabe 0001, Hye-Jin Shim, Hemlata Tak, Siddhant Arora, Junichi Yamagishi, Joon Son Chung |
INTERSPEECH | 5 |
| 2023 | Towards Single Integrated Spoofing-aware Speaker Verification Embeddings
Sung Hwan Mun, Hye-Jin Shim, Hemlata Tak, Xin Wang 0037, Xuechen Liu 0001, Md. Sahidullah, Myeonghun Jeong, Min Hyun Han, Massimiliano Todisco, Kong-Aik Lee, Junichi Yamagishi, Nicholas W. D. Evans, Tomi Kinnunen, Nam Soo Kim, Jee-Weon Jung |
INTERSPEECH | 2 |
| 2023 | How to Construct Perfect and Worse-than-Coin-Flip Spoofing Countermeasures: A Word of Warning on Shortcut Learning
Hye-Jin Shim, Rosa González Hautamäki, Md. Sahidullah, Tomi Kinnunen |
INTERSPEECH | 1 |
| 2023 | Multi-Dataset Co-Training with Sharpness-Aware Optimization for Audio Anti-spoofing
Hye-Jin Shim, Jee-Weon Jung, Tomi Kinnunen |
INTERSPEECH | 1 |
| 2022 | AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention NetworksabstractArtefacts that differentiate spoofed from bona-fide utterances can reside in specific temporal or spectral intervals. Their reliable detection usually depends upon computationally demanding ensemble systems where each subsystem is tuned to some specific artefacts. We seek to develop an efficient, single system that can detect a broad range of different spoofing attacks without score-level ensembles. We propose a novel heterogeneous stacking graph attention layer that models artefacts spanning heterogeneous temporal and spectral intervals with a heterogeneous attention mechanism and a stack node. With a new max graph operation that involves a competitive mechanism and a new readout scheme, our approach, named AASIST, outperforms the current state-of-the-art by 20% relative. Even a lightweight variant, AASIST-L, with only 85k parameters, outperforms all competing systems. Jee-Weon Jung, Hee-Soo Heo, Hemlata Tak, Hye-Jin Shim, Joon Son Chung, Bong-Jin Lee, Ha-Jin Yu, Nicholas W. D. Evans |
ICASSP | 4 |
| 2022 | RawNeXt: Speaker Verification System For Variable-Duration Utterances With Deep Layer Aggregation And Extended Dynamic Scaling PoliciesabstractDespite achieving satisfactory performance in speaker verification using deep neural networks, variable-duration utterances remain a challenge that threatens the robustness of systems. To deal with this issue, we propose a speaker verification system called RawNeXt that can handle input raw waveforms of arbitrary length by employing the following two components: (1) A deep layer aggregation strategy enhances speaker information by iteratively and hierarchically aggregating features of various time scales and spectral channels output from blocks. (2) An extended dynamic scaling policy flexibly processes features according to the length of the utterance by selectively merging the activations of different resolution branches in each block. Owing to these two components, our proposed model can extract speaker embeddings rich in time-spectral information and operate dynamically on length variations. Experimental results on the VoxCeleb1 test set consisting of various duration utterances demonstrate that RawNeXt achieves state-of-the-art performance compared to the recently proposed systems. Our code and trained model weights are available at https://github.com/wngh1187/RawNeXt. Ju-ho Kim, Hye-Jin Shim, Jungwoo Heo, Ha-Jin Yu |
ICASSP | 2 |
| 2022 | Graph Attentive Feature Aggregation for Text-Independent Speaker VerificationabstractThe objective of this paper is to combine multiple frame-level features into a single utterance-level representation considering pair-wise relationships. For this purpose, we propose a novel graph attentive feature aggregation module by interpreting each frame-level feature as a node of a graph. The inter-relationship between all possible pairs of features, typically exploited indirectly, can be directly modeled using a graph. The module comprises a graph attention layer and a graph pooling layer followed by a readout operation. The graph attention layer first models the non-Euclidean data manifold between different nodes. Then, the graph pooling layer discards less informative nodes considering the significance of the nodes. Finally, the readout operation combines the remaining nodes into a single representation. We employ two recent systems, SE-ResNet and RawNet2, with different input features and architectures and demonstrate that the proposed feature aggregation module consistently shows a relative improvement over 10%, compared to the baseline. Hye-Jin Shim, Jungwoo Heo, Jae-Han Park, Ga-Hui Lee, Ha-Jin Yu |
ICASSP | 1 |
| 2022 | Attentive Max Feature Map and Joint Training for Acoustic Scene ClassificationabstractVarious attention mechanisms are being widely applied to acoustic scene classification. However, we empirically found that the attention mechanism can excessively discard potentially valuable information, despite improving performance. We propose the attentive max feature map that combines two effective techniques, attention and a max feature map, to further elaborate the attention mechanism and mitigate the above-mentioned phenomenon. We also explore various joint training methods, including multi-task learning, that allocate additional abstract labels for each audio recording. Our proposed system demonstrates competitive performance with much larger state-of-the-art systems for single systems on Subtask A of the DCASE 2020 challenge by applying the two proposed techniques using relatively fewer parameters. Furthermore, adopting the proposed attentive max feature map, our team placed fourth in the recent DCASE 2021 challenge. Hye-Jin Shim, Jee-Weon Jung, Ju-ho Kim, Ha-Jin Yu |
ICASSP | 1 |
| 2022 | SASV 2022: The First Spoofing-Aware Speaker Verification ChallengeabstractThe first spoofing-aware speaker verification (SASV) challenge aims to integrate research efforts in speaker verification and anti-spoofing.We extend the speaker verification scenario by introducing spoofed trials to the usual set of target and impostor trials.In contrast to the established ASVspoof challenge where the focus is upon separate, independently optimised spoofing detection and speaker verification sub-systems, SASV targets the development of integrated and jointly optimised solutions.Pre-trained spoofing detection and speaker verification models are provided as open source and are used in two baseline SASV solutions.Both models and baselines are freely available to participants and can be used to develop back-end fusion approaches or end-to-end solutions.Using the provided common evaluation protocol, 23 teams submitted SASV solutions.When assessed with target, bona fide non-target and spoofed non-target trials, the top-performing system reduces the equal error rate of a conventional speaker verification system from 23.83% to 0.13%.SASV challenge results are a testament to the reliability of today's state-of-the-art approaches to spoofing detection and speaker verification. Jee-Weon Jung, Hemlata Tak, Hye-Jin Shim, Hee-Soo Heo, Bong-Jin Lee, Soo-Whan Chung, Ha-Jin Yu, Nicholas W. D. Evans, Tomi Kinnunen |
INTERSPEECH | 3 |
| 2022 | Extended U-Net for Speaker Verification in Noisy EnvironmentsabstractBackground noise is a well-known factor that deteriorates the accuracy and reliability of speaker verification (SV) systems by blurring speech intelligibility.Various studies have used separate pretrained enhancement models as the front-end module of the SV system in noisy environments, and these methods effectively remove noises.However, the denoising process of independent enhancement models not tailored to the SV task can also distort the speaker information included in utterances.We argue that the enhancement network and speaker embedding extractor should be fully jointly trained for SV tasks under noisy conditions to alleviate this issue.Therefore, we proposed a U-Net-based integrated framework that simultaneously optimizes speaker identification and feature enhancement losses.Moreover, we analyzed the structural limitations of using U-Net directly for noise SV tasks and further proposed Extended U-Net to reduce these drawbacks.We evaluated the models on the noise-synthesized VoxCeleb1 test set and VOiCES development set recorded in various noisy scenarios.The experimental results demonstrate that the U-Net-based fully joint training framework is more effective than the baseline, and the extended U-Net exhibited state-of-the-art performance versus the recently proposed compensation systems. Ju-ho Kim, Jungwoo Heo, Hye-Jin Shim, Ha-Jin Yu |
INTERSPEECH | 3 |
| 2021 | DCASENET: An Integrated Pretrained Deep Neural Network for Detecting and Classifying Acoustic Scenes and EventsabstractAlthough acoustic scenes and events include many related tasks, their combined detection and classification have been scarcely investigated. We propose three architectures of deep neural networks that are integrated to simultaneously perform acoustic scene classification, audio tagging, and sound event detection. The first two architectures are inspired by human cognitive processes. The first architecture resembles the short-term perception for scene classification of adults, who can detect various sound events that are then used to identify the acoustic scene. The second architecture resembles the long-term learning of babies, being also the concept under-lying self-supervised learning. Babies first observe the effects of abstract notions such as gravity and then learn specific tasks using such perceptions. The third architecture adds a few layers to the second one that solely perform a single task before its corresponding output layer. The aim is to build an integrated system that can serve as a pretrained model to perform the three abovementioned tasks. Experiments on three datasets demonstrate that the proposed architecture, called DcaseNet, can be either directly used for any of the tasks while providing suitable results or fine-tuned to improve the performance of one task. The code and pretrained DcaseNet weights are available at https://github.com/Jungjee/DcaseNet. Jee-Weon Jung, Hye-Jin Shim, Ju-ho Kim, Ha-Jin Yu |
ICASSP | 2 |
| 2020 | Improved RawNet with Feature Map Scaling for Text-Independent Speaker Verification Using Raw WaveformsabstractRecent advances in deep learning have facilitated the design of speaker verification systems that directly input raw waveforms.For example, RawNet [1] extracts speaker embeddings from raw waveforms, which simplifies the process pipeline and demonstrates competitive performance.In this study, we improve RawNet by scaling feature maps using various methods.The proposed mechanism utilizes a scale vector that adopts a sigmoid non-linear function.It refers to a vector with dimensionality equal to the number of filters in a given feature map.Using a scale vector, we propose to scale the feature map multiplicatively, additively, or both.In addition, we investigate replacing the first convolution layer with the sinc-convolution layer of SincNet.Experiments performed on the VoxCeleb1 evaluation dataset demonstrate the effectiveness of the proposed methods, and the best performing system reduces the equal error rate by half compared to the original RawNet.Expanded evaluation results obtained using the VoxCeleb1-E and VoxCeleb-H protocols marginally outperform existing state-ofthe-art systems. Jee-Weon Jung, Seung-bin Kim, Hye-Jin Shim, Ju-ho Kim, Ha-Jin Yu |
INTERSPEECH | 3 |
| 2020 | Acoustic Scene Classification Using Audio TaggingabstractAcoustic scene classification systems using deep neural networks classify given recordings into pre-defined classes. In this study, we propose a novel scheme for acoustic scene classification which adopts an audio tagging system inspired by the human perception mechanism. When humans identify an acoustic scene, the existence of different sound events provides discriminative information which affects the judgement. The proposed framework mimics this mechanism using various approaches. Firstly, we employ three methods to concatenate tag vectors extracted using an audio tagging system with an intermediate hidden layer of an acoustic scene classification system. We also explore the multi-head attention on the feature map of an acoustic scene classification system using tag vectors. Experiments conducted on the detection and classification of acoustic scenes and events 2019 task 1-a dataset demonstrate the effectiveness of the proposed scheme. Concatenation and multi-head attention show a classification accuracy of 75.66 % and 75.58 %, respectively, compared to 73.63 % accuracy of the baseline. The system with the proposed two approaches combined demonstrates an accuracy of 76.75 %. Jee-Weon Jung, Hye-Jin Shim, Ju-ho Kim, Seung-bin Kim, Ha-Jin Yu |
INTERSPEECH | 2 |
| 2020 | Segment Aggregation for Short Utterances Speaker Verification Using Raw WaveformsabstractMost studies on speaker verification systems focus on longduration utterances, which are composed of sufficient phonetic information.However, the performances of these systems are known to degrade when short-duration utterances are inputted due to the lack of phonetic information as compared to the long utterances.In this paper, we propose a method that compensates for the performance degradation of speaker verification for short utterances, referred to as "segment aggregation".The proposed method adopts an ensemble-based design to improve the stability and accuracy of speaker verification systems.The proposed method segments an input utterance into several short utterances and then aggregates the segment embeddings extracted from the segmented inputs to compose a speaker embedding.Then, this method simultaneously trains the segment embeddings and the aggregated speaker embedding.In addition, we also modified the teacher-student learning method for the proposed method.Experimental results on different input duration using the VoxCeleb1 test set demonstrate that the proposed technique improves speaker verification performance by about 45.37% relatively compared to the baseline system with 1-second test utterance condition. Seung-bin Kim, Jee-Weon Jung, Hye-Jin Shim, Ju-ho Kim, Ha-Jin Yu |
INTERSPEECH | 3 |
| 2020 | Self-Supervised Pre-Training with Acoustic Configurations for Replay Spoofing DetectionabstractConstructing a dataset for replay spoofing detection requires a physical process of playing an utterance and re-recording it, presenting a challenge to the collection of large-scale datasets.In this study, we propose a self-supervised framework for pretraining acoustic configurations using datasets published for other tasks, such as speaker verification.Here, acoustic configurations refer to the environmental factors generated during the process of voice recording but not the voice itself, including microphone types, place and ambient noise levels.Specifically, we select pairs of segments from utterances and train deep neural networks to determine whether the acoustic configurations of the two segments are identical.We validate the effectiveness of the proposed method based on the ASVspoof 2019 physical access dataset utilizing two well-performing systems.The experimental results demonstrate that the proposed method outperforms the baseline approach by 30%. Hye-Jin Shim, Hee-Soo Heo, Jee-Weon Jung, Ha-Jin Yu |
INTERSPEECH | 1 |
| 2019 | Short Utterance Compensation in Speaker Verification via Cosine-Based Teacher-Student Learning of Speaker EmbeddingsabstractThe short duration of an input utterance is one of the most critical threats that degrade the performance of speaker verification systems. This study aimed to develop an integrated text-independent speaker verification system that inputs utterances with short duration of 2 seconds or less. We propose an approach using a teacher-student learning framework for this goal, applied to short utterance compensation for the first time in our knowledge. The core concept of the proposed system is to conduct the compensation throughout the network that extracts the speaker embedding, mainly in phonetic-level, rather than compensating via a separate system after extracting the speaker embedding. In the proposed architecture, phonetic-level features where each feature represents a segment of 130 ms are extracted using convolutional layers. A layer of gated recurrent units extracts an utterance-level feature using phonetic-level features. The proposed approach also adopts a new objective function for teacher-student learning that considers both Kullback-Leibler divergence of output layers and cosine distance of speaker embeddings layers. Experiments were conducted using deep neural networks that take raw waveforms as input, and output speaker embeddings on VoxCelebl dataset. The proposed model showed 16.6 % relative improvement compared to a baseline approach. Jee-Weon Jung, Hee-Soo Heo, Hye-Jin Shim, Ha-Jin Yu |
ASRU | 3 |
| 2019 | Acoustic Scene Classification Using Teacher-Student Learning with Soft-LabelsabstractAcoustic scene classification identifies an input segment into one of the pre-defined classes using spectral information. The spectral information of acoustic scenes may not be mutually exclusive due to common acoustic properties across different classes, such as babble noises included in both airports and shopping malls. However, conventional training procedure based on one-hot labels does not consider the similarities between different acoustic scenes. We exploit teacher-student learning with the purpose to derive soft-labels that consider common acoustic properties among different acoustic scenes. In teacher-student learning, the teacher network produces soft-labels, based on which the student network is trained. We investigate various methods to extract soft-labels that better represent similarities across different scenes. Such attempts include extracting soft-labels from multiple audio segments that are defined as an identical acoustic scene. Experimental results demonstrate the potential of our approach, showing a classification accuracy of 77.36 % on the DCASE 2018 task 1 validation set. Hee-Soo Heo, Jee-Weon Jung, Hye-Jin Shim, Ha-Jin Yu |
INTERSPEECH | 3 |
| 2019 | End-to-End Losses Based on Speaker Basis Vectors and All-Speaker Hard Negative Mining for Speaker VerificationabstractIn recent years, speaker verification has primarily performed using deep neural networks that are trained to output embeddings from input features such as spectrograms or Mel-filterbank energies.Studies that design various loss functions, including metric learning have been widely explored.In this study, we propose two end-to-end loss functions for speaker verification using the concept of speaker bases, which are trainable parameters.One loss function is designed to further increase the interspeaker variation, and the other is designed to conduct the identical concept with hard negative mining.Each speaker basis is designed to represent the corresponding speaker in the process of training deep neural networks.In contrast to the conventional loss functions that can consider only a limited number of speakers included in a mini-batch, the proposed loss functions can consider all the speakers in the training set regardless of the mini-batch composition.In particular, the proposed loss functions enable hard negative mining and calculations of betweenspeaker variations with consideration of all speakers.Through experiments on VoxCeleb1 and VoxCeleb2 datasets, we confirmed that the proposed loss functions could supplement conventional softmax and center loss functions. Hee-Soo Heo, Jee-Weon Jung, Il-Ho Yang, Sung-Hyun Yoon, Hye-Jin Shim, Ha-Jin Yu |
INTERSPEECH | 5 |
| 2019 | RawNet: Advanced End-to-End Deep Neural Network Using Raw Waveforms for Text-Independent Speaker VerificationabstractRecently, direct modeling of raw waveforms using deep neural networks has been widely studied for a number of tasks in audio domains.In speaker verification, however, utilization of raw waveforms is in its preliminary phase, requiring further investigation.In this study, we explore end-to-end deep neural networks that input raw waveforms to improve various aspects: front-end speaker embedding extraction including model architecture, pre-training scheme, additional objective functions, and back-end classification.Adjustment of model architecture using a pre-training scheme can extract speaker embeddings, giving a significant improvement in performance.Additional objective functions simplify the process of extracting speaker embeddings by merging conventional two-phase processes: extracting utterance-level features such as i-vectors or x-vectors and the feature enhancement phase, e.g., linear discriminant analysis.Effective back-end classification models that suit the proposed speaker embedding are also explored.We propose an end-toend system that comprises two deep neural networks, one frontend for utterance-level speaker embedding extraction and the other for back-end classification.Experiments conducted on the VoxCeleb1 dataset demonstrate that the proposed model achieves state-of-the-art performance among systems without data augmentation.The proposed system is also comparable to the state-of-the-art x-vector system that adopts data augmentation. Jee-Weon Jung, Hee-Soo Heo, Ju-ho Kim, Hye-Jin Shim, Ha-Jin Yu |
INTERSPEECH | 4 |
| 2019 | Replay Attack Detection with Complementary High-Resolution Information Using End-to-End DNN for the ASVspoof 2019 ChallengeabstractIn this study, we concentrate on replacing the process of extracting hand-crafted acoustic feature with end-to-end DNN using complementary high-resolution spectrograms.As a result of advance in audio devices, typical characteristics of a replayed speech based on conventional knowledge alter or diminish in unknown replay configurations.Thus, it has become increasingly difficult to detect spoofed speech with a conventional knowledge-based approach.To detect unrevealed characteristics that reside in a replayed speech, we directly input spectrograms into an end-to-end DNN without knowledge-based intervention.Explorations dealt in this study that differentiates from existing spectrogram-based systems are twofold: complementary information and high-resolution.Spectrograms with different information are explored, and it is shown that additional information such as the phase information can be complementary.High-resolution spectrograms are employed with the assumption that the difference between a bona-fide and a replayed speech exists in the details.Additionally, to verify whether other features are complementary to spectrograms, we also examine raw waveform and an i-vector based system.Experiments conducted on the ASVspoof 2019 physical access challenge show promising results, where t-DCF and equal error rates are 0.0570 and 2.45 % for the evaluation set, respectively. Jee-Weon Jung, Hye-Jin Shim, Hee-Soo Heo, Ha-Jin Yu |
INTERSPEECH | 2 |
| 2018 | A Complete End-to-End Speaker Verification System Using Deep Neural Networks: From Raw Signals to Verification ResultabstractEnd-to-end systems using deep neural networks have been widely studied in the field of speaker verification. Raw audio signal processing has also been widely studied in the fields of automatic music tagging and speech recognition. However, as far as we know, end-to-end systems using raw audio signals have not been explored in speaker verification. In this paper, a complete end-to-end speaker verification system is proposed, which inputs raw audio signals and outputs the verification results. A pre-processing layer and the embedded speaker feature extraction models were mainly investigated. The proposed pre-emphasis layer was combined with a strided convolution layer for pre-processing at the first two hidden layers. In addition, speaker feature extraction models using convolutionallayer and long short-term memory are proposed to be embedded in the proposed end-to-end system. Jee-Weon Jung, Hee-Soo Heo, Il-Ho Yang, Hye-Jin Shim, Ha-Jin Yu |
ICASSP | 4 |
| 2018 | Avoiding Speaker Overfitting in End-to-End DNNs Using Raw Waveform for Text-Independent Speaker Verification
Jee-Weon Jung, Hee-Soo Heo, Il-Ho Yang, Hye-Jin Shim, Ha-Jin Yu |
INTERSPEECH | 4 |
| 2004 | Effective Digital Watermarking Algorithm by Contour Detection
Won-Hyuck Choi, Hye-Jin Shim, Jung-Sun Kim |
ICCSA (4) | 2 |