Hee-Soo Heo

dblp:180/2735 · also Hee Soo Heo, HeeSoo Heo, Heesoo Heo · DBLP profile ↗
← Back
34ranked-venue papers
7as first author
22since 2021 · last 2024
0000-0003-1567-123XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 7 first-author · 22 since 2021Artificial intelligence and machine learning · 19 · 4 first-author · 10 since 2021
YearPublicationVenuePosition
2024 Rethinking Session Variability: Leveraging Session Embeddings for Session Robustness in Speaker Verification
abstract
In the field of speaker verification, session or channel variability poses a significant challenge. While many contemporary methods aim to disentangle session information from speaker embeddings, we introduce a novel approach using an additional embedding to represent the session information. This is achieved by training an auxiliary network appended to the speaker embedding extractor which remains fixed in this training process. This results in two similarity scores: one for the speakers information and one for the session information. The latter score acts as a compensator for the former that might be skewed due to session variations. Our extensive experiments demonstrate that session information can be effectively compensated without retraining of the embedding extractor.
Hee-Soo Heo, Kihyun Nam, Bong-Jin Lee, Youngki Kwon, You Jin Kim, Joon Son Chung
ICASSP1
2024 Self-supervised speaker verification with relational mask prediction
Ju-ho Kim, Hee-Soo Heo, Bong-Jin Lee, Youngki Kwon, Ha-Jin Yu
INTERSPEECH2
2024 Disentangled Representation Learning for Environment-agnostic Speaker Recognition
Kihyun Nam, Hee-Soo Heo, Jee-Weon Jung, Joon Son Chung
INTERSPEECH2
2023 High-Resolution Embedding Extractor for Speaker Diarisation
abstract
Speaker embedding extractors significantly influence the performance of clustering-based speaker diarisation systems. Conventionally, only one embedding is extracted from each speech segment. However, because of the sliding window approach, a segment easily includes two or more speakers owing to speaker change points. This study proposes a novel embedding extractor architecture, referred to as a high-resolution embedding extractor (HEE), which extracts multiple high-resolution embeddings from each speech segment. Hee consists of a feature-map extractor and an enhancer, where the enhancer with the self-attention mechanism is the key to success. The enhancer of HEE replaces the aggregation process; instead of a global pooling layer, the enhancer combines relative information to each frame via attention leveraging the global context. Extracted dense frame-level embeddings can each represent a speaker. Thus, multiple speakers can be represented by different frame-level features in each segment. We also propose an artificially generating mixture data training framework to train the proposed HEE. Through experiments on five evaluation sets, including four public datasets, the proposed HEE demonstrates at least 10% improvement on each evaluation set, except for one dataset, which we analyse that rapid speaker changes less exist.
Hee-Soo Heo, Youngki Kwon, Bong-Jin Lee, You Jin Kim, Jee-Weon Jung
ICASSP1
2023 In Search of Strong Embedding Extractors for Speaker Diarisation
abstract
Speaker embedding extractors (EEs), which map input audio to a speaker discriminant latent space, are of paramount importance in speaker diarisation. However, there are several challenges when adopting EEs for diarisation, from which we tackle two key problems. First, the evaluation is not straightforward because the required features differ between speaker verification and diarisation. We show that better performance on widely adopted speaker verification evaluation protocols does not lead to better diarisation performance. Second, embedding extractors have not seen utterances in which multiple speakers exist. These inputs are inevitably present in speaker diarisation because of overlapped speech and speaker changes; they degrade the performance. To mitigate the first problem, we generate speaker verification evaluation protocols that better mimic the diarisation scenario. We propose two data augmentation techniques to alleviate the second problem, making embedding extractors aware of overlapped speech or speaker change input. One technique generates overlapped speech segments, and the other generates segments where two speakers utter sequentially. Extensive experimental results using three state-of-the-art speaker embedding extractors demonstrate that both proposed approaches are effective.
Jee-Weon Jung, Hee-Soo Heo, Bong-Jin Lee, Jaesung Huh, Andrew Brown 0006, Youngki Kwon, Shinji Watanabe 0001, Joon Son Chung
ICASSP2
2023 Advancing the Dimensionality Reduction of Speaker Embeddings for Speaker Diarisation: Disentangling Noise and Informing Speech Activity
abstract
The objective of this work is to train noise-robust speaker embeddings adapted for speaker diarisation. Speaker embeddings play a crucial role in the performance of diarisation systems, but they often capture spurious information such as noise, adversely affecting performance. Our previous work has proposed an auto-encoder-based dimensionality reduction module to help remove the redundant information. However, they do not explicitly separate such information and have also been found to be sensitive to hyper-parameter values. To this end, we propose two contributions to overcome these issues: (i) a novel dimensionality reduction framework that can disentangle spurious information from the speaker embeddings; (ii) the use of speech activity vector to prevent the speaker code from representing the background noise. Through a range of experiments conducted on four datasets, our approach consistently demonstrates the state-of-the-art performance among models without system fusion.
You Jin Kim, Hee-Soo Heo, Jee-Weon Jung, Youngki Kwon, Bong-Jin Lee, Joon Son Chung
ICASSP2
2023 Absolute Decision Corrupts Absolutely: Conservative Online Speaker Diarisation
abstract
Our focus lies in developing an online speaker diarisation framework which demonstrates robust performance across diverse domains. In online speaker diarisation, outputs generated in real-time are irreversible, and a few misjudgements in the early phase of an input session can lead to catastrophic results. We hypothesise that cautiously increasing the number of estimated speakers is of paramount importance among many other factors. Thus, our proposed framework includes decreasing the number of speakers by one when the system judges that an increase in the past was faulty. We also adopt dual buffers, checkpoints and centroids, where checkpoints are combined with silhouette coefficients to estimate the number of speakers and centroids represent speakers. Again, we believe that more than one centroid can be generated from one speaker. Thus we design a clustering-based label matching technique to assign labels in realtime. The resulting system is lightweight yet surprisingly effective. The system demonstrates state-of-the-art performance on DIHARD II and III datasets, where it is also competitive in AMI and VoxConverse test sets.
Youngki Kwon, Hee-Soo Heo, Bong-Jin Lee, You Jin Kim, Jee-Weon Jung
ICASSP2
2023 Curriculum Learning for Self-supervised Speaker Verification
Hee-Soo Heo, Jee-Weon Jung, Youngki Kwon, Bong-Jin Lee, You Jin Kim, Joon Son Chung
INTERSPEECH1
2023 Encoder-decoder Multimodal Speaker Change Detection
Jee-Weon Jung, Soonshin Seo, Hee-Soo Heo, Geonmin Kim, You Jin Kim, Youngki Kwon, Bong-Jin Lee
INTERSPEECH3
2023 Disentangled Representation Learning for Multilingual Speaker Recognition
Kihyun Nam, Youkyum Kim, Jaesung Huh, Hee-Soo Heo, Jee-Weon Jung, Joon Son Chung
INTERSPEECH4
2022 AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks
abstract
Artefacts that differentiate spoofed from bona-fide utterances can reside in specific temporal or spectral intervals. Their reliable detection usually depends upon computationally demanding ensemble systems where each subsystem is tuned to some specific artefacts. We seek to develop an efficient, single system that can detect a broad range of different spoofing attacks without score-level ensembles. We propose a novel heterogeneous stacking graph attention layer that models artefacts spanning heterogeneous temporal and spectral intervals with a heterogeneous attention mechanism and a stack node. With a new max graph operation that involves a competitive mechanism and a new readout scheme, our approach, named AASIST, outperforms the current state-of-the-art by 20% relative. Even a lightweight variant, AASIST-L, with only 85k parameters, outperforms all competing systems.
Jee-Weon Jung, Hee-Soo Heo, Hemlata Tak, Hye-Jin Shim, Joon Son Chung, Bong-Jin Lee, Ha-Jin Yu, Nicholas W. D. Evans
ICASSP2
2022 Multi-Scale Speaker Embedding-Based Graph Attention Networks For Speaker Diarisation
abstract
The objective of this work is effective speaker diarisation using multi-scale speaker embeddings. Typically, there is a trade-off between the ability to recognise short speaker segments and the discriminative power of the embedding, according to the segment length used for embedding extraction. To this end, recent works have proposed the use of multi-scale embeddings where segments with varying lengths are used. However, the scores are combined using a weighted summation scheme where the weights are fixed after the training phase, whereas the importance of segment lengths can differ within a single session.To address this issue, we present three key contributions in this paper: (1) we propose graph attention networks for multi-scale speaker diarisation; (2) we design scale indicators to utilise scale information of each embedding; (3) we adapt the attention-based aggregation to utilise a pre-computed affinity matrix from multi-scale embeddings.We demonstrate the effectiveness of our method in various datasets where the speaker confusion which constitutes the primary metric drops over 10% in average relative compared to the baseline.
Youngki Kwon, Hee-Soo Heo, Jee-Weon Jung, You Jin Kim, Bong-Jin Lee, Joon Son Chung
ICASSP2
2022 Pushing the limits of raw waveform speaker recognition
abstract
In recent years, speaker recognition systems based on raw waveform inputs have received increasing attention.However, the performance of such systems are typically inferior to the state-of-the-art handcrafted feature-based counterparts, which demonstrate equal error rates under 1% on the popular VoxCeleb1 test set.This paper proposes a novel speaker recognition model based on raw waveform inputs.The model incorporates recent advances in machine learning and speaker verification, including the Res2Net backbone module and multi-layer feature aggregation.Our best model achieves an equal error rate of 0.89%, which is competitive with the state-of-the-art models based on handcrafted features, and outperforms the best model based on raw waveform inputs by a large margin.We also explore the application of the proposed model in the context of self-supervised learning framework.Our self-supervised model outperforms single phase-based existing works in this line of research.Finally, we show that self-supervised pre-training is effective for the semi-supervised scenario where we only have a small set of labelled training data, along with a larger set of unlabelled examples.
Jee-Weon Jung, You Jin Kim, Hee-Soo Heo, Bong-Jin Lee, Youngki Kwon, Joon Son Chung
INTERSPEECH3
2022 SASV 2022: The First Spoofing-Aware Speaker Verification Challenge
abstract
The first spoofing-aware speaker verification (SASV) challenge aims to integrate research efforts in speaker verification and anti-spoofing.We extend the speaker verification scenario by introducing spoofed trials to the usual set of target and impostor trials.In contrast to the established ASVspoof challenge where the focus is upon separate, independently optimised spoofing detection and speaker verification sub-systems, SASV targets the development of integrated and jointly optimised solutions.Pre-trained spoofing detection and speaker verification models are provided as open source and are used in two baseline SASV solutions.Both models and baselines are freely available to participants and can be used to develop back-end fusion approaches or end-to-end solutions.Using the provided common evaluation protocol, 23 teams submitted SASV solutions.When assessed with target, bona fide non-target and spoofed non-target trials, the top-performing system reduces the equal error rate of a conventional speaker verification system from 23.83% to 0.13%.SASV challenge results are a testament to the reliability of today's state-of-the-art approaches to spoofing detection and speaker verification.
Jee-Weon Jung, Hemlata Tak, Hye-Jin Shim, Hee-Soo Heo, Bong-Jin Lee, Soo-Whan Chung, Ha-Jin Yu, Nicholas W. D. Evans, Tomi Kinnunen
INTERSPEECH4
2021 Graph Attention Networks for Speaker Verification
abstract
This work presents a novel back-end framework for speaker verification using graph attention networks. Segment-wise speaker embeddings extracted from multiple crops within an utterance are interpreted as node representations of a graph. The proposed framework inputs segment-wise speaker embeddings from an enrollment and a test utterance and directly outputs a similarity score. We first construct a graph using segment-wise speaker embeddings and then input these to graph attention networks. After a few graph attention layers with residual connections, each node is projected into a one-dimensional space using affine transform, followed by a readout operation resulting in a scalar similarity score. To enable successful adaptation for speaker verification, we propose techniques such as separating trainable weights for attention map calculations between segment-wise speaker embeddings from different utterances. The effectiveness of the proposed framework is validated using three different speaker embedding extractors trained with different architectures and objective functions. Experimental results demonstrate consistent improvement over various baseline back-end classifiers, with an average equal error rate improvement of 20% over the cosine similarity back-end without test time augmentation.
Jee-Weon Jung, Hee-Soo Heo, Ha-Jin Yu, Joon Son Chung
ICASSP2
2021 The ins and outs of speaker recognition: lessons from VoxSRC 2020
abstract
The VoxCeleb Speaker Recognition Challenge (VoxSRC) at Interspeech 2020 offers a challenging evaluation for speaker recognition systems, which includes celebrities playing different parts in movies. The goal of this work is robust speaker recognition of utterances recorded in these challenging environments. We utilise variants of the popular ResNet architecture for speaker recognition and perform extensive experiments using a range of loss functions and training parameters. To this end, we optimise an efficient training framework that allows powerful models to be trained with limited time and resources. Our trained models demonstrate improvements over most existing works with lighter models and a simple pipeline. The paper shares the lessons learned from our participation in the challenge.
Yoohwan Kwon, Hee-Soo Heo, Bong-Jin Lee, Joon Son Chung
ICASSP2
2021 Three-Class Overlapped Speech Detection Using a Convolutional Recurrent Neural Network
abstract
In this work, we propose an overlapped speech detection system trained as a three-class classifier. Unlike conventional systems that perform binary classification as to whether or not a frame contains overlapped speech, the proposed approach classifies into three classes: non-speech, single speaker speech, and overlapped speech. By training a network with the more detailed label definition, the model can learn a better notion on deciding the number of speakers included in a given frame. A convolutional recurrent neural network architecture is explored to benefit from both convolutional layer's capability to model local patterns and recurrent layer's ability to model sequential information. The proposed overlapped speech detection model establishes a state-of-the-art performance with a precision of 0.6648 and a recall of 0.3222 on the DIHARD II evaluation set, showing a 20% increase in recall along with higher precision. In addition, we also introduce a simple approach to utilize the proposed overlapped speech detection model for speaker diarization which ranked third place in the Track 1 of the DIHARD III challenge.
Jee-Weon Jung, Hee-Soo Heo, Youngki Kwon, Joon Son Chung, Bong-Jin Lee
Interspeech2
2021 Look Who's Talking: Active Speaker Detection in the Wild
abstract
In this work, we present a novel audio-visual dataset for active speaker detection in the wild.A speaker is considered active when his or her face is visible and the voice is audible simultaneously.Although active speaker detection is a crucial pre-processing step for many audio-visual tasks, there is no existing dataset of natural human speech to evaluate the performance of active speaker detection.We therefore curate the Active Speakers in the Wild (ASW) dataset which contains videos and co-occurring speech segments with dense speech activity labels.Videos and timestamps of audible segments are parsed and adopted from VoxConverse, an existing speaker diarisation dataset that consists of videos in the wild.Face tracks are extracted from the videos and active segments are annotated based on the timestamps of VoxConverse in a semi-automatic way.Two reference systems, a self-supervised system and a fully supervised one, are evaluated on the dataset to provide the baseline performances of ASW.Cross-domain evaluation is conducted in order to show the negative effect of dubbed videos in the training data.
You Jin Kim, Hee-Soo Heo, Soyeon Choe, Soo-Whan Chung, Yoohwan Kwon, Bong-Jin Lee, Youngki Kwon, Joon Son Chung
Interspeech2
2021 Adapting Speaker Embeddings for Speaker Diarisation
abstract
The goal of this paper is to adapt speaker embeddings for solving the problem of speaker diarisation. The quality of speaker embeddings is paramount to the performance of speaker diarisation systems. Despite this, prior works in the field have directly used embeddings designed only to be effective on the speaker verification task. In this paper, we propose three techniques that can be used to better adapt the speaker embeddings for diarisation: dimensionality reduction, attention-based embedding aggregation, and non-speech clustering. A wide range of experiments is performed on various challenging datasets. The results demonstrate that all three techniques contribute positively to the performance of the diarisation system achieving an average relative improvement of 25.07% in terms of diarisation error rate over the baseline.
Youngki Kwon, Jee-Weon Jung, Hee-Soo Heo, You Jin Kim, Bong-Jin Lee, Joon Son Chung
Interspeech3
2021 Metric Learning for Keyword Spotting
abstract
The goal of this work is to train effective representations for keyword spotting via metric learning. Most existing works address keyword spotting as a closed-set classification problem, where both target and non-target keywords are predefined. Therefore, prevailing classifier-based keyword spot-ting systems perform poorly on non-target sounds which are unseen during the training stage, causing high false alarm rates in real-world scenarios. In reality, keyword spotting is a detection problem where predefined target keywords are detected from a variety of unknown sounds. This shares many similarities to metric learning problems in that the unseen and unknown non-target sounds must be clearly differentiated from the target keywords. However, a key difference is that the target keywords are known and predefined. To this end, we propose a new method based on metric learning that maximises the distance between target and non-target key-words, but also learns per-class weights for target keywords as in classification objectives. Experiments on the Google Speech Commands dataset show that our method significantly reduces false alarms to unseen non-target keywords, while maintaining the overall classification accuracy.
Jaesung Huh, Hee-Soo Heo, Seongkyu Mun, Joon Son Chung
SLT3
2021 End-To-End Lip Synchronisation Based on Pattern Classification
abstract
The goal of this work is to synchronise audio and video of a talking face using deep neural network models. Existing works have trained networks on proxy tasks such as cross-modal similarity learning, and then computed similarities between audio and video frames using a sliding window approach. While these methods demonstrate satisfactory performance, the networks are not trained directly on the task. To this end, we propose an end-to-end trained network that can directly predict the offset between an audio stream and the corresponding video stream. The similarity matrix between the two modalities is first computed from the features, then the inference of the offset can be considered to be a pattern recognition problem where the matrix is considered equivalent to an image. The feature extractor and the classifier are trained jointly. We demonstrate that the proposed approach outperforms the previous work by a large margin on LRS2 and LRS3 datasets.
You Jin Kim, Hee-Soo Heo, Soo-Whan Chung, Bong-Jin Lee
SLT2
2021 Look Who's Not Talking
abstract
The objective of this work is speaker diarisation of speech recordings `in the wild'. The ability to determine speech segments is a crucial part of diarisation systems, accounting for a large proportion of errors. In this paper, we present a simple but effective solution for speech activity detection based on the speaker embeddings. In particular, we discover that the norm of the speaker embedding is an extremely effective indicator of speech activity. The method does not require an independent model for speech activity detection, therefore allows speaker diarisation to be performed using a unified representation for both speaker modelling and speech activity detection. We perform a number of experiments on in-house and public datasets, in which our method outperforms popular baselines.
Youngki Kwon, Hee-Soo Heo, Jaesung Huh, Bong-Jin Lee, Joon Son Chung
SLT2
2020 In Defence of Metric Learning for Speaker Recognition
abstract
The objective of this paper is 'open-set' speaker recognition of unseen speakers, where ideal embeddings should be able to condense information into a compact utterance-level representation that has small intra-speaker and large inter-speaker distance. A popular belief in speaker recognition is that networks trained with classification objectives outperform metric learning methods. In this paper, we present an extensive evaluation of most popular loss functions for speaker recognition on the VoxCeleb dataset. We demonstrate that the vanilla triplet loss shows competitive performance compared to classification-based losses, and those trained with our proposed metric learning objective outperform state-of-the-art methods.
Joon Son Chung, Jaesung Huh, Seongkyu Mun, Hee-Soo Heo, Soyeon Choe, Chiheon Ham, Sunghwan Jung, Bong-Jin Lee, Icksang Han
INTERSPEECH5
2020 Self-Supervised Pre-Training with Acoustic Configurations for Replay Spoofing Detection
abstract
Constructing a dataset for replay spoofing detection requires a physical process of playing an utterance and re-recording it, presenting a challenge to the collection of large-scale datasets.In this study, we propose a self-supervised framework for pretraining acoustic configurations using datasets published for other tasks, such as speaker verification.Here, acoustic configurations refer to the environmental factors generated during the process of voice recording but not the voice itself, including microphone types, place and ambient noise levels.Specifically, we select pairs of segments from utterances and train deep neural networks to determine whether the acoustic configurations of the two segments are identical.We validate the effectiveness of the proposed method based on the ASVspoof 2019 physical access dataset utilizing two well-performing systems.The experimental results demonstrate that the proposed method outperforms the baseline approach by 30%.
Hye-Jin Shim, Hee-Soo Heo, Jee-Weon Jung, Ha-Jin Yu
INTERSPEECH2
2019 Short Utterance Compensation in Speaker Verification via Cosine-Based Teacher-Student Learning of Speaker Embeddings
abstract
The short duration of an input utterance is one of the most critical threats that degrade the performance of speaker verification systems. This study aimed to develop an integrated text-independent speaker verification system that inputs utterances with short duration of 2 seconds or less. We propose an approach using a teacher-student learning framework for this goal, applied to short utterance compensation for the first time in our knowledge. The core concept of the proposed system is to conduct the compensation throughout the network that extracts the speaker embedding, mainly in phonetic-level, rather than compensating via a separate system after extracting the speaker embedding. In the proposed architecture, phonetic-level features where each feature represents a segment of 130 ms are extracted using convolutional layers. A layer of gated recurrent units extracts an utterance-level feature using phonetic-level features. The proposed approach also adopts a new objective function for teacher-student learning that considers both Kullback-Leibler divergence of output layers and cosine distance of speaker embeddings layers. Experiments were conducted using deep neural networks that take raw waveforms as input, and output speaker embeddings on VoxCelebl dataset. The proposed model showed 16.6 % relative improvement compared to a baseline approach.
Jee-Weon Jung, Hee-Soo Heo, Hye-Jin Shim, Ha-Jin Yu
ASRU2
2019 Acoustic Scene Classification Using Teacher-Student Learning with Soft-Labels
abstract
Acoustic scene classification identifies an input segment into one of the pre-defined classes using spectral information. The spectral information of acoustic scenes may not be mutually exclusive due to common acoustic properties across different classes, such as babble noises included in both airports and shopping malls. However, conventional training procedure based on one-hot labels does not consider the similarities between different acoustic scenes. We exploit teacher-student learning with the purpose to derive soft-labels that consider common acoustic properties among different acoustic scenes. In teacher-student learning, the teacher network produces soft-labels, based on which the student network is trained. We investigate various methods to extract soft-labels that better represent similarities across different scenes. Such attempts include extracting soft-labels from multiple audio segments that are defined as an identical acoustic scene. Experimental results demonstrate the potential of our approach, showing a classification accuracy of 77.36 % on the DCASE 2018 task 1 validation set.
Hee-Soo Heo, Jee-Weon Jung, Hye-Jin Shim, Ha-Jin Yu
INTERSPEECH1
2019 End-to-End Losses Based on Speaker Basis Vectors and All-Speaker Hard Negative Mining for Speaker Verification
abstract
In recent years, speaker verification has primarily performed using deep neural networks that are trained to output embeddings from input features such as spectrograms or Mel-filterbank energies.Studies that design various loss functions, including metric learning have been widely explored.In this study, we propose two end-to-end loss functions for speaker verification using the concept of speaker bases, which are trainable parameters.One loss function is designed to further increase the interspeaker variation, and the other is designed to conduct the identical concept with hard negative mining.Each speaker basis is designed to represent the corresponding speaker in the process of training deep neural networks.In contrast to the conventional loss functions that can consider only a limited number of speakers included in a mini-batch, the proposed loss functions can consider all the speakers in the training set regardless of the mini-batch composition.In particular, the proposed loss functions enable hard negative mining and calculations of betweenspeaker variations with consideration of all speakers.Through experiments on VoxCeleb1 and VoxCeleb2 datasets, we confirmed that the proposed loss functions could supplement conventional softmax and center loss functions.
Hee-Soo Heo, Jee-Weon Jung, Il-Ho Yang, Sung-Hyun Yoon, Hye-Jin Shim, Ha-Jin Yu
INTERSPEECH1
2019 RawNet: Advanced End-to-End Deep Neural Network Using Raw Waveforms for Text-Independent Speaker Verification
abstract
Recently, direct modeling of raw waveforms using deep neural networks has been widely studied for a number of tasks in audio domains.In speaker verification, however, utilization of raw waveforms is in its preliminary phase, requiring further investigation.In this study, we explore end-to-end deep neural networks that input raw waveforms to improve various aspects: front-end speaker embedding extraction including model architecture, pre-training scheme, additional objective functions, and back-end classification.Adjustment of model architecture using a pre-training scheme can extract speaker embeddings, giving a significant improvement in performance.Additional objective functions simplify the process of extracting speaker embeddings by merging conventional two-phase processes: extracting utterance-level features such as i-vectors or x-vectors and the feature enhancement phase, e.g., linear discriminant analysis.Effective back-end classification models that suit the proposed speaker embedding are also explored.We propose an end-toend system that comprises two deep neural networks, one frontend for utterance-level speaker embedding extraction and the other for back-end classification.Experiments conducted on the VoxCeleb1 dataset demonstrate that the proposed model achieves state-of-the-art performance among systems without data augmentation.The proposed system is also comparable to the state-of-the-art x-vector system that adopts data augmentation.
Jee-Weon Jung, Hee-Soo Heo, Ju-ho Kim, Hye-Jin Shim, Ha-Jin Yu
INTERSPEECH2
2019 Replay Attack Detection with Complementary High-Resolution Information Using End-to-End DNN for the ASVspoof 2019 Challenge
abstract
In this study, we concentrate on replacing the process of extracting hand-crafted acoustic feature with end-to-end DNN using complementary high-resolution spectrograms.As a result of advance in audio devices, typical characteristics of a replayed speech based on conventional knowledge alter or diminish in unknown replay configurations.Thus, it has become increasingly difficult to detect spoofed speech with a conventional knowledge-based approach.To detect unrevealed characteristics that reside in a replayed speech, we directly input spectrograms into an end-to-end DNN without knowledge-based intervention.Explorations dealt in this study that differentiates from existing spectrogram-based systems are twofold: complementary information and high-resolution.Spectrograms with different information are explored, and it is shown that additional information such as the phase information can be complementary.High-resolution spectrograms are employed with the assumption that the difference between a bona-fide and a replayed speech exists in the details.Additionally, to verify whether other features are complementary to spectrograms, we also examine raw waveform and an i-vector based system.Experiments conducted on the ASVspoof 2019 physical access challenge show promising results, where t-DCF and equal error rates are 0.0570 and 2.45 % for the evaluation set, respectively.
Jee-Weon Jung, Hye-Jin Shim, Hee-Soo Heo, Ha-Jin Yu
INTERSPEECH3
2018 A Complete End-to-End Speaker Verification System Using Deep Neural Networks: From Raw Signals to Verification Result
abstract
End-to-end systems using deep neural networks have been widely studied in the field of speaker verification. Raw audio signal processing has also been widely studied in the fields of automatic music tagging and speech recognition. However, as far as we know, end-to-end systems using raw audio signals have not been explored in speaker verification. In this paper, a complete end-to-end speaker verification system is proposed, which inputs raw audio signals and outputs the verification results. A pre-processing layer and the embedded speaker feature extraction models were mainly investigated. The proposed pre-emphasis layer was combined with a strided convolution layer for pre-processing at the first two hidden layers. In addition, speaker feature extraction models using convolutionallayer and long short-term memory are proposed to be embedded in the proposed end-to-end system.
Jee-Weon Jung, Hee-Soo Heo, Il-Ho Yang, Hye-Jin Shim, Ha-Jin Yu
ICASSP2
2018 Avoiding Speaker Overfitting in End-to-End DNNs Using Raw Waveform for Text-Independent Speaker Verification
Jee-Weon Jung, Hee-Soo Heo, Il-Ho Yang, Hye-Jin Shim, Ha-Jin Yu
INTERSPEECH2
2017 Applying compensation techniques on i-vectors extracted from short-test utterances for speaker verification using deep neural network
abstract
We propose a method to improve speaker verification performance when a test utterance is very short. In some situations with short test utterances, performance of ivector/probabilistic linear discriminant analysis systems degrades. The proposed method transforms short-utterance feature vectors to adequate vectors using a deep neural network, which compensate for short utterances. To reduce the dimensionality of the search space, we extract several principal components from the residual vectors between every long utterance i-vector in a development set and its truncated short utterance i-vector. Then an input i-vector of the network is transformed by linear combination of these directions. In this case, network outputs correspond to weights for linear combination of principal components. We use public speech databases to evaluate the method. The experimental results on short2-10sec condition (det6, male portion) of the NIST 2008 speaker recognition evaluation corpus show that the proposed method reduces the minimum detection cost relative to the baseline system, which uses linear discriminant analysis transformed i-vectors as features.
Il-Ho Yang, Hee-Soo Heo, Sung-Hyun Yoon, Ha-Jin Yu
ICASSP2
2017 Joint Training of Expanded End-to-End DNN for Text-Dependent Speaker Verification
Hee-Soo Heo, Jee-Weon Jung, Il-Ho Yang, Sung-Hyun Yoon, Ha-Jin Yu
INTERSPEECH1
2016 Advanced b-vector system based deep neural network as classifier for speaker verification
abstract
Few studies on speaker verification have directly used a deep neural network (DNN) as a classifier. It is difficult to directly apply a DNN as a discriminative model to speaker-verification tasks because the training data for each speaker are very limited. Therefore, a b-vector has been proposed to solve the problem. However, the DNN with the b-vectors showed lower performance than the conventional i-vector probabilistic linear-discriminant analysis (PLDA) system. In this paper, we propose an improved version of the b-vector DNN system, which incorporates the background speakers' information into the DNN. In this study, each input feature is paired with a representative background speaker's feature vectors, and a b-vector is extracted from each pair; thus, feeding background information into the DNN. We confirmed that the performance improvements of the proposed system compensate for the shortcomings of conventional b-vectors in experiments carried out using the National Institute of Standards and Technology 2008 Speaker-Recognition Evaluation tests.
Hee-Soo Heo, Il-Ho Yang, Myung-Jae Kim, Sung-Hyun Yoon, Ha-Jin Yu
ICASSP1