Seongkyu Mun

dblp:170/0707 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 AdaptVC: High Quality Voice Conversion with Adaptive Learning
abstract
The goal of voice conversion is to transform the speech of a source speaker to sound like that of a reference speaker while preserving the original content. A key challenge is to extract disentangled linguistic content from the source and voice style from the reference. While existing approaches leverage various methods to isolate the two, a generalization still requires further attention, especially for robustness in zero-shot scenarios. In this paper, we achieve successful disentanglement of content and speaker features by tuning self-supervised speech features with adapters. The adapters are trained to dynamically encode nuanced features from rich self-supervised features, and the decoder fuses them to produce speech that accurately resembles the reference with minimal loss of content. Moreover, we leverage a conditional flow matching decoder with cross-attention speaker conditioning to further boost the synthesis quality and efficiency. Subjective and objective evaluations in a zero-shot scenario demonstrate that the proposed method outperforms existing models in speech quality and similarity to the reference speech.
Jaehun Kim, Yeunju Choi, Tan Dat Nguyen, Seongkyu Mun, Joon Son Chung
ICASSP5
2025 Boundary-Conscious Pruning: Hard Set-Aware Model Compression for Efficient Speaker Recognition
Seongkyu Mun, Jubum Han
INTERSPEECH1
2024 Latent Filling: Latent Space Data Augmentation for Zero-Shot Speech Synthesis
abstract
Previous works in zero-shot text-to-speech (ZS-TTS) have attempted to enhance its systems by enlarging the training data through crowd-sourcing or augmenting existing speech data. However, the use of low-quality data has led to a decline in the overall system performance. To avoid such degradation, instead of directly augmenting the input data, we propose a latent filling (LF) method that adopts simple but effective latent space data augmentation in the speaker embedding space of the ZS-TTS system. By incorporating a consistency loss, LF can be seamlessly integrated into existing ZS-TTS systems without the need for additional training stages. Experimental results show that LF significantly improves speaker similarity while preserving speech quality.
Jae-Sung Bae, Joun Yeop Lee, Ji-Hyun Lee, Seongkyu Mun, Taehwa Kang, Hoonyoung Cho, Chanwoo Kim 0001
ICASSP4
2024 Mels-Tts : Multi-Emotion Multi-Lingual Multi-Speaker Text-To-Speech System Via Disentangled Style Tokens
abstract
This paper proposes a multi-emotion, multi-lingual, and multi-speaker text-to-speech (MELS-TTS) system, employing disentangled style tokens for effective emotion transfer. In speech encompassing various attributes, such as emotional state, speaker identity, and linguistic style, disentangling these elements is crucial for an efficient multi-emotion, multi-lingual, and multi-speaker TTS system. To accomplish this purpose, we propose to utilize separate style tokens to disentangle emotion, language, speaker, and residual information, inspired by the global style tokens (GSTs). Through the attention mechanism, each style token learns its respective speech attribute from the target speech. Our proposed approach yields improved performance in both objective and subjective evaluations, demonstrating the ability to generate cross-lingual speech with diverse emotions, even from a neutral source speaker, while preserving the speaker’s identity.
Heejin Choi, Jae-Sung Bae, Joun Yeop Lee, Seongkyu Mun, Hoonyoung Cho, Chanwoo Kim 0001
ICASSP4
2024 VoxSim: A perceptual voice similarity dataset
Junseok Ahn, Youkyum Kim, Yeunju Choi, Doyeop Kwak, Seongkyu Mun, Joon Son Chung
INTERSPEECH6
2023 Hierarchical Timbre-Cadence Speaker Encoder for Zero-shot Speech Synthesis
Joun Yeop Lee, Jae-Sung Bae, Seongkyu Mun, Ji-Hyun Lee, Hoonyoung Cho, Chanwoo Kim 0001
INTERSPEECH3
2022 Prototypical speaker-interference loss for target voice separation using non-parallel audio samples
Seongkyu Mun, Dhananjaya Gowda, Dokyun Lee, Chanwoo Kim 0001
INTERSPEECH1
2021 Streaming End-to-End Speech Recognition with Jointly Trained Neural Feature Enhancement
abstract
In this paper, we present a streaming end-to-end speech recognition model based on Monotonic Chunkwise Attention (MoCha) jointly trained with enhancement layers. Even though the MoCha attention enables streaming speech recognition with recognition accuracy comparable to a full attention-based approach, training this model is sensitive to various factors such as the difficulty of training examples, hyper-parameters, and so on. Because of these issues, speech recognition accuracy of a MoCha-based model for clean speech drops significantly when a multi-style training approach is applied. Inspired by Curriculum Learning [1], we introduce two training strategies: Gradual Application of Enhanced Features (GAEF) and Gradual Reduction of Enhanced Loss (GREL). With GAEF, the model is initially trained using clean features. Subsequently, the portion of outputs from the enhancement layers gradually increases. With GREL, the portion of the Mean Squared Error (MSE) loss for the enhanced output gradually reduces as training proceeds. In experimental results on the LibriSpeech corpus and noisy far-field test sets, the proposed model with GAEF-GREL training strategies shows significantly better results than the conventional multi-style training approach.
Chanwoo Kim 0001, Abhinav Garg, Dhananjaya Gowda, Seongkyu Mun
ICASSP4
2021 Metric Learning for Keyword Spotting
abstract
The goal of this work is to train effective representations for keyword spotting via metric learning. Most existing works address keyword spotting as a closed-set classification problem, where both target and non-target keywords are predefined. Therefore, prevailing classifier-based keyword spot-ting systems perform poorly on non-target sounds which are unseen during the training stage, causing high false alarm rates in real-world scenarios. In reality, keyword spotting is a detection problem where predefined target keywords are detected from a variety of unknown sounds. This shares many similarities to metric learning problems in that the unseen and unknown non-target sounds must be clearly differentiated from the target keywords. However, a key difference is that the target keywords are known and predefined. To this end, we propose a new method based on metric learning that maximises the distance between target and non-target key-words, but also learns per-class weights for target keywords as in classification objectives. Experiments on the Google Speech Commands dataset show that our method significantly reduces false alarms to unseen non-target keywords, while maintaining the overall classification accuracy.
Jaesung Huh, Hee-Soo Heo, Seongkyu Mun, Joon Son Chung
SLT4
2020 The Sound of My Voice: Speaker Representation Loss for Target Voice Separation
abstract
Content and style representations have been widely studied in the field of style transfer. In this paper, we propose a new loss function using speaker content representation for audio source separation, and we call it speaker representation loss. The objective is to extract the target speaker voice from the noisy input and also remove it from the residual components. Compared to the conventional spectral reconstruction, our proposed framework maximizes the use of target speaker information by minimizing the distance between the speaker representations of reference and source separation output. We also propose triplet speaker representation loss as an additional criterion to remove the target speaker information from residual spectrogram output. VoiceFilter framework is adopted to evaluate source separation performance using the VCTK database, and we achieved improved performances compared to the baseline loss function without any additional network parameters.
Seongkyu Mun, Soyeon Choe, Jaesung Huh, Joon Son Chung
ICASSP1
2020 In Defence of Metric Learning for Speaker Recognition
abstract
The objective of this paper is 'open-set' speaker recognition of unseen speakers, where ideal embeddings should be able to condense information into a compact utterance-level representation that has small intra-speaker and large inter-speaker distance. A popular belief in speaker recognition is that networks trained with classification objectives outperform metric learning methods. In this paper, we present an extensive evaluation of most popular loss functions for speaker recognition on the VoxCeleb dataset. We demonstrate that the vanilla triplet loss shows competitive performance compared to classification-based losses, and those trained with our proposed metric learning objective outperform state-of-the-art methods.
Joon Son Chung, Jaesung Huh, Seongkyu Mun, Hee-Soo Heo, Soyeon Choe, Chiheon Ham, Sunghwan Jung, Bong-Jin Lee, Icksang Han
INTERSPEECH3
2019 Domain Mismatch Robust Acoustic Scene Classification Using Channel Information Conversion
abstract
In recent acoustic scene classification (ASC) research field, training and test device channel mismatch have become an issue for the real world implementation. To address the issue, this paper proposes a channel domain conversion using factorized hierarchical variational autoencoder. Proposed method adapts both the source and target domain to a pre-defined specific domain. Unlike the conventional approach, the relationship between the target and source domain and information of each domain are not required in the adaptation process. Based on the experimental results using the IEEE Detection and Classification of Acoustic Scenes and Event 2018 task 1-B dataset and the baseline system, it is shown that the proposed approach can mitigate the channel mismatching issue of different recording devices.
Seongkyu Mun, Suwon Shon
ICASSP1
2018 Image fusion and influence function for performance improvement of ATM vandalism action recognition
abstract
Rising rate of vandalism against Automatic Teller Machines (ATMs) is a serious issue within banking industries, prompting needs of a technology to autonomously recognize such events. A vision based fusion method proposed here for classifying these incidents is rooted on visually recognizing heavy or sharp objects potentially used for detecting vandalism actions inferred from optical flow. The recognition performance has been improved chiefly by a novel employment of influence functions in selecting data points of each class useful in learning. We show that the tool recognition performance can be improved when the training data is selected from the ImageNet data set as guided by the influence function.
Jeongseop Yun, Junyeop Lee, Seongkyu Mun, Chul Jin Cho, David K. Han, Hanseok Ko
AVSS3
2017 Deep Neural Network based learning and transferring mid-level audio features for acoustic scene classification
abstract
Deep Neural Network (DNN) based transfer learning has been shown to be effective in Visual Object Classification (VOC) for complementing the deficit of target domain training samples by adapting classifiers that have been pre-trained for other large-scaled DataBase (DB). Although there exists an abundance of acoustic data, it can also be said that datasets of specific acoustic scenes are sparse for training Acoustic Scene Classification (ASC) models. By exploiting VOC DNN's ability of learning beyond its pre-trained environments, this paper proposes DNN based transfer learning for ASC. Effectiveness of the proposed method is demonstrated on the database of IEEE DCASE Challenge 2016 Task 1 and home surveillance environment via representative experiments. Its improved performance is verified by comparing it to prominent conventional methods.
Seongkyu Mun, Suwon Shon, Wooil Kim, David K. Han, Hanseok Ko
ICASSP1
2017 Recursive Whitening Transformation for Speaker Recognition on Language Mismatched Condition
abstract
Recently in speaker recognition, performance degradation due to the channel domain mismatched condition has been actively addressed. However, the mismatches arising from language is yet to be sufficiently addressed. This paper proposes an approach which employs recursive whitening transformation to mitigate the language mismatched condition. The proposed method is based on the multiple whitening transformation, which is intended to remove un-whitened residual components in the dataset associated with i-vector length normalization. The experiments were conducted on the Speaker Recognition Evaluation 2016 trials of which the task is non-English speaker recognition using development dataset consist of both a large scale out-of-domain (English) dataset and an extremely low-quantity in-domain (non-English) dataset. For performance comparison, we develop a state-of- the-art system using deep neural network and bottleneck feature, which is based on a phonetically aware model. From the experimental results, along with other prior studies, effectiveness of the proposed method on language mismatched condition is validated.
Suwon Shon, Seongkyu Mun, Hanseok Ko
INTERSPEECH2
2017 Autoencoder Based Domain Adaptation for Speaker Recognition Under Insufficient Channel Information
abstract
In real-life conditions, mismatch between development and test domain degrades speaker recognition performance. To solve the issue, many researchers explored domain adaptation approaches using matched in-domain dataset. However, adaptation would be not effective if the dataset is insufficient to estimate channel variability of the domain. In this paper, we explore the problem of performance degradation under such a situation of insufficient channel information. In order to exploit limited in-domain dataset effectively, we propose an unsupervised domain adaptation approach using Autoencoder based Domain Adaptation (AEDA). The proposed approach combines an autoencoder with a denoising autoencoder to adapt resource-rich development dataset to test domain. The proposed technique is evaluated on the Domain Adaptation Challenge 13 experimental protocols that is widely used in speaker recognition for domain mismatched condition. The results show significant improvements over baselines and results from other prior studies.
Suwon Shon, Seongkyu Mun, Wooil Kim, Hanseok Ko
INTERSPEECH2
2016 Deep Neural Network Bottleneck Features for Acoustic Event Recognition
Seongkyu Mun, Suwon Shon, Wooil Kim, Hanseok Ko
INTERSPEECH1
2015 Maximum likelihood Linear Dimension Reduction of heteroscedastic feature for robust Speaker Recognition
abstract
This paper analyzes heteroscedasticity in i-vector for robust forensics and surveillance speaker recognition system. Linear Discriminant Analysis (LDA), a widely-used linear dimension reduction technique, assumes that classes are homoscedastic within a same covariance. In this paper it is assumed that general speech utterances contain both homoscedastic and heteroscedastic elements. We show the validity of this assumption by employing several analyses and also demonstrate that dimension reduction using principal components is feasible. To effectively handle the presence of heteroscedastic and homoscedastic elements, we propose a fusion approach of applying both LDA and Heteroscedastic-LDA (HLDA). The experiments are conducted to show its effectiveness and compare to other methods using the telephone database of National Institute of Standards and Technology (NIST) Speaker Recognition Evaluation (SRE) 2010 extended.
Suwon Shon, Seongkyu Mun, David K. Han, Hanseok Ko
AVSS2