EDBT 2026 Demo / reviewers in the wild / expert
Hoonyoung Cho
dblp:359/7732 · also Hoon-Young Cho
· DBLP profile ↗
27ranked-venue papers
3as first author
19since 2021 · last 2025
0000-0002-6850-6580ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 2 first-author · 18 since 2021Artificial intelligence and machine learning · 18 · 2 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Single-Channel Distance-Based Source Separation for Mobile GPU in Outdoor and Indoor EnvironmentsabstractThis study emphasizes the significance of exploring distance-based source separation (DSS) in outdoor environments. Unlike existing studies that primarily focus on indoor settings, the proposed model is designed to capture the unique characteristics of outdoor audio sources. It incorporates advanced techniques, including a two-stage conformer block, a linear relation-aware self-attention (RSA), and a TensorFlow Lite GPU delegate. While the linear RSA may not capture physical cues as explicitly as the quadratic RSA, the linear RSA enhances the model’s context awareness, leading to improved performance on the DSS that requires an understanding of physical cues in outdoor and indoor environments. The experimental results demonstrated that the proposed model overcomes the limitations of existing approaches and considerably enhances energy efficiency and real-time inference speed on mobile devices. Hanbin Bae, Byungjun Kang, Jaeyong Hwang, Hosang Sung, Hoonyoung Cho |
ICASSP | 6 |
| 2025 | Text-Aware Adapter for Few-Shot Keyword SpottingabstractRecent advances in flexible keyword spotting (KWS) with text enrollment allow users to personalize keywords without uttering them during enrollment. However, there is still room for improvement in target keyword performance. In this work, we propose a novel few-shot transfer learning method, called text-aware adapter (TA-adapter), designed to enhance a pre-trained flexible KWS model for specific keywords with limited speech samples. To adapt the acoustic encoder, we leverage a jointly pre-trained text encoder to generate a text embedding that acts as a representative vector for the keyword. By fine-tuning only a small portion of the network while keeping the core components’ weights intact, the TA-adapter proves highly efficient for few-shot KWS, enabling a seamless return to the original pre-trained model. In our experiments, the TA-adapter demonstrated significant performance improvements across 35 distinct keywords from the Google Speech Commands V2 dataset, with only a 0.14% increase in the total number of parameters. Youngmoon Jung, Myunghun Jung, Yong-Hyeok Lee, Hoonyoung Cho |
ICASSP | 6 |
| 2025 | Adversarial Deep Metric Learning for Cross-Modal Audio-Text Alignment in Open-Vocabulary Keyword Spotting
Youngmoon Jung, Yong-Hyeok Lee, Myunghun Jung, Jaeyoung Roh, Chang Woo Han, Hoonyoung Cho |
INTERSPEECH | 6 |
| 2025 | Quadruple Path Modeling with Latent Feature Transfer for Permutation-free Continuous Speech Separation
Hyewon Han, Jonguk Yoo, Chang Woo Han, Jeongook Song, Hoonyoung Cho, Hong-Goo Kang |
INTERSPEECH | 8 |
| 2025 | Efficient Streaming TTS Acoustic Model with Depthwise RVQ Decoding Strategies in a Mamba Framework
Joun Yeop Lee, Byoung Jin Choi, Ji-Hyun Lee, Min-Kyung Kim 0005, Hoonyoung Cho |
INTERSPEECH | 6 |
| 2024 | Latent Filling: Latent Space Data Augmentation for Zero-Shot Speech SynthesisabstractPrevious works in zero-shot text-to-speech (ZS-TTS) have attempted to enhance its systems by enlarging the training data through crowd-sourcing or augmenting existing speech data. However, the use of low-quality data has led to a decline in the overall system performance. To avoid such degradation, instead of directly augmenting the input data, we propose a latent filling (LF) method that adopts simple but effective latent space data augmentation in the speaker embedding space of the ZS-TTS system. By incorporating a consistency loss, LF can be seamlessly integrated into existing ZS-TTS systems without the need for additional training stages. Experimental results show that LF significantly improves speaker similarity while preserving speech quality. Jae-Sung Bae, Joun Yeop Lee, Ji-Hyun Lee, Seongkyu Mun, Taehwa Kang, Hoonyoung Cho, Chanwoo Kim 0001 |
ICASSP | 6 |
| 2024 | Mels-Tts : Multi-Emotion Multi-Lingual Multi-Speaker Text-To-Speech System Via Disentangled Style TokensabstractThis paper proposes a multi-emotion, multi-lingual, and multi-speaker text-to-speech (MELS-TTS) system, employing disentangled style tokens for effective emotion transfer. In speech encompassing various attributes, such as emotional state, speaker identity, and linguistic style, disentangling these elements is crucial for an efficient multi-emotion, multi-lingual, and multi-speaker TTS system. To accomplish this purpose, we propose to utilize separate style tokens to disentangle emotion, language, speaker, and residual information, inspired by the global style tokens (GSTs). Through the attention mechanism, each style token learns its respective speech attribute from the target speech. Our proposed approach yields improved performance in both objective and subjective evaluations, demonstrating the ability to generate cross-lingual speech with diverse emotions, even from a neutral source speaker, while preserving the speaker’s identity. Heejin Choi, Jae-Sung Bae, Joun Yeop Lee, Seongkyu Mun, Hoonyoung Cho, Chanwoo Kim 0001 |
ICASSP | 6 |
| 2024 | Speech Boosting: Low-Latency Live Speech Enhancement for TWS EarbudsabstractThis paper introduces a speech enhancement solution tailored for true wireless stereo (TWS) earbuds on-device usage.The solution was specifically designed to support conversations in noisy environments, with active noise cancellation (ANC) activated.The primary challenges for speech enhancement models in this context arise from computational complexity that limits on-device usage and latency that must be less than 3 ms to preserve a live conversation.To address these issues, we evaluated several crucial design elements, including the network architecture and domain, design of loss functions, pruning method, and hardware-specific optimization.Consequently, we demonstrated substantial improvements in speech enhancement quality compared with that in baseline models, while simultaneously reducing the computational complexity and algorithmic latency. Hanbin Bae, Pavel Andreev, Azat Saginbaev, Nicholas Babaev, Won-Jun Lee, Hosang Sung, Hoonyoung Cho |
INTERSPEECH | 7 |
| 2024 | CTC-aligned Audio-Text Embedding for Streaming Open-vocabulary Keyword Spotting
Sichen Jin, Youngmoon Jung, Jaeyoung Roh, Hoonyoung Cho |
INTERSPEECH | 6 |
| 2024 | Relational Proxy Loss for Audio-Text based Keyword Spotting
Youngmoon Jung, Joon-Young Yang, Jaeyoung Roh, Chang Woo Han, Hoonyoung Cho |
INTERSPEECH | 6 |
| 2024 | High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
Joun Yeop Lee, Myeonghun Jeong, Ji-Hyun Lee, Hoonyoung Cho, Nam Soo Kim |
INTERSPEECH | 5 |
| 2024 | FINALLY: fast and universal speech enhancement with studio-like qualityabstractIn this paper, we address the challenge of speech enhancement in real-world recordings, which often contain various forms of distortion, such as background noise, reverberation, and microphone artifacts.
We revisit the use of Generative Adversarial Networks (GANs) for speech enhancement and theoretically show that GANs are naturally inclined to seek the point of maximum density within the conditional clean speech distribution, which, as we argue, is essential for speech enhancement task.
We study various feature extractors for perceptual loss to facilitate the stability of adversarial training, developing a methodology for probing the structure of the feature space.
This leads us to integrate WavLM-based perceptual loss into MS-STFT adversarial training pipeline, creating an effective and stable training procedure for the speech enhancement model.
The resulting speech enhancement model, which we refer to as FINALLY, builds upon the HiFi++ architecture, augmented with a WavLM encoder and a novel training pipeline.
Empirical results on various datasets confirm our model's ability to produce clear, high-quality speech at 48 kHz, achieving state-of-the-art performance in the field of speech enhancement. Demo page: https://samsunglabs.github.io/FINALLY-page/ Nicholas Babaev, Kirill Tamogashev, Azat Saginbaev, Ivan Shchekotov, Hanbin Bae, Hosang Sung, Won-Jun Lee, Hoonyoung Cho, Pavel Andreev |
NeurIPS | 8 |
| 2023 | Randmasking Augment: A Simple and Randomized Data Augmentation For Acoustic Scene ClassificationabstractIn this work, we describe RandMasking Augment as an effective data augmentation method for acoustic scene classification research. We concentrate on both time and frequency domains masking augmentation introduced in SpecAugment, and apply various transformations that can maintain time and frequency information of the original spectrogram to the masking region. Because acoustic feature is transformed into various forms without distortion of frequency and time information, the proposed augmentation can capture unique characteristics of the input audio in detail. Moreover, RandMasking Augment can be extended by mixing other audio samples and applying different weights on frequency bands in the randomized masking region. We evaluate the suggested augmentation on the DCASE 2018 Task1A dataset and the DCASE 2019 Task1A dataset, and it is compared with other augmentation methods. The proposed augmentation shows outstanding performances with various popular convolutional neural networks. Jubum Han, Mateusz Matuszewski, Olaf Sikorski, Hosang Sung, Hoonyoung Cho |
ICASSP | 5 |
| 2023 | Hierarchical Timbre-Cadence Speaker Encoder for Zero-shot Speech Synthesis
Joun Yeop Lee, Jae-Sung Bae, Seongkyu Mun, Ji-Hyun Lee, Hoonyoung Cho, Chanwoo Kim 0001 |
INTERSPEECH | 6 |
| 2021 | A Neural Text-to-Speech Model Utilizing Broadcast Data Mixed with Background MusicabstractRecently, it has become easier to obtain speech data from various media such as the internet or YouTube, but directly utilizing them to train a neural text-to-speech (TTS) model is difficult. The proportion of clean speech is insufficient and the remainder includes background music. Even with the global style token (GST). Therefore, we propose the following method to successfully train an end-to-end TTS model with limited broadcast data. First, the background music is removed from the speech by introducing a music filter. Second, the GST-TTS model with an auxiliary quality classifier is trained with the filtered speech and a small amount of clean speech. In particular, the quality classifier makes the embedding vector of the GST layer focus on representing the speech quality (filtered or clean) of the input speech. The experimental results verified that the proposed method synthesized much more high-quality speech than conventional methods. Hanbin Bae, Jae-Sung Bae, Young-Sun Joo, Young-Ik Kim, Hoonyoung Cho |
ICASSP | 5 |
| 2021 | Hierarchical Context-Aware Transformers for Non-Autoregressive Text to SpeechabstractIn this paper, we propose methods for improving the modeling performance of a Transformer-based non-autoregressive textto-speech (TNA-TTS) model.Although the text encoder and audio decoder handle different types and lengths of data (i.e., text and audio), the TNA-TTS models are not designed considering these variations.Therefore, to improve the modeling performance of the TNA-TTS model we propose a hierarchical Transformer structure-based text encoder and audio decoder that are designed to accommodate the characteristics of each module.For the text encoder, we constrain each self-attention layer so the encoder focuses on a text sequence from the local to the global scope.Conversely, the audio decoder constrains its self-attention layers to focus in the reverse direction, i.e., from global to local scope.Additionally, we further improve the pitch modeling accuracy of the audio decoder by providing sentence and word-level pitch as conditions.Various objective and subjective evaluations verified that the proposed method outperformed the baseline TNA-TTS. Jae-Sung Bae, Taejun Bak, Young-Sun Joo, Hoonyoung Cho |
Interspeech | 4 |
| 2021 | FastPitchFormant: Source-Filter Based Decomposed Modeling for Speech SynthesisabstractMethods for modeling and controlling prosody with acoustic features have been proposed for neural text-to-speech (TTS) models. Prosodic speech can be generated by conditioning acoustic features. However, synthesized speech with a large pitch-shift scale suffers from audio quality degradation, and speaker characteristics deformation. To address this problem, we propose a feed-forward Transformer based TTS model that is designed based on the source-filter theory. This model, called FastPitchFormant, has a unique structure that handles text and acoustic features in parallel. With modeling each feature separately, the tendency that the model learns the relationship between two features can be mitigated. Taejun Bak, Jae-Sung Bae, Hanbin Bae, Young-Ik Kim, Hoonyoung Cho |
Interspeech | 5 |
| 2021 | N-Singer: A Non-Autoregressive Korean Singing Voice Synthesis System for Pronunciation EnhancementabstractRecently, end-to-end Korean singing voice systems have been designed to generate realistic singing voices.However, these systems still suffer from a lack of robustness in terms of pronunciation accuracy.In this paper, we propose N-Singer, a nonautoregressive Korean singing voice system, to synthesize accurate and pronounced Korean singing voices in parallel.N-Singer consists of a Transformer-based mel-generator, a convolutional network-based postnet, and voicing-aware discriminators.It can contribute in the following ways.First, for accurate pronunciation, N-Singer separately models linguistic and pitch information without other acoustic features.Second, to achieve improved mel-spectrograms, N-Singer uses a combination of Transformer-based modules and convolutional networkbased modules.Third, in adversarial training, voicing-aware conditional discriminators are used to capture the harmonic features of voiced segments and noise components of unvoiced segments.The experimental results prove that N-Singer can synthesize a natural singing voice in parallel with a more accurate pronunciation than the baseline model. Gyeong-Hoon Lee, Hanbin Bae, Min-Ji Lee, Young-Ik Kim, Hoonyoung Cho |
Interspeech | 6 |
| 2021 | GANSpeech: Adversarial Training for High-Fidelity Multi-Speaker Speech SynthesisabstractRecent advances in neural multi-speaker text-to-speech (TTS) models have enabled the generation of reasonably good speech quality with a single model and made it possible to synthesize the speech of a speaker with limited training data.Finetuning to the target speaker data with the multi-speaker model can achieve better quality, however, there still exists a gap compared to the real speech sample and the model depends on the speaker.In this work, we propose GANSpeech, which is a high-fidelity multi-speaker TTS model that adopts the adversarial training method to a non-autoregressive multi-speaker TTS model.In addition, we propose simple but efficient automatic scaling methods for feature matching loss used in adversarial training.In the subjective listening tests, GANSpeech significantly outperformed the baseline multi-speaker FastSpeech and FastSpeech2 models, and showed a better MOS score than the speaker-specific fine-tuned FastSpeech2. Jinhyeok Yang, Jae-Sung Bae, Taejun Bak, Young-Ik Kim, Hoonyoung Cho |
Interspeech | 5 |
| 2020 | Detecting Mismatch Between Text Script and Voice-Over Using Utterance Verification Based on Phoneme Recognition RankingabstractThe purpose of this study is to detect the mismatch between text script and voice-over. For this, we present a novel utterance verification (UV) method, which calculates the degree of correspondence between a voice-over and the phoneme sequence of a script. We found that the phoneme recognition probabilities of exaggerated voice-overs decrease compared to ordinary utterances, but their rankings do not demonstrate any significant change. The proposed method, therefore, uses the recognition ranking of each phoneme segment corresponding to a phoneme sequence for measuring the confidence of a voice-over utterance for its corresponding script. The experimental results show that the proposed UV method outperforms a state-of-the-art approach using cross modal attention used for detecting mismatch between speech and transcription. Yoonjae Jeong, Hoonyoung Cho |
ICASSP | 2 |
| 2020 | Speaking Speed Control of End-to-End Speech Synthesis Using Sentence-Level ConditioningabstractThis paper proposes a controllable end-to-end text-to-speech (TTS) system to control the speaking speed (speed-controllable TTS; SCTTS) of synthesized speech with sentence-level speaking-rate value as an additional input. The speaking-rate value, the ratio of the number of input phonemes to the length of input speech, is adopted in the proposed system to control the speaking speed. Furthermore, the proposed SCTTS system can control the speaking speed while retaining other speech attributes, such as the pitch, by adopting the global style token-based style encoder. The proposed SCTTS does not require any additional well-trained model or an external speech database to extract phoneme-level duration information and can be trained in an end-to-end manner. In addition, our listening tests on fast-, normal-, and slow-speed speech showed that the SCTTS can generate more natural speech than other phoneme duration control approaches which increase or decrease duration at the same rate for the entire sentence, especially in the case of slow-speed speech. Jae-Sung Bae, Hanbin Bae, Young-Sun Joo, Gyeong-Hoon Lee, Hoonyoung Cho |
INTERSPEECH | 6 |
| 2020 | VocGAN: A High-Fidelity Real-Time Vocoder with a Hierarchically-Nested Adversarial NetworkabstractWe present a novel high-fidelity real-time neural vocoder called VocGAN.A recently developed GAN-based vocoder, MelGAN, produces speech waveforms in real-time.However, it often produces a waveform that is insufficient in quality or inconsistent with acoustic characteristics of the input mel spectrogram.VocGAN is nearly as fast as MelGAN, but it significantly improves the quality and consistency of the output waveform.VocGAN applies a multi-scale waveform generator and a hierarchically-nested discriminator to learn multiple levels of acoustic properties in a balanced way.It also applies the joint conditional and unconditional objective, which has shown successful results in high-resolution image synthesis.In experiments, VocGAN synthesizes speech waveforms 416.7x faster on a GTX 1080Ti GPU and 3.24x faster on a CPU than realtime.Compared with MelGAN, it also exhibits significantly improved quality in multiple evaluation metrics including mean opinion score (MOS) with minimal additional overhead.Additionally, compared with Parallel WaveGAN, another recently developed high-fidelity vocoder, VocGAN is 6.98x faster on a CPU and exhibits higher MOS. Jinhyeok Yang, Young-Ik Kim, Hoonyoung Cho, Injung Kim 0001 |
INTERSPEECH | 4 |
| 2011 | Zero-Crossing-Based Channel Attentive Weighting of Cepstral Features for Robust Speech Recognition: The ETRI 2011 CHiME Challenge System
Young-Ik Kim, Hoonyoung Cho, Sang-Hun Kim |
INTERSPEECH | 2 |
| 2007 | Data-Driven Subvector Clustering using the Cross-Entropy MethodabstractAutomatic speech recognition (ASR) systems are limited in the computational power and memory resources, especially in low-memory/low-power environments such as personal digital assistants. The parameter quantization is the one of the ways to achieve these conditions. In this work, we compare various subvector clustering procedures for the parameter quantization in the ASR system and propose a data-driven subvector clustering technique based on the entropy minimization. The cross-entropy(CE) method is a good choice for the combinatorial optimization problems. We compare the ASR performance on resource management (RM) speech recognition task and show that the proposed technique produces better performance than previous heuristic techniques. Cue Jun Jung, Hoonyoung Cho, Yung-Hwan Oh |
ICASSP (4) | 2 |
| 2004 | Emotion verification for emotion detection and unknown emotion rejectionabstractThis paper focuses on detection of a single emotion and verification of a specific emotion type in a test utterance. To utilize a probabilistic output of a classifier as well as to exploit various long term acoustic features, we built a probabilistic output SVM and applied several approximated log likelihood ratio tests for emotion verification. Experimental results on SUSAS and AIBO emotion database show that anger and sadness are easier emotions to be detected than boredom and happiness. Results also verify the efficacy of applying log likelihood ratio with respect to neutral emotion as a measure for emotion verification. Hoonyoung Cho, Kaisheng Yao, Te-Won Lee |
INTERSPEECH | 1 |
| 2004 | On the use of channel-attentive MFCC for robust recognition of partially corrupted speechabstractThis letter proposes a channel-attentive mel frequency cepstral coefficient (CAMFCC) method to improve the utilization of uncorrupted or more reliable frequency bands for robust speech recognition. This method obtains a channel attention matrix by reliability estimation of mel filter bank channels, and both the input mel frequency cepstral coefficients and the mean vectors of hidden Markov models are corrected using the channel attention matrix at the output probability calculation of the Viterbi decoding. Experimental results on the TIDIGITS database corrupted by various band-selective noises indicated that the proposed CAMFCC method utilizes the uncorrupted partial frequency bands better than a multiband method, resolving the limitation of noise localization caused by the fixed boundaries of the multiband approach. Hoonyoung Cho, Yung-Hwan Oh |
IEEE Signal Process. Lett. | 1 |
| 1998 | A Robust Front-End for Telephone Speech Recognition
Hoonyoung Cho, Sang-Mun Chi, Yung-Hwan Oh |
PRICAI | 1 |