Hyungchan Yoon

dblp:317/6951 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-6331-2437ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 APTTS: Adversarial Post-training in Latent Flow Matching for Fast and High-fidelity Text-to-Speech
Hyungchan Yoon, Chanwoo Lee, Hoodong Lee, Stanley Jungkyu Choi
INTERSPEECH1
2023 HappyQuokka System for ICASSP 2023 Auditory EEG Challenge
abstract
This report describes our submission to Task 2 of the Auditory EEG Decoding Challenge at ICASSP 2023 Signal Processing Grand Challenge (SPGC). Task 2 is a regression problem that focuses on reconstructing a speech envelope from an EEG signal. For the task, we propose a pre-layer normalized feedforward transformer (FFT) architecture. For within-subjects generation, we additionally utilize an auxiliary global conditioner which provides our model with additional information about seen individuals. Experimental results show that our proposed method outperforms the VLAAI baseline and all other submitted systems. Notably, it demonstrates significant improvements on the within-subjects task, likely thanks to our use of the auxiliary global conditioner. In terms of evaluation metrics set by the challenge, we obtain Pearson correlation values of 0.1895 ± 0.0869 for the within-subjects generation test and 0.0976 ± 0.0444 for the heldout-subjects test. We release the training code for our model online.1
Zhenyu Piao, Miseul Kim, Hyungchan Yoon, Hong-Goo Kang
ICASSP3
2023 Pruning Self-Attention for Zero-Shot Multi-Speaker Text-to-Speech
abstract
For personalized speech generation, a neural text-to-speech (TTS) model must be successfully implemented with limited data from a target speaker. To this end, the baseline TTS model needs to be amply generalized to out-of-domain data (i.e., target speaker's speech). However, approaches to address this out-of-domain generalization problem in TTS have yet to be thoroughly studied. In this work, we propose an effective pruning method for a transformer known as sparse attention, to improve the TTS model's generalization abilities. In particular, we prune off redundant connections from self-attention layers whose attention weights are below the threshold. To flexibly determine the pruning strength for searching optimal degree of generalization, we also propose a new differentiable pruning method that allows the model to automatically learn the thresholds. Evaluations on zero-shot multi-speaker TTS verify the effectiveness of our method in terms of voice quality and speaker similarity.
Hyungchan Yoon, Eunwoo Song, Hyun-Wook Yoon, Hong-Goo Kang
INTERSPEECH1
2023 Adversarial Learning of Intermediate Acoustic Feature for End-to-End Lightweight Text-to-Speech
abstract
To simplify the generation process, several text-to-speech (TTS) systems implicitly learn intermediate latent representations instead of relying on predefined features (e.g., mel-spectrogram).However, their generation quality is unsatisfactory as these representations lack speech variances.In this paper, we improve TTS performance by adding prosody embeddings to the latent representations.During training, we extract reference prosody embeddings from mel-spectrograms, and during inference, we estimate these embeddings from text using generative adversarial networks (GANs).Using GANs, we reliably estimate the prosody embeddings in a fast way, which have complex distributions due to the dynamic nature of speech.We also show that the prosody embeddings work as efficient features for learning a robust alignment between text and acoustic features.Our proposed model surpasses several publicly available models with less parameters and computational complexity in comparative experiments.
Hyungchan Yoon, Seyun Um, Hong-Goo Kang
INTERSPEECH1
2023 SC-CNN: Effective Speaker Conditioning Method for Zero-Shot Multi-Speaker Text-to-Speech Systems
abstract
This letter proposes an effective speaker-conditioning method that is applicable to zero-shot multi-speaker text-to-speech (ZSM-TTS) systems. Based on the inductive bias in the speech generation task, in which local context information in text/phoneme sequences heavily affect the speaker characteristics of the output speech, we propose a Speaker-Conditional Convolutional Neural Network (SC-CNN) for the ZSM-TTS task. SC-CNN first predicts convolutional kernels from each learned speaker embedding, then applies 1-D convolutions to phoneme sequences with the predicted kernels. It utilizes the aforementioned inductive bias and effectively models the characteristic of speech by providing the speaker-specific local context in phonetic domain. We also build both FastSpeech2 and VITS-based ZSM-TTS systems to verify its superiority over conventional speaker conditioning methods. The results confirm that the models with SC-CNN outperform the recent ZSM-TTS models in terms of both subjective and objective measurements.
Hyungchan Yoon, Seyun Um, Hyun-Wook Yoon, Hong-Goo Kang
IEEE Signal Process. Lett.1
2022 FluentTTS: Text-dependent Fine-grained Style Control for Multi-style TTS
Seyun Um, Hyungchan Yoon, Hong-Goo Kang
INTERSPEECH3