EDBT 2026 Demo / reviewers in the wild / expert
Seyun Um
dblp:252/4964 · also Se-Yun Um
· DBLP profile ↗
11ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-2229-6741ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Content-Aware Style Augmentation for Zero-Shot Voice Conversion With Short Target SpeechabstractIn this letter, we propose a neural zero-shot voice conversion (ZS-VC) system that simultaneously achieves high speaker similarity and speech intelligibility by incorporating a content-aware style generation module. Although recent neural ZS-VC systems have shown strong performance in either speaker similarity or speech intelligibility, attaining high performance in both remains challenging, especially when only a short target speech sample is available. We attribute this limitation to the insufficient content problem—where the linguistic content of the target speech fails to fully cover that of the source speech. To address this issue, we introduce a method that augments the target speaker's style features for underrepresented content using self-supervised feature generation. Experimental results demonstrate that the proposed system, when integrated with the feature matching-based approach kNN-VC, outperforms existing methods in both key metrics. Demo samples are available at https://hyeonjincha.github.io/. Hyeonjin Cha, Seyun Um, Miseul Kim, Seungshin Lee, Hong-Goo Kang |
IEEE Signal Process. Lett. | 2 |
| 2025 | The Text-to-speech in the Wild (TITW) Database
Jee-Weon Jung, Wangyou Zhang, Soumi Maiti, Yihan Wu 0008, Xin Wang 0037, Yuta Matsunaga, Seyun Um, Jinchuan Tian, Hye-Jin Shim, Nicholas W. D. Evans, Joon Son Chung, Shinnosuke Takamichi, Shinji Watanabe 0001 |
INTERSPEECH | 8 |
| 2025 | SpeechMLC: Speech Multi-label Classification
Miseul Kim, Seyun Um, Hyeonjin Cha, Hong-Goo Kang |
INTERSPEECH | 2 |
| 2024 | PARAN: Variational Autoencoder-based End-to-End Articulation-to-Speech System for Speech Intelligibility
Seyun Um, Hong-Goo Kang |
INTERSPEECH | 1 |
| 2024 | UNIQUE : Unsupervised Network for Integrated Speech Quality Evaluation
Juhwan Yoon, WooSeok Ko, Seyun Um, Sungwoong Hwang, Soojoong Hwang, Hong-Goo Kang |
INTERSPEECH | 3 |
| 2023 | Adversarial Learning of Intermediate Acoustic Feature for End-to-End Lightweight Text-to-SpeechabstractTo simplify the generation process, several text-to-speech (TTS) systems implicitly learn intermediate latent representations instead of relying on predefined features (e.g., mel-spectrogram).However, their generation quality is unsatisfactory as these representations lack speech variances.In this paper, we improve TTS performance by adding prosody embeddings to the latent representations.During training, we extract reference prosody embeddings from mel-spectrograms, and during inference, we estimate these embeddings from text using generative adversarial networks (GANs).Using GANs, we reliably estimate the prosody embeddings in a fast way, which have complex distributions due to the dynamic nature of speech.We also show that the prosody embeddings work as efficient features for learning a robust alignment between text and acoustic features.Our proposed model surpasses several publicly available models with less parameters and computational complexity in comparative experiments. Hyungchan Yoon, Seyun Um, Hong-Goo Kang |
INTERSPEECH | 2 |
| 2023 | SC-CNN: Effective Speaker Conditioning Method for Zero-Shot Multi-Speaker Text-to-Speech SystemsabstractThis letter proposes an effective speaker-conditioning method that is applicable to zero-shot multi-speaker text-to-speech (ZSM-TTS) systems. Based on the inductive bias in the speech generation task, in which local context information in text/phoneme sequences heavily affect the speaker characteristics of the output speech, we propose a Speaker-Conditional Convolutional Neural Network (SC-CNN) for the ZSM-TTS task. SC-CNN first predicts convolutional kernels from each learned speaker embedding, then applies 1-D convolutions to phoneme sequences with the predicted kernels. It utilizes the aforementioned inductive bias and effectively models the characteristic of speech by providing the speaker-specific local context in phonetic domain. We also build both FastSpeech2 and VITS-based ZSM-TTS systems to verify its superiority over conventional speaker conditioning methods. The results confirm that the models with SC-CNN outperform the recent ZSM-TTS models in terms of both subjective and objective measurements. Hyungchan Yoon, Seyun Um, Hyun-Wook Yoon, Hong-Goo Kang |
IEEE Signal Process. Lett. | 3 |
| 2022 | Light-Weight Speaker Verification with Global Context Information
Miseul Kim, Zhenyu Piao, Seyun Um, Ran Lee, Jaemin Joh, Seungshin Lee, Hong-Goo Kang |
INTERSPEECH | 3 |
| 2022 | FluentTTS: Text-dependent Fine-grained Style Control for Multi-style TTS
Seyun Um, Hyungchan Yoon, Hong-Goo Kang |
INTERSPEECH | 2 |
| 2021 | LiteTTS: A Lightweight Mel-Spectrogram-Free Text-to-Wave Synthesizer Based on Generative Adversarial Networks
Huu-Kim Nguyen, Kihyuk Jeong, Seyun Um, Min-Jae Hwang, Eunwoo Song, Hong-Goo Kang |
Interspeech | 3 |
| 2020 | Emotional Speech Synthesis with Rich and Granularized ControlabstractThis paper proposes an effective emotion control method for an end-to-end text-to-speech (TTS) system. To flexibly control the distinct characteristic of a target emotion category, it is essential to determine embedding vectors representing the TTS input. We introduce an inter-to-intra emotional distance ratio algorithm to the embedding vectors that can minimize the distance to the target emotion category while maximizing its distance to the other emotion categories. To further enhance the expressiveness of a target speech, we also introduce an effective interpolation technique that enables the intensity of a target emotion to be gradually changed to that of neutral speech. Subjective evaluation results in terms of emotional expressiveness and controllability show the superiority of the proposed algorithm to the conventional methods. Seyun Um, Sangshin Oh, Kyungguen Byun, Inseon Jang, Chunghyun Ahn, Hong-Goo Kang |
ICASSP | 1 |