VLDB 2026 Research / reviewers in the wild / expert
Young-Sun Joo
dblp:120/1230
· DBLP profile ↗
10ranked-venue papers
1as first author
6since 2021 · last 2024
0000-0002-7428-5868ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SYNTHE-SEES: Face Based Text-to-Speech for Virtual SpeakerabstractRecent virtual voice generation researches have limitations in that they results in low-quality voice and generate inconsistent voice from the same speaker’s different facial images. To handle this, we propose a facial encoder module for the pre-trained multi-speaker TTS system called SYNTHE-SEES, which utilizes face embeddings as speaker embeddings by sharing the embedding space of the pre-trained speech embeddings using cross-modal contrastive learning. We trained the facial encoder in two ways: 1) for consistent embeddings, we use the dataset supervision to capture discriminative speaker attributes; 2) we leverage internal structure of the speech embedding to generate diverse and high-quality voices. Experimental results demonstrate that our method generates more distinct, consistent, and high-quality speaker embeddings than other state-of-the-art methods in both quantitative and qualitative evaluations. Especially, the result of cluster-level evaluation verifies that our method shows the highest distinction performance of diverse speaker embedding. Our demo is available at ${\color{Cyan}{\text{Demo}}}$. Jaehyun Park 0012, Joon-Gyu Maeng, Taejun Bak, Young-Sun Joo |
ICASSP | 4 |
| 2023 | Avocodo: Generative Adversarial Network for Artifact-Free VocoderabstractNeural vocoders based on the generative adversarial neural network (GAN) have been widely used due to their fast inference speed and lightweight networks while generating high-quality speech waveforms. Since the perceptually important speech components are primarily concentrated in the low-frequency bands, most GAN-based vocoders perform multi-scale analysis that evaluates downsampled speech waveforms. This multi-scale analysis helps the generator improve speech intelligibility. However, in preliminary experiments, we discovered that the multi-scale analysis which focuses on the low-frequency bands causes unintended artifacts, e.g., aliasing and imaging artifacts, which degrade the synthesized speech waveform quality. Therefore, in this paper, we investigate the relationship between these artifacts and GAN-based vocoders and propose a GAN-based vocoder, called Avocodo, that allows the synthesis of high-fidelity speech with reduced artifacts. We introduce two kinds of discriminators to evaluate speech waveforms in various perspectives: a collaborative multi-band discriminator and a sub-band discriminator. We also utilize a pseudo quadrature mirror filter bank to obtain downsampled multi-band speech waveforms while avoiding aliasing. According to experimental results, Avocodo outperforms baseline GAN-based vocoders, both objectively and subjectively, while reproducing speech with fewer artifacts. Taejun Bak, Hanbin Bae, Jinhyeok Yang, Jae-Sung Bae, Young-Sun Joo |
AAAI | 6 |
| 2022 | Enhancement of Pitch Controllability using Timbre-Preserving Pitch Augmentation in FastPitchabstractThe recently developed pitch-controllable text-to-speech (TTS) model, i.e.FastPitch, was conditioned for the pitch contours.However, the quality of the synthesized speech degraded considerably for pitch values that deviated significantly from the average pitch; i.e. the ability to control pitch was limited.To address this issue, we propose two algorithms to improve the robustness of FastPitch.First, we propose a novel timbrepreserving pitch-shifting algorithm for natural pitch augmentation.Pitch-shifted speech samples sound more natural when using the proposed algorithm because the speaker's vocal timbre is maintained.Moreover, we propose a training algorithm that defines FastPitch using pitch-augmented speech datasets with different pitch ranges for the same sentence.The experimental results demonstrate that the proposed algorithms improve the pitch controllability of FastPitch. Hanbin Bae, Young-Sun Joo |
INTERSPEECH | 2 |
| 2022 | Hierarchical and Multi-Scale Variational Autoencoder for Diverse and Natural Non-Autoregressive Text-to-SpeechabstractThis paper proposes a hierarchical and multi-scale variational autoencoder-based non-autoregressive text-to-speech model (HiMuV-TTS) to generate natural speech with diverse speaking styles. Recent advances in non-autoregressive TTS (NAR-TTS) models have significantly improved the inference speed and robustness of synthesized speech. However, the diversity of speaking styles and naturalness are needed to be improved. To solve this problem, we propose the HiMuV-TTS model that first determines the global-scale prosody and then determines the local-scale prosody via conditioning on the global-scale prosody and the learned text representation. In addition, we improve the quality of speech by adopting the adversarial training technique. Experimental results verify that the proposed HiMuV-TTS model can generate more diverse and natural speech as compared to TTS models with single-scale variational autoencoders, and can represent different prosody information in each scale. Jae-Sung Bae, Jinhyeok Yang, Taejun Bak, Young-Sun Joo |
INTERSPEECH | 4 |
| 2021 | A Neural Text-to-Speech Model Utilizing Broadcast Data Mixed with Background MusicabstractRecently, it has become easier to obtain speech data from various media such as the internet or YouTube, but directly utilizing them to train a neural text-to-speech (TTS) model is difficult. The proportion of clean speech is insufficient and the remainder includes background music. Even with the global style token (GST). Therefore, we propose the following method to successfully train an end-to-end TTS model with limited broadcast data. First, the background music is removed from the speech by introducing a music filter. Second, the GST-TTS model with an auxiliary quality classifier is trained with the filtered speech and a small amount of clean speech. In particular, the quality classifier makes the embedding vector of the GST layer focus on representing the speech quality (filtered or clean) of the input speech. The experimental results verified that the proposed method synthesized much more high-quality speech than conventional methods. Hanbin Bae, Jae-Sung Bae, Young-Sun Joo, Young-Ik Kim, Hoonyoung Cho |
ICASSP | 3 |
| 2021 | Hierarchical Context-Aware Transformers for Non-Autoregressive Text to SpeechabstractIn this paper, we propose methods for improving the modeling performance of a Transformer-based non-autoregressive textto-speech (TNA-TTS) model.Although the text encoder and audio decoder handle different types and lengths of data (i.e., text and audio), the TNA-TTS models are not designed considering these variations.Therefore, to improve the modeling performance of the TNA-TTS model we propose a hierarchical Transformer structure-based text encoder and audio decoder that are designed to accommodate the characteristics of each module.For the text encoder, we constrain each self-attention layer so the encoder focuses on a text sequence from the local to the global scope.Conversely, the audio decoder constrains its self-attention layers to focus in the reverse direction, i.e., from global to local scope.Additionally, we further improve the pitch modeling accuracy of the audio decoder by providing sentence and word-level pitch as conditions.Various objective and subjective evaluations verified that the proposed method outperformed the baseline TNA-TTS. Jae-Sung Bae, Taejun Bak, Young-Sun Joo, Hoonyoung Cho |
Interspeech | 3 |
| 2020 | Speaking Speed Control of End-to-End Speech Synthesis Using Sentence-Level ConditioningabstractThis paper proposes a controllable end-to-end text-to-speech (TTS) system to control the speaking speed (speed-controllable TTS; SCTTS) of synthesized speech with sentence-level speaking-rate value as an additional input. The speaking-rate value, the ratio of the number of input phonemes to the length of input speech, is adopted in the proposed system to control the speaking speed. Furthermore, the proposed SCTTS system can control the speaking speed while retaining other speech attributes, such as the pitch, by adopting the global style token-based style encoder. The proposed SCTTS does not require any additional well-trained model or an external speech database to extract phoneme-level duration information and can be trained in an end-to-end manner. In addition, our listening tests on fast-, normal-, and slow-speed speech showed that the SCTTS can generate more natural speech than other phoneme duration control approaches which increase or decrease duration at the same rate for the entire sentence, especially in the case of slow-speed speech. Jae-Sung Bae, Hanbin Bae, Young-Sun Joo, Gyeong-Hoon Lee, Hoonyoung Cho |
INTERSPEECH | 3 |
| 2015 | Improved time-frequency trajectory excitation modeling for a statistical parametric speech synthesis systemabstractThis paper proposes an improved time-frequency trajectory excitation (TFTE) modeling method for a statistical parametric speech synthesis system. The proposed approach overcomes the dimensional variation problem of the training process caused by the inherent nature of the pitch-dependent analysis paradigm. By reducing the redundancies of the parameters using predicted average block coefficients (PABC), the proposed algorithm efficiently models excitation, even if its dimension is varied. Objective and subjective test results verify that the proposed algorithm provides not only robustness to the training process but also naturalness to the synthesized speech. Eunwoo Song, Young-Sun Joo, Hong-Goo Kang |
ICASSP | 2 |
| 2013 | Enhancement of spectral clarity for HMM-based text-to-speech systemsabstractThis paper proposes a method to enhance the spectral clarity of hidden Markov model (HMM)-based text-to-speech (TTS) systems. A simple way of enhancing spectral clarity is increasing the order of spectral parameters in the speech analysis/synthesis stage, but the method has an inherent statistical modeling problem. The proposed algorithm takes a low-to-high-order spectral parameter mapping approach that adopts low-order parameters for HMM training but does high-order parameters for the actual synthesis step. Various ways of mapping criterion to find appropriate high-order parameters are investigated to further enhance the quality of synthesized speech. Performance evaluation results verify the superiority of the proposed method compared to the conventional one. Young-Sun Joo, Chi-Sang Jung, Hong-Goo Kang |
ICASSP | 1 |
| 2012 | Waveform Interpolation-Based Speech Analysis/Synthesis for HMM-Based TTS SystemsabstractThis letter proposes an HMM-based Text-to-Speech (TTS) system using waveform interpolation (WI)-based speech analysis and synthesis. The synthesized speech quality of the proposed system is significantly improved due to adopting an enhanced excitation modeling technique. The decomposition of characteristic waveform (CW) into slowly evolving waveform (SEW) and rapidly evolving waveform (REW) is efficient not only for excitation modeling but also for training process of HMMs. Objective and subjective test results verify the superiority of the proposed approach to conventional ones. Chi-Sang Jung, Young-Sun Joo, Hong-Goo Kang |
IEEE Signal Process. Lett. | 2 |