Taejun Bak

dblp:296/2072 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2024 SYNTHE-SEES: Face Based Text-to-Speech for Virtual Speaker
abstract
Recent virtual voice generation researches have limitations in that they results in low-quality voice and generate inconsistent voice from the same speaker’s different facial images. To handle this, we propose a facial encoder module for the pre-trained multi-speaker TTS system called SYNTHE-SEES, which utilizes face embeddings as speaker embeddings by sharing the embedding space of the pre-trained speech embeddings using cross-modal contrastive learning. We trained the facial encoder in two ways: 1) for consistent embeddings, we use the dataset supervision to capture discriminative speaker attributes; 2) we leverage internal structure of the speech embedding to generate diverse and high-quality voices. Experimental results demonstrate that our method generates more distinct, consistent, and high-quality speaker embeddings than other state-of-the-art methods in both quantitative and qualitative evaluations. Especially, the result of cluster-level evaluation verifies that our method shows the highest distinction performance of diverse speaker embedding. Our demo is available at ${\color{Cyan}{\text{Demo}}}$.
Jaehyun Park 0012, Joon-Gyu Maeng, Taejun Bak, Young-Sun Joo
ICASSP3
2023 Avocodo: Generative Adversarial Network for Artifact-Free Vocoder
abstract
Neural vocoders based on the generative adversarial neural network (GAN) have been widely used due to their fast inference speed and lightweight networks while generating high-quality speech waveforms. Since the perceptually important speech components are primarily concentrated in the low-frequency bands, most GAN-based vocoders perform multi-scale analysis that evaluates downsampled speech waveforms. This multi-scale analysis helps the generator improve speech intelligibility. However, in preliminary experiments, we discovered that the multi-scale analysis which focuses on the low-frequency bands causes unintended artifacts, e.g., aliasing and imaging artifacts, which degrade the synthesized speech waveform quality. Therefore, in this paper, we investigate the relationship between these artifacts and GAN-based vocoders and propose a GAN-based vocoder, called Avocodo, that allows the synthesis of high-fidelity speech with reduced artifacts. We introduce two kinds of discriminators to evaluate speech waveforms in various perspectives: a collaborative multi-band discriminator and a sub-band discriminator. We also utilize a pseudo quadrature mirror filter bank to obtain downsampled multi-band speech waveforms while avoiding aliasing. According to experimental results, Avocodo outperforms baseline GAN-based vocoders, both objectively and subjectively, while reproducing speech with fewer artifacts.
Taejun Bak, Hanbin Bae, Jinhyeok Yang, Jae-Sung Bae, Young-Sun Joo
AAAI1
2022 Hierarchical and Multi-Scale Variational Autoencoder for Diverse and Natural Non-Autoregressive Text-to-Speech
abstract
This paper proposes a hierarchical and multi-scale variational autoencoder-based non-autoregressive text-to-speech model (HiMuV-TTS) to generate natural speech with diverse speaking styles. Recent advances in non-autoregressive TTS (NAR-TTS) models have significantly improved the inference speed and robustness of synthesized speech. However, the diversity of speaking styles and naturalness are needed to be improved. To solve this problem, we propose the HiMuV-TTS model that first determines the global-scale prosody and then determines the local-scale prosody via conditioning on the global-scale prosody and the learned text representation. In addition, we improve the quality of speech by adopting the adversarial training technique. Experimental results verify that the proposed HiMuV-TTS model can generate more diverse and natural speech as compared to TTS models with single-scale variational autoencoders, and can represent different prosody information in each scale.
Jae-Sung Bae, Jinhyeok Yang, Taejun Bak, Young-Sun Joo
INTERSPEECH3
2021 Hierarchical Context-Aware Transformers for Non-Autoregressive Text to Speech
abstract
In this paper, we propose methods for improving the modeling performance of a Transformer-based non-autoregressive textto-speech (TNA-TTS) model.Although the text encoder and audio decoder handle different types and lengths of data (i.e., text and audio), the TNA-TTS models are not designed considering these variations.Therefore, to improve the modeling performance of the TNA-TTS model we propose a hierarchical Transformer structure-based text encoder and audio decoder that are designed to accommodate the characteristics of each module.For the text encoder, we constrain each self-attention layer so the encoder focuses on a text sequence from the local to the global scope.Conversely, the audio decoder constrains its self-attention layers to focus in the reverse direction, i.e., from global to local scope.Additionally, we further improve the pitch modeling accuracy of the audio decoder by providing sentence and word-level pitch as conditions.Various objective and subjective evaluations verified that the proposed method outperformed the baseline TNA-TTS.
Jae-Sung Bae, Taejun Bak, Young-Sun Joo, Hoonyoung Cho
Interspeech2
2021 FastPitchFormant: Source-Filter Based Decomposed Modeling for Speech Synthesis
abstract
Methods for modeling and controlling prosody with acoustic features have been proposed for neural text-to-speech (TTS) models. Prosodic speech can be generated by conditioning acoustic features. However, synthesized speech with a large pitch-shift scale suffers from audio quality degradation, and speaker characteristics deformation. To address this problem, we propose a feed-forward Transformer based TTS model that is designed based on the source-filter theory. This model, called FastPitchFormant, has a unique structure that handles text and acoustic features in parallel. With modeling each feature separately, the tendency that the model learns the relationship between two features can be mitigated.
Taejun Bak, Jae-Sung Bae, Hanbin Bae, Young-Ik Kim, Hoonyoung Cho
Interspeech1
2021 GANSpeech: Adversarial Training for High-Fidelity Multi-Speaker Speech Synthesis
abstract
Recent advances in neural multi-speaker text-to-speech (TTS) models have enabled the generation of reasonably good speech quality with a single model and made it possible to synthesize the speech of a speaker with limited training data.Finetuning to the target speaker data with the multi-speaker model can achieve better quality, however, there still exists a gap compared to the real speech sample and the model depends on the speaker.In this work, we propose GANSpeech, which is a high-fidelity multi-speaker TTS model that adopts the adversarial training method to a non-autoregressive multi-speaker TTS model.In addition, we propose simple but efficient automatic scaling methods for feature matching loss used in adversarial training.In the subjective listening tests, GANSpeech significantly outperformed the baseline multi-speaker FastSpeech and FastSpeech2 models, and showed a better MOS score than the speaker-specific fine-tuned FastSpeech2.
Jinhyeok Yang, Jae-Sung Bae, Taejun Bak, Young-Ik Kim, Hoonyoung Cho
Interspeech3