VLDB 2026 Research / reviewers in the wild / expert
Jinhyeok Yang
dblp:182/5262
· DBLP profile ↗
8ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0002-0182-5114ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
Jinhyeok Yang, Junhyeok Lee 0001, Hyeong-Seok Choi, Seunghoon Ji, Hyeongju Kim, Juheon Lee |
INTERSPEECH | 1 |
| 2023 | Avocodo: Generative Adversarial Network for Artifact-Free VocoderabstractNeural vocoders based on the generative adversarial neural network (GAN) have been widely used due to their fast inference speed and lightweight networks while generating high-quality speech waveforms. Since the perceptually important speech components are primarily concentrated in the low-frequency bands, most GAN-based vocoders perform multi-scale analysis that evaluates downsampled speech waveforms. This multi-scale analysis helps the generator improve speech intelligibility. However, in preliminary experiments, we discovered that the multi-scale analysis which focuses on the low-frequency bands causes unintended artifacts, e.g., aliasing and imaging artifacts, which degrade the synthesized speech waveform quality. Therefore, in this paper, we investigate the relationship between these artifacts and GAN-based vocoders and propose a GAN-based vocoder, called Avocodo, that allows the synthesis of high-fidelity speech with reduced artifacts. We introduce two kinds of discriminators to evaluate speech waveforms in various perspectives: a collaborative multi-band discriminator and a sub-band discriminator. We also utilize a pseudo quadrature mirror filter bank to obtain downsampled multi-band speech waveforms while avoiding aliasing. According to experimental results, Avocodo outperforms baseline GAN-based vocoders, both objectively and subjectively, while reproducing speech with fewer artifacts. Taejun Bak, Hanbin Bae, Jinhyeok Yang, Jae-Sung Bae, Young-Sun Joo |
AAAI | 4 |
| 2023 | NANSY++: Unified Voice Synthesis with Neural Analysis and Synthesis
Hyeong-Seok Choi, Jinhyeok Yang, Juheon Lee, Hyeongju Kim |
ICLR | 2 |
| 2022 | Varianceflow: High-Quality and Controllable Text-to-Speech using Variance Information via Normalizing FlowabstractThere are two types of methods for non-autoregressive text-to-speech models to learn the one-to-many relationship between text and speech effectively. The first one is to use an advanced generative framework such as normalizing flow (NF). The second one is to use variance information such as pitch or energy together when generating speech. For the second type, it is also possible to control the variance factors by adjusting the variance values provided to a model. In this paper, we propose a novel model called VarianceFlow combining the advantages of the two types. By modeling the variance with NF, VarianceFlow predicts the variance information more precisely with improved speech quality. Also, the objective function of NF makes the model use the variance information and the text in a disentangled manner resulting in more precise variance control. In experiments, VarianceFlow shows superior performance over other state-of-the-art TTS models both in terms of speech quality and controllability. Yoonhyung Lee, Jinhyeok Yang, Kyomin Jung |
ICASSP | 2 |
| 2022 | Hierarchical and Multi-Scale Variational Autoencoder for Diverse and Natural Non-Autoregressive Text-to-SpeechabstractThis paper proposes a hierarchical and multi-scale variational autoencoder-based non-autoregressive text-to-speech model (HiMuV-TTS) to generate natural speech with diverse speaking styles. Recent advances in non-autoregressive TTS (NAR-TTS) models have significantly improved the inference speed and robustness of synthesized speech. However, the diversity of speaking styles and naturalness are needed to be improved. To solve this problem, we propose the HiMuV-TTS model that first determines the global-scale prosody and then determines the local-scale prosody via conditioning on the global-scale prosody and the learned text representation. In addition, we improve the quality of speech by adopting the adversarial training technique. Experimental results verify that the proposed HiMuV-TTS model can generate more diverse and natural speech as compared to TTS models with single-scale variational autoencoders, and can represent different prosody information in each scale. Jae-Sung Bae, Jinhyeok Yang, Taejun Bak, Young-Sun Joo |
INTERSPEECH | 2 |
| 2021 | GANSpeech: Adversarial Training for High-Fidelity Multi-Speaker Speech SynthesisabstractRecent advances in neural multi-speaker text-to-speech (TTS) models have enabled the generation of reasonably good speech quality with a single model and made it possible to synthesize the speech of a speaker with limited training data.Finetuning to the target speaker data with the multi-speaker model can achieve better quality, however, there still exists a gap compared to the real speech sample and the model depends on the speaker.In this work, we propose GANSpeech, which is a high-fidelity multi-speaker TTS model that adopts the adversarial training method to a non-autoregressive multi-speaker TTS model.In addition, we propose simple but efficient automatic scaling methods for feature matching loss used in adversarial training.In the subjective listening tests, GANSpeech significantly outperformed the baseline multi-speaker FastSpeech and FastSpeech2 models, and showed a better MOS score than the speaker-specific fine-tuned FastSpeech2. Jinhyeok Yang, Jae-Sung Bae, Taejun Bak, Young-Ik Kim, Hoonyoung Cho |
Interspeech | 1 |
| 2020 | VocGAN: A High-Fidelity Real-Time Vocoder with a Hierarchically-Nested Adversarial NetworkabstractWe present a novel high-fidelity real-time neural vocoder called VocGAN.A recently developed GAN-based vocoder, MelGAN, produces speech waveforms in real-time.However, it often produces a waveform that is insufficient in quality or inconsistent with acoustic characteristics of the input mel spectrogram.VocGAN is nearly as fast as MelGAN, but it significantly improves the quality and consistency of the output waveform.VocGAN applies a multi-scale waveform generator and a hierarchically-nested discriminator to learn multiple levels of acoustic properties in a balanced way.It also applies the joint conditional and unconditional objective, which has shown successful results in high-resolution image synthesis.In experiments, VocGAN synthesizes speech waveforms 416.7x faster on a GTX 1080Ti GPU and 3.24x faster on a CPU than realtime.Compared with MelGAN, it also exhibits significantly improved quality in multiple evaluation metrics including mean opinion score (MOS) with minimal additional overhead.Additionally, compared with Parallel WaveGAN, another recently developed high-fidelity vocoder, VocGAN is 6.98x faster on a CPU and exhibits higher MOS. Jinhyeok Yang, Young-Ik Kim, Hoonyoung Cho, Injung Kim 0001 |
INTERSPEECH | 1 |
| 2019 | HanFont: large-scale adaptive Hangul font recognizer using CNN and font clustering
Jinhyeok Yang, Heebeom Kim, Hyobin Kwak, Injung Kim 0001 |
Int. J. Document Anal. Recognit. | 1 |