Yu-Huai Peng

dblp:214/2303 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
5since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2022 SVSNet: An End-to-End Speaker Voice Similarity Assessment Model
abstract
Neural evaluation metrics derived for numerous speech generation tasks have recently attracted great attention. In this paper, we propose SVSNet, the first end-to-end neural network model to assess the speaker voice similarity between converted speech and natural speech for voice conversion tasks. Unlike most neural evaluation metrics that use hand-crafted features, SVSNet directly takes the raw waveform as input to more completely utilize speech information for prediction. SVSNet consists of encoder, co-attention, distance calculation, and prediction modules and is trained in an end-to-end manner. The experimental results on the Voice Conversion Challenge 2018 and 2020 (VCC2018 and VCC2020) datasets show that SVSNet outperforms well-known baseline systems in the assessment of speaker similarity at the utterance and system levels.
Cheng-Hung Hu, Yu-Huai Peng, Junichi Yamagishi, Yu Tsao 0001, Hsin-Min Wang
IEEE Signal Process. Lett.2
2021 Mandarin Electrolaryngeal Speech Voice Conversion with Sequence-to-Sequence Modeling
abstract
The electrolaryngeal speech (EL speech) is typically spoken with an electrolarynx device that generates excitation signals to substitute human vocal fold vibrations. Because the excitation signals cannot perfectly characterize sound sources generated by vocal folds, the naturalness and intelligibility of the EL speech are inevitably worse than that of the natural speech (NL speech). To improve speech naturalness, statistical models, such as Gaussian mixture models and deep-learning-based models, have been employed for EL speech voice conversion (ELVC). The ELVC task aims to convert EL speech into NL speech through an ELVC model. To implement a frame-wise ELVC system, accurate feature alignment is crucial for model training. However, the abnormal acoustic characteristics of the EL speech cause misalignments and accordingly limit the ELVC performance. To address this issue, we propose a novel ELVC system based on sequence-to-sequence (seq2seq) modeling with text-to-speech (TTS) pretraining. The seq2seq model involves an attention mechanism to concurrently perform representation learning and alignment. Meanwhile, TTS pretraining provides efficient training with limited data. Experimental results show that the proposed ELVC system yields notable improvements in terms of standardized evaluation metrics and subjective listening tests over a well-known frame-wise ELVC system.
Ming-Chi Yen, Wen-Chin Huang, Kazuhiro Kobayashi, Yu-Huai Peng, Shu-Wei Tsai, Yu Tsao 0001, Tomoki Toda, Jyh-Shing Roger Jang, Hsin-Min Wang
ASRU4
2021 A Preliminary Study of a Two-Stage Paradigm for Preserving Speaker Identity in Dysarthric Voice Conversion
abstract
We propose a new paradigm for maintaining speaker identity in dysarthric voice conversion (DVC). The poor quality of dysarthric speech can be greatly improved by statistical VC, but as the normal speech utterances of a dysarthria patient are nearly impossible to collect, previous work failed to recover the individuality of the patient. In light of this, we suggest a novel, two-stage approach for DVC, which is highly flexible in that no normal speech of the patient is required. First, a powerful parallel sequence-to-sequence model converts the input dysarthric speech into a normal speech of a reference speaker as an intermediate product, and a nonparallel, frame-wise VC model realized with a variational autoencoder then converts the speaker identity of the reference speech back to that of the patient while assumed to be capable of preserving the enhanced quality. We investigate several design options. Experimental evaluation results demonstrate the potential of our approach to improving the quality of the dysarthric speech while maintaining the speaker identity.
Wen-Chin Huang, Kazuhiro Kobayashi, Yu-Huai Peng, Ching-Feng Liu, Yu Tsao 0001, Hsin-Min Wang, Tomoki Toda
Interspeech3
2021 Dual-Path Filter Network: Speaker-Aware Modeling for Speech Separation
abstract
Speech separation has been extensively studied to deal with the cocktail party problem in recent years.All related approaches can be divided into two categories: time-frequency domain methods and time domain methods.In addition, some methods try to generate speaker vectors to support source separation.In this study, we propose a new model called dualpath filter network (DPFN).Our model focuses on the postprocessing of speech separation to improve speech separation performance.DPFN is composed of two parts: the speaker module and the separation module.First, the speaker module infers the identities of the speakers.Then, the separation module uses the speakers' information to extract the voices of individual speakers from the mixture.DPFN constructed based on DPRNN-TasNet is not only superior to DPRNN-TasNet, but also avoids the problem of permutation-invariant training (PIT).
Fan-Lin Wang, Yu-Huai Peng, Hung-Shin Lee, Hsin-Min Wang
Interspeech2
2021 Relational Data Selection for Data Augmentation of Speaker-Dependent Multi-Band MelGAN Vocoder
abstract
Nowadays, neural vocoders can generate very high-fidelity speech when a bunch of training data is available.Although a speaker-dependent (SD) vocoder usually outperforms a speaker-independent (SI) vocoder, it is impractical to collect a large amount of data of a specific target speaker for most realworld applications.To tackle the problem of limited target data, a data augmentation method based on speaker representation and similarity measurement of speaker verification is proposed in this paper.The proposed method selects utterances that have similar speaker identity to the target speaker from an external corpus, and then combines the selected utterances with the limited target data for SD vocoder adaptation.The evaluation results show that, compared with the vocoder adapted using only limited target data, the vocoder adapted using augmented data improves both the quality and similarity of synthesized speech.
Yi-Chiao Wu, Cheng-Hung Hu, Hung-Shin Lee, Yu-Huai Peng, Wen-Chin Huang, Yu Tsao 0001, Hsin-Min Wang, Tomoki Toda
Interspeech4
2020 ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
Xin Wang 0037, Junichi Yamagishi, Massimiliano Todisco, Héctor Delgado, Andreas Nautsch, Nicholas W. D. Evans, Md. Sahidullah, Ville Vestman, Tomi Kinnunen, Kong-Aik Lee, Lauri Juvela, Paavo Alku, Yu-Huai Peng, Hsin-Te Hwang, Yu Tsao 0001, Hsin-Min Wang, Sébastien Le Maguer, Zhen-Hua Ling
Comput. Speech Lang.13
2018 Exemplar-Based Spectral Detail Compensation for Voice Conversion
Yu-Huai Peng, Hsin-Te Hwang, Yi-Chiao Wu, Yu Tsao 0001, Hsin-Min Wang
INTERSPEECH1