Cheng-Hung Hu

dblp:156/0822 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-6112-9380ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
Cheng-Hung Hu, Yusuke Yasuda, Akifumi Yoshimoto, Tomoki Toda
INTERSPEECH1
2025 E2EPref: An end-to-end preference-based framework for speech quality assessment to alleviate bias in direct assessment scores
abstract
In speech quality assessment (SQA), direct assessment (DA) scores are frequently used as the objective of model training. However, because the DA scores themselves have listener-wise bias and equal range bias, the scores predicted by models trained with DA scores do not always reflect the true quality score. In this study, we utilize preference-based learning for SQA by transforming the DA score prediction framework into a preference prediction framework. Our proposed End-to-End Preference-based framework (E2EPref) for SQA is designed for predicting system-level quality scores directly. It contains four proposed components: pair generation, preference function, threshold selection, and preference aggregation. Through these functions of E2EPref, we aim to mitigate biases introduced by directly using DA scores for training. In experiments, we show that this framework helps the SQA model alleviate biases, resulting in higher system-level Spearman’s rank correlation coefficient and linear correlation coefficient. Additionally, we evaluate the quality prediction capability of the framework in a zero-shot out-of-domain scenario. Finally, we collect subjective preference scores on a dataset already containing DA scores and analyze the advantages and disadvantages of using DA scores versus subjective preference scores as the ground truth or for model training. • Proposed E2EPref: an end-to-end preference-based SQA framework. • Pair generation, preference function, and aggregation enhance model performance. • E2EPref outperforms baselines in in-domain, out-of-domain, and cross-dataset tests. • Collected preference score dataset for advancing preference-based SQA research.
Cheng-Hung Hu, Yusuke Yasuda, Tomoki Toda
Comput. Speech Lang.1
2024 Embedding Learning for Preference-based Speech Quality Assessment
Cheng-Hung Hu, Yusuke Yasuda, Tomoki Toda
INTERSPEECH1
2023 Preference-based training framework for automatic speech quality assessment using deep neural network
abstract
One objective of Speech Quality Assessment (SQA) is to estimate the ranks of synthetic speech systems. However, recent SQA models are typically trained using low-precision direct scores such as mean opinion scores (MOS) as the training objective, which is not straightforward to estimate ranking. Although it is effective for predicting quality scores of individual sentences, this approach does not account for speech and system preferences when ranking multiple systems. We propose a training framework of SQA models that can be trained with only preference scores derived from pairs of MOS to improve ranking prediction. Our experiment reveals conditions where our framework works the best in terms of pair generation, aggregation functions to derive system score from utterance preferences, and threshold functions to determine preference from a pair of MOS. Our results demonstrate that our proposed method significantly outperforms the baseline model in Spearman's Rank Correlation Coefficient.
Cheng-Hung Hu, Yusuke Yasuda, Tomoki Toda
INTERSPEECH1
2022 NASTAR: Noise Adaptive Speech Enhancement with Target-Conditional Resampling
abstract
For deep learning-based speech enhancement (SE) systems, the training-test acoustic mismatch can cause notable performance degradation.To address the mismatch issue, numerous noise adaptation strategies have been derived.In this paper, we propose a novel method, called noise adaptive speech enhancement with target-conditional resampling (NASTAR), which reduces mismatches with only one sample (one-shot) of noisy speech in the target environment.NASTAR uses a feedback mechanism to simulate adaptive training data via a noise extractor and a retrieval model.The noise extractor estimates the target noise from the noisy speech, called pseudo-noise.The noise retrieval model retrieves relevant noise samples from a pool of noise signals according to the noisy speech, called relevant-cohort.The pseudo-noise and the relevant-cohort set are jointly sampled and mixed with the source speech corpus to prepare simulated training data for noise adaptation.Experimental results show that NASTAR can effectively use one noisy speech sample to adapt an SE model to a target condition.Moreover, both the noise extractor and the noise retrieval model contribute to model adaptation.To our best knowledge, NASTAR is the first work to perform one-shot noise adaptation through noise extraction and retrieval.
Chi-Chang Lee, Cheng-Hung Hu, Yuchen Lin 0003, Chu-Song Chen, Hsin-Min Wang, Yu Tsao 0001
INTERSPEECH2
2022 SVSNet: An End-to-End Speaker Voice Similarity Assessment Model
abstract
Neural evaluation metrics derived for numerous speech generation tasks have recently attracted great attention. In this paper, we propose SVSNet, the first end-to-end neural network model to assess the speaker voice similarity between converted speech and natural speech for voice conversion tasks. Unlike most neural evaluation metrics that use hand-crafted features, SVSNet directly takes the raw waveform as input to more completely utilize speech information for prediction. SVSNet consists of encoder, co-attention, distance calculation, and prediction modules and is trained in an end-to-end manner. The experimental results on the Voice Conversion Challenge 2018 and 2020 (VCC2018 and VCC2020) datasets show that SVSNet outperforms well-known baseline systems in the assessment of speaker similarity at the utterance and system levels.
Cheng-Hung Hu, Yu-Huai Peng, Junichi Yamagishi, Yu Tsao 0001, Hsin-Min Wang
IEEE Signal Process. Lett.1
2021 Relational Data Selection for Data Augmentation of Speaker-Dependent Multi-Band MelGAN Vocoder
abstract
Nowadays, neural vocoders can generate very high-fidelity speech when a bunch of training data is available.Although a speaker-dependent (SD) vocoder usually outperforms a speaker-independent (SI) vocoder, it is impractical to collect a large amount of data of a specific target speaker for most realworld applications.To tackle the problem of limited target data, a data augmentation method based on speaker representation and similarity measurement of speaker verification is proposed in this paper.The proposed method selects utterances that have similar speaker identity to the target speaker from an external corpus, and then combines the selected utterances with the limited target data for SD vocoder adaptation.The evaluation results show that, compared with the vocoder adapted using only limited target data, the vocoder adapted using augmented data improves both the quality and similarity of synthesized speech.
Yi-Chiao Wu, Cheng-Hung Hu, Hung-Shin Lee, Yu-Huai Peng, Wen-Chin Huang, Yu Tsao 0001, Hsin-Min Wang, Tomoki Toda
Interspeech2
2017 A Two-Stage Latent Variable Estimation Procedure for Time-Censored Accelerated Degradation Tests
abstract
Parallel constant-stress accelerated degradation testing (PCSADT) is widely used to assess the reliability of highly reliable products in a timely manner when the products' degradation can be measured. Under a time-censored PCSADT, several groups of units are tested simultaneously, but under different stress levels, until a prespecified censoring time is reached. At this time, degradation values from the censored units, and failure times of the failed units are obtained. When the degradation follows a Wiener process where the parameters depend on the stress level through a life-stress model containing an unknown nuisance parameter, estimating this parameter often biases the maximum likelihood and least-squares estimators of the lifetime parameters. In this paper, we propose a two-stage procedure to address this problem. In the first stage, we transform the data under the different stress levels of a PCSADT so that the resulting data can be considered to have been obtained under normal stress. In the second stage, we introduce a latent variable for the unobserved degradation after the failure time for each failed unit to obtain a pseudodegradation value at the censoring time. We then use all degradation values (pseudo or observed) at the censoring time to develop latent variable estimators for all model parameters. Unlike other existing estimators, the proposed estimators are shown to be s-consistent, have closed-form expressions, and are easy to interpret. We use a real example of light-emitting diodes to illustrate the proposed method. In addition to proving s-consistencies, we conduct a simulation study to demonstrate that the proposed estimators also perform well in finite samples.
Ming-Yung Lee, Cheng-Hung Hu, Jen Tang
IEEE Trans. Reliab.2