Huayun Zhang

dblp:70/3736 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
14since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
abstract
Current emotional text-to-speech (TTS) models pre-dominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a single emotion per text-speech pair. These models only learn the correct emotional outputs without fully comprehending other emotion characteristics, which limits their capabilities of capturing the nuances between different emotions. We propose a controllable Emo-DPO approach, which employs direct preference optimization to differentiate subtle emotional nuances between emotions through optimizing towards preferred emotions over less preferred emotional ones. Instead of relying on traditional neural architectures used in existing emotional TTS models, we propose utilizing the emotion-aware LLM-TTS neural architecture to leverage LLMs’ in-context learning and instruction-following capabilities. Comprehensive experiments confirm that our proposed method outperforms the existing baselines.
Xiaoxue Gao, Chen Zhang 0020, Yiming Chen 0010, Huayun Zhang, Nancy F. Chen
ICASSP4
2025 Prompt-Unseen-Emotion: Mixed Emotional Speech Synthesis With Prompt-LLM Contextual Knowledge
abstract
Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse emotional speech generation for more natural interactions. To bridge this gap, this paper proposes a novel prompt-unseen-emotion (PUE) approach to generate unseen emotional speech via emotion-guided prompt learning.PUEis trained utilizing an LLM-TTS architecture to ensure emotional consistency between categorical emotion-relevant prompts and emotional speech, allowing the model to quantitatively capture different emotion weightings per utterance. During inference, mixed emotional speech can be generated by flexibly adjusting emotion proportions and leveraging LLM contextual knowledge, enabling the model to quantify different emotional styles. Our proposedPUEsuccessfully facilitates expressive speech synthesis of unseen emotions in a zero-shot setting.
Xiaoxue Gao, Huayun Zhang, Nancy F. Chen
IEEE Signal Process. Lett.2
2025 PRESENT: Zero-Shot Text-to-Prosody Control
abstract
Current strategies for achieving fine-grained prosody control in speech synthesis entail extracting additional style embeddings or adopting more complex architectures. To enable zero-shot application of pretrained text-to-speech (TTS) models, we present PRESENT (PRosody Editing without Style Embeddings or New Training), which exploits explicit prosody prediction in FastSpeech2-based models by modifying the inference process directly. We apply our text-to-prosody framework to zero-shot language transfer using a JETS model exclusively trained on English LJSpeech data. We obtain character error rates (CER) of 12.8%, 18.7% and 5.9% for German, Hungarian and Spanish respectively, beating the previous state-of-the-art CER by over 2× for all three languages. Furthermore, we allow subphoneme-level control, a first in this field. To evaluate its effectiveness, we show that PRESENT can improve the prosody of questions, and use it to generate Mandarin, a tonal language where vowel pitch varies at subphoneme level. We attain 25.3% hanzi CER and 13.0% pinyin CER with the JETS model. All our code and audio samples11https://github.com/iamanigeeit/presentandhttps://present2024.web.app/are available online.
Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman, Dorien Herremans
IEEE Signal Process. Lett.2
2024 Semi-Supervised Learning for Robust Speech Evaluation
abstract
Speech evaluation measures a learner’s oral proficiency using automatic models. Corpora for training such models often pose sparsity challenges given that there often is limited scored data from teachers, in addition to the score distribution across proficiency levels being often imbalanced among student cohorts. Automatic scoring is thus not robust when faced with under-represented samples or out-of-distribution samples, which inevitably exist in real-world deployment scenarios. This paper proposes to address such challenges by exploiting semi-supervised pre-training and objective regularization to approximate subjective evaluation criteria. In particular, normalized mutual information is used to quantify the speech characteristics from the learner and the reference. An anchor model is trained using pseudo labels to predict the correctness of pronunciation. An interpolated loss function is proposed to minimize not only the prediction error with respect to ground-truth scores but also the divergence between two probability distributions estimated by the speech evaluation model and the anchor model. Compared to other state-of-the-art methods on a public data-set, this approach not only achieves high performance while evaluating the entire test-set as a whole, but also brings the most evenly distributed prediction error across distinct proficiency levels. Furthermore, empirical results show the model accuracy on out-of-distribution data also compares favorably with competitive baselines.
Huayun Zhang, Jeremy H. M. Wong, Geyu Lin, Nancy F. Chen
SLT1
2024 SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
abstract
Text-to-speech (TTS) models have achieved remarkable naturalness in recent years, yet like most deep neural models, they have more parameters than necessary. Sparse TTS models can improve on dense models via pruning and extra retraining, or converge faster than dense models with some performance loss. Thus, we propose training TTS models using decaying sparsity, i.e. a high initial sparsity to accelerate training first, followed by a progressive rate reduction to obtain better eventual performance. This decremental approach differs from current methods of incrementing sparsity to a desired target, which costs significantly more time than dense training. We call our method SNIPER training: Single-shot Initialization Pruning Evolving-Rate training. Our experiments on FastSpeech2 show that we were able to obtain better losses in the first few training epochs with SNIPER, and that the final SNIPER-trained models outperformed constant-sparsity models and edged out dense models, with negligible difference in training time. Our code is available on Github11https://github.com/iamanigeeit/sniper.
Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman, Dorien Herremans
TENCON2
2023 Variational Gaussian Process Data Uncertainty
abstract
A Gaussian process (GP) computes an explainable distributional uncertainty by hypothesising higher output uncertainty for query inputs far from the training inputs. However, a GP may not capture data uncertainty well. Accurate data uncertainty estimation may be important for subjective tasks, such as Spoken Language Assessment (SLA), where human expert raters may disagree on the output scores. This paper shows that a variational approximation of a GP has capacity to learn data uncertainty from the training data. However, standard training criteria tune only a scalar noise hyper-parameter toward the standard deviation of the output reference, thereby limiting the learning of this uncertainty. A training criterion is proposed to explicitly encourage the GP posterior to emulate the distribution of scores from multiple raters. Experiments on the speechocean762 SLA task show that this allows the GP to better express data uncertainty and improves the modelling of inter-rater disagreements.
Jeremy H. M. Wong, Huayun Zhang, Nancy F. Chen
ASRU2
2023 Student Engagement Detection: Case Study on Using Peer-to-Peer Emotion Comparison with Context Regularization
abstract
This paper describes a method to automatically assess participants' engagement in online education. Similar to emotion recognition, student engagement can be subjective. Hence, it is challenging to obtain large-enough and consistent ground-truth engagement labels for automatic student engagement. We propose an unsupervised method that could detect abnormal engagement states using peer-to-peer emotion correlation analysis in different modalities. Without any human engagement labeling, this zero-shot method accurately pinpoints the abnormal student engagements in our experiment. Modality-dependent engagement prediction also suggests possible distractions on the student's device.
Geyu Lin, Manas Gupta, Cheryl Sze Yin Wong, Huayun Zhang
ICCE4
2023 Distilling knowledge from Gaussian process teacher to neural network student
Jeremy H. M. Wong, Huayun Zhang, Nancy F. Chen
INTERSPEECH2
2023 Modelling Inter-Rater Uncertainty in Spoken Language Assessment
abstract
In a subjective task, such as Spoken Language Assessment (SLA), the reference scores provided by different human raters may vary. A collection of annotated scores from multiple raters can be interpreted as an expression of data uncertainty. Previous studies often treat SLA as classification or regression tasks, and train and evaluate models against scalar reference scores that were computed from the multiple rater scores, for example by majority voting. However, a scalar representation may not adequately capture information about uncertainty that is expressed by the multiple rater scores. This paper proposes to reformulate this subjective task as a distribution fitting problem, where the model should aim to emulate the uncertainty expressed by the multiple raters. Toward this aim, the model is trained and evaluated by computing a distance between the model's output posterior and the distribution of reference scores from the multiple raters. Different methods to infer a scalar score from the model's output posterior are also considered. This paper also proposes to improve the match between the model and the SLA task, by interpreting the model's outputs as parameters of a beta density function, to capture both uncertainty and score monotonicity. Finally, ensemble combination is investigated and a novel combination method is proposed, to marginalise out model uncertainty from the combined output distribution. These approaches are evaluated on the speechocean762 dataset and an in-house Tamil dataset.
Jeremy H. M. Wong, Huayun Zhang, Nancy F. Chen
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 EPIC TTS Models: Empirical Pruning Investigations Characterizing Text-To-Speech Models
abstract
Neural models are known to be over-parameterized, and recent work has shown that sparse text-to-speech (TTS) models can outperform dense models. Although a plethora of sparse methods has been proposed for other domains, such methods have rarely been applied in TTS. In this work, we seek to answer the question: what are the characteristics of selected sparse techniques on the performance and model complexity? We compare a Tacotron2 baseline and the results of applying five techniques. We then evaluate the performance via the factors of naturalness, intelligibility and prosody, while reporting model size and training time. Complementary to prior research, we find that pruning before or during training can achieve similar performance to pruning after training and can be trained much faster, while removing entire neurons degrades performance much more than removing parameters. To our best knowledge, this is the first work that compares sparsity paradigms in text-to-speech synthesis.
Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman
INTERSPEECH2
2022 Variations of multi-task learning for spoken language assessment
Jeremy H. M. Wong, Huayun Zhang, Nancy F. Chen
INTERSPEECH2
2021 WittyKiddy: Multilingual Spoken Language Learning for Kids
Ke Shi 0001, Kye Min Tan, Huayun Zhang, Siti Umairah Md. Salleh, Shikang Ni, Nancy F. Chen
Interspeech3
2021 Multilingual Speech Evaluation: Case Studies on English, Malay and Tamil
abstract
Speech evaluation is an essential component in computerassisted language learning (CALL).While speech evaluation on English has been popular, automatic speech scoring on low resource languages remains challenging.Work in this area has focused on monolingual specific designs and handcrafted features stemming from resource-rich languages like English.Such approaches are often difficult to generalize to other languages, especially if we also want to consider suprasegmental qualities such as rhythm.In this work, we examine three different languages that possess distinct rhythm patterns: English (stresstimed), Malay (syllable-timed), and Tamil (mora-timed).We exploit robust feature representations inspired by music processing and vector representation learning.Empirical validations show consistent gains for all three languages when predicting pronunciation, rhythm and intonation performance.
Huayun Zhang, Ke Shi 0001, Nancy F. Chen
Interspeech1
2021 Online-Semisupervised Neural Anomaly Detector to Identify MQTT-Based Attacks in Real Time
abstract
Industry 4.0 focuses on continuous interconnection services, allowing for the continuous and uninterrupted exchange of signals or information between related parties. The application of messaging protocols for transferring data to remote locations must meet specific specifications such as asynchronous communication, compact messaging, operating in conditions of unstable connection of the transmission line of data, limited network bandwidth operation, support multilevel Quality of Service (QoS), and easy integration of new devices. The Message Queue Telemetry Transport (MQTT) protocol is used in software applications that require asynchronous communication. It is a light and simplified protocol based on publish-subscribe messaging and is placed functionally over the TCP/IP protocol. It is designed to minimize the required communication bandwidth and system requirements increasing reliability and probability of successful message transmission, making it ideal for use in Machine-to-Machine (M2M) communication or networks where bandwidth is limited, delays are long, coverage is not reliable, and energy consumption should be as low as possible. Despite the fact that the advantage that MQTT offers its way of operating does not provide a serious level of security in how to achieve its interconnection, as it does not require protocol dependence on one intermediate third entity, the interface is dependent on each application. This paper presents an innovative real-time anomaly detection system to detect MQTT-based attacks in cyber-physical systems. This is an online-semisupervised learning neural system based on a small number of sampled patterns that identify crowd anomalies in the MQTT protocol related to specialized attacks to undermine cyber-physical systems.
Xingwei Wang 0001, Huayun Zhang, Zengrong Xu
Secur. Commun. Networks4
2006 Pattern-Based Dynamic Compensation Towards Robust Speech Recognition in Mobile Environments
abstract
Today, the high mobility provided by wireless networks places users in a wild variety of noise and channel conditions, which poses serious challenge to telephone-base Acoustic Speech Recognition (ASR). In this paper, we propose a Pattern-based Dynamic Compensation (PDC) scheme to improve the robustness of ASR in mobile environments. In PDC, a distortion pattern-set is employed to normalize the environmental variations in training data according to a set of pre-defined application scenarios. At recognition time, instantaneous distortion is calculated as a linear combination of several possible patterns. To online estimate the combination weights robustly, a Bayesian learning process with Speech-conditioned Prior Evolution is introduced into PDC (PDC-SPE). In outdoor experiments, the PDC-SPE method outperforms other commonly used compensation/adaptation methods and leads to 20∼25% relative reduction in Word Error Rate (WER) over a well-trained baseline system.
Huayun Zhang
ICASSP (1)1
2003 A vector statistical piecewise polynomial approximation algorithm for environment compensation in telephone LVCSR
abstract
A vector statistical piecewise polynomial (VPP) approximation algorithm is proposed for environment compensation in speech signals that are degraded by both additive and convolutive noise. By investigating the model of the telephone environment, we address a piecewise polynomial, namely two linear polynomials and a quadratic polynomial, to approximate the environment function precisely. The VPP is applied either to stationary noise, or to non-stationary noise. In the first case, batch EM is used in the log-spectral domain; in the second case, recursive EM with iterative stochastic approximation is developed in the cepstral domain. Both approaches are based on the minimum mean squared error (MMSE) sense. Experimental results are presented on the application of this approach in improving the performance of Mandarin large vocabulary continuous speech recognition (LVCSR) in background noise and different transmission channels (such as fixed telephone line and GSM). The method can reduce the average character error rate (CER) by about 18%.
Zhaobing Han, Shuwu Zhang, Huayun Zhang, Bo Xu 0002
ICASSP (2)3
2003 Dynamic channel compensation based on maximum a posteriori estimation
Huayun Zhang, Zhaobing Han, Bo Xu 0002
INTERSPEECH1
2003 Geometric constrained maximum likelihood linear regression on Mandarin dialect adaptation
Huayun Zhang, Bo Xu 0002
INTERSPEECH1
2002 Pitch and tone's modeling in parametric trajectory model
abstract
It is described in this paper for the application of pitch/tone information in the parametric trajectory model. Pitch as a dynamic feature and its contours—tone as a segmental-level feature are deserved their own particular characteristics, which match case of parametric trajectory model better compared with MFCC and energy. Here we give an improved pitch extraction algorithm and especially the “total” and “parallel” integration methods to combine these information with the base model. In the experiment of Mandarin connected digit recognition, we achieve 22.87% and 33.54% error reduction respectively for them, moreover when combined with these two methods, 38.72% error reduction is obtained.
Bo Xu 0002, Huayun Zhang
ICASSP4
2002 Codebook dependent dynamic channel estimation for Mandarin speech recognition over telephone
Huayun Zhang, Zhaobing Han, Bo Xu 0002
INTERSPEECH1
2002 Improving parametric trajectory modeling by integration of pitch and tone information
Bo Xu 0002, Huayun Zhang
INTERSPEECH4