Hao Huang 0009

dblp:04/5616-9 · DBLP profile ↗
← Back
62ranked-venue papers
3as first author
56since 2021 · last 2026
0000-0001-6604-0951ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 46 · 1 first-author · 43 since 2021Artificial intelligence and machine learning · 30 · 1 first-author · 26 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding
abstract
Spoken Language Understanding (SLU) consists of two sub-tasks: intent detection (ID) and slot filling (SF). Given its broad range of real-world applications, enhancing SLU for practical deployment is increasingly critical. Profile-based SLU addresses ambiguous user utterances by incorporating context awareness (CA), user profiles (UP), and knowledge graphs (KG) to support disambiguation, thereby advancing SLU research toward real-world applicability. However, existing SLU datasets still fall short in representing real-world scenarios. Specifically, (1) CA uses one-hot vectors for representation, which is overly idealized, and (2) models typically focuses solely on predicting intents and slot labels, neglecting the reasoning process that could enhance performance and interpretability. To overcome these limitations, we introduce VRSLU, a novel SLU dataset that integrates both Visual images and explicit Reasoning. For over-idealized CA, we use GPT-4o and FLUX.1-dev to generate images reflecting users’ environments and statuses, followed by human verification to ensure quality. For reasoning, GPT-4o is employed to generate explanations for predicted labels, which are then refined by human annotators to ensure accuracy and coherence. Additionally, we propose an instructional template, LR-Instruct, which first predicts labels and then generates corresponding reasoning. This two-step approach helps mitigate the influence of reasoning bias on label prediction. Experimental results confirm the effectiveness of incorporating visual information and highlight the promise of explicit reasoning in advancing SLU.
Di Wu 0088, Liting Jiang, Ruiyu Fang, Bianjing, Hongyan Xie, Haoxiang Su, Hao Huang 0009, Zhongjiang He, Shuangyong Song, Xuelong Li 0001
AAAI7
2026 LCMA-SRT: Language-Conditional Mixture-of-Experts Adapters for Joint Multilingual Speech Recognition and Translation
abstract
Neural transducers offer an alignment-free framework for speech-to-text modeling, and hierarchical transducer architectures further improve multilingual joint automatic speech recognition (ASR) and speech translation (ST) by stacking a translation-focused encoder on top of an ASR encoder.However, extending hierarchical transducers to multilingual manyto-many settings remains challenging: fully shared models often suffer from negative transfer and unstable target-language generation, while training separate models for each direction is computationally prohibitive.We propose LCMA-SRT (Language-Conditional Mixtureof-Experts Adapters for Speech Recognition and Translation), which augments a hierarchical transducer with language-conditional Mixture-of-Experts (MoE) adapters.A sourceconditioned MoE adapter (SRC-MoE) uses source-language embeddings to reduce crosslanguage interference and improve multilingual ASR.A target-conditioned MoE adapter (TGT-MoE) uses the desired target language to reduce cross-target interference and stabilize targetlanguage generation in many-to-many ST.Experiments on Europarl-ST (9 languages, 72 directions) show that LCMA-SRT improves both ASR and ST within a single joint model, reducing average WER and improving BLEU and COMET over strong hierarchical transducer baselines.We release our code and models at https://github.com/linanjie0820/ LCMA-SRT.
Nanjie Li, Xiaoyong Guo, Hao Huang 0009, Haihua Xu 0001
ACL (1)3
2025 Improved Cross-Lingual Speaker Verification Using Speaker Sensitive Feature Guidance and Fine-grained Phonetic Information
abstract
Speaker verification performance significantly degrades when there exists a language mismatch between training and evaluation. Domain Adversarial Training (DAT) has shown to be effective in mitigating this gap by incorporating adversarial training with domain information (language id). Inspired by recent research of DAT, we propose to decompose and locate features sensitive to speaker identity, so that domain adaptation can better serve the speaker classification task. Additionally, we further promote DAT by replacing alignment based on language identification with alignment of fine-grained phonetic information, where pre-trained speech recognition models are utilized to provide frame-level phonetic labels. The speaker verification model is trained on VoxCeleb, while CnCeleb is used for adversarial training and evaluation. Results show the methods effectively mitigate the performance degradation caused by language mismatch.
Yongtai Ji, Guangxing Li, Hao Huang 0009, Yanbing Li, Wushour Slamu
ICASSP3
2025 Utterance as A Bridge: Few-shot Joint Learning of Empathy Detection and Empathy Intent Classification
abstract
Empathy detection (ED) and empathy intent classification (EIC) aim to identify the empathy direction expressed in user utterances and the underlying empathy intent behind them. Previous studies show that facilitating information transfer between tasks can enhance model performance. However, the interaction between ED and EIC in few-shot learning remains underexplored. To this end, we identify the challenges in jointly training ED and EIC in a few-shot setting: establishing effective information transfer between them and improving the model’s generalization capability. We propose a novel model called USB. For information transfer, the interactive module maps empathy and empathy intent labels through utterances to model task correlations. For generalization capability, after capturing empathy and empathy intent representations with an adaptive fusion module, we introduce a multi-level contrastive learning strategy to optimize representations at task and label levels, enhancing generalization. Experimental results on two public datasets show that our model outperforms all baselines.
Liting Jiang, Di Wu 0088, Shuangyong Song, Yanbing Li, Hao Huang 0009
ICASSP6
2025 Robust and Efficient Text-based Speech Editing using Noise Conditioning and Rectified Flow
abstract
Significant advancements have been made in text-based speech editing (TSE) for clear speech, but effectively editing the noise-contaminated speech remains a challenge. Background noise degrades the quality of generated speech, and edited speech that fails to maintain noise context consistency often sounds unnatural. We propose Reflow-TSE, a robust and efficient TSE model for noise-consistent speech editing. For noise robustness, 1) we design a noise condition module to extract frame-level noise sequences as the conditioning information, and 2) we introduce an enhanced context-conditioned prediction module that predicts masked noise sequences along with conventional duration and pitch, using these conditions to guide generation. For boosting efficiency, 3) we introduce the rectified flow model, leveraging speech context and predicted conditions to achieve high-quality editing with limited sampling steps. Experimental results show that with just two steps of sampling, Reflow-TSE achieves context-consistent noisy speech editing, a capability absent in other TSE models. Additionally, for clean speech, Reflow-TSE also matches or surpasses baseline models. Audio samples are available at https://hw-su.github.io.
Haowen Yin, Hao Huang 0009, Wushour Slamu
ICASSP4
2025 Multi-Segment Soft Robot Control Via Deep Koopman-Based Model Predictive Control
abstract
Soft robots, compared to regular rigid robots, as their multiple segments with soft materials bring flexibility and compliance, have the advantages of safe interaction and dexterous operation in the environment. However, due to its characteristics of high dimensional, nonlinearity, time-varying nature, and infinite degree of freedom, it has been challenges in achieving precise and dynamic control such as trajectory tracking and position reaching. To address these challenges, we propose a framework of Deep Koopman-based Model Predictive Control (DK-MPC) for handling multi-segment soft robots. We first employ a deep learning approach with sampling data to approximate the Koopman operator, which therefore linearizes the high-dimensional nonlinear dynamics of the soft robots into a finite-dimensional linear representation. Secondly, this linearized model is utilized within a model predictive control framework to compute optimal control inputs that minimize the tracking error between the desired and actual state trajectories. The real-world experiments on the soft robot “Chordata” demonstrate that DK-MPC could achieve highprecision control, showing the potential of DK-MPC for future applications to soft robots. More visualization results can be found at https://pinkmoon-io.github.io/DKMPC/.
Lei Lv, Lei Liu 0076, Fuchun Sun 0001, Jiahong Dong, Jianwei Zhang 0001, Xuemei Shan, Kai Sun 0014, Hao Huang 0009, Yu Luo 0021
ICRA9
2025 VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge
Zijing Zhao 0008, Hao Huang 0009, Ying Hu 0005, Liang He 0003
INTERSPEECH3
2025 A Joint Network for Singing Melody Extraction from Polyphonic Music with Attention Aggregation and Self-Consistency Training
Jiabo Jing, Ying Hu 0005, Hao Huang 0009, Liang He 0003, Zhijian Ou
INTERSPEECH3
2025 Lightweight and Robust Multi-Channel End-to-End Speech Recognition with Spherical Harmonic Transform
Xiangzhu Kong, Hao Huang 0009, Zhijian Ou
INTERSPEECH2
2025 LLM-based phoneme-to-grapheme for phoneme-based speech recognition
Te Ma, Min Bi, Saierdaer Yusuyin, Hao Huang 0009, Zhijian Ou
INTERSPEECH4
2025 Language-Aware Prompt Tuning for Parameter-Efficient Seamless Language Expansion in Multilingual ASR
abstract
Recent advancements in multilingual automatic speech recognition (ASR) have been driven by large-scale end-to-end models like Whisper. However, challenges such as language interference and expanding to unseen languages (language expansion) without degrading performance persist. This paper addresses these with three contributions: 1) Entire Soft Prompt Tuning (Entire SPT), which applies soft prompts to both the encoder and decoder, enhancing feature extraction and decoding; 2) Language-Aware Prompt Tuning (LAPT), which leverages cross-lingual similarities to encode shared and language-specific features using lightweight prompt matrices; 3) SPT-Whisper, a toolkit that integrates SPT into Whisper and enables efficient continual learning. Experiments across three languages from FLEURS demonstrate that Entire SPT and LAPT outperform Decoder SPT by 5.0% and 16.0% in language expansion tasks, respectively, providing an efficient solution for dynamic, multilingual ASR models with minimal computational overhead.
Sheng Li 0010, Hao Huang 0009, Ayiduosi Tuohan, Yizhou Peng
INTERSPEECH3
2025 Adapting Whisper for Parameter-efficient Code-Switching Speech Recognition via Soft Prompt Tuning
abstract
Large-scale multilingual ASR models like Whisper excel in high-resource settings but face challenges in low-resource scenarios, such as rare languages and code-switching (CS), due to computational costs and catastrophic forgetting. We explore Soft Prompt Tuning (SPT), a parameter-efficient method to enhance CS ASR while preserving prior knowledge. We evaluate two strategies: (1) full fine-tuning (FFT) of both soft prompts and the entire Whisper model, demonstrating improved cross-lingual capabilities compared to traditional methods, and (2) adhering to SPT's original design by freezing model parameters and only training soft prompts. Additionally, we introduce SPT4ASR, a combination of different SPT variants. Experiments on the SEAME and ASRU2019 datasets show that deep prompt tuning is the most effective SPT approach, and our SPT4ASR methods achieve further error reductions in CS ASR, maintaining parameter efficiency similar to LoRA, without degrading performance on existing languages.
Yizhou Peng, Hao Huang 0009, Sheng Li 0010
INTERSPEECH3
2025 RAICL-DSC: Retrieval-Augmented In-Context Learning for Dialogue State Correction
Haoxiang Su, Hongyan Xie, Di Wu 0088, Liting Jiang, Hao Huang 0009, Zhongjiang He, Ruiyu Fang, Shuangyong Song
Knowl. Based Syst.6
2025 A classifier expansion framework with dual knowledge distillation and dynamic weighting for continual relation extraction
Aonan Mao, Di Wu 0088, Liting Jiang, Shuangyong Song, Yanbing Li, Hao Huang 0009, Wushour Slamu
J. Supercomput.6
2024 Image-Text Sentiment Analysis Based on Cross-Modal Interactive Attention
Wushouer Mairidan, Gulanbaier Tuerhong, Hao Huang 0009, Liwei Tian, Suping Liu, Longqing Zhang
GPC4
2024 Fact-Aware Summarization with Contrastive Learning for Few-Shot Dialogue State Tracking
abstract
Dialogue state tracking (DST) is a crucial component of task-oriented dialogue systems, as it aims to accurately track the user’s goals throughout the dialogue history. However, DST models struggle with new domains due to limited annotated data, leading to poor performance. To solve this key challenge in DST, we propose a model called Fact-aware Summarization model for few-shot DST (FaS-DST), which introduces a "Summarize, Extract, and Select" pattern. Specifically, we decompose DST into three sub-tasks: generating candidate summaries, extracting dialogue states, and scoring the candidates to select the most accurate one. Contrastive learning is incorporated to train a candidate scorer, which improves faithfulness and factuality in dialogue summarization. Additionally, we employ two strategies namely data augmentation and summary & state concatenation to improve the model’s training effectiveness. Experimental results demonstrate that FaS-DST outperforms state-of-the-art models on both MultiWOZ 2.0 and MultiWOZ 2.1 datasets in few-shot settings.
Sijie Feng, Haoxiang Su, Hongyan Xie, Di Wu 0088, Hao Huang 0009, Wushour Slamu
ICASSP5
2024 SMMA-Net: An Audio Clue-Based Target Speaker Extraction Network with Spectrogram Matching and Mutual Attention
abstract
We propose a deep neural network with spectrogram matching and mutual attention (SMMA-Net) for audio clue-based target speaker extraction (TSE). To effectively use the auxiliary speech, we proposed spectrogram matching (SM) strategy and mutual attention (MA) block. We conducted all experiments on the WSJ0-2mix-extr dataset. The ablation and comparison studies verified the effectiveness of SM strategy and MA block. The experimental results show that our proposed method outperforms the state-of-the-art methods by a sizable margin of 1.3 dB on the metric of scale-invariant signal-to-distortion ratio improvement. Additionally, SMMA-Net achieved that the performance of model for TSE task exceeds that for speaker separation task under the similar architecture. The main code will be available at https://github.com/Ht-Xu/SMMA-Net.
Ying Hu 0005, Zhongcun Guo, Hao Huang 0009, Liang He 0003
ICASSP4
2024 Introducing Multilingual Phonetic Information to Speaker Embedding for Speaker Verification
abstract
Incorporating frame-level phonetic information during the extraction of speaker embeddings has been shown to enhance the performance of speaker verification systems. However, previous studies have primarily relied on phonetic information obtained from pre-trained models of monolingual automatic speech recognition (ASR). Considering that speaker verification datasets typically consist of multiple languages, there are instances where speakers are proficient in multiple languages, resulting in discrepancies between the languages used in the enrolled and test utterances. To address these challenges, we employ a pre-trained multilingual ASR Conformer encoder to initialize the MFA-Conformer network for speaker verification. Experimental results on the VoxCeleb dataset demonstrate a significant improvement in the performance of the system that incorporates multilingual phonetic information across different evaluation sets, including VoxCeleb1-O, E, and H, as well as the VoxSRC21 validation set, which focuses on multilingual verification. The source code is released at https://github.com/zds-potato/multilingual-phonetic-sv.
Zhida Song, Liang He 0003, Ying Hu 0005, Hao Huang 0009
ICASSP5
2024 Domain-Slot Aware Contrastive Learning for Improved Dialogue State Tracking
abstract
Large-scale pre-trained neural language model has facilitated to achieve the state-of-the-art performance on Dialogue State Tracking (DST) tasks. One of the existing works models the semantic correlation between the dialogue context and (domain, slot) pair encoded by BERT and make the prediction. Despite the effectiveness, they ignore the fact that there is no perfect semantic correspondence between (domain, slot) pair and the dialogue context. In this paper, we propose a domain-slot aware contrastive learning framework to solve this problem, which proposes three methods to bridge the semantic gap between the dialogue context and the (domain, slot) by constructing training sample pairs to fine-tune the BERT model and use it for base DST model. The experiments demonstrate that our proposed method has improved the performance of the baseline model on the MultiWOZ2.1 and MultiWOZ2.4 datasets, yielding competitive results.
Haoxiang Su, Sijie Feng, Hongyan Xie, Di Wu 0088, Hao Huang 0009, Zhongjiang He, Shuangyong Song, Ruiyu Fang, Xiaomeng Huang, Wushour Slamu
ICASSP5
2024 Phase Continuity-Aware Self-Attentive Recurrent Network with Adaptive Feature Selection for Robust VAD
abstract
Deep neural network (DNN) applications have significantly progressed in voice activity detection (VAD). Most current DNN-based VAD methods ignore the rich audio information in the phase domain. Therefore, applying this auxiliary information rationally and coping with low signal-to-noise ratio (SNR) background noise environments remains one of the challenges for VAD. To address this problem, we propose a VAD model robust to noise called phase continuity-aware self-attentive recurrent network (PC-ARN). For the input of PC-ARN, we draw inspiration from recent speech enhancement research by introducing phase-related features and further employing an adaptive feature selection module (AFSM) to combine magnitude features with it efficiently. The backbone network is an ARN module combining the attention mechanism and recurrent neural network (RNN), which can consider the relationship between local and global information to improve VAD performance competently. Experimental results show that our method has remarkable generalization ability and robustness compared to the traditional VAD techniques.
Minjie Tang, Hao Huang 0009, Liang He 0003
ICASSP2
2024 Dual Level Intent-Slot Interaction for Improved Multi-Intent Spoken Language Understanding
abstract
Multi-intent spoken language understanding consists of two typical subtasks: multi-intent detection and slot filling. Existing approach suffers from two limitations: (1) It fails to explicitly model the information transfer between slots associated within the same intent clause; (2) Using a co-occurrence matrix of both label encodings introduces needless slot positional information such as the prefix ‘B-’ or ‘I-’. For (1), we propose a Gaussian Graph Attention Network that allows interaction to focus not only on the connection between slots within the current intent clause, but also on the connection between intent, and between intent and slot. For (2), we use a co-occurrence matrix of intent categories and slot types to model the knowledge transfer between the two subtasks in the corpus-level interaction, bypassing the introduction of slot positional information. Our framework achieves significant accuracy gains on both the MixATIS and MixSNIPS datasets.
Di Wu 0088, Liting Jiang, Haoxiang Su, Hao Huang 0009
ICASSP7
2024 A State of the Art Review on Artificial Intelligence-Enabled Cyber Security in Smart Grid
Hao Huang 0009, Weidong Fang 0002, Wei Chen 0036, Andrew W. H. Ip, Kai-Leung Yung
ICIC (9)1
2024 Speaker Recognition Based on Pre-Trained Model and Deep Clustering
abstract
In this paper, we propose a novel loss by integrating a deep clustering (DC) loss at the frame-level and a speaker recognition loss at the segment-level into a single network without additional data requirements and exhaustive computation. The DC loss implicitly generates soft pseudo-phoneme labels for each frame-level feature, which facilitates extracting more discriminant speaker representation by suppressing phonetic content information. We study the DC loss not only on the acoustic feature, but also on the features extracted by the pre-trained models, such as wav2vec 2.0, HuBERT and WavLM. Experimental results on the VoxCeleb dataset shows that the overall system performance based on the pre-trained model features are better than the one on the acoustic feature. The proposed loss is significantly effective for systems on the acoustic feature and has a marginal improvement for systems on the pre-trained model feature.
Liang He 0003, Zhida Song, Shuanghong Liu, Mengqi Niu, Ying Hu 0005, Hao Huang 0009
ICME6
2024 Improving Pointer Network based Dialogue State Tracking via Dual Hierarchical Selective Augmentation
abstract
Dialogue state tracking is responsible for predicting the user’s dialogue state during the whole dialogue process. In practical applications, values for different slots exist in individual utterances of the dialog history. With the accumulation of the dialogue history, it becomes extremely difficult to accurately predict slots and corresponding values from the lengthy dialogue history. To solve the problem of the interference caused by lengthy dialogue history, we propose a dual hierarchical selective augmentation method, which makes use of two hierarchical level information selection strategy to generate slot values. In the encoding phase, we first extract word-level matching features between the slot and each dialogue turn, and then build turn-level context relevance. In the decoding phase, first of all, from a global perspective, the dialogue turn information is selected multiple according to the dialogue context and slot, so that the model focuses more on the turn containing slot value. Secondly, our model performs weighted context attention to capture the critical words of dialogue turn from the local view. This dual hierarchical context selection alleviates the interference caused by excessive redundant information in the dialogue history and enhances the judgment ability of the model for vital turns and words. Furthermore, to enhance the copying ability of the model, we use the turn selection-guided pointer network to copy slot values from the dialogue. Experimental results show that our model significantly outperforms multiple baselines on the released MultiWOZ benchmark.
Shuangyong Song, Hongyan Xie, Haoxiang Su, Hao Huang 0009, Mengxiang Li, Zhongjiang He, Ruiyu Fang
IJCNN4
2024 Graph-based Dynamic Domain Selection for Dialogue State Tracking
abstract
The Dialogue State Tracking (DST) module tracks the user’s intent by populating multiple predefined slots related to the dialogue task. In recent years, various graph neural network-based DST methods have been proposed to establish graph structures capturing the correlations between domains and slots, thereby enhancing model performance. However, these methods may involve redundant connections in the graph structure. To better construct relationships between domains and slots, we introduce a graph neural network-based dialogue state tracking method called Dynamic Domain Selection Graph DST (DDSG-DST). Specifically, (1) we employ Graphormer to establish hierarchical relationships between domains and slots; (2) we propose an additional domain prediction auxiliary task to predict the domain relevant to the dialogue context; (3) based on the predicted relevant domain from the auxiliary task, we dynamically select domain node information in the graph and perform dialogue state prediction. Experimental results demonstrate that we effectively establish hierarchical relationships between domains and slots, mitigate the negative impact of redundant connections in the graph structure, and enhance model performance.
Shuangyong Song, Hao Huang 0009, Hongyan Xie, Haoxiang Su, Mengxiang Li, Zhongjiang He, Ruiyu Fang
IJCNN3
2024 Cross-modal Features Interaction-and-Aggregation Network with Self-consistency Training for Speech Emotion Recognition
Ying Hu 0005, Hao Huang 0009, Liang He 0003
INTERSPEECH3
2024 An Uyghur Extension to the MASSIVE Multi-lingual Spoken Language Understanding Corpus with Comprehensive Evaluations
Ainikaerjiang Aimaiti, Di Wu 0088, Liting Jiang, Gulinigeer Abudouwaili, Hao Huang 0009, Wushour Slamu
INTERSPEECH5
2024 Synthesizing Long-Form Speech merely from Sentence-Level Corpus with Content Extrapolation and LLM Contextual Enrichment
Shijie Lai, Minglu He, Zijing Zhao 0008, Hao Huang 0009
INTERSPEECH5
2024 YOLOPitch: A Time-Frequency Dual-Branch YOLO Model for Pitch Estimation
Hao Huang 0009, Ying Hu 0005, Liang He 0003, Yuyi Wang 0007
INTERSPEECH2
2024 CEA-Net: a co-interactive external attention network for joint intent detection and slot filling
Di Wu 0088, Liting Jiang, Hao Huang 0009
Neural Comput. Appl.5
2024 IIFC-Net: A Monaural Speech Enhancement Network With High-Order Information Interaction and Feature Calibration
abstract
Recently, many Transformer-style dual-path models have achieved impressive performance for speech enhancement. However, their high parameters and computational complexity hinder their practical application. In this letter, we propose a monaural speech enhancement network with lower parameter count and complexity based on high-order information interaction and feature calibration (IIFC-Net). The network includes high-order information interaction Transformer (HOIIFormer) with high-order information interaction (HOII) block instead of a multi-head self-attention (MHSA) in Transformer. IIFC-Net leverages dual-path HOIIFormer (DPH) to model the distant dependency relation along time and frequency dimensions, respectively, and effectively captures deep-level information through the HOII block. We also design a feature calibration (FC) block to enhance the frequency components of target speech, which can be verified by a visualization analysis. The outcomes of experiments conducted on the VoiceBank+DEMAND and WHAMR! Datasets demonstrate that IIFC-Net achieves comparable performance in terms of denoising, dereverberation, and simultaneous denoising & dereverberation.
Wenbing Wei, Ying Hu 0005, Hao Huang 0009, Liang He 0003
IEEE Signal Process. Lett.3
2023 Hierarchical Softmax for End-To-End Low-Resource Multilingual Speech Recognition
abstract
Low-resource speech recognition has been long-suffering from insufficient training data. In this paper, we propose an approach that leverages neighboring languages to improve low-resource scenario performance, founded on the hypothesis that similar linguistic units in neighboring languages exhibit comparable term frequency distributions, which enables us to construct a Huffman tree for performing multilingual hierarchical Softmax decoding. This hierarchical structure enables cross-lingual knowledge sharing among similar tokens, thereby enhancing low-resource training outcomes. Empirical analyses demonstrate that our method is effective in improving the accuracy and efficiency of low-resource speech recognition.
Qianying Liu, Zhuo Gong, Zhengdong Yang, Sheng Li 0010, Chenchen Ding, Nobuaki Minematsu, Hao Huang 0009, Fei Cheng 0002, Chenhui Chu, Sadao Kurohashi
ICASSP8
2023 SRTNET: Time Domain Speech Enhancement via Stochastic Refinement
abstract
Diffusion model, as a new generative model which is very popular in image generation and audio synthesis, is rarely used in speech enhancement. In this paper, we use the diffusion model as a module for stochastic refinement. We propose SRTNet, a novel method for speech enhancement via Stochastic Refinement in complete Time b domain. Specifically, we design a joint network consisting of a deterministic module and a stochastic module, which makes up the "enhance-and-refine" paradigm. We theoretically demonstrate the feasibility of our method and experimentally prove that our method achieves faster training, faster sampling and higher quality. Our code is available at https://github.com/zhibinQiu/SRTNet.git
Zhibin Qiu, Mengfan Fu, Yinfeng Yu, Fuchun Sun 0001, Hao Huang 0009
ICASSP6
2023 Speakeraugment: Data Augmentation for Generalizable Source Separation via Speaker Parameter Manipulation
abstract
Existing speech separation models based on deep learning typically generalize poorly due to domain mismatch. In this paper, we propose SpeakerAugment (SA), a data augmentation method for generalizable speech separation that aims to increase the diversity of speaker identity in training data, to mitigate speaker mismatch of domain mismatch. The SA consists of two sub-policies: (1) SA-Vocoder, which uses a vocoder to manipulate pitch and formants parameters of speakers. (2) SA-Spectrum, which directly performs pitch-shift and time-stretch on the spectrum of each speech signal. The SA is simple and effective. Experimental results show that using SA can significantly improve the generalization ability of models, especially for: 1) The training set with fewer speakers, e.g., WSJ0-2mix, or 2) The target test set with complex linguistic conditions, e.g., the TIMIT based test set. Moreover, as a data augmentation method, SA has good potential to be applicable to other speech related tasks. We validate this by applying SA in speech recognition, and experimental results show that the generalization ability is also improved.
Hao Huang 0009, Ying Hu 0005, Sheng Li 0010
ICASSP3
2023 Speech-Text Based Multi-Modal Training with Bidirectional Attention for Improved Speech Recognition
abstract
To let the state-of-the-art end-to-end ASR model enjoy data efficiency, as well as much more unpaired text data by multi-modal training, one needs to address two problems: 1) the synchronicity of feature sampling rates between speech and language (aka text data); 2) the homogeneity of the learned representations from two encoders. In this paper we propose to employ a novel bidirectional attention mechanism (BiAM) to jointly learn both ASR encoder (bottom layers) and text encoder with a multi-modal learning method. The BiAM is to facilitate feature sampling rate exchange, realizing the quality of the transformed features for the one kind to be measured in another space, with diversified objective functions. As a result, the speech representations are enriched with more linguistic information, while the representations generated by the text encoder are more similar to corresponding speech ones, and therefore the shared ASR models are more amenable for unpaired text data pretraining. To validate the efficacy of the proposed method, we perform two categories of experiments with or without extra unpaired text data. Experimental results on Librispeech corpus show it can achieve up to 6.15% word error rate reduction (WERR) with only paired data learning, while 9.23% WERR when more unpaired text data is employed1.
Haihua Xu 0001, Hao Huang 0009, Chng Eng Siong, Sheng Li 0010
ICASSP3
2023 Mitigating Domain Dependency for Improved Speech Enhancement Via SNR Loss Boosting
abstract
Current supervised speech enhancement methods based on deep learning typically utilize amplitude-based loss functions for optimization, such as Mean Absolute Error (MAE) or Mean Square Error (MSE) loss, which measures the difference between the amplitudes of the estimated and clean speech signals. However, models trained with these losses heavily depend on specific domain properties, i.e. speaker, noise type, and signal-to-noise ratio (SNR). In this paper, we first validate this assumption by visually analyzing the model’s internal representation, and these dependencies result in severe performance degradation in unseen situations. Given that the SNR is irrelevant to speakers and noise types, we propose a simple but effective novel objective function by minimizing the discrepancy between indirectly estimated SNR and true SNR over time-frequency units to alleviate the model’s reliance on those domain properties. Experimental results demonstrate that our proposed method outperforms other prevalent loss functions in terms of both performance gain and generalization capability.
Di Wu 0088, Zhibin Qiu, Hao Huang 0009
ICASSP4
2023 A Joint Network Based on Interactive Attention for Speech Emotion Recognition
abstract
Speech emotion recognition (SER) has played a vital role in human-machine interaction. In this paper, we propose a separate spectrum-based SER model and a joint network combining pre-trained and spectrum-based models. In the joint network, we design an interactive attention module to effectively fuse the intermediate features from two models. Our proposed separate spectrum-based model is superior to four compared spectrum-based methods under the speaker-dependent setting. For the application in real scenarios, we compared our proposed joint network with six methods utilizing the pre-trained model under the speaker-independent setting. Experimental results show that our proposed joint network achieves the best performance among four unimodal models on the unweighted accuracy (UA) of 73.32 % and weighted accuracy (WA) of 72.48 %, respectively.
Ying Hu 0005, Shijing Hou, Hao Huang 0009, Liang He 0003
ICME4
2023 Speech Topic Classification Based on Pre-trained and Graph Networks
abstract
Speech Topic Classification (STC) automatically classifies audio clips into predefined categories, which is widely used in short video, personalized recommendation and other fields. At present, the common system is composed of two parts: first, the speech is converted into text by automatic speech recognition (ASR), and then the text topic is classified by natural language processing (NLP). Most of them have problems such as error propagation and lack of global structure. So in this paper, we propose a new end-to-end framework based on a pre-trained model and graph network. The pre-trained model is used to extract the semantic features with sequential structure instead of acoustic features, and the combination with the global features of conversational context constructed by graph network has achieved good results on the Fisher dataset.
Fangjing Niu, Ying Hu 0005, Hao Huang 0009, Liang He 0003
ICME4
2023 CRA-DIFFUSE: Improved Cross-Domain Speech Enhancement Based on Diffusion Model with T-F Domain Pre-Denoising
abstract
Speech enhancement (SE) methods in both the Time-Frequency (T-F) domain and time-domain domains have their own advantages. Leveraging both T-F domain and time-domain (cross domain) inputs has shown to be successful in the speech enhancement task. Recent SE methods based on diffusion models have shown promising results. However, little research effort has been made in the cross-domain speech enhancement using a diffusion model. We propose CRA-DiffuSE, a cross-domain SE model that uses a diffusion-based enhancement model as a refinement module after initial enhancement to achieve better results. For pre-enhance stage, we design CRANet, a T-F domain enhancement model combining channel attention and spatial attention. For the post-enhance stage, we design DiffuNet, a conditional generation model based on Denoising Diffusion Implicit Model (DDIM) for speech enhancement. Experiments demonstrate that the proposed CRA-DiffuSE is significantly superior to the baselines.
Zhibin Qiu, Yachao Guo, Mengfan Fu, Hao Huang 0009, Ying Hu 0005, Liang He 0003, Fuchun Sun 0001
ICME4
2023 Improved Keyword Recognition Based on Aho-Corasick Automaton
abstract
The recognition of out-of-vocabulary (OOV) words in many state-of-art automatic speech recognition (ASR) systems, which need the to recognize a word that has never been seen before during training (named entities mostly), is very challenging. Current keyword boosted beam search method solves this problem to some extent, However, underperformed on rare long-tail words and ignore the mismatch rate in the trie search. To reduce the mismatch, we propose a effective and fast method based on Aho-Corasick algorithm. So that the ASR can reduce the matching failure and be more efficient and fast in the decoding. We also assign different reward sizes to different lengths of OOV words in order to alleviate the insertion error of rare long-tail words that ASR can identify more easily in decoding. And we also improve the original Contextual Biasing method in Wenet, reduced the number of matching failures in the context graph, and increase the efficiency of matching. We also combine our proposed method with the internal language model estimation method, and experimentally demonstrate that our proposed method not only has a significant improvement in the performance of intra-domain OOV words recognition, but also improves the accuracy of OOV words in cross-domain speech recognition.
Yachao Guo, Zhibin Qiu, Hao Huang 0009, Chng Eng Siong
IJCNN3
2023 MTANet: Multi-band Time-frequency Attention Network for Singing Melody Extraction from Polyphonic Music
Ying Hu 0005, Liusong Wang, Hao Huang 0009, Liang He 0003
INTERSPEECH4
2023 Self-supervised Learning Representation based Accent Recognition with Persistent Accent Memory
Zhiwei Xie 0007, Haihua Xu 0001, Yizhou Peng, Hexin Liu, Hao Huang 0009, Chng Eng Siong
INTERSPEECH6
2023 Reprogramming Self-supervised Learning-based Speech Representations for Speaker Anonymization
abstract
Current speaker anonymization methods, especially with self-supervised learning (SSL) models, require massive computational resources when hiding speaker identity. This paper proposes an effective and parameter-efficient speaker anonymization method based on recent End-to-End model reprogramming technology. To improve the anonymization performance, we first extract speaker representation from large SSL models as the speaker identifies. To hide the speaker’s identity, we reprogram the speaker representation by adapting the speaker to a pseudo domain. Extensive experiments are carried out on the VoicePrivacy Challenge (VPC) 2022 datasets to demonstrate the effectiveness of our proposed parameter-efficient learning anonymization methods. Additionally, while achieving comparable performance with the VPC 2022 strong baseline 1.b, our approach also consumes less computational resources during anonymization.
Sheng Li 0010, Jiyi Li, Hao Huang 0009, Yang Cao 0011, Liang He 0003
MMAsia4
2023 GhostVec: A New Threat to Speaker Privacy of End-to-End Speech Recognition System
abstract
Speaker adaptation systems face privacy concerns, for such systems are trained on private datasets and often overfitting. This paper demonstrates that an attacker can extract speaker information by querying speaker-adapted speech recognition (ASR) systems. We focus on the speaker information of a transformer-based ASR and propose GhostVec, a simple and efficient attack method to extract the speaker information from an encoder-decoder-based ASR system without any external speaker verification system or natural human voice as a reference. To make our results quantitative, we pre-process GhostVec using singular value decomposition (SVD) and synthesize it into waveform. Experiment results show that the synthesized audio of GhostVec reaches 10.83% EER and 0.47 minDCF with target speakers, which suggests the effectiveness of the proposed method. We hope the preliminary discovery in this study to catalyze future speech recognition research on privacy-preserving topics.
Sheng Li 0010, Jiyi Li, Yang Cao 0011, Hao Huang 0009, Liang He 0003
MMAsia5
2022 Minimum Word Error Training For Non-Autoregressive Transformer-Based Code-Switching ASR
abstract
Non-autoregressive end-to-end ASR framework might be potentially appropriate for code-switching recognition task thanks to its inherent property that present output token being independent of historical ones. However, it still under-performs the state-of-the-art autoregressive ASR frameworks. In this paper, we propose various approaches to boosting the performance of a CTC-mask-based non-autoregressive Transformer under code-switching ASR scenario. To begin with, we attempt diversified masking method that are closely related with code-switching point, yielding an improved baseline model. More importantly, we employ Minimum Word Error (MWE) criterion to train the model. One of the challenges is how to generate a diversified hypothetical space, so as to obtain the average loss for a given ground truth. To address such a challenge, we explore different approaches to yielding desired N-best-based hypothetical space. We demonstrate the efficacy of the proposed methods on SEAME corpus, a challenging English-Mandarin code-switching corpus for Southeast Asia community. Compared with the cross-entropy-trained strong baseline, the proposed MWE training method achieves consistent performance improvement on the test sets.
Yizhou Peng, Haihua Xu 0001, Hao Huang 0009, Chng Eng Siong
ICASSP4
2022 Mining Hard Samples Locally And Globally For Improved Speech Separation
abstract
Speech separation dataset typically consists of hard and non-hard samples, and the former is minority and latter majority. The data imbalance problem biases the model towards non-hard samples and weakens the generalization capability. Given that the average separation performance is sufficiently good, improving hard samples may contribute more to back-end tasks. In this paper, we propose two methods to alleviate data imbalance in speech separation task, based on local and global hard sample mining. For the local, we propose weighted loss to compensate for hard samples by increasing their weights in each batch. For the global, we perform global hard sample mining and re-sample to increase the proportion of hard samples in the training set. Because hard sample mining using objective loss in dynamic mixing leads to local results, we propose an indirect method using speaker-specific parameters, based on the fact that pitch median difference and x-vector cosine distance of two speakers in a mixture are closely correlated with separation SI-SNRi. Experimental results show that both methods decrease the percentage of hard samples in the test set than using dynamic mixing only while keeping the average SI-SNRi comparable, and the global method shows more promising results than the local one.
Yizhou Peng, Hao Huang 0009, Ying Hu 0005, Sheng Li 0010
ICASSP3
2022 GhostVec: Directly Extracting Speaker Embedding from End-to-End Speech Recognition Model Using Adversarial Examples
Sheng Li 0010, Hao Huang 0009
ICONIP (6)3
2022 Investigating Effective Domain Adaptation Method for Speaker Verification Task
Guangxing Li, Wangjin Zhou, Sheng Li 0010, Yi Zhao 0006, Hao Huang 0009
ICONIP (6)6
2022 A Graph Isomorphism Network with Weighted Multiple Aggregators for Speech Emotion Recognition
abstract
Speech emotion recognition (SER) is an essential part of human-computer interaction. In this paper, we propose an SER network based on a Graph Isomorphism Network with Weighted Multiple Aggregators (WMA-GIN), which can effectively handle the problem of information confusion when neighbour nodes' features are aggregated together in GIN structure. Moreover, a Full-Adjacent (FA) layer is adopted for alleviating the over-squashing problem, which is existed in all Graph Neural Network (GNN) structures, including GIN. Furthermore, a multi-phase attention mechanism and multi-loss training strategy are employed to avoid missing the useful emotional information in the stacked WMA-GIN layers. We evaluated the performance of our proposed WMA-GIN on the popular IEMOCAP dataset. The experimental results show that WMA-GIN outperforms other GNN-based methods and is comparable to some advanced non-graph-based methods by achieving 72.48% of weighted accuracy (WA) and 67.72% of unweighted accuracy (UA).
Ying Hu 0005, Yuwu Tang, Hao Huang 0009, Liang He 0003
INTERSPEECH3
2022 A Multi-grained based Attention Network for Semi-supervised Sound Event Detection
abstract
Sound event detection (SED) is an interesting but challenging task due to the scarcity of data and diverse sound events in real life. This paper presents a multi-grained based attention network (MGA-Net) for semi-supervised sound event detection. To obtain the feature representations related to sound events, a residual hybrid convolution (RH-Conv) block is designed to boost the vanilla convolution's ability to extract the time-frequency features. Moreover, a multi-grained attention (MGA) module is designed to learn temporal resolution features from coarse-level to fine-level. With the MGA module,the network could capture the characteristics of target events with short- or long-duration, resulting in more accurately determining the onset and offset of sound events. Furthermore, to effectively boost the performance of the Mean Teacher (MT) method, a spatial shift (SS) module as a data perturbation mechanism is introduced to increase the diversity of data. Experimental results show that the MGA-Net outperforms the published state-of-the-art competitors, achieving 53.27% and 56.96% event-based macro F1 (EB-F1) score, 0.709 and 0.739 polyphonic sound detection score (PSDS) on the validation and public set respectively.
Ying Hu 0005, Xiujuan Zhu, Hao Huang 0009, Liang He 0003
INTERSPEECH4
2022 Multi-stage music separation network with dual-branch attention and hybrid convolution
Yadong Chen 0003, Ying Hu 0005, Liang He 0003, Hao Huang 0009
J. Intell. Inf. Syst.4
2022 A bimodal network based on Audio-Text-Interactional-Attention with ArcFace loss for speech emotion recognition
Yuwu Tang, Ying Hu 0005, Liang He 0003, Hao Huang 0009
Speech Commun.4
2022 Hierarchic Temporal Convolutional Network With Cross-Domain Encoder for Music Source Separation
abstract
Recently, the time-domain-based methods (i.e., the method of modeling the raw waveform directly) for audio source separation have shown tremendous potential. In this paper, we propose a model which combines the complexed spectrogram domain feature and time-domain feature by a cross-domain encoder (CDE) and adopts the hierarchic temporal convolutional network (HTCN) for multiple music sources separation. The CDE is designed to enable the network to code the interactive information of the time-domain and complexed spectrogram domain features. HTCN enables it to learn the long-time series dependence effectively. We also designed a feature calibration unit (FCU) to be applied in the HTCN and adopted the multi-stage training strategy during the training stage. The ablation study demonstrates the effectiveness of each designed component in the model. We conducted the experiments on the MUSDB18 dataset. The experimental results indicate that our proposed CDE-HTCN model outperforms the top-of-the-line methods and, compared with the state-of-the-art method, DEMUCS, achieves the improvement of the average SDR score of 0.61 dB. Significantly, the improvement of the SDR score for the$\ bass$source has a sizable margin of 0.91 dB.
Ying Hu 0005, Yadong Chen 0003, Wenzhong Yang, Liang He 0003, Hao Huang 0009
IEEE Signal Process. Lett.5
2021 Encoder-Decoder Based Pitch Tracking and Joint Model Training for Mandarin Tone Classification
abstract
We pursue an interpretable pitch tracking model and a jointly trained tone model for Mandarin tone classification. For pitch tracking, present deep learning based pitch model structure seldom considers the Viterbi decoding commonly implemented in prevalent manually designed pitch tracking algorithms. We propose RNN based Encoder-Decoder framework with gating mechanism which underlying models both the state cost estimation and Viterbi back-tracing pass implemented in the RAPT algorithm. Then we apply the pitch extractor to a down-stream Mandarin tone classification task. The basic motivation is to combine together the two conventional components in tone classification (i.e., the pitch extractor and tone classifier) and then the whole network are trained simultaneously in an end-to-end fashion. Various cascade methods are evaluated. We carry out pitch extraction and tone classification experiments on Mandarin continuous speech database to show the superiority of the proposed models. Experimental results on pitch extraction show proposed pitch tracking model outperforms the DNN-RNN and bi-directional variants. Tone classification experimental results show the composite model outperforms the traditional cascade tone classification framework which makes use of pitch related feature and a back-end classifier.
Hao Huang 0009, Ying Hu 0005, Sheng Li 0010
ICASSP1
2021 End-to-End Speech Separation Using Orthogonal Representation in Complex and Real Time-Frequency Domain
Hao Huang 0009, Ying Hu 0005, Sheng Li 0010
Interspeech2
2021 E2E-Based Multi-Task Learning Approach to Joint Speech and Accent Recognition
abstract
In this paper, we propose a single multi-task learning framework to perform End-to-End (E2E) speech recognition (ASR) and accent recognition (AR) simultaneously.The proposed framework is not only more compact but can also yield comparable or even better results than standalone systems.Specifically, we found that the overall performance is predominantly determined by the ASR task, and the E2E-based ASR pretraining is essential to achieve improved performance, particularly for the AR task.Additionally, we conduct several analyses of the proposed method.First, though the objective loss for the AR task is much smaller compared with its counterpart of ASR task, a smaller weighting factor with the AR task in the joint objective function is necessary to yield better results for each task.Second, we found that sharing only a few layers of the encoder yields better AR results than sharing the overall encoder.Experimentally, the proposed method produces WER results close to the best standalone E2E ASR ones, while it achieves 7.7% and 4.2% relative improvement over standalone and single-task-based joint recognition methods on test set for accent recognition respectively.
Yizhou Peng, Van Tung Pham, Haihua Xu 0001, Hao Huang 0009, Chng Eng Siong
Interspeech5
2020 Monolingual Data Selection Analysis for English-Mandarin Hybrid Code-Switching Speech Recognition
abstract
In this paper, we conduct data selection analysis in building an English-Mandarin code-switching (CS) speech recognition (CSSR) system, which is aimed for a real CSSR contest in China.The overall training sets have three subsets, i.e., a codeswitching data set, an English (LibriSpeech) and a Mandarin data set respectively.The code-switching data are Mandarin dominated.First of all, it is found using the overall data yields worse results, and hence data selection study is necessary.Then to exploit monolingual data, we find data matching is crucial.Mandarin data is closely matched with the Mandarin part in the code-switching data, while English data is not.However, Mandarin data only helps on those utterances that are significantly Mandarin-dominated.Besides, there is a balance point, over which more monolingual data will divert the CSSR system, degrading results.Finally, we analyze the effectiveness of combining monolingual data to train a CSSR system with the HMM-DNN hybrid framework.The CSSR system can perform within-utterance code-switch recognition, but it still has a margin with the one trained on code-switching data.
Haihua Xu 0001, Van Tung Pham, Hao Huang 0009, Chng Eng Siong
INTERSPEECH4
2020 A Lightweight Model Based on Separable Convolution for Speech Emotion Recognition
Ying Hu 0005, Hao Huang 0009, Wushour Slamu
INTERSPEECH3
2016 Monaural Singing Voice Separation by Non-negative Matrix Partial Co-Factorization with Temporal Continuity and Sparsity Criteria
Ying Hu 0005, Hao Huang 0009
ICIC (3)3
2016 Semi-Supervised and Cross-Lingual Knowledge Transfer Learnings for DNN Hybrid Acoustic Models Under Low-Resource Conditions
Haihua Xu 0001, Chongjia Ni, Hao Huang 0009, Chng Eng Siong, Haizhou Li 0001
INTERSPEECH5
2015 Maximum F1-Score Discriminative Training Criterion for Automatic Mispronunciation Detection
abstract
We carry out an in-depth investigation on a newly proposed Maximum F1-score Criterion (MFC) discriminative training objective function for Goodness of Pronunciation (GOP) based automatic mispronunciation detection that makes use of Gaussian Mixture Model-hidden Markov model (GMM-HMM) as acoustic models. The formulation of MFC seeks to directly optimize F1-score by converting the non-differentiable F1-score function into a continuous objective function to facilitate optimization. We present model-space training algorithm according to MFC using extended Baum–Welch form like update equations based on the weak-sense auxiliary function method. We then present MFC based feature-space discriminative training. We train a matrix projecting from posteriors of Gaussians to a normal size feature space, and add the projected features to traditional spectral features. Mispronunciation detection experiments show MFC based model-space training and feature-space training are effective in improving F1-score and other commonly used evaluation metrics. It is also shown MFC training in both the feature-space and model-space outperforms either model-space training or feature-space training alone, and is about 11.6% better than the maximum likelihood (ML) trained baseline in terms of F1-score. Further, we review and compare mispronunciation detection results with the use of MFC and some traditional training criteria that minimize word error rate in speech recognition. The experimental analysis and comparison provide useful insight into the correlations between F1-score maximization and optimization of these training criteria.
Hao Huang 0009, Haihua Xu 0001, Wushour Slamu
IEEE ACM Trans. Audio Speech Lang. Process.1
2009 Minimum tag error for discriminative training of conditional random fields
Jie Zhu 0006, Hao Huang 0009, Haihua Xu 0001
Inf. Sci.3