Yanhua Long

dblp:64/8052 · DBLP profile ↗
← Back
34ranked-venue papers
8as first author
21since 2021 · last 2027
0000-0003-0924-408XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 7 first-author · 20 since 2021Artificial intelligence and machine learning · 16 · 5 first-author · 9 since 2021
YearPublicationVenuePosition
2027 Lightweight speech enhancement guided target speech extraction in noisy scenarios
Ziling Huang, Junnan Wu, Lichun Fan, Haixin Guan, Yanhua Long
Comput. Speech Lang.5
2026 PRSE: A two-stage joint optimization approach for lightweight speech enhancement
Haixin Guan, Guanyong Wang, Yanhua Long, Jiaen Liang, Xiaobin Tan
Speech Commun.3
2025 SEF-PNet: Speaker Encoder-Free Personalized Speech Enhancement with Local and Global Contexts Aggregation
abstract
Personalized speech enhancement (PSE) methods typically rely on pre-trained speaker verification models or self-designed speaker encoders to extract target speaker clues, guiding the PSE model in isolating the desired speech. However, these approaches suffer from significant model complexity and often underutilize enrollment speaker information, limiting the potential performance of the PSE model. To address these limitations, we propose a novel Speaker Encoder-Free PSE network, termed SEF-PNet, which fully exploits the information present in both the enrollment speech and noisy mixtures. SEF-PNet incorporates two key innovations: Interactive Speaker Adaptation (ISA) and Local-Global Context Aggregation (LCA). ISA dynamically modulates the interactions between enrollment and noisy signals to enhance the speaker adaptation, while LCA employs advanced channel attention within the PSE encoder to effectively integrate local and global contextual information, thus improving feature learning. Experiments on the Libri2Mix dataset demonstrate that SEF-PNet significantly outperforms baseline models, achieving state-of-the-art PSE performance. Our source code is available at https://github.com/isHuangZiling/SEF-PNet.
Ziling Huang, Haixin Guan, Yanhua Long
ICASSP4
2025 Leveraging Out-of-Domain Noise for Unsupervised Domain Adaptation in Speech Enhancement
abstract
When there’s a mismatch between the training and test domains, supervised speech enhancement (SE) models trained on synthetic paired noisy-clean data often struggle in real-world scenarios, highlighting the industry’s strong demand for unsupervised training and domain adaptation methods. In this study, we introduce PHA-ReMixIT, a novel approach for leveraging out-of-domain (OOD) noise signals to enhance unsupervised domain adaptation in SE. Our method builds upon the state-of-the-art ReMixIT by introducing a paired unsupervised remixing technique, which augments the diversity of target domain training data with OOD noise signals. We further propose a heterogeneous noise invariant training to align the OOD augmented noisy mixtures with their paired heterogeneous counterparts, encouraging the model to output cleaner speech. Additionally, an adaptive focal weighting mechanism is also introduced to dynamically emphasize the data importance of both in-domain and OOD noisy mixtures during model adaptation. Experiments on CHiME-7 unsupervised domain adaptation for conversational speech enhancement (UDASE) task demonstrate that PHA-ReMixIT significantly outperforms the ReMixIT baseline, boosting SE performance on both real and synthesized test sets.
Yu Liao, Haixin Guan, Yanhua Long
ICASSP4
2025 Personalized Speech Enhancement without User Enrollment for Real-World Audio Replay Scenarios
abstract
Many speech enhancement (SE) approaches have been proposed to deal with cocktail party problem. Personalized speech enhancement (PSE) approaches improve SE performance by utilizing user enrollment speech. However, PSE requires users to record additional clean audio for registration, which can be redundant works or impractical for many real-world scenarios. For instance, in personal devices and Vloggers’ audio playback scenarios, there are already many video/audios available recorded under different noise types and SNR conditions, but without any target speaker pre-registered speech. To better utilize information from existing video/audio stock, this paper propose a novel speech enhancement approach that integrates PSE methods without requiring pre-registered user speech. With user adaptation and noise adaptation training modules, the proposed approach automatically selects high-quality speech segments to assist in denoising low-quality speech segments. Additionally, two test sets were collected to evaluate the performance in the aforementioned scenarios. Experimental results demonstrate that the proposed approach outperforms the corresponding SE methods in both objective and subjective evaluation metrics.
Shi-Lin Wang, Yanhua Long
ICASSP3
2025 Enhanced cross-modal parallel training for improving end-to-end accented speech recognition
Renchang Dong, Yanhua Long, Dongxing Xu
Speech Commun.3
2024 Cross-Modal Parallel Training for Improving end-to-end Accented Speech Recognition
abstract
Multi-accent speech recognition is a key challenge in current speech recognition due to the pronunciation variations of different accents. In this study, we propose a Cross-modal Parallel Training (CPT) approach for improving the accent robustness of state-of-the-art Conformer-Transducer (Conformer-T) ASR system. Specifically, in CPT, a novel cross-modal attention and fusion module is first designed as a frontend to align the low-level acoustic speech representations with phonetic embeddings, and thus normalizing accent variations into a shared standard pronunciation latent space; Then, a parallel training mechanism is proposed to simultaneously model both the acoustic and accent normalized multi-modal features for improving accented ASR performance. Moreover, different multi-objective training losses with text-induced and phonetic target units are also investigated. Our experiments are performed on the public CommonVoice English accented ASR tasks, results show that the proposed CPT outperforms the strong baseline by relative 9.3%-13.4% WER reductions on six evaluation sets, all without increasing any model parameters or computational costs during ASR inference.
Renchang Dong, Yijie Li 0001, Dongxing Xu, Yanhua Long
ICASSP4
2024 Accent-Specific Vector Quantization for Joint Unsupervised and Supervised Training in Accent Robust Speech Recognition
abstract
How to effectively use limited supervised accent data to improve the accented ASR is of paramount importance. In this work, we propose an accent-specific quantization for joint unsupervised and supervised training (AQ-JUST) of end-to-end ASR models to address this issue. Specifically, two variants of AQ-JUST are investigated, namely BAQ-JUST and SAQ-JUST, by employing different model structures and training methods to capture the distinctiveness and commonalities between diverse accents, thus enhancing the performance of accented ASR systems. Our experiments are performed on both accented English and Mandarin ASR tasks. Results show that the proposed methods outperform the strong JUST baseline by relative 3.9% to 9.4% word/character error rate reductions on accented test sets.
Yijie Li 0001, Dongxing Xu, Yanhua Long
ICASSP5
2024 Score Calibration Based on Consistency Measure Factor for Speaker Verification
abstract
This paper proposes a new scoring calibration method named "Consistency-Aware Score Calibration", which introduces a Consistency Measure Factor (CMF) to measure the stability of audio voiceprints in similarity scores for speaker verification. The CMF is inspired by the limitations in segment scoring, where the segments with shorter length are not friendly to calculate the similarity score. By using CMF as a scale to calibrate scores calculated on the whole audio length, the method improves the performance of different state-of-the-art systems significantly, including ResNet with 34 to 518 layers and RepVGG. Experimental results show that the CMF method is better than segment scoring and shows excellent complementary information with other normalization or calibration methods. The proposed method was first proposed in a system description for the VoxCeleb speaker recognition challenge 2023, where it achieved the 1st place in Track1 of the challenge.
Chuanying Niu, Yibin Zhan, Yanhua Long, Dongxing Xu
ICASSP5
2024 QMixCAT: Unsupervised Speech Enhancement Using Quality-guided Signal Mixing and Competitive Alternating Model Training
Shi-Lin Wang, Haixin Guan, Yanhua Long
INTERSPEECH3
2023 FEW-Shot Continual Learning with Weight Alignment and Positive Enhancement for Bioacoustic Event Detection
abstract
In this paper, we propose a new continual learning framework for few-shot bioacoustic event detection (BED). First, we modify the recently proposed dynamic few-shot learning (DFSL) and generalize it to the BED task. Then, we introduce a weight alignment loss to enhance the weight generator of modified DFSL for detecting novel events. Furthermore, to augment the few positive samples of each target bioacoustic event, a positive enhancement approach is proposed to select high-confidence pseudo positives using the cumulative distribution of initial detection posterior probabilities. All experiments are performed on DCASE 2022 Task5 Challenge dataset, results show that the proposed methods significantly outperform the prototypical network (PN) baseline, it brings the overall F-measure of validation set from 52.2% to 55.9%. Moreover, the proposed framework shows great complementarity with the conventional PN, the F-measure is improved to 60.6% after applying a simple score fusion.
Sissi Xiaoxiao Wu, Dongxing Xu, Yanhua Long
ICASSP4
2023 Advanced RawNet2 with Attention-based Channel Masking for Synthetic Speech Detection
Yanhua Long, Yijie Li 0001, Dongxing Xu
INTERSPEECH2
2023 Phonetic-assisted Multi-Target Units Modeling for Improving Conformer-Transducer ASR system
Dongxing Xu, Yanhua Long
INTERSPEECH4
2023 Multi-pass Training and Cross-information Fusion for Low-resource End-to-end Accented Speech Recognition
Yanhua Long, Yijie Li 0001
INTERSPEECH2
2023 Dual-model self-regularization and fusion for domain adaptation of robust speaker verification
Yibo Duan, Yanhua Long, Jiaen Liang
Speech Commun.2
2022 DPCCN: Densely-Connected Pyramid Complex Convolutional Network for Robust Speech Separation and Extraction
abstract
In recent years, a number of time-domain speech separation methods have been proposed. However, most of them are very sensitive to the environments and wide domain coverage tasks. In this paper, from the time-frequency domain perspective, we propose a densely-connected pyramid complex convolutional network, termed DPCCN, to improve the robustness of speech separation under complicated conditions. Furthermore, we generalize the DPCCN to target speech extraction (TSE) by integrating a new specially designed speaker encoder. Moreover, we also investigate the robustness of DPCCN to unsupervised cross-domain TSE tasks. A Mixture-Remix approach is proposed to adapt the target domain acoustic characteristics for fine-tuning the source model. We evaluate the proposed methods not only under noisy and reverberant in-domain condition, but also in clean but cross-domain conditions. Results show that for both speech separation and extraction, the DPCCN-based systems achieve significantly better performance and robustness than the currently dominating time-domain methods, especially for the cross-domain tasks. Particularly, we find that the Mixture-Remix fine-tuning with DPCCN significantly outperforms the TD-SpeakerBeam for unsupervised cross-domain TSE, with around 3.5 dB SISNR improvement on target domain test set, without any source domain performance degradation.
Jiangyu Han, Yanhua Long, Lukás Burget, Jan Cernocký
ICASSP2
2022 PercepNet+: A Phase and SNR Aware PercepNet for Real-Time Speech Enhancement
Xiaofeng Ge, Jiangyu Han, Yanhua Long, Haixin Guan
INTERSPEECH3
2022 Selective Pseudo-labeling and Class-wise Discriminative Fusion for Sound Event Detection
abstract
In recent years, exploring effective sound separation (SSep) techniques to improve overlapping sound event detection (SED) attracts more and more attention.Creating accurate separation signals to avoid the catastrophic error accumulation during SED model training is very important and challenging.In this study, we first propose a novel selective pseudo-labeling approach, termed SPL, to produce high confidence separated target events from blind sound separation outputs.These target events are then used to fine-tune the original SED model that pre-trained on the sound mixtures in a multi-objective learning style.Then, to further leverage the SSep outputs, a class-wise discriminative fusion is proposed to improve the final SED performances, by combining multiple frame-level event predictions of both sound mixtures and their separated signals.All experiments are performed on the public DCASE 2021 Task 4 dataset, and results show that our approaches significantly outperforms the official baseline, the collar-based F 1, PSDS1 and PSDS2 performances are improved from 44.3%, 37.3% and 54.9% to 46.5%, 44.5% and 75.4%, respectively.
Yunhao Liang, Yanhua Long, Yijie Li 0001, Jiaen Liang
INTERSPEECH2
2021 Attention-Based Scaling Adaptation for Target Speech Extraction
abstract
The target speech extraction has attracted widespread attention in recent years. In this work, we focus on investigating the dynamic interaction between different mixtures and the target speaker to exploit the discriminative target speaker clues. We propose a special attention mechanism without introducing any additional parameters in a scaling adaptation layer to better adapt the network towards extracting the target speech. Furthermore, by introducing a mixture embedding matrix pooling method, our proposed attention-based scaling adaptation (ASA) can exploit the target speaker clues in a more efficient way. Experimental results on the spatialized reverberant WSJ0 2-mix dataset demonstrate that the proposed method can improve the performance of the target speech extraction effectively. Furthermore, we find that under the same network configurations, the ASA in a single-channel condition can achieve competitive performance gains as that achieved from two-channel mixtures with inter-microphone phase difference (IPD) features.
Jiangyu Han, Yanhua Long, Jiaen Liang
ASRU3
2021 Multi-Channel Target Speech Extraction with Channel Decorrelation and Target Speaker Adaptation
abstract
The end-to-end approaches for single-channel target speech extraction have attracted widespread attention. However, the studies for end-to-end multi-channel target speech extraction are still relatively limited. In this work, we propose two methods for exploiting the multi-channel spatial information to extract the target speech. The first one is using a target speech adaptation layer in a parallel encoder architecture. The second one is designing a channel decorrelation mechanism to extract the inter-channel differential information to enhance the multi-channel encoder representation. We compare the proposed methods with two strong state-of-the-art baselines. Experimental results on the multi-channel reverberant WSJ0 2-mix dataset demonstrate that our proposed methods achieve up to 11.2% and 11.5% relative improvements in SDR and SiSDR respectively, which are the best reported results on this task to the best of our knowledge.
Jiangyu Han, Xinyuan Zhou, Yanhua Long, Yijie Li 0001
ICASSP3
2021 Improving Channel Decorrelation for Multi-Channel Target Speech Extraction
abstract
Target speech extraction has attracted widespread attention. When microphone arrays are available, the additional spatial information can be helpful in extracting the target speech. We have recently proposed a channel decorrelation (CD) mechanism to extract the inter-channel differential information to enhance the reference channel encoder representation. Although the proposed mechanism has shown promising results for extracting the target speech from mixtures, the extraction performance is still limited by the nature of the original decorrelation theory. In this paper, we propose two methods to broaden the horizon of the original channel decorrelation, by replacing the original softmax-based inter-channel similarity between encoder representations, using an unrolled probability and a normalized cosine-based similarity at the dimensional-level. Moreover, new combination strategies of the CD-based spatial information and target speaker adaptation of parallel encoder outputs are also investigated. Experiments on the reverberant WSJ0 2-mix show that the improved CD can result in more discriminative differential information and the new adaptation strategy is also very effective to improve the target speech extraction.
Jiangyu Han, Wei Rao 0002, Yannan Wang, Yanhua Long
Interspeech4
2020 Self-and-Mixed Attention Decoder with Deep Acoustic Structure for Transformer-Based LVCSR
abstract
Transformer has shown impressive performance in automatic speech recognition.It uses an encoder-decoder structure with self-attention to learn the relationship between high-level representation of source inputs and embedding of target outputs.In this paper, we propose a novel decoder structure that features a self-and-mixed attention decoder (SMAD) with a deep acoustic structure (DAS) to improve the acoustic representation of Transformer-based LVCSR.Specifically, we introduce a self-attention mechanism to learn a multi-layer deep acoustic structure for multiple levels of acoustic abstraction.We also design a mixed attention mechanism that learns the alignment between different levels of acoustic abstraction and its corresponding linguistic information simultaneously in a shared embedding space.The ASR experiments on Aishell-1 show that the proposed structure achieves CERs of 4.8% on the dev set and 5.1% on the test set, which are the best reported results on this task to the best of our knowledge.
Xinyuan Zhou, Grandee Lee, Emre Yilmaz 0001, Yanhua Long, Jiaen Liang, Haizhou Li 0001
INTERSPEECH4
2020 Multi-Encoder-Decoder Transformer for Code-Switching Speech Recognition
abstract
Code-switching (CS) occurs when a speaker alternates words of two or more languages within a single sentence or across sentences.Automatic speech recognition (ASR) of CS speech has to deal with two or more languages at the same time.In this study, we propose a Transformer-based architecture with two symmetric language-specific encoders to capture the individual language attributes, that improve the acoustic representation of each language.These representations are combined using a language-specific multi-head attention mechanism in the decoder module.Each encoder and its corresponding attention module in the decoder are pre-trained using a large monolingual corpus aiming to alleviate the impact of limited CS training data.We call such a network a multi-encoder-decoder (MED) architecture.Experiments on the SEAME corpus show that the proposed MED architecture achieves 10.2% and 10.8% relative error rate reduction on the CS evaluation sets with Mandarin and English as the matrix language respectively.
Xinyuan Zhou, Emre Yilmaz 0001, Yanhua Long, Yijie Li 0001, Haizhou Li 0001
INTERSPEECH3
2018 Active Learning for LF-MMI Trained Neural Networks in ASR
Yanhua Long, Yijie Li 0001, Jiaen Liang
INTERSPEECH1
2018 Offline to online speaker adaptation for real-time deep neural network based LVCSR systems
Yanhua Long, Yijie Li 0001
Multim. Tools Appl.1
2017 Domain compensation based on phonetically discriminative features for speaker verification
Yanhua Long, Jifeng Ni
Comput. Speech Lang.1
2013 Improving lightly supervised training for broadcast transcription
abstract
This paper investigates improving lightly supervised acoustic model training for an archive of broadcast data. Standard lightly supervised training uses automatically derived decoding hypotheses using a biased language model. However, as the actual speech can deviate significantly from the original programme scripts that are supplied, the quality of standard lightly supervised hypotheses can be poor. To address this issue, word and segment level combination approaches are used between the lightly supervised transcripts and the original programme scripts which yield improved transcriptions. Experimental results show that systems trained using these improved transcriptions consistently outperform those trained using only the original lightly supervised decoding hypotheses. This is shown to be the case for both the maximum likelihood and minimum phone error trained systems.
Yanhua Long, Mark J. F. Gales, Pierre Lanchantin, Xunying Liu, Matthew Stephen Seigel, Philip C. Woodland
INTERSPEECH1
2012 Transcription of multi-genre media archives using out-of-domain data
abstract
We describe our work on developing a speech recognition system for multi-genre media archives. The high diversity of the data makes this a challenging recognition task, which may benefit from systems trained on a combination of in-domain and out-of-domain data. Working with tandem HMMs, we present Multi-level Adaptive Networks (MLAN), a novel technique for incorporating information from out-of-domain posterior features using deep neural networks. We show that it provides a substantial reduction in WER over other systems, with relative WER reductions of 15% over a PLP baseline, 9% over in-domain tandem features and 8% over the best out-of-domain tandem features.
Peter Bell 0001, Mark J. F. Gales, Pierre Lanchantin, Xunying Liu, Yanhua Long, Steve Renals, Pawel Swietojanski, Philip C. Woodland
SLT5
2011 Speaker characterization using spectral subband energy ratio based on Harmonic plus Noise Model
abstract
This paper proposes a feature extraction for speaker characterization by exploring the relationship between the two distinct components of the speech signal, one is harmonics accounting for the periodicity of the signal and the other is modulated noise accounting for the turbulences of the glottal airflow. The harmonic and noise parts of the speech signal are decomposed based on the Harmonic plus Noise Model approach. We estimate the spectral subband energy ratios (SSERs) as the speaker characteristic features, which are expected to reflect the interaction property of the vocal tract and glottal airflow of individual speakers for speaker verification. The speaker verification experiments based on a GMM-UBM system have shown the efficiency of the SSER features, reducing the error equal rate by 27.2% by combining with the conventional MFCC features.
Yanhua Long, Zhijie Yan, Frank K. Soong, Li-Rong Dai 0001, Wu Guo
ICASSP1
2011 Improvements in Speaker Characterization Using Spectral Subband Energy Based on Harmonic plus Noise Model
Yanhua Long, Zhijie Yan, Frank K. Soong, Li-Rong Dai 0001, Wu Guo
INTERSPEECH1
2010 N-gram nearest neighbor algorithm for voice password system
abstract
A specific issue in the voice password system is addressed in this paper: When the text content of target speaker's enrollment password has been already known by imposters, they can do a well-behaved impersonation using the same text content as the target speaker. This results in a much higher false acceptance than the traditional voice password system. N-gram based nearest neighbor algorithm is proposed here to improve the speaker detection accuracy. Furthermore, correlation coefficient is adopted as the distance measurement between two acoustic features instead of the traditional Euclidean distance. Experimental results show that the proposed method outperforms the DTW and GMM-UBM algorithms.
Wu Guo, Yanhua Long, Li-Rong Dai 0001
ICASSP3
2010 Effects of the phonological relevance in speaker verification
Yanhua Long, Li-Rong Dai 0001, Bin Ma 0001, Wu Guo
INTERSPEECH1
2009 iFLY system for the NIST 2008 speaker recognition evaluation
abstract
The description of iFLY system submitted for NIST 2008 speaker recognition evaluation (SRE), which has achieved excellent performance in the 2008 SRE evaluation, is presented in this paper. Our primary system is a fusion of two subsystems GMM-UBM and GMM-SVM. For each sub-system, two kinds of short-time acoustic features PLP and LPCC are adopted. We focus on three key issues in this evaluation: channel compensation, multi-lingual or bi-lingual cues and the voice activity detection. We also point out that data selection and factor analysis play key roles in the system improvement.
Wu Guo, Yanhua Long, Yijie Li 0001, Eryu Wang, Li-Rong Dai 0001
ICASSP2
2009 Exploiting prosodic information for Speaker Recognition
abstract
In this paper, we study speaker characterization using prosodic supervectors with negative within-class covariance normalization (NWCCN) projection and speaker modeling with support vector regression (SVR). We also propose a segmental weight fusion (SWF) technique that combines acoustic and prosodic subsystems effectively, despite the big performance gap between the subsystems. We validate the effectiveness of our proposed techniques on the NIST 2006 Speaker Recognition Evaluation (SRE) in comparison with other prominent solutions. The experiments have reported competitive results of 17.72% Equal Error Rate for the prosodic subsystem alone and 4.50% for the fusion system on NIST 2006 SRE core test condition.
Yanhua Long, Bin Ma 0001, Haizhou Li 0001, Wu Guo, Chng Eng Siong, Li-Rong Dai 0001
ICASSP1