Dongxing Xu

dblp:94/5872 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0001-7445-1398ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021
YearPublicationVenuePosition
2025 Enhanced cross-modal parallel training for improving end-to-end accented speech recognition
Renchang Dong, Yanhua Long, Dongxing Xu
Speech Commun.5
2024 Cross-Modal Parallel Training for Improving end-to-end Accented Speech Recognition
abstract
Multi-accent speech recognition is a key challenge in current speech recognition due to the pronunciation variations of different accents. In this study, we propose a Cross-modal Parallel Training (CPT) approach for improving the accent robustness of state-of-the-art Conformer-Transducer (Conformer-T) ASR system. Specifically, in CPT, a novel cross-modal attention and fusion module is first designed as a frontend to align the low-level acoustic speech representations with phonetic embeddings, and thus normalizing accent variations into a shared standard pronunciation latent space; Then, a parallel training mechanism is proposed to simultaneously model both the acoustic and accent normalized multi-modal features for improving accented ASR performance. Moreover, different multi-objective training losses with text-induced and phonetic target units are also investigated. Our experiments are performed on the public CommonVoice English accented ASR tasks, results show that the proposed CPT outperforms the strong baseline by relative 9.3%-13.4% WER reductions on six evaluation sets, all without increasing any model parameters or computational costs during ASR inference.
Renchang Dong, Yijie Li 0001, Dongxing Xu, Yanhua Long
ICASSP3
2024 Accent-Specific Vector Quantization for Joint Unsupervised and Supervised Training in Accent Robust Speech Recognition
abstract
How to effectively use limited supervised accent data to improve the accented ASR is of paramount importance. In this work, we propose an accent-specific quantization for joint unsupervised and supervised training (AQ-JUST) of end-to-end ASR models to address this issue. Specifically, two variants of AQ-JUST are investigated, namely BAQ-JUST and SAQ-JUST, by employing different model structures and training methods to capture the distinctiveness and commonalities between diverse accents, thus enhancing the performance of accented ASR systems. Our experiments are performed on both accented English and Mandarin ASR tasks. Results show that the proposed methods outperform the strong JUST baseline by relative 3.9% to 9.4% word/character error rate reductions on accented test sets.
Yijie Li 0001, Dongxing Xu, Yanhua Long
ICASSP3
2024 Score Calibration Based on Consistency Measure Factor for Speaker Verification
abstract
This paper proposes a new scoring calibration method named "Consistency-Aware Score Calibration", which introduces a Consistency Measure Factor (CMF) to measure the stability of audio voiceprints in similarity scores for speaker verification. The CMF is inspired by the limitations in segment scoring, where the segments with shorter length are not friendly to calculate the similarity score. By using CMF as a scale to calibrate scores calculated on the whole audio length, the method improves the performance of different state-of-the-art systems significantly, including ResNet with 34 to 518 layers and RepVGG. Experimental results show that the CMF method is better than segment scoring and shows excellent complementary information with other normalization or calibration methods. The proposed method was first proposed in a system description for the VoxCeleb speaker recognition challenge 2023, where it achieved the 1st place in Track1 of the challenge.
Chuanying Niu, Yibin Zhan, Yanhua Long, Dongxing Xu
ICASSP6
2023 FEW-Shot Continual Learning with Weight Alignment and Positive Enhancement for Bioacoustic Event Detection
abstract
In this paper, we propose a new continual learning framework for few-shot bioacoustic event detection (BED). First, we modify the recently proposed dynamic few-shot learning (DFSL) and generalize it to the BED task. Then, we introduce a weight alignment loss to enhance the weight generator of modified DFSL for detecting novel events. Furthermore, to augment the few positive samples of each target bioacoustic event, a positive enhancement approach is proposed to select high-confidence pseudo positives using the cumulative distribution of initial detection posterior probabilities. All experiments are performed on DCASE 2022 Task5 Challenge dataset, results show that the proposed methods significantly outperform the prototypical network (PN) baseline, it brings the overall F-measure of validation set from 52.2% to 55.9%. Moreover, the proposed framework shows great complementarity with the conventional PN, the F-measure is improved to 60.6% after applying a simple score fusion.
Sissi Xiaoxiao Wu, Dongxing Xu, Yanhua Long
ICASSP2
2023 Advanced RawNet2 with Attention-based Channel Masking for Synthetic Speech Detection
Yanhua Long, Yijie Li 0001, Dongxing Xu
INTERSPEECH4
2023 Phonetic-assisted Multi-Target Units Modeling for Improving Conformer-Transducer ASR system
Dongxing Xu, Yanhua Long
INTERSPEECH2
2010 A fast implementation of factor analysis for speaker verification
Dongxing Xu, Hongbin Cai, Beiqian Dai
INTERSPEECH3