EDBT 2026 Demo / reviewers in the wild / expert
Xinkang Xu
dblp:277/3578
· DBLP profile ↗
12ranked-venue papers
0as first author
11since 2021 · last 2025
0009-0003-2771-1398ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 11 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Speaker Normalization and Content Restoration for Zero-Shot Voice Conversion with Attention-Enhanced Discriminator
Desheng Hu, Xinhui Hu, Xinkang Xu |
INTERSPEECH | 5 |
| 2025 | RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
Haoqin Sun, Jingguang Tian, Jiaming Zhou 0001, Hui Wang 0075, Jiabei He 0001, Shiwan Zhao, Xiangyu Kong 0001, Desheng Hu, Xinkang Xu, Xinhui Hu |
INTERSPEECH | 9 |
| 2025 | Discrete Audio Representations for Automated Audio Captioning
Jingguang Tian, Haoqin Sun, Xinhui Hu, Xinkang Xu |
INTERSPEECH | 4 |
| 2024 | Learning Emotion-Invariant Speaker Representations for Speaker VerificationabstractIn recent years, the rapid progress in speaker verification (SV) technology has been driven by the extraction of speaker representations based on deep learning. However, such representations are still vulnerable to emotion variability. To address this issue, we propose multiple improvements to train speaker encoders to increase emotion robustness. Firstly, we utilize CopyPaste-based data augmentation to gather additional parallel data, which includes different emotional expressions from the same speaker. Secondly, we apply cosine similarity loss to restrict parallel sample pairs and minimize intraclass variation of speaker representations to reduce their correlation with emotional information. Finally, we use emotion-aware masking (EM) based on the speech signal energy on the input parallel samples to further strengthen the speaker representation and make it emotion-invariant. We conduct a comprehensive ablation study to demonstrate the effectiveness of these various components. Experimental results show that our proposed method achieves a relative 19.29% drop in EER compared to the baseline system. Jingguang Tian, Xinhui Hu, Xinkang Xu |
ICASSP | 3 |
| 2024 | A Deep Representation Learning-Based Speech Enhancement Method Using Complex Convolution Recurrent Variational AutoencoderabstractGenerally, the performance of deep neural networks (DNNs) heavily depends on the quality of data representation learning. Our preliminary work has emphasized the significance of deep representation learning (DRL) in the context of speech enhancement (SE) applications. Specifically, our initial SE algorithm employed a gated recurrent unit variational autoencoder (VAE) with a Gaussian distribution to enhance the performance of certain existing SE systems. Building upon our preliminary framework, this paper introduces a novel approach for SE using deep complex convolutional recurrent networks with a VAE (DCCRN-VAE). DCCRN-VAE assumes that the latent variables of signals follow complex Gaussian distributions that are modeled by DCCRN, as these distributions can better capture the behaviors of complex signals. Additionally, we propose the application of a residual loss in DCCRN-VAE to further improve the quality of the enhanced speech. Compared to our preliminary work, DCCRN-VAE introduces a more sophisticated DCCRN structure and probability distribution for DRL. Furthermore, in comparison to DCCRN, DCCRN-VAE employs a more advanced DRL strategy. The experimental results demonstrate that the proposed SE algorithm outperforms both our preliminary SE framework and the state-of-the-art DCCRN SE method in terms of scale-invariant signal-to-distortion ratio, speech quality, and speech intelligibility. Jingguang Tian, Xinhui Hu, Xinkang Xu, Zhaohui Yin |
ICASSP | 4 |
| 2024 | SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR
Shuaishuai Ye, Shunfei Chen, Xinhui Hu, Xinkang Xu |
INTERSPEECH | 4 |
| 2023 | LE-SSL-MOS: Self-Supervised Learning MOS Prediction with Listener EnhancementabstractRecently, researchers have shown an increasing interest in automatically predicting the subjective evaluation for speech synthesis systems. This prediction is a challenging task, especially on the out-of-domain test set. In this paper, we proposed a novel fusion model for MOS prediction that combines both supervised and unsupervised approaches. In the supervised aspect, we developed a SSL-based predictor called LE-SSL-MOS. The LE-SSL-MOS utilizes pre-trained self-supervised learning models and further improves prediction accuracy by utilizing the opinion scores of each utterance in the listener enhancement branch. In the unsupervised aspect, two steps are contained: one is that we fine-tuned unit language model (ULM) using highly-intelligible domain data to improve the correlation of an unsupervised metric SpeechLMScore. Another is that we utilized ASR confidence as a new metric with the help of ensemble learning. To the best of our knowledge, this is the first architecture that fuses supervised and unsupervised methods for MOS prediction.With these approaches, our experimental results on the VoiceMOS Challenge 2023 show that LE-SSL-MOS performs better than the baseline. Our fusion system achieved an absolute improvement of 13 % over LE-SSL-MOS on the noisy and enhanced speech track. And our system ranked 1st and 2 nd respectively in the French speech synthesis track and the noisy and enhanced speech track of the challenge. Zili Qi, Xinhui Hu, Wangjin Zhou, Sheng Li 0010, Xinkang Xu |
ASRU | 7 |
| 2022 | The Royalflush System of Speech Recognition for M2met ChallengeabstractThis paper describes our RoyalFlush system for the track of multi-speaker automatic speech recognition (ASR) in the M2MeT challenge. We adopted the serialized output training (SOT) based multi-speakers ASR system with large-scale simulation data. Firstly, we investigated a set of front-end methods, including multi-channel weighted predicted error (WPE), beamforming, speech separation, speech enhancement, etc., to process training, evaluation, and test sets. However, according to their experimental results, we only selected the WPE and beamforming approach as our front-end methods. Secondly, we made great efforts in the data augmentation for multi-speaker ASR, including adding noise and reverberation, over-lapped speech simulation, multi-channel speech simulation, speed perturbation, front-end processing, etc., which brought us a significant performance improvement. Finally, to make full use of the performance complementary of different model architecture, we trained the standard conformer based joint CTC/Attention (Conformer) and U2++ ASR model with a bidirectional attention decoder, a modification of Conformer, to fuse their results. Compared with the official baseline system, our system got a 12.22% absolute Character Error Rate (CER) reduction on the evaluation set and 12.11% on the test set. Shuaishuai Ye, Shunfei Chen, Xinhui Hu, Xinkang Xu |
ICASSP | 5 |
| 2022 | Multiple Enhancements to LSTM for Learning Emotion-Salient Features in Speech Emotion Recognition
Desheng Hu, Xinhui Hu, Xinkang Xu |
INTERSPEECH | 3 |
| 2021 | An Investigation of Using Hybrid Modeling Units for Improving End-to-End Speech Recognition SystemabstractThe acoustic modeling unit is crucial for an end-to-end speech recognition system, especially for the Mandarin language. Until now, most of the studies on Mandarin speech recognition focused on individual units, and few of them paid attention to using a combination of these units. This paper uses a hybrid of the syllable, Chinese character, and subword as the modeling units for the end-to-end speech recognition system based on the CTC/attention multi-task learning. In this approach, the character-subword unit is assigned to train the transformer model in the main task learning stage. In contrast, the syllable unit is assigned to enhance the transformer’s shared encoder in the auxiliary task stage with the Connectionist Temporal Classification (CTC) loss function. The recognition experiments were conducted on AISHELL-1 and an open data set of 1200-hour Mandarin speech corpus collected from the OpenSLR, respectively. The experimental results demonstrated that using the syllable-char-subword hybrid modeling unit can achieve better performances than the conventional units of char-subword, and 6.6% relative CER reduction on our 1200-hour data. The substitution error also achieves a considerable reduction. Shunfei Chen, Xinhui Hu, Sheng Li 0010, Xinkang Xu |
ICASSP | 4 |
| 2021 | An End-to-End Dialect Identification System with Transfer Learning from a Multilingual Automatic Speech Recognition Model
Shuaishuai Ye, Xinhui Hu, Sheng Li 0010, Xinkang Xu |
Interspeech | 5 |
| 2020 | Data Augmentation for Code-Switch Language Modeling by Fusing Multiple Text Generation Methods
Xinhui Hu, Binbin Gu, Xinkang Xu |
INTERSPEECH | 5 |