VLDB 2026 Research / reviewers in the wild / expert
Qingyang Hong
dblp:39/4831
· DBLP profile ↗
60ranked-venue papers
4as first author
43since 2021 · last 2026
0000-0001-7380-8690ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 4 first-author · 41 since 2021Artificial intelligence and machine learning · 34 · 2 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ZoneSep: A Lightweight End-to-End Neural Beamformer With Post-Mask Decoder for In-Vehicle Multi-Zone Speech SeparationabstractRecently, the task of in-vehicle multi-zone speech separation (IMSS) has attracted significant interest. However, a persistent issue in this area is the cross-zone leakage problem. To address this, we present ZoneSep, a lightweight, fully end-to-end neural beamformer equipped with a post-mask decoder to suppress leakage effectively. Additionally, we propose a novel loss function, SC-SI-SDR loss, which enables simultaneous optimization of both non-silent and silent zones. Experiments on the IMSS dataset demonstrate the effectiveness of our approaches, showing that ZoneSep significantly outperforms recent advanced methods while requiring only 0.65M parameters and 4.02 GMACs. The source code is available athttps://github.com/EeLLJ/ZoneSep/. Longjie Luo, Junnan Wu, Lichun Fan, Zhenbo Luo, Jian Luan 0001, Qingyang Hong, Lin Li 0032 |
IEEE Signal Process. Lett. | 6 |
| 2025 | Dynamic Language Group-based MoE: Enhancing Code-Switching Speech Recognition with Hierarchical RoutingabstractThe Mixture of Experts (MoE) model is a promising approach for handling code-switching speech recognition (CS-ASR) tasks. However, the existing CS-ASR work on MoE has yet to leverage the advantages of MoE’s parameter scaling ability fully. This work proposes DLG-MoE, a Dynamic Language Group-based MoE, which can effectively handle the CS-ASR task and leverage the advantages of parameter scaling. DLG-MoE operates based on a hierarchical routing mechanism. First, the language router explicitly models the language attribute and dispatches the representations to the corresponding language expert groups. Subsequently, the unsupervised router within each language group implicitly models attributes beyond language and coordinates expert routing and collaboration. DLG-MoE outperforms the existing MoE methods on CS-ASR tasks while demonstrating great flexibility. It supports different top-k inference and streaming capabilities and can also prune the model parameters flexibly to obtain a monolingual sub-model. Hukai Huang, Shenghui Lu, Yahui Shan, He Qu, Fengrun Zhang, Wenhao Guan, Qingyang Hong |
ICASSP | 7 |
| 2025 | SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified FlowabstractRecently, flow matching based speech synthesis has significantly enhanced the quality of synthesized speech while reducing the number of inference steps. In this paper, we introduce SlimSpeech, a lightweight and efficient speech synthesis system based on rectified flow. We have built upon the existing speech synthesis method utilizing the rectified flow model, modifying its structure to reduce parameters and serve as a teacher model. By refining the reflow operation, we directly derive a smaller model with a more straight sampling trajectory from the larger model, while utilizing distillation techniques to further enhance the model performance. Experimental results demonstrate that our proposed method, with significantly reduced model parameters, achieves comparable performance to larger models through one-step sampling. Kaidi Wang 0001, Wenhao Guan, Shenghui Lu, Jianglong Yao, Qingyang Hong |
ICASSP | 6 |
| 2025 | DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec
Peijie Chen, Wenhao Guan, Weijie Wu, Hukai Huang, Qingyang Hong |
INTERSPEECH | 6 |
| 2025 | Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm
Wangjie Li, Longjie Luo, Qingyang Hong |
INTERSPEECH | 7 |
| 2025 | Cross-attention and Self-attention for Audio-visual Speaker Diarization in MISP-Meeting Challenge
Haodong Zhou, Longjie Luo, Qingyang Hong |
INTERSPEECH | 7 |
| 2025 | A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement
Shenghui Lu, Hukai Huang, Jinanglong Yao, Qingyang Hong |
INTERSPEECH | 5 |
| 2025 | SuPseudo: A Pseudo-supervised Learning Method for Neural Speech Enhancement in Far-field Speech Recognition
Longjie Luo, Qingyang Hong |
INTERSPEECH | 3 |
| 2025 | Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge
Longjie Luo, Shenghui Lu, Qingyang Hong |
INTERSPEECH | 4 |
| 2025 | ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
Wenhao Guan, Peijie Chen, Qingyang Hong |
INTERSPEECH | 5 |
| 2025 | Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
Kaidi Wang 0001, Wenhao Guan, Ziyue Jiang 0001, Hukai Huang, Peijie Chen, Weijie Wu, Qingyang Hong |
INTERSPEECH | 7 |
| 2024 | MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech SynthesisabstractThe style transfer task in Text-to-Speech (TTS) refers to the process of transferring style information into text content to generate corresponding speech with a specific style. However, most existing style transfer approaches are either based on fixed emotional labels or reference speech clips, which cannot achieve flexible style transfer. Recently, some methods have adopted text descriptions to guide style transfer. In this paper, we propose a more flexible multi-modal and style controllable TTS framework named MM-TTS. It can utilize any modality as the prompt in unified multi-modal prompt space, including reference speech, emotional facial images, and text descriptions, to control the style of the generated speech in a system. The challenges of modeling such a multi-modal style controllable TTS mainly lie in two aspects: 1) aligning the multi-modal information into a unified style space to enable the input of arbitrary modality as the style prompt in a single system, and 2) efficiently transferring the unified style representation into the given text content, thereby empowering the ability to generate prompt style-related voice. To address these problems, we propose an aligned multi-modal prompt encoder that embeds different modalities into a unified style space, supporting style transfer for different modalities. Additionally, we present a new adaptive style transfer method named Style Adaptive Convolutions (SAConv) to achieve a better style representation. Furthermore, we design a Rectified Flow based Refiner to solve the problem of over-smoothing Mel-spectrogram and generate audio of higher fidelity. Since there is no public dataset for multi-modal TTS, we construct a dataset named MEAD-TTS, which is related to the field of expressive talking head. Our experiments on the MEAD-TTS dataset and out-of-domain datasets demonstrate that MM-TTS can achieve satisfactory results based on multi-modal prompts. The audio samples and constructed dataset are available at https://multimodal-tts.github.io. Wenhao Guan, Yishuang Li, Hukai Huang, Jiayan Lin, Lingyan Huang, Qingyang Hong |
AAAI | 9 |
| 2024 | Reflow-TTS: A Rectified Flow Model for High-Fidelity Text-to-SpeechabstractThe diffusion models including Denoising Diffusion Probabilistic Models (DDPM) and score-based generative models have demonstrated excellent performance in speech synthesis tasks. However, its effectiveness comes at the cost of numerous sampling steps, resulting in prolonged sampling time required to synthesize high-quality speech. This drawback hinders its practical applicability in real-world scenarios. In this paper, we introduce ReFlow-TTS, a novel rectified flow based method for speech synthesis with high-fidelity. Specifically, our ReFlow-TTS is simply an Ordinary Differential Equation (ODE) model that transports Gaussian distribution to the ground-truth Mel-spectrogram distribution by straight line paths as much as possible. Furthermore, our proposed approach enables high-quality speech synthesis with a single sampling step and eliminates the need for training a teacher model. Our experiments on LJSpeech Dataset show that our ReFlow-TTS method achieves the best performance compared with other diffusion based models. And the ReFlow-TTS with one step sampling achieves competitive performance compared with existing one-step TTS models. Wenhao Guan, Haodong Zhou, Shiyu Miao, Xingjia Xie, Qingyang Hong |
ICASSP | 7 |
| 2024 | SR-HuBERT : An Efficient Pre-Trained Model for Speaker VerificationabstractRecently, pre-trained models (PTMs) have been extensively applied in speaker verification (SV) and greatly boosted system performance. However, mainstream PTMs currently concentrate on using frame-level universal representations. In this paper, we propose a novel pre-training framework that jointly models speaker information — Speaker Related HuBERT, abbreviated as SR-HuBERT. This framework aims to further explore speaker-related information inherent in speech universal representations. The proposed SR-HuBERT utilizes an unsupervised clustering algorithm based on graph structures to generate speaker pseudo-labels and promotes the learning of segment-level speaker-related representations through a multi-task pre-training framework. Experimental results on VoxCeleb1 test set demonstrate the effectiveness of the proposed SR-HuBERT. Even in the scenarios of limited fine-tuning data, SR-HuBERT outperforms the other existing PTMs on SV tasks. Additionally, SR-HuBERT also performs well on speaker-related tasks of SUPERB benchmark. Yishuang Li, Hukai Huang, Zhicong Chen, Wenhao Guan, Jiayan Lin, Qingyang Hong |
ICASSP | 7 |
| 2024 | Improving Multi-Speaker ASR With Overlap-Aware Encoding And Monotonic AttentionabstractEnd-to-end (E2E) multi-speaker speech recognition with the serialized output training (SOT) strategy demonstrates good performance in modeling diverse speaker scenarios. However, the E2E architecture doesn’t explicitly address the modeling of overlapping speech areas, potentially limiting the model’s ability to generalize. To tackle this issue, we introduce two approaches: overlap-aware encoding method and monotonic attention loss. The former enables the model to acquire knowledge about overlapping speech through multitask learning, while the latter encourages the model to learn specific attention patterns associated with overlap by constraining the attention of adjacent text time steps. Our experimental results on the AliMeeting dataset show that the combination of these two methods effectively enhances the model’s performance. Wenhao Guan, Lingyan Huang, Qingyang Hong |
ICASSP | 5 |
| 2024 | LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
Wenhao Guan, Wangjin Zhou, Feng Deng, Qingyang Hong |
INTERSPEECH | 8 |
| 2024 | Efficient Integrated Features Based on Pre-trained Models for Speaker Verification
Yishuang Li, Wenhao Guan, Hukai Huang, Shiyu Miao, Qingyang Hong |
INTERSPEECH | 7 |
| 2024 | MinSpeech: A Corpus of Southern Min Dialect for Automatic Speech Recognition
Jiayan Lin, Shenghui Lu, Hukai Huang, Wenhao Guan, Hui Bu, Qingyang Hong |
INTERSPEECH | 7 |
| 2024 | Enhancing Code-Switching Speech Recognition With LID-Based Collaborative Mixture of Experts ModelabstractDue to the inherent difficulty in modeling phonetic similarities across different languages, code-switching speech recognition presents a formidable challenge. This study proposes a Collaborative-MoE, a Mixture of Experts (MoE) model that leverages a collaborative mechanism among expert groups. Initially, a preceding routing network explicitly learns Language Identification (LID) tasks and selects experts based on acquired LID weights. This process ensures robust routing information to the MoE layer, mitigating interference from diverse language domains on expert network parameter updates. The LID weights are also employed to facilitate inter-group collaboration, enabling the integration of language-specific representations. Furthermore, within each language expert group, a gating network operates unsupervised to foster collaboration on attributes beyond language. Extensive experiments demonstrate the efficacy of our approach, achieving significant performance enhancements compared to alternative methods. Importantly, our method preserves the efficient inference capabilities characteristic of MoE models without necessitating additional pre-training. Hukai Huang, Jiayan Lin, Yishuang Li, Wenhao Guan, Qingyang Hong |
SLT | 7 |
| 2023 | Unsupervised Speaker Verification Using Pre-Trained Model and Label CorrectionabstractRecently, the fine-tuning pre-trained model framework has emerged as a promising paradigm for speech-processing tasks. In this study, we present a novel strategy for unsupervised speaker verification using the Sub-structure of Pre-Trained Model (Sub-PTM), which consists of a CNN-based feature extractor and several Transformer blocks. To obtain the initial pseudo labels, we utilize Infomap to perform clustering on the representations extracted from the Sub-PTM. The generated pseudo labels are then leveraged to train a speaker verification model containing a Sub-PTM and a downstream network. We also propose an Online and Offline Label Correction (OAO-LC) method to alleviate the effects of incorrect pseudo labels. By incorporating these techniques, our system achieves competitive results compared to the supervised baseline. Zhicong Chen, Wenxuan Hu, Lin Li 0032, Qingyang Hong |
ICASSP | 5 |
| 2023 | The XMU System for Audio-Visual Diarization and Recognition in MISP Challenge 2022abstractIn this paper, we present our work in track 2 of the Multi-modal Information based Speech Processing (MISP) 2022 Challenge. We built a cascaded system and explored different acoustic front-ends and end-to-end speech recognition back-ends based on multimodal. To promote effective fusion between the different modalities, we introduced a multi-level feature fusion network. By utilizing several additional strategies, our system achieved 31.88% in the concatenated minimum permutation character error rate (cpCER) on the evaluation set, achieving the 3th place ranking in the competition. Haodong Zhou, Qingyang Hong |
ICASSP | 4 |
| 2023 | Towards A Unified Conformer Structure: from ASR to ASV TaskabstractTransformer has achieved extraordinary performance in Natural Language Processing and Computer Vision tasks thanks to its powerful self-attention mechanism, and its variant Conformer has become a state-of-the-art architecture in the field of Automatic Speech Recognition (ASR). However, the main-stream architecture for Automatic Speaker Verification (ASV) is convolutional Neural Networks, and there is still much room for research on the Conformer based ASV. In this paper, firstly, we modify the Conformer architecture from ASR to ASV with very minor changes. Length-Scaled Attention (LSA) method and Sharpness-Aware Minimization (SAM) are adopted to improve model generalization. Experiments conducted on VoxCeleb and CN-Celeb show that our Conformer based ASV achieves competitive performance com-pared with the popular ECAPA-TDNN. Secondly, inspired by the transfer learning strategy, ASV Conformer is natural to be initialized from the pretrained ASR model. Via parameter transferring, self-attention mechanism could better focus on the relationship between sequence features, brings about 11% relative improvement in EER on test set of VoxCeleb and CN-Celeb, which reveals the potential of Conformer to unify ASV and ASR task. Finally, we provide a runtime in ASV-Subtools to evaluate its inference speed in production scenario. Our code is released at https://github.com/Snowdar/asv-subtools/tree/master/doc/papers/conformer.md. Dexin Liao, Tao Jiang 0033, Lin Li 0032, Qingyang Hong |
ICASSP | 5 |
| 2023 | Community Detection Graph Convolutional Network for Overlap-Aware Speaker DiarizationabstractThe clustering algorithm plays a crucial role in speaker diarization systems. However, traditional clustering algorithms suffer from the complex distribution of speaker embeddings and lack of digging potential relationships between speakers in a session. We propose a novel graph-based clustering approach called Community Detection Graph Convolutional Network (CDGCN) to improve the performance of the speaker diarization system. The CDGCN-based clustering method consists of graph generation, sub-graph detection, and Graph-based Overlapped Speech Detection (Graph-OSD). Firstly, the graph generation refines the local linkages among speech segments. Secondly the sub-graph detection finds the optimal global partition of the speaker graph. Finally, we view speaker clustering for overlap-aware speaker diarization as an overlapped community detection task and design a Graph-OSD component to output overlap-aware labels. By capturing local and global information, the speaker diarization system with CDGCN clustering outperforms the traditional Clustering-based Speaker Diarization (CSD) systems on the DIHARD III corpus. Zhicong Chen, Haodong Zhou, Qingyang Hong |
ICASSP | 5 |
| 2023 | Meta Learning with Adaptive Loss Weight for Low-Resource Speech RecognitionabstractModel Agnostic Meta-Learning (MAML) is an effective meta-learning algorithm for low-resource automatic speech recognition (ASR). It uses gradient descent to learn the initialization parameters of the model through various languages, making the model quickly adapt to unseen low-resource languages. But MAML is unstable due to its unique bilevel loss backward structure, which significantly affects the stability and generalization of the model. Since various languages have different contributions to the target language, the loss weights corresponding to the effects of diverse languages require costly manual adjustment in the training stage. Proper selection of these weights will influence the performance of the entire model. In this paper, we propose to apply a loss weight adaption method to MAML using Convolutional Neural Network (CNN) with Homoscedastic Uncertainty. The results of experiments showed that the proposed method outperformed previous gradient-based meta-learning methods and other loss weights adaption methods, and it further improved the stability and effectiveness of MAML. Qiulin Wang, Wenxuan Hu, Lin Li 0032, Qingyang Hong |
ICASSP | 4 |
| 2023 | Interpretable Style Transfer for Text-to-Speech with ControlVAE and Diffusion Bridge
Wenhao Guan, Yishuang Li, Hukai Huang, Qingyang Hong |
INTERSPEECH | 5 |
| 2023 | Cross-Modal Semantic Alignment before Fusion for Two-Pass End-to-End Spoken Language Understanding
Lingyan Huang, Haodong Zhou, Qingyang Hong |
INTERSPEECH | 4 |
| 2023 | Conformer-based Language Embedding with Self-Knowledge Distillation for Spoken Language Identification
Lingyan Huang, Qingyang Hong, Lin Li 0032 |
INTERSPEECH | 4 |
| 2022 | Graph Convolutional Network Based Semi-Supervised Learning on Multi-Speaker Meeting DataabstractUnsupervised clustering on speakers is becoming increasingly important for its potential uses in semi-supervised learning. In reality, we are often presented with enormous amounts of unlabeled data from multi-party meetings and discussions. An effective unsupervised clustering approach would allow us to significantly increase the amount of training data without additional costs for annotations. Recently, methods based on graph convolutional networks (GCN) have received growing attention for unsupervised clustering, as these methods exploit the connectivity patterns between nodes to improve learning performance. In this work, we present a GCN-based approach for semi-supervised learning. Given a pre-trained embedding extractor, a graph convolutional network is trained on the labeled data and clusters unlabeled data with "pseudo-labels". We present a self-correcting training mechanism that iteratively runs the cluster-train-correct process on pseudo-labels. We show that this proposed approach effectively uses unlabeled data and improves speaker recognition accuracy. Fuchuan Tong, Yafeng Chen, Hongbin Suo, Qingyang Hong, Lin Li 0032 |
ICASSP | 6 |
| 2022 | Spatial-aware Speaker Diarizaiton for Multi-channel Multi-party MeetingabstractThis paper describes a spatial-aware speaker diarization system for the multi-channel multi-party meeting. The diarization system obtains direction information of speaker by microphone array. Speaker spatial embedding is generated by xvector and s-vector derived from superdirective beamforming (SDB) which makes the embedding more robust. Specifically, we propose a novel multi-channel sequence-to-sequence neural network architecture named discriminative multi-stream neural network (DMSNet) which consists of attention superdirective beamforming (ASDB) block and Conformer encoder. The proposed ASDB is a self-adapted channel-wise block that extracts the latent spatial features of array audios by modeling interdependencies between channels. We explore DMSNet to address overlapped speech problem on multi-channel audio and achieve 93.53% accuracy on evaluation set. By performing DMSNet based overlapped speech detection (OSD) module, the diarization error rate (DER) of cluster-based diarization system decrease significantly from 13.45% to 7.64%. Yuji Liu, Binling Wang, Yiming Zhi, Shipeng Xia, Feng Tong, Lin Li 0032, Qingyang Hong |
INTERSPEECH | 10 |
| 2022 | Oriental Language Recognition (OLR) 2021: Summary and AnalysisabstractThe fifth Oriental Language Recognition (OLR) Challenge focuses on language recognition in a variety of complex environments to promote its development.The OLR 2020 Challenge includes three tasks: (1) cross-channel language identification, (2) dialect identification, and (3) noisy language identification.We choose Cavg as the principle evaluation metric, and the Equal Error Rate (EER) as the secondary metric.There were 58 teams participating in this challenge and one third of the teams submitted valid results.Compared with the best baseline, the Cavg values of Top 1 system for the three tasks were relatively reduced by 82%, 62% and 48%, respectively.This paper describes the three tasks, the database profile, and the final results.We also outline the novel approaches that improve the performance of language recognition systems most significantly, such as the utilization of auxiliary information. Binling Wang, Wenxuan Hu, Qiulin Wang, Dong Wang 0013, Lin Li 0032, Qingyang Hong |
INTERSPEECH | 8 |
| 2022 | When Speaker Recognition Meets Noisy Labels: Optimizations for Front-Ends and Back-EndsabstractA typical speaker recognition system often involves two modules: a feature extractor front-end and a speaker identification back-end. Despite the superior performance that deep neural networks have achieved for the front-end, their success benefits from the availability of large-scale, correctly labeled datasets. While label noise is unavoidable in speaker recognition datasets, both the front-end and back-end are affected by label noise, which degrades speaker recognition performance. In this paper, we first conduct comprehensive experiments to help improve our understanding of the effects of label noise on both the front-end and back-end. Then, we propose a simple yet effective training paradigm and loss correction method to handle label noise in the front-end. We combine our proposed method with the recently proposed Bayesian estimation of PLDA for noisy labels, and the whole system shows strong robustness to label noise. Furthermore, we show two practical applications of the improved system: one application corrects noisy labels based on an utterance’s chunk-level predictions, and the other algorithmically filters out high-confidence noisy samples within a dataset. By applying the second application to the NIST SRE04–10 dataset and verifying filtered utterances by human validation, we identify that approximately 1% of the NIST SRE04–10 dataset is made up of label errors. Lin Li 0032, Fuchuan Tong, Qingyang Hong |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Light-TTS: Lightweight Multi-Speaker Multi-Lingual Text-to-SpeechabstractWith the development of deep learning, end-to-end neural text-to-speech (TTS) systems have achieved significant improvements in high-quality speech synthesis. However, most of these systems are attention-based autoregressive models, resulting in slow synthesis speed and large model parameters. In addition, speech in different languages is usually synthesized using different models, which increases the complexity of the speech synthesis systems. In this paper, we propose a new lightweight multi-speaker multi-lingual speech synthesis system, named LightTTS, which can quickly synthesize the Chinese, English or code-switch speech of multiple speakers in a non-autoregressive generation manner using only one model. Moreover, compared to FastSpeech with the same number of neural network layers and nodes, our LightTTS achieves a 2.50x Mel-spectrum generation acceleration on CPU, and the parameters are compressed by 12.83x. Beibei Ouyang, Lin Li 0032, Qingyang Hong |
ICASSP | 4 |
| 2021 | End-To-End Multi-Accent Speech Recognition with Unsupervised Accent ModellingabstractEnd-to-end speech recognition has achieved good recognition performance on standard English pronunciation datasets. However, one prominent problem with end-to-end speech recognition systems is that non-native English speakers tend to have complex and varied accents, which reduces the accuracy of English speech recognition in different countries. In order to grapple with such an issue, we first investigate and improve the current mainstream end-to-end multi-accent speech recognition technologies. In addition, we propose two unsupervised accent modelling methods, which convert accent information into a global embedding, and use it to improve the performance of the end-to-end multi-accent speech recognition systems. Experimental results on accented English datasets of eight countries (AESRC2020) show that, compared with the Transformer baseline, our proposed methods achieve relative 14.8% and 15.4% average word error rate (WER) reduction in the development set and evaluation set, respectively. Beibei Ouyang, Dexin Liao, Shipeng Xia, Lin Li 0032, Qingyang Hong |
ICASSP | 6 |
| 2021 | ASV-SUBTOOLS: Open Source Toolkit for Automatic Speaker VerificationabstractIn this paper, we introduce a new open source toolkit for automatic speaker verification (ASV), named ASV-Subtools. Adopting PyTorch as main deep learning engine and Kaldi toolkit for data processing, ASV-Subtools allows users to develop modern speaker recognizers flexibly and efficiently. The toolkit prioritizes efficiency, modularity, and extensibility with the goal of supporting the state-of-the-art technologies in speaker recognition. In addition to including the commonly used networks, such as the time delay neural networks (TDNN), factorized TDNN (F-TDNN) and ResNet, ASV-Subtools also integrates an upgraded version of SpecAugment data augmentation method, named Inverted SpecAugment, with focus on making it more appropriate for speaker recognition subtasks. Besides, for alleviating the domain mismatch between training and test data, ASV-Subtools provides multiple domain adaptation methods of Probabilistic Linear Discriminant Analysis (PLDA). Experimental results show that state-of-the-art techniques implemented on ASV-Subtools could achieve competitive performance compared to other implementations. Fuchuan Tong, Miao Zhao, Lin Li 0032, Qingyang Hong |
ICASSP | 7 |
| 2021 | Additive Phoneme-Aware Margin Softmax Loss for Language RecognitionabstractThis paper proposes an additive phoneme-aware margin softmax (APM-Softmax) loss to train the multi-task learning network with phonetic information for language recognition.In additive margin softmax (AM-Softmax) loss, the margin is set as a constant during the entire training for all training samples, and that is a suboptimal method since the recognition difficulty varies in training samples.In additive angular margin softmax (AAM-Softmax) loss, the additional angular margin is set as a costant as well.In this paper, we propose an APM-Softmax loss for language recognition with phoneitc multi-task learning, in which the additive phoneme-aware margin is automatically tuned for different training samples.More specifically, the margin of language recognition is adjusted according to the results of phoneme recognition.Experiments are reported on Oriental Language Recognition (OLR) datasets, and the proposed method improves AM-Softmax loss and AAM-Softmax loss in different language recognition testing conditions. Lin Li 0032, Qingyang Hong |
Interspeech | 4 |
| 2021 | Real-Time End-to-End Monaural Multi-Speaker Speech Recognition
Beibei Ouyang, Fuchuan Tong, Dexin Liao, Lin Li 0032, Qingyang Hong |
Interspeech | 6 |
| 2021 | Oriental Language Recognition (OLR) 2020: Summary and AnalysisabstractThe fifth Oriental Language Recognition (OLR) Challenge focuses on language recognition in a variety of complex environments to promote its development. The OLR 2020 Challenge includes three tasks: (1) cross-channel language identification, (2) dialect identification, and (3) noisy language identification. We choose Cavg as the principle evaluation metric, and the Equal Error Rate (EER) as the secondary metric. There were 58 teams participating in this challenge and one third of the teams submitted valid results. Compared with the best baseline, the Cavg values of Top 1 system for the three tasks were relatively reduced by 82%, 62% and 48%, respectively. This paper describes the three tasks, the database profile, and the final results. We also outline the novel approaches that improve the performance of language recognition systems most significantly, such as the utilization of auxiliary information. Binling Wang, Yiming Zhi, Lin Li 0032, Qingyang Hong, Dong Wang 0013 |
Interspeech | 6 |
| 2021 | An Integrated Framework for Two-Pass Personalized Voice TriggerabstractIn this paper, we present the XMUSPEECH system for Task 1 of 2020 Personalized Voice Trigger Challenge (PVTC2020).Task 1 is a joint wake-up word detection with speaker verification on close talking data.The whole system consists of a keyword spotting (KWS) sub-system and a speaker verification (SV) sub-system.For the KWS system, we applied a Temporal Depthwise Separable Convolution Residual Network (TDSC-ResNet) to improve the system's performance.For the SV system, we proposed a multi-task learning network, where phonetic branch is trained with the character label of the utterance, and speaker branch is trained with the label of the speaker.Phonetic branch is optimized with connectionist temporal classification (CTC) loss, which is treated as an auxiliary module for speaker branch.Experiments show that our system gets significant improvements compared with baseline system. Dexin Liao, Yiming Zhi, Qingyang Hong, Lin Li 0032 |
Interspeech | 5 |
| 2021 | Phoneme-Aware and Channel-Wise Attentive Learning for Text Dependent Speaker Verification
Lin Li 0032, Qingyang Hong |
Interspeech | 4 |
| 2021 | Automatic Error Correction for Speaker Embedding Learning with Noisy Labels
Fuchuan Tong, Lin Li 0032, Qingyang Hong |
Interspeech | 6 |
| 2021 | Lightspeech: Lightweight Non-Autoregressive Multi-Speaker Text-To-SpeechabstractWith the development of deep learning, end-to-end neural text-to-speech systems have achieved significant improvements on high-quality speech synthesis. However, most of these systems are attention-based autoregressive models, resulting in slow synthesis speed and large model parameters. In this paper, we propose a new lightweight non-autoregressive multi-speaker speech synthesis system, named LightSpeech, which utilizes the lightweight feedforward neural networks to accelerate synthesis and reduce the amount of parameters. With the speaker embedding, LightSpeech achieves multi-speaker speech synthesis extremely quickly. Experiments on the LibriTTS dataset show that, compared with FastSpeech, our smallest LightSpeech model achieves a 9.27x Mel-spectrogram generation acceleration on CPU, and the model size and parameters are compressed by 37.06x and 37.36x, respectively. Beibei Ouyang, Lin Li 0032, Qingyang Hong |
SLT | 4 |
| 2021 | Multi-Feature Learning with Canonical Correlation Analysis Constraint for Text-Independent Speaker VerificationabstractIn order to improve the performance and robustness of text-independent speaker verification systems, various speaker embedding representation learning algorithms have been developed. Typically, exploring manifold kinds of features to describe speaker-related embeddings is a common approach, such as introducing more acoustic features or different resolution scale features. In this paper, a new multi-feature learning strategy with canonical correlation analysis (CCA) constraint is proposed to learn the instinct speaker embeddings, which maximizes the correlation between two features from the same utterance. Based on the multi-feature learning structure, the CCA constraint layer and the CCA loss are utilized to explore the correlation representation between the two kinds of features and alleviate the redundancy. Therefore, two multi-feature learning strategies are studied, using the pairwise acoustic features, and the pair of short-term and long-term features. Furthermore, we improve the long short-term feature learning structure by replacing the LSTM block with the Bidirectional-GRU (B-GRU) block and introducing more dense layers. The effectiveness of these improvements are shown on the VoxCeleb 1 evaluation set, the noisy Vox-Celeb 1 evaluation set and the SITW evaluation set. Miao Zhao, Lin Li 0032, Qingyang Hong |
SLT | 4 |
| 2021 | Deep joint learning for language recognition
Lin Li 0032, Qingyang Hong |
Neural Networks | 4 |
| 2020 | XMU-TS Systems for NIST SRE19 CTS Challenge
Miao Zhao, Wendian Lei, Qingyang Hong, Lin Li 0032 |
ICASSP | 5 |
| 2020 | The XMUSPEECH System for Short-Duration Speaker Verification Challenge 2020
Tao Jiang 0033, Miao Zhao, Lin Li 0032, Qingyang Hong |
INTERSPEECH | 4 |
| 2020 | Improving Transformer-Based Speech Recognition with Unsupervised Pre-Training and Multi-Task Semantic Knowledge Learning
Lin Li 0032, Qingyang Hong, Lingling Liu |
INTERSPEECH | 3 |
| 2020 | On the Usage of Multi-Feature Integration for Speaker Verification and Language Identification
Miao Zhao, Lin Li 0032, Qingyang Hong |
INTERSPEECH | 5 |
| 2020 | The XMUSPEECH System for the AP19-OLR Challenge
Miao Zhao, Yiming Zhi, Lin Li 0032, Qingyang Hong |
INTERSPEECH | 6 |
| 2019 | Training Multi-task Adversarial Network for Extracting Noise-robust Speaker EmbeddingabstractUnder noisy environments, to achieve the robust performance of speaker recognition is still a challenging task. Motivated by the promising performance of multi-task training in a variety of image processing tasks, we explore the potential of multitask adversarial training for learning a noise-robust speaker embedding. In this paper, we present a novel framework that consists of three components: an encoder that extracts the noise-robust speaker embeddings; a classifier that classifies the speakers; a discriminator that discriminates the noise type of the speaker embeddings. Additionally, we propose a training strategy using the training accuracy as an indicator to stabilize the multi-class adversarial optimization process. We conduct our experiments on the English and Mandarin corpuses and the experimental results demonstrate that our proposed multi-task adversarial training method could greatly outperform the other methods without adversarial training in noisy environments. Furthermore, the experiments indicate that our method is also able to improve the speaker verification performance under the clean condition. Tao Jiang 0033, Lin Li 0032, Qingyang Hong, Bingyin Xia |
ICASSP | 4 |
| 2019 | Anti-Spoofing Speaker Verification System with Multi-Feature Integration and Multi-Task Learning
Rongjin Li, Miao Zhao, Lin Li 0032, Qingyang Hong |
INTERSPEECH | 5 |
| 2019 | Deep Speaker Embedding Extraction with Channel-Wise Feature Responses and Additive Supervision Softmax Loss Function
Tao Jiang 0033, Lin Li 0032, Qingyang Hong |
INTERSPEECH | 5 |
| 2019 | Systematic analysis and prediction of type IV secreted effector proteins by machine learning approachesabstractIn the course of infecting their hosts, pathogenic bacteria secrete numerous effectors, namely, bacterial proteins that pervert host cell biology. Many Gram-negative bacteria, including context-dependent human pathogens, use a type IV secretion system (T4SS) to translocate effectors directly into the cytosol of host cells. Various type IV secreted effectors (T4SEs) have been experimentally validated to play crucial roles in virulence by manipulating host cell gene expression and other processes. Consequently, the identification of novel effector proteins is an important step in increasing our understanding of host-pathogen interactions and bacterial pathogenesis. Here, we train and compare six machine learning models, namely, Naïve Bayes (NB), K-nearest neighbor (KNN), logistic regression (LR), random forest (RF), support vector machines (SVMs) and multilayer perceptron (MLP), for the identification of T4SEs using 10 types of selected features and 5-fold cross-validation. Our study shows that: (1) including different but complementary features generally enhance the predictive performance of T4SEs; (2) ensemble models, obtained by integrating individual single-feature models, exhibit a significantly improved predictive performance and (3) the 'majority voting strategy' led to a more stable and accurate classification performance when applied to predicting an ensemble learning model with distinct single features. We further developed a new method to effectively predict T4SEs, Bastion4 (Bacterial secretion effector predictor for T4SS), and we show our ensemble classifier clearly outperforms two recent prediction tools. In summary, we developed a state-of-the-art T4SE predictor by conducting a comprehensive performance evaluation of different machine learning algorithms along with a detailed analysis of single- and multi-feature selections. Jiawei Wang 0002, Bingjiao Yang, Yi An, Tatiana T. Marquez-Lago, André Leier, Jonathan Wilksch, Qingyang Hong, Yang Zhang 0010, Morihiro Hayashida, Tatsuya Akutsu, Geoffrey I. Webb, Richard A. Strugnell, Jiangning Song, Trevor Lithgow |
Briefings Bioinform. | 7 |
| 2018 | Electroencephalogram-based brain-computer interface for the Chinese spelling system: a surveyabstractElectroencephalogram (EEG) based brain-computer interfaces allow users to communicate with the external environment by means of their EEG signals, without relying on the brain’s usual output pathways such as muscles. A popular application for EEGs is the EEG-based speller, which translates EEG signals into intentions to spell particular words, thus benefiting those suffering from severe disabilities, such as amyotrophic lateral sclerosis. Although the EEG-based English speller (EEGES) has been widely studied in recent years, few studies have focused on the EEG-based Chinese speller (EEGCS). The EEGCS is more difficult to develop than the EEGES, because the English alphabet contains only 26 letters. By contrast, Chinese contains more than 11 000 logographic characters. The goal of this paper is to survey the literature on EEGCS systems. First, the taxonomy of current EEGCS systems is discussed to get the gist of the paper. Then, a common framework unifying the current EEGCS and EEGES systems is proposed, in which the concept of EEG-based choice acts as a core component. In addition, a variety of current EEGCS systems are investigated and discussed to highlight the advances, current problems, and future directions for EEGCS. Minghui Shi, Changle Zhou, Jun Xie 0002, Shaozi Li, Qingyang Hong, Min Jiang 0005, Fei Chao 0001, Weifeng Ren, Xiangqian Liu, Dajun Zhou |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2017 | Transfer learning for PLDA-based speaker verification
Qingyang Hong, Lin Li 0032, Lihong Wan, Huiyang Guo |
Speech Commun. | 1 |
| 2016 | A transfer learning method for PLDA-based speaker verificationabstractCurrently, the state-of-the-art speaker verification system is based on i-vector and PLDA. However, PLDA requires tens of thousands of development data from many speakers. This makes it difficult to learn the PLDA parameters for a domain with scarce data. In this paper, we propose an effective transfer learning method based on Bayesian joint probability in which Kullback-Leibler (KL) divergence between the source domain and the target domain is added as a regularization factor. This hypothesis could utilize the development data of source domain to help find a better optimal solution of PLDA parameters for the target domain. Experimental results based on the NIST SRE and Switchboard corpus demonstrate that our proposed method could produce the largest gain of performance compared with the traditional PLDA and the other adaptation approach. Qingyang Hong, Lin Li 0032, Lihong Wan, Feng Tong |
ICASSP | 1 |
| 2016 | Transfer Learning for Speaker Verification on Short Utterances
Qingyang Hong, Lin Li 0032, Lihong Wan, Feng Tong |
INTERSPEECH | 1 |
| 2015 | Duration dependent covariance regularization in PLDA modeling for speaker verification
Weicheng Cai, Ming Li 0026, Lin Li 0032, Qingyang Hong |
INTERSPEECH | 4 |
| 2015 | Modified-prior PLDA and score calibration for duration mismatch compensation in speaker recognition systemabstractTo deal with the performance degradation of speaker recognition due to duration mismatch between enrollment and test utterances, a novel strategy to modify the standard normal prior distribution of the i-vector during probabilistic linear discriminant analysis (PLDA) modeling is employed. This new modified-prior PLDA model incorporates the covariance matrix scaled with duration of each utterance for each speaker, which achieves more discriminative characteristics by learning the duration variability as well as session variation in the i-vector space. Furthermore, an efficient Quality Measure Function (QMF) method which adopts duration variation as a compensation technique is employed to eliminate the linear shift in the score domain. To evaluate the robustness of the proposed approach, experiments were conducted on the NIST SRE10 core-core task in condition-5 with varying test utterance duration, in which the i-vectors of test utterances were extracted from full segment and randomly truncated segments of duration 10s and 20s. The results demonstrated the efficiency of modified-prior PLDA in different duration conditions, and the combined score calibration further improved the performance of speaker recognition. Qingyang Hong, Lin Li 0032, Ming Li 0026, Lihong Wan |
INTERSPEECH | 1 |
| 2007 | Translation Memory Sharing Models in XMCATabstractIn this paper, two Translation Memory (TM) sharing models adopted in XMCAT, a Computer Assisted Translation tool (CAT) supporting cooperated work in machine translation, was described in detail. One is Center-based TM sharing model, which is only fit for users in a local area network (LAN) and the other is a novel model called P2P-based TM sharing model, which could be used through Internet by geographically distributed users. With the two TM sharing models, a user may share data with other users through network, so that he/she may reduce the repeated work further and cooperate with others more easily. Besides, the methods used in XMCA T to deal with the problem of multi-translations arose in the cooperated memory sharing models, were also proposed in this paper. XMCAT system has been adopted and approved by some translation companies. Yidong Chen 0001, Xiaodong Shi, Changle Zhou, Tangqiu Li, Qingyang Hong |
CSCWD | 5 |
| 2001 | A hybrid method for syntactic and semantic structure disambiguation for ChineseabstractThis paper presents a method of syntactic and semantic structure disambiguation for Chinese. The method finds the most plausible interpretation of a phrase or a sentence in Chinese by evaluation of the similarity between the structure and examples in related word entries in a knowledge base, Hownet. First we put forward the general idea of the method and briefly introduce its semantic knowledge resource-the Hownet Dictionary. The main algorithm of the method is then proposed with detail. The experimental result shows that the method is effective. Tangqiu Li, Qingyang Hong, Shaozi Li |
SMC | 3 |