EDBT 2026 Demo / reviewers in the wild / expert
Jiaming Zhou 0001
dblp:250/0527-1
· DBLP profile ↗
27ranked-venue papers
6as first author
27since 2021 · last 2026
0009-0002-4819-4572ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 22 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio ModelsabstractText-to-Audio (TTA) generation has made rapid progress, but current evaluation methods remain narrow, focusing mainly on perceptual quality while overlooking robustness, generalization, and ethical concerns. We present TTA-Bench, a comprehensive benchmark for evaluating TTA models across functional performance, reliability, and social responsibility. It covers seven dimensions including accuracy, robustness, fairness, and toxicity, and includes 2,999 diverse prompts generated through automated and manual methods. We introduce a unified evaluation protocol that combines objective metrics with over 118,000 human annotations from both experts and general users. Ten state-of-the-art models are benchmarked under this framework, offering detailed insights into their strengths and limitations. TTA-Bench establishes a new standard for holistic evaluation of TTA systems. Hui Wang 0075, Haoze Liu, Yuhang Jia, Shiwan Zhao, Jiaming Zhou 0001, Haoqin Sun, Hui Bu |
AAAI | 7 |
| 2026 | DIFFA: Large Language Diffusion Models Can Listen and UnderstandabstractRecent advances in large language models (LLMs) have shown remarkable capabilities across textual and multimodal domains. In parallel, large language diffusion models have emerged as a promising alternative to the autoregressive paradigm, offering improved controllability, bidirectional context modeling, and robust generation. However, their application to the audio modality remains underexplored. In this work, we introduce DIFFA, the first diffusion-based large audio-language model designed to perform spoken language understanding. DIFFA integrates a frozen diffusion language model with a lightweight dual-adapter architecture that bridges speech understanding and natural language reasoning. We employ a two-stage training pipeline: first, aligning semantic representations via an ASR objective; then, learning instruction-following abilities through synthetic audio-caption pairs automatically generated by prompting LLMs. Despite being trained on only 960 hours of ASR and 127 hours of synthetic instruction data, DIFFA demonstrates competitive performance on major benchmarks, including MMSU, MMAU, and VoiceBench, outperforming several autoregressive open-source baselines. Our results reveal the potential of large language diffusion models for efficient and scalable audio understanding, opening a new direction for speech-driven AI. Jiaming Zhou 0001, Hongjie Chen 0001, Shiwan Zhao, Jian Kang 0006, Jie Li 0001, Enzhi Wang, Haoqin Sun, Hui Wang 0075, Aobo Kong, Xuelong Li 0001 |
AAAI | 1 |
| 2026 | RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal AnalysisabstractRecent advances in speech large language models (e.g., GPT-4o) have enabled end-to-end spoken interactions, yet their robustness in realworld applications remains unclear, where systems must assist users in completing specific tasks under complex conditions such as multiturn, ambiguous, and often spontaneous speech, as well as natural alternation between speech and text.Task-oriented dialogue (TOD) offers a realistic scenario to evaluate whether models can effectively help users accomplish such task-oriented goals, but existing benchmarks are mainly text-based, and the few speech datasets are limited to English and often neglect spontaneous disfluencies and speaker diversity.To address this gap, we introduce RealTalk-CN, the first Chinese multi-turn, multi-domain speech-text TOD dataset, containing 5.4k dialogues (60K turns, ~150 hours) of real human-to-human recordings with detailed annotations for dialogue states, disfluency types, and speaker characteristics.Based on this dataset, we propose a cross-modal interaction task supporting dynamic speech-text switching and a comprehensive evaluation protocol assessing robustness to disfluencies, sensitivity to speaker variation, and cross-domain generalization.Experiments on state-of-the-art models demonstrate the challenges posed by RealTalk-CN and establish its value as a benchmark for developing reliable and fair Speech LLMs in real-world deployments.The dataset and evaluation framework are available 1 to encourage further research. Enzhi Wang, Jiaming Zhou 0001, Yuhang Jia, Aobo Kong, Qicheng Li |
ACL (1) | 2 |
| 2026 | SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality EvaluationabstractHui Wang, Jinghua Zhao, Yifan Yang, Shujie Liu, Junyang Chen, Yanzhe Zhang, Shiwan Zhao, Jinyu Li, Jiaming Zhou, Haoqin Sun, Yan Lu, Yong Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hui Wang 0075, Jinghua Zhao 0004, Yifan Yang 0005, Shujie Liu 0001, Shiwan Zhao, Jinyu Li 0001, Jiaming Zhou 0001, Haoqin Sun, Yan Lu 0001 |
ACL (1) | 9 |
| 2025 | ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5abstractAutomatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children’s speech remains challenging due to differences in pronunciation, tone, and pace compared to adult speech. In this paper, we introduce a new Mandarin speech dataset focused on children aged 3 to 5, addressing the scarcity of resources in this area. The dataset comprises 41.25 hours of speech with carefully crafted manual transcriptions, collected from 397 speakers across various provinces in China, with balanced gender representation. We provide a comprehensive analysis of speaker demographics, speech duration distribution and geographic coverage. Additionally, we evaluate ASR performance on models trained from scratch, such as Conformer, as well as fine-tuned pre-trained models like HuBERT and Whisper, where fine-tuning demonstrates significant performance improvements. Furthermore, we assess speaker verification (SV) on our dataset, showing that, despite the challenges posed by the unique vocal characteristics of young children, the dataset effectively supports both ASR and SV tasks. This dataset is a valuable contribution to Mandarin child speech research and holds potential for applications in educational technology and child-computer interaction. It will be open-source and freely available for all academic purposes. Jiaming Zhou 0001, Shiwan Zhao, Jiabei He 0001, Haoqin Sun, Hui Wang 0075, Aobo Kong, Xi Yang 0023, Yequan Wang, Yonghua Lin |
ACL (1) | 1 |
| 2025 | Emotion-Preserving Prosody Anonymization Network for Voice Privacy ProtectionabstractBalancing emotion preservation and privacy protection in voice anonymization presents a significant challenge, particularly due to the difficulty of effectively handling prosody, a key feature in speech. While preserving prosodic features in anonymized speech enhances emotional expression, it also increases the risk of leaking speaker information. To address this conflict, we propose a lightweight Emotion-Preserving Prosody Anonymization (EPPA) network, which extracts speaker-independent prosodic features to preserve speech emotion while converting them into another speaker’s style for anonymization. By combining EPPA with timbre cloning for anonymization while retaining speech content, we achieve a more balanced voice conversion. Evaluated using the Voice Privacy Challenge (VPC) 2024 metrics, our proposed EPPA, utilizing the closest center distance (CCD) anonymization strategy, demonstrates strong performance across emotional expression, content clarity, and privacy protection, achieving the highest ranking in both average and weighted ranks compared to the six baseline solutions. Jiabei He 0001, Shiwan Zhao, Jiaming Zhou 0001, Haoqin Sun, Hui Wang 0075 |
ICASSP | 3 |
| 2025 | MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music EvaluationabstractThe technology for generating music from textual descriptions has seen rapid advancements. However, evaluating text-to-music (TTM) systems remains a significant challenge, primarily due to the difficulty of balancing performance and cost with existing objective and subjective evaluation methods. In this paper, we propose an automatic assessment task for TTM models to align with human perception. To address the TTM evaluation challenges posed by the professional requirements of music evaluation and the complexity of the relationship between text and music, we collect MusicEval, the first generative music assessment dataset. This dataset contains 2,748 music clips generated by 31 advanced and widely used models in response to 384 text prompts, along with 13,740 ratings from 14 music experts. Furthermore, we design a CLAP-based assessment model built on this dataset, and our experimental results validate the feasibility of the proposed task, providing a valuable reference for future development in TTM evaluation. The dataset is available at https://www.aishelltech.com/AISHELL_7A. Hui Wang 0075, Jinghua Zhao 0004, Shiwan Zhao, Hui Bu, Jiaming Zhou 0001, Haoqin Sun |
ICASSP | 7 |
| 2025 | Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement FrameworkabstractMultimodal emotion recognition systems rely heavily on the full availability of modalities, suffering significant performance declines when modal data is incomplete. To tackle this issue, we present the Cross-Modal Alignment, Reconstruction, and Refinement (CM-ARR) framework, an innovative approach that sequentially engages in cross-modal alignment, reconstruction, and refinement phases to handle missing modalities and enhance emotion recognition. This framework utilizes unsupervised distribution-based contrastive learning to align heterogeneous modal distributions, reducing discrepancies and modeling semantic uncertainty effectively. The reconstruction phase applies normalizing flow models to transform these aligned distributions and recover missing modalities. The refinement phase employs supervised point-based contrastive learning to disrupt semantic correlations and accentuate emotional traits, thereby enriching the affective content of the reconstructed representations. Extensive experiments confirm the superior performance of CM-ARR. Notably, averaged across six scenarios of missing modalities, CM-ARR achieves absolute improvements of 2.11%/2.12% (WAR/UAR), and 1.71%/1.96% (WAR/UAR), respectively, on IEMOCAP and MSP-IMPROV datasets. Haoqin Sun, Shiwan Zhao, Shaokai Li, Xiangyu Kong 0001, Xuechen Wang, Jiaming Zhou 0001, Aobo Kong, Wenjia Zeng |
ICASSP | 6 |
| 2025 | Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal AlignmentabstractMultimodal emotion recognition (MER), leveraging speech and text, has emerged as a pivotal domain within human-computer interaction, demanding sophisticated methods for effective multimodal integration. The challenge of aligning features across these modalities is significant, with most existing approaches adopting a singular alignment strategy. Such a narrow focus not only limits model performance but also fails to address the complexity and ambiguity inherent in emotional expressions. In response, this paper introduces a Multi-Granularity Cross-Modal Alignment (MGCMA) framework, distinguished by its comprehensive approach encompassing distribution-based, instance-based, and token-based alignment modules. This framework enables a multi-level perception of emotional information across modalities. Our experiments on IEMOCAP demonstrate that our proposed method outperforms current state-of-the-art techniques. Xuechen Wang, Shiwan Zhao, Haoqin Sun, Hui Wang 0075, Jiaming Zhou 0001 |
ICASSP | 5 |
| 2025 | M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing WhisperabstractState-of-the-art models like OpenAI’s Whisper exhibit strong performance in multilingual automatic speech recognition (ASR), but they still face challenges in accurately recognizing diverse subdialects. In this paper, we propose M2R-Whisper, a novel multi-stage and multi-scale retrieval augmentation approach designed to enhance ASR performance in low-resource settings. Building on the principles of in-context learning (ICL) and retrieval-augmented techniques, our method employs sentence-level ICL in the pre-processing stage to harness contextual information, while integrating token-level k-Nearest Neighbors (kNN) retrieval as a post-processing step to further refine the final output distribution. By synergistically combining sentence-level and token-level retrieval strategies, M2R-Whisper effectively mitigates various types of recognition errors. Experiments conducted on Mandarin and subdialect datasets, including AISHELL-1 and KeSpeech, demonstrate substantial improvements in ASR accuracy, all achieved without any parameter updates. Jiaming Zhou 0001, Shiwan Zhao, Jiabei He 0001, Hui Wang 0075, Wenjia Zeng, Haoqin Sun, Aobo Kong |
ICASSP | 1 |
| 2025 | Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual DatastoresabstractThe kNN-CTC model has proven to be effective for monolingual automatic speech recognition (ASR). However, its direct application to multilingual scenarios like code-switching, presents challenges. Although there is potential for performance improvement, a kNN-CTC model utilizing a single bilingual datastore can inadvertently introduce undesirable noise from the alternative language. To address this, we propose a novel kNN-CTC-based code-switching ASR (CS-ASR) framework that employs dual monolingual datastores and a gated datastore selection mechanism to reduce noise interference. Our method selects the appropriate datastore for decoding each frame, ensuring the injection of language-specific information into the ASR process. We apply this framework to cutting-edge CTC-based models, developing an advanced CS-ASR system. Extensive experiments demonstrate the remarkable effectiveness of our gated datastore mechanism in enhancing the performance of zero-shot Chinese-English CS-ASR. Jiaming Zhou 0001, Shiwan Zhao, Hui Wang 0075, Tian-Hao Zhang, Haoqin Sun, Xuechen Wang |
ICASSP | 1 |
| 2025 | Chinese-LiPS: A Chinese Audio-Visual Speech Recognition Dataset with Lip-Reading and Presentation SlidesabstractIncorporating visual modalities to assist Automatic Speech Recognition (ASR) tasks has led to significant improvements. However, existing Audio-Visual Speech Recognition (AVSR) datasets and methods typically rely solely on lip-reading information or speaking contextual video, neglecting the potential of combining these different valuable visual cues within the speaking context. In this paper, we release a multimodal Chinese AVSR dataset, Chinese-LiPS, comprising 100 hours of speech, video, and corresponding manual transcription, with the visual modality encompassing both lip-reading information and the presentation slides used by the speaker. Based on Chinese-LiPS, we develop a simple yet effective pipeline, LiPS-AVSR, which leverages both lip-reading and presentation slide information as visual modalities for AVSR tasks. Experiments show that lip-reading and presentation slide information improve ASR performance by approximately 8% and 25%, respectively, with a combined performance improvement of about 35%. The dataset is available at https://kiri0824.github.io/Chinese-LiPS/ Jinghua Zhao 0004, Yuhang Jia, Jiaming Zhou 0001, Hui Wang 0075 |
ICME | 4 |
| 2025 | RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
Haoqin Sun, Jingguang Tian, Jiaming Zhou 0001, Hui Wang 0075, Jiabei He 0001, Shiwan Zhao, Xiangyu Kong 0001, Desheng Hu, Xinkang Xu, Xinhui Hu |
INTERSPEECH | 3 |
| 2025 | A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
Jiaming Zhou 0001, Shiwan Zhao |
INTERSPEECH | 2 |
| 2025 | FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow MatchingabstractTo advance continuous token modeling and temporal-coherence enforcement, we propose FELLE, an autoregressive model that integrates language modeling with token-wise flow matching. By leveraging the autoregressive nature of language models and the generative efficacy of flow matching, FELLE effectively predicts continuous-valued tokens (mel-spectrograms). For each continuous-valued token, FELLE modifies the general prior distribution in flow matching by incorporating information from the previous step, improving coherence and stability. Furthermore, to enhance synthesis quality, FELLE introduces a coarse-to-fine flow-matching mechanism, generating continuous-valued tokens hierarchically, conditioned on the language model's output. Experimental results demonstrate the potential of incorporating flow-matching techniques in autoregressive mel-spectrogram modeling, leading to significant improvements in TTS generation quality, as shown in https://aka.ms/felle. Hui Wang 0075, Shujie Liu 0001, Lingwei Meng, Jinyu Li 0001, Yifan Yang 0005, Shiwan Zhao, Haiyang Sun 0004, Haoqin Sun, Jiaming Zhou 0001, Yan Lu 0001 |
ACM Multimedia | 10 |
| 2025 | SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged SeniorsabstractWhile voice technologies increasingly serve aging populations, current systems exhibit significant performance gaps due to inadequate training data capturing elderly-specific vocal characteristics like presbyphonia and dialectal variations. The limited data available on super-aged individuals in existing elderly speech datasets, coupled with overly simple recording styles and annotation dimensions, exacerbates this issue. To address the critical scarcity of speech data from individuals aged 75 and above, we introduce SeniorTalk, a carefully annotated Chinese spoken dialogue dataset. This dataset contains 55.53 hours of speech from 101 natural conversations involving 202 participants, ensuring a strategic balance across gender, region, and age. Through detailed annotation across multiple dimensions, it can support a wide range of speech tasks. We perform extensive experiments on speaker verification, speaker diarization, speech recognition, and speech editing tasks, offering crucial insights for the development of speech technologies targeting this age group. Code is available at https://github.com/flageval-baai/SeniorTalk and data at https://huggingface.co/datasets/evan0617/seniortalk. Yang Chen 0034, Hui Wang 0075, Jiabei He 0001, Jiaming Zhou 0001, Xi Yang 0023, Yequan Wang, Yonghua Lin |
NeurIPS | 6 |
| 2025 | Self-Prompt Tuning: Enable Autonomous Role-Playing in LLMs
Aobo Kong, Shiwan Zhao, Qicheng Li, Jiaming Zhou 0001, Haoqin Sun |
NLPCC (1) | 6 |
| 2025 | StreamMel: Real-Time Zero-Shot Text-to-Speech Via Interleaved Continuous Autoregressive ModelingabstractRecent advances in zero-shot text-to-speech (TTS) synthesis have achieved high-quality speech generation for unseen speakers, but most systems remain unsuitable for real-time applications because of their offline design. Current streaming TTS paradigms often rely on multi-stage pipelines and discrete representations, leading to increased computational cost and suboptimal system performance. In this work, we propose StreamMel, a pioneering single-stage streaming TTS framework that models continuous mel-spectrograms. By interleaving text tokens with acoustic frames, StreamMel enables low-latency, autoregressive synthesis while preserving high speaker similarity and naturalness. Experiments on LibriSpeech demonstrate that StreamMel outperforms existing streaming TTS baselines in both quality and latency. It even achieves performance comparable to offline systems while supporting efficient real-time generation, showcasing broad prospects for integration with real-time speech large language models. Audio samples are available at:https://aka.ms/StreamMel. Hui Wang 0075, Yifan Yang 0005, Shujie Liu 0001, Jinyu Li 0001, Lingwei Meng, Tie-Yan Liu, Jiaming Zhou 0001, Haoqin Sun, Yan Lu 0001 |
IEEE Signal Process. Lett. | 7 |
| 2024 | CIF-T: A Novel CIF-Based Transducer Architecture for Automatic Speech RecognitionabstractRNN-T models are widely used in ASR, which rely on the RNN-T loss to achieve length alignment between input audio and target sequence. However, the implementation complexity and the alignment-based optimization target of RNN-T loss lead to computational redundancy and a reduced role for predictor network, respectively. In this paper, we propose a novel model named CIF-Transducer (CIF-T) which incorporates the Continuous Integrate-and-Fire (CIF) mechanism with the RNN-T model to achieve efficient alignment. In this way, the RNN-T loss is abandoned, thus bringing a computational reduction and allowing the predictor network a more significant role. We also introduce Funnel-CIF, Context Blocks, Unified Gating and Bilinear Pooling joint network, and auxiliary training strategy to further improve performance. Experiments on the 178-hour AISHELL-1 and 10000-hour WenetSpeech datasets show that CIF-T achieves state-of-the-art results with lower computational overhead compared to RNN-T models. Tian-Hao Zhang, Dinghao Zhou, Guiping Zhong, Jiaming Zhou 0001, Baoxiang Li |
ICASSP | 4 |
| 2024 | KNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo LabelsabstractThe success of retrieval-augmented language models in various natural language processing (NLP) tasks has been constrained in automatic speech recognition (ASR) applications due to challenges in constructing fine-grained audio-text datastores. This paper presents kNN-CTC, a novel approach that overcomes these challenges by leveraging Connectionist Temporal Classification (CTC) pseudo labels to establish frame-level audio-text key-value pairs, circumventing the need for precise ground truth alignments. We further introduce a "skip-blank" strategy, which strategically ignores CTC blank frames, to reduce datastore size. By incorporating a k-nearest neighbors retrieval mechanism into pretrained CTC ASR systems and leveraging a fine-grained, pruned datastore, kNN-CTC consistently achieves substantial improvements in performance under various experimental settings. Our code is available at https://github.com/NKU-HLT/KNN-CTC. Jiaming Zhou 0001, Shiwan Zhao, Wenjia Zeng |
ICASSP | 1 |
| 2024 | AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
Rong Gong, Hongfei Xue, Lezhi Wang, Qisheng Li, Lei Xie 0001, Hui Bu, Shaomei Wu, Jiaming Zhou 0001, Jun Du 0002, Jia Bin, Ming Li 0026 |
INTERSPEECH | 9 |
| 2024 | Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
Haoqin Sun, Shiwan Zhao, Xiangyu Kong 0001, Xuechen Wang, Hui Wang 0075, Jiaming Zhou 0001 |
INTERSPEECH | 6 |
| 2024 | Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based AdaptationabstractDysarthric speech recognition (DSR) presents a formidable challenge due to inherent inter-speaker variability, leading to severe performance degradation when applying DSR models to new dysarthric speakers.Traditional speaker adaptation methodologies typically involve fine-tuning models for each speaker, but this strategy is cost-prohibitive and inconvenient for disabled users, requiring substantial data collection.To address this issue, we introduce a prototype-based approach that markedly improves DSR performance for unseen dysarthric speakers without additional fine-tuning.Our method employs a feature extractor trained with HuBERT to produce per-word prototypes that encapsulate the characteristics of previously unseen speakers.These prototypes serve as the basis for classification.Additionally, we incorporate supervised contrastive learning to refine feature extraction.By enhancing representation quality, we further improve DSR performance, enabling effective personalized DSR. Shiwan Zhao, Jiaming Zhou 0001, Aobo Kong |
INTERSPEECH | 3 |
| 2024 | Uncertainty-Aware Mean Opinion Score Prediction
Hui Wang 0075, Shiwan Zhao, Jiaming Zhou 0001, Xiguang Zheng, Haoqin Sun, Xuechen Wang |
INTERSPEECH | 3 |
| 2024 | PB-LRDWWS System For the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting ChallengeabstractFor the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting (LRDWWS) Challenge, we introduce the PB-LRDWWS system. This system combines a dysarthric speech content feature extractor for prototype construction with a prototype-based classification method. The feature extractor is a fine-tuned HuBERT model obtained through a three-stage fine-tuning process using cross-entropy loss. This fine-tuned HuBERT extracts features from the target dysarthric speaker’s enrollment speech to build prototypes. Classification is achieved by calculating the cosine similarity between the HuBERT features of the target dysarthric speaker’s evaluation speech and prototypes. Despite its simplicity, our method demonstrates effectiveness through experimental results. Our system achieves second place in the final Test-B of the LRDWWS Challenge. Jiaming Zhou 0001, Shiwan Zhao |
SLT | 2 |
| 2024 | Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition ChallengeabstractThe StutteringSpeech Challenge focuses on advancing speech technologies for people who stutter, specifically targeting Stuttering Event Detection (SED) and Automatic Speech Recognition (ASR) in Mandarin. The challenge comprises three tracks: (1) SED, which aims to develop systems for detection of stuttering events; (2) ASR, which focuses on creating robust systems for recognizing stuttered speech; and (3) Research track for innovative approaches utilizing the provided dataset. We utilizes an open-source Mandarin stuttering dataset AS-70, which has been split into new training and test sets for the challenge. This paper presents the dataset, details the challenge tracks, and analyzes the performance of the top systems, highlighting improvements in detection accuracy and reductions in recognition error rates. Our findings underscore the potential of specialized models and augmentation strategies in developing stuttered speech technologies. Hongfei Xue, Rong Gong, Mingchen Shao, Lezhi Wang, Lei Xie 0001, Hui Bu, Jiaming Zhou 0001, Jun Du 0002, Ming Li 0026 |
SLT | 8 |
| 2023 | MADI: Inter-Domain Matching and Intra-Domain Discrimination for Cross-Domain Speech RecognitionabstractEnd-to-end automatic speech recognition (ASR) usually suffers from performance degradation when applied to a new domain due to domain shift. Unsupervised domain adaptation (UDA) aims to improve the performance on the unlabeled target domain by transferring knowledge from the source to the target domain. To improve transferability, existing UDA approaches mainly focus on matching the distributions of the source and target domains globally and/or locally, while ignoring the model discriminability. In this paper, we propose a novel UDA approach for ASR via inter-domain MAtching and intra-domain DIscrimination (MADI), which improves the model transferability by fine-grained inter-domain matching and discriminability by intra-domain contrastive discrimination simultaneously. Evaluations on the Libri-Adapt dataset demonstrate the effectiveness of our approach. MADI reduces the relative word error rate (WER) on cross-device and cross-environment ASR by 17.7% and 22.8%, respectively. Jiaming Zhou 0001, Shiwan Zhao |
ICASSP | 1 |