EDBT 2026 Demo / reviewers in the wild / expert
Ziyang Ma 0001
dblp:138/5915-1
· DBLP profile ↗
46ranked-venue papers
9as first author
46since 2021 · last 2026
0000-0002-8195-3262ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 8 first-author · 34 since 2021Artificial intelligence and machine learning · 29 · 7 first-author · 29 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream QuantizationabstractWenxi Chen, Ruiqi Yan, Yushen Chen, Zhikang Niu, Ziyang Ma, Xiquan Li, Yuzhe Liang, Wenhanlin, Shunshun Yin, Ming Tao, Xinsheng Wang, Xie Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ruiqi Yan, Yushen Chen, Zhikang Niu, Ziyang Ma 0001, Xiquan Li, Yuzhe Liang, Wenhan Lin, Shunshun Yin, Xie Chen 0001 |
ACL (1) | 5 |
| 2026 | FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio PretrainingabstractContrastively pretrained audio–language models (e.g., CLAP) excel at clip-level understanding but struggle with frame-level tasks.Existing extensions fail to exploit the varying granularity of real-world audio–text data, where massive clip-level textual descriptions coexist with limited frame-level annotations. This paper proposes Fine-grained Language-Audio Pretraining (FineLAP), a novel training paradigm that advances both clip- and frame-level alignment in CLAP with heterogeneous data.FineLAP introduces a dual-stream sigmoid loss with a cluster-based sampling strategy to jointly learn from clip- and frame-level supervision. To capture both global semantics and local details, FineLAP uses a decoupled audio projector on top of a self-supervised encoder.To alleviate the scarcity of temporally annotated data, we present FineLAP-100k, a large-scale synthetic SED dataset constructed through a scalable curation pipeline.Extensive experiments demonstrate that FineLAP achieves SOTA performance across multiple audio understanding tasks, including retrieval, classification, sound event detection, and text-to-audio grounding. Ablation studies further show that coarse- and fine-grained alignment are mutually beneficial, providing insights for building better audio-language models (ALMs). Xiquan Li, Xuenan Xu, Ziyang Ma 0001, Haolin He, Qiuqiang Kong, Xie Chen 0001 |
ACL (1) | 3 |
| 2026 | Evaluating the Expressive Appropriateness of Speech in Rich ContextsabstractTianrui Wang, Ziyang Ma, Yizhou Peng, Haoyu Wang, Zhikang Niu, Zikang Huang, Yihao Wu, Yi-Wen Chao, Yu Jiang, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Cheng Gong, Yifan Yang, Tianchi Liu, Junyu Wang, Nana Hou, Meng Ge, Fuming You, Yang Wei, Zhongqian Sun, Hu Haifeng, Xiaobao Wang, Eng Siong Chng, Xie Chen, Longbiao Wang, Jianwu Dang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianrui Wang, Ziyang Ma 0001, Yizhou Peng, Zhikang Niu, Zikang Huang, Yi-Wen Chao, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Yifan Yang 0005, Tianchi Liu 0004, Nana Hou, Meng Ge, Fuming You, Zhongqian Sun, Haifeng Hu 0009, Xiaobao Wang, Chng Eng Siong, Xie Chen 0001, Longbiao Wang, Jianwu Dang 0001 |
ACL (1) | 2 |
| 2026 | Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-trainingabstractYifan Yang, Bing Han, Hui Wang, Wei Wang, Ziyang Ma, Long Zhou, Zengrui Jin, Guanrou Yang, Tianrui Wang, Xu Tan, Xie Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yifan Yang 0005, Bing Han 0008, Hui Wang 0075, Wei Wang 0010, Ziyang Ma 0001, Zengrui Jin, Guanrou Yang, Tianrui Wang, Xu Tan 0003, Xie Chen 0001 |
ACL (1) | 5 |
| 2025 | VQTalker: Towards Multilingual Talking Avatars Through Facial Motion TokenizationabstractWe present VQTalker, a Vector Quantization-based framework for multilingual talking head generation that addresses the challenges of lip synchronization and natural motion across diverse languages. Our approach is grounded in the phonetic principle that human speech comprises a finite set of distinct sound units (phonemes) and corresponding visual articulations (visemes), which often share commonalities across languages. We introduce a facial motion tokenizer based on Group Residual Finite Scalar Quantization (GRFSQ), which creates a discretized representation of facial features. This method enables comprehensive capture of facial movements while improving generalization to multiple languages, even with limited training data. Building on this quantized representation, we implement a coarse-to-fine motion generation process that progressively refines facial animations. Extensive experiments demonstrate that VQTalker achieves state-of-the-art performance in both video-driven and speech-driven scenarios, particularly in multilingual settings. Notably, our method achieves high-quality results at a resolution of 512 × 512 pixels while maintaining a lower bitrate of approximately 11 kbps. Our work opens new possibilities for cross-lingual talking face generation. Tao Liu 0068, Ziyang Ma 0001, Qi Chen 0009, Shuai Fan 0005, Xie Chen 0001, Kai Yu 0004 |
AAAI | 2 |
| 2025 | Language Model Can Listen While SpeakingabstractDialogue serves as the most natural manner of human-computer interaction (HCI). Recent advancements in speech language models (SLM), have significantly enhanced speech-based conversational AI. However, these models are limited to turn-based conversation, lacking the ability to interact with humans in real-time spoken scenarios, for example, being interrupted when the generated content is not satisfactory. To address these limitations, we explore full duplex modeling (FDM) in interactive speech language models (iSLM), focusing on enhancing real-time interaction and, more explicitly, exploring the quintessential ability of interruption. We introduce a novel model design, namely listening-while-speaking language model (LSLM), an end-to-end system equipped with both listening and speaking channels. Our LSLM employs a token-based decoder-only TTS for speech generation and a streaming self-supervised learning (SSL) encoder for real-time audio input. LSLM fuses both channels for autoregressive generation and detects turn-taking in real time. Three fusion strategies—early fusion, middle fusion, and late fusion—are explored, with middle fusion achieving an optimal balance between speech generation and real-time interaction. Two experimental settings, command-based FDM and voice-based FDM, demonstrate LSLM’s robustness to noise and sensitivity to diverse instructions. Our results highlight LSLM’s capability to achieve duplex communication with minimal impact on existing systems. This study aims to advance the development of interactive speech dialogue systems, enhancing their applicability in real-world contexts. Ziyang Ma 0001, Yakun Song, Chenpeng Du, Jian Cong, Zhuo Chen 0006, Yuping Wang 0005, Yuxuan Wang 0002, Xie Chen 0001 |
AAAI | 1 |
| 2025 | Speech Recognition Meets Large Language Model: Benchmarking, Models, and ExplorationabstractIn this paper, we focus on prompting one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and large language models (LLM). Despite the growing body of research in this area, we find that many crucial design decisions in LLM-based ASR systems are often inadequately justified. This lack of clarity impedes the field's progress, making it challenging to pinpoint which design choices truly improve model performance. To address these challenges, we conduct a comprehensive series of experiments that explore various aspects, leading to the optimal LLM-based ASR system. We found that delicate designs are not necessary, while a clean setup with little task-specific design is competent. The models achieve strong performance on the Librispeech and Gigaspeech datasets, compared to both LLM-based models and non-LLM-based models. Finally, we explore the capability emergence of LLM-based ASR in the process of modal alignment. We hope that our study can facilitate the research on extending LLM with cross-modality capacity and shed light on the LLM-based ASR community. Ziyang Ma 0001, Guanrou Yang, Yifan Yang 0005, Zhifu Gao, Jiaming Wang 0004, Zhihao Du, Fan Yu 0002, Qian Chen 0003, Shiliang Zhang, Xie Chen 0001 |
AAAI | 1 |
| 2025 | ELLA-V: Stable Neural Codec Language Modeling with Alignment-Guided Sequence ReorderingabstractThe language model (LM) approach based on acoustic and linguistic prompts, such as VALL-E, has achieved remarkable progress in the field of zero-shot audio generation. However, existing methods still have some limitations: 1) repetitions, transpositions, and omissions in the output synthesized speech due to limited alignment constraints between audio and phoneme tokens; 2) challenges of fine-grained control over the synthesized speech with autoregressive (AR) language model; 3) infinite silence generation due to the nature of AR-based decoding, especially under the greedy strategy. To alleviate these issues, we propose ELLA-V, a simple but efficient LM-based zero-shot text-to-speech (TTS) framework, which enables fine-grained control over synthesized audio at the phoneme level. The key to ELLA-V is interleaving sequences of acoustic and phoneme tokens, where phoneme tokens appear ahead of the corresponding acoustic tokens. The experimental findings reveal that our model outperforms baselines in terms of accuracy and delivers more stable results using both greedy and sampling-based decoding strategies. Yakun Song, Zhuo Chen 0006, Ziyang Ma 0001, Xie Chen 0001 |
AAAI | 4 |
| 2025 | GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and RefinementabstractYifan Yang, Zheshu Song, Jianheng Zhuo, Mingyu Cui, Jinpeng Li, Bo Yang, Yexing Du, Ziyang Ma, Xunying Liu, Ziyuan Wang, Ke Li, Shuai Fan, Kai Yu, Wei-Qiang Zhang, Guoguo Chen, Xie Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yifan Yang 0005, Zheshu Song, Jianheng Zhuo, Bo Yang 0006, Yexing Du, Ziyang Ma 0001, Xunying Liu, Ke Li 0018, Shuai Fan 0005, Kai Yu 0004, Weiqiang Zhang 0001, Guoguo Chen, Xie Chen 0001 |
ACL (1) | 8 |
| 2025 | F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow MatchingabstractThis paper introduces F5-TTS, a fully non-autoregressive text-to-speech system based on flow matching with Diffusion Transformer (DiT). Without requiring complex designs such as duration model, text encoder, and phoneme alignment, the text input is simply padded with filler tokens to the same length as input speech, and then the denoising is performed for speech generation, which was originally proved feasible by E2 TTS. However, the original design of E2 TTS makes it hard to follow due to its slow convergence and low robustness. To address these issues, we first model the input with ConvNeXt to refine the text representation, making it easy to align with the speech. We further propose an inference-time Sway Sampling strategy, which significantly improves our model’s performance and efficiency. This sampling strategy for flow step can be easily applied to existing flow matching based models without retraining. Our design allows faster training and achieves an inference RTF of 0.15, which is greatly improved compared to state-of-the-art diffusion-based TTS models. Trained on a public 100K hours multilingual dataset, our F5-TTS exhibits highly natural and expressive zero-shot ability, seamless code-switching capability, and speed control efficiency. We have released all codes and checkpoints to promote community development, at https://SWivid.github.io/F5-TTS/. Yushen Chen, Zhikang Niu, Ziyang Ma 0001, Keqi Deng, Kai Yu 0004, Xie Chen 0001 |
ACL (1) | 3 |
| 2025 | Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum LearningabstractYexing Du, Youcheng Pan, Ziyang Ma, Bo Yang, Yifan Yang, Keqi Deng, Xie Chen, Yang Xiang, Ming Liu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yexing Du, Youcheng Pan, Ziyang Ma 0001, Bo Yang 0006, Yifan Yang 0005, Keqi Deng, Xie Chen 0001, Yang Xiang 0003, Ming Liu 0004, Bing Qin 0001 |
ACL (1) | 3 |
| 2025 | Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language ModelabstractLarge Audio-Language Models (LALMs) have demonstrated remarkable performance in tasks involving audio perception and understanding, such as speech recognition and audio captioning. However, their reasoning capabilities-critical for solving complex real-world problems-remain underexplored. In this work, we conduct the first exploration into integrating Chain-of-Thought (CoT) reasoning into LALMs to enhance their reasoning ability across auditory modalities. We evaluate representative CoT methods, analyzing their performance in both information extraction and reasoning tasks across sound, music, and speech domains. Our findings reveal that CoT methods significantly improve performance on easy and medium tasks but encounter challenges with hard tasks, where reasoning chains can confuse the model rather than improve accuracy. Additionally, we identify a positive correlation between reasoning path length and accuracy, demonstrating the potential of scaling inference for advanced instruction-following and reasoning. This study not only highlights the promise of CoT in enhancing LALM reasoning capabilities but also identifies key limitations and provides actionable directions for future research. Index Terms-Chain-of-Thought (CoT), Reasoning, Large Audio Language Model (LALM) Ziyang Ma 0001, Zhuo Chen 0006, Yuping Wang 0005, Chng Eng Siong, Xie Chen 0001 |
ASRU | 1 |
| 2025 | SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMsabstractAutomated Audio Captioning (AAC) aims to generate natural textual descriptions for input audio signals. Recent progress in audio pre-trained models and large language models (LLMs) has significantly enhanced audio understanding and textual reasoning capabilities, making improvements in AAC possible. In this paper, we propose SLAM-AAC to further enhance AAC with paraphrasing augmentation and CLAP-Refine through LLMs. Our approach uses the self-supervised EAT model to extract fine-grained audio representations, which are then aligned with textual embeddings via lightweight linear layers. The caption generation LLM is efficiently fine-tuned using the LoRA adapter. Drawing inspiration from the back-translation method in machine translation, we implement paraphrasing augmentation to expand the Clotho dataset during pre-training. This strategy helps alleviate the limitation of scarce audio-text pairs and generates more diverse captions from a small set of audio clips. During inference, we introduce the plug-and-play CLAP-Refine strategy to fully exploit multiple decoding outputs, akin to the n-best rescoring strategy in speech recognition. Using the CLAP model for audio-text similarity calculation, we could select the textual descriptions generated by multiple searching beams that best match the input audio. Experimental results show that SLAM-AAC achieves state-of-the-art performance on Clotho V2 and AudioCaps, surpassing previous mainstream models. Ziyang Ma 0001, Xiquan Li, Xuenan Xu, Yuzhe Liang, Zhisheng Zheng, Kai Yu 0004, Xie Chen 0001 |
ICASSP | 2 |
| 2025 | DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio CaptioningabstractWhile automated audio captioning (AAC) has made notable progress, traditional fully supervised AAC models still face two critical challenges: the need for expensive audio-text pair data for training and performance degradation when transferring across domains. To overcome these limitations, we present DRCap, a data-efficient and flexible zero-shot audio captioning system that requires text-only data for training and can quickly adapt to new domains without additional fine-tuning. DRCap integrates a contrastive language-audio pre-training (CLAP) model and a large language model (LLM) as its backbone. During training, the model predicts the ground-truth caption with a fixed text encoder from CLAP, whereas, during inference, the text encoder is replaced with the audio encoder to generate captions for audio clips in a zero-shot manner. To mitigate the modality gap of the CLAP model, we use both the projection strategy from the encoder side and the retrieval-augmented generation strategy from the decoder side. Specifically, audio embeddings are first projected onto a text embedding support to absorb extensive semantic information within the joint multi-modal space of CLAP. At the same time, similar captions retrieved from a datastore are fed as prompts to instruct the LLM, incorporating external knowledge to take full advantage of its strong generative capability. Conditioned on both the projected CLAP embedding and the retrieved similar captions, the model is able to produce a more accurate and semantically rich textual description. By tailoring the text embedding support and the caption datastore to the target domain, DRCap acquires a robust ability to adapt to new domains in a training-free manner. Experimental results demonstrate that DRCap outperforms all other zero-shot models in in-domain scenarios and achieves state-of-the-art performance in cross-domain scenarios. Xiquan Li, Ziyang Ma 0001, Xuenan Xu, Yuzhe Liang, Zhisheng Zheng, Qiuqiang Kong, Xie Chen 0001 |
ICASSP | 3 |
| 2025 | Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data GapabstractWhile automatic speech recognition (ASR) systems have achieved remarkable performance with large-scale datasets, their efficacy remains inadequate in low-resource settings, encompassing dialects, accents, minority languages, and long-tail hotwords, domains with significant practical relevance. With the advent of versatile and powerful text-to-speech (TTS) models, capable of generating speech with human-level naturalness, expressiveness, and diverse speaker profiles, leveraging TTS for ASR data augmentation provides a cost-effective and practical approach to enhancing ASR performance. Comprehensive experiments on an unprecedentedly rich variety of low-resource datasets demonstrate consistent and substantial performance improvements, proving that the proposed method of enhancing low-resource ASR through a versatile TTS model is highly effective and has broad application prospects. Furthermore, we delve deeper into key characteristics of synthesized speech data that contribute to ASR improvement, examining factors such as text diversity, speaker diversity, and the volume of synthesized data, with text diversity being studied for the first time in this work. We hope our findings provide helpful guidance and reference for the practical application of TTS-based data augmentation and push the advancement of low-resource ASR one step further. Guanrou Yang, Fan Yu 0002, Ziyang Ma 0001, Zhihao Du, Zhifu Gao, Shiliang Zhang, Xie Chen 0001 |
ICASSP | 3 |
| 2025 | MuPT: A Generative Symbolic Music Pretrained TransformerabstractIn this paper, we explore the application of Large Language Models (LLMs) to the pre-training of music. While the prevalent use of MIDI in music modeling is well-established, our findings suggest that LLMs are inherently more compatible with ABC Notation, which aligns more closely with their design and strengths, thereby enhancing the model's performance in musical composition.
To address the challenges associated with misaligned measures from different tracks during generation, we propose the development of a $\underline{S}$ynchronized $\underline{M}$ulti-$\underline{T}$rack ABC Notation ($\textbf{SMT-ABC Notation}$), which aims to preserve coherence across multiple musical tracks.
Our contributions include a series of models capable of handling up to 8192 tokens, covering 90\% of the symbolic music data in our training set. Furthermore, we explore the implications of the $\underline{S}$ymbolic $\underline{M}$usic $\underline{S}$caling Law ($\textbf{SMS Law}$) on model performance. The results indicate a promising research direction in music generation, offering extensive resources for further research through our open-source contributions. Xingwei Qu, Yuelin Bai, Yinghao Ma, Ziya Zhou, Ka Man Lo, Ruibin Yuan, Lejun Min, Xueling Liu 0001, Xeron Du, Shuyue Guo, Yiming Liang, Shangda Wu, Junting Zhou, Tianyu Zheng, Ziyang Ma 0001, Fengze Han, Wei Xue 0002, Gus Xia, Emmanouil Benetos, Xiang Yue, Chenghua Lin 0002, Xu Tan 0003, Wenhao Huang 0001, Jie Fu 0001, Ge Zhang 0009 |
ICLR | 18 |
| 2025 | A Progressive Generation Framework with Speech Pre-trained Model for Expressive Voice ConversionabstractExpressive voice conversion (EVC) aims to modify the speaker identity and emotional style of speech while preserving its content. Existing approaches often focus on disentangling speaker, emotion, and content information but overlook the progressive generation mechanisms in human speech production. To address this, we propose a three-stage framework that includes a speech disentanglement module, a progressive generator, and an acoustic refiner. This framework enables speech pre-trained models to parse linguistic content, emotional style, and speaker identity, which are then progressively integrated into the speech reconstruction branch to generate high-quality speech with replaceable emotional style and speaker identity. Experiments with six different pre-trained models show that our framework activates their disentanglement capabilities, surpassing baseline performance in EVC, and supports speaker and emotion control from different target samples. This framework also provides a valuable reference for evaluating the disentanglement capabilities of speech pre-training models. Tianrui Wang, Meng Ge, Zhikang Niu, Chunyu Qiang, Zikang Huang, Ziyang Ma 0001, Xiaobao Wang, Xie Chen 0001, Longbiao Wang, Jianwu Dang 0001 |
ICME | 8 |
| 2025 | k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation LearningabstractSelf-supervised learning (SSL) has achieved great success in speech-related tasks. While Transformer and Conformer architectures have dominated SSL backbones, encoders like Zipformer, which excel in automatic speech recognition (ASR), remain unexplored in SSL. Concurrently, inefficiencies in data processing within existing SSL training frameworks, such as fairseq, pose challenges in managing the growing volumes of training data. To address these issues, we propose k2SSL, an open-source framework that offers faster, more memory-efficient, and better-performing self-supervised speech representation learning, focusing on downstream ASR tasks. The optimized HuBERT and proposed Zipformer-based SSL systems exhibit substantial reductions in both training time and memory usage during SSL training. Experiments on LibriSpeech demonstrate that Zipformer Base significantly outperforms HuBERT and WavLM, achieving up to a 34.8% relative WER reduction compared to HuBERT Base after fine-tuning, along with a 3.5x pre-training speedup in GPU hours. When scaled to 60k hours of LibriLight data, Zipformer Large exhibits remarkable efficiency, matching HuBERT Large’s performance while requiring only 5/8 pre-training steps. Yifan Yang 0005, Jianheng Zhuo, Zengrui Jin, Ziyang Ma 0001, Xiaoyu Yang 0005, Zengwei Yao, Liyong Guo, Wei Kang 0006, Long Lin, Daniel Povey, Xie Chen 0001 |
ICME | 4 |
| 2025 | Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
Qixi Zheng, Yushen Chen, Zhikang Niu, Ziyang Ma 0001, Kai Yu 0004, Xie Chen 0001 |
INTERSPEECH | 4 |
| 2025 | MER 2025: When Affective Computing Meets Large Language ModelsabstractMER2025 is the third year of our MER series of challenges. Previously, MER2023 (http://merchallenge.cn/mer2023) focused on multi-label learning, noise robustness, and semi-supervised learning, while MER2024 (https://zeroqiaoba.github.io/MER2024-website) introduced a new track dedicated to open-vocabulary emotion recognition. This year, MER2025 centers on the theme ''When Affective Computing Meets Large Language Models (LLMs)''. We aim to shift the paradigm from traditional categorical frameworks reliant on predefined emotion taxonomies to LLM-driven generative methods, offering innovative solutions for more accurate and reliable emotion understanding. The challenge contains four tracks: MER-SEMI focuses on fixed categorical emotion recognition enhanced by semi-supervised learning; MER-FG explores fine-grained emotions, expanding recognition from basic to nuanced emotional states; MER-DES incorporates multimodal cues (beyond emotion words) into predictions to enhance model interpretability; MER-PR reveals whether emotion prediction results can improve personality recognition performance. For the first three tracks, the baseline code is available at MERTools (https://github.com/zeroQiaoba/MERTools) and datasets can be accessed via Hugging Face (https://huggingface.co/datasets/MERChallenge/MER2025). For the last track, the dataset and baseline code are available on GitHub (https://github.com/cai-cong/MER25_personality). Zheng Lian 0004, Rui Liu 0008, Kele Xu, Bin Liu 0041, Xuefei Liu, Yazhou Zhang 0001, Xin Liu 0012, Yong Li 0032, Zebang Cheng, Haolin Zuo, Ziyang Ma 0001, Xiaojiang Peng, Xie Chen 0001, Ya Li 0001, Erik Cambria, Guoying Zhao 0001, Björn W. Schuller, Jianhua Tao 0001 |
ACM Multimedia | 11 |
| 2025 | EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingabstractHuman speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made huge progress, they still face challenges in controlling the emotional expression in the generated speech. In this work, we propose EmoVoice, a novel emotion-controllable TTS model that exploits large language models (LLMs) to enable fine-grained freestyle natural language emotion control, and a phoneme boost variant design that makes the model output phoneme tokens and audio tokens in parallel to enhance content consistency, inspired by chain-of-thought (CoT) and modality-of-thought (CoM) techniques. Besides, we introduce EmoVoice-DB, a high-quality 40-hour English emotion dataset featuring expressive speech and fine-grained emotion labels with natural language descriptions. EmoVoice achieves state-of-the-art performance on the English EmoVoice-DB test set using only synthetic training data, and on the Chinese Secap test set using our in-house data. We further investigate the reliability of existing emotion evaluation metrics and their alignment with human perceptual preferences, and explore using SOTA multimodal LLMs GPT-4o-audio and Gemini to assess emotional speech. Dataset, code, checkpoints and demo samples are available at https://github.com/yanghaha0908/EmoVoice. Guanrou Yang, Qian Chen 0003, Ziyang Ma 0001, Wen Wang 0019, Tianrui Wang, Yifan Yang 0005, Zhikang Niu, Wenrui Liu 0003, Fan Yu 0002, Zhihao Du, Zhifu Gao, Shiliang Zhang, Xie Chen 0001 |
ACM Multimedia | 4 |
| 2025 | MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their MixabstractWe introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through iterative error corrections and quality checks to ensure high quality. Unlike existing benchmarks that are limited to specific domains of sound, music, or speech, MMAR extends them to a broad spectrum of real-world audio scenarios, including mixed-modality combinations of sound, music, and speech. Each question in MMAR is hierarchically categorized across four reasoning layers: Signal, Perception, Semantic, and Cultural, with additional sub-categories within each layer to reflect task diversity and complexity. To further foster research in this area, we annotate every question with a Chain-of-Thought (CoT) rationale to promote future advancements in audio reasoning. Each item in the benchmark demands multi-step deep reasoning beyond surface-level understanding. Moreover, a part of the questions requires graduate-level perceptual and domain-specific knowledge, elevating the benchmark's difficulty and depth. We evaluate MMAR using a broad set of models, including Large Audio-Language Models (LALMs), Large Audio Reasoning Models (LARMs), Omni Language Models (OLMs), Large Language Models (LLMs), and Large Reasoning Models (LRMs), with audio caption inputs. The performance of these models on MMAR highlights the benchmark's challenging nature, and our analysis further reveals critical limitations of understanding and reasoning capabilities among current models. These findings underscore the urgent need for greater research attention in audio-language reasoning, including both data and algorithm innovation. We hope MMAR will serve as a catalyst for future advances in this important but little-explored area. Ziyang Ma 0001, Yinghao Ma, Yanqiao Zhu 0003, Yi-Wen Chao, Yuanzhe Chen, Zhuo Chen 0006, Jian Cong, Keliang Li, Siyou Li, Xinfeng Li, Xiquan Li, Zheng Lian 0004, Yuzhe Liang, Minghao Liu 0003, Zhikang Niu, Tianrui Wang, Yuping Wang 0005, Yuxuan Wang 0002, Guanrou Yang, Jianwei Yu 0001, Ruibin Yuan, Zhisheng Zheng, Ziya Zhou, Haina Zhu, Wei Xue 0002, Emmanouil Benetos, Kai Yu 0004, Chng Eng Siong, Xie Chen 0001 |
NeurIPS | 1 |
| 2025 | Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech SynthesisabstractWhile emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level control. Achieving word-level expressive control poses fundamental challenges, primarily due to the complexity of modeling multi-emotion transitions and the scarcity of annotated datasets that capture intra-sentence emotional and prosodic variation. In this paper, we propose WeSCon, the first self-training framework that enables word-level control of both emotion and speaking rate in a pretrained zero-shot TTS model, without relying on datasets containing intra-sentence emotion or speed transitions.
Our method introduces a transition-smoothing strategy and a dynamic speed control mechanism to guide the pretrained TTS model in performing word-level expressive synthesis through a multi-round inference process. To further simplify the inference, we incorporate a dynamic emotional attention bias mechanism and fine-tune the model via self-training, thereby activating its ability for word-level expressive control in an end-to-end manner. Experimental results show that WeSCon effectively overcomes data scarcity, achieving state-of-the-art performance in word-level emotional expression control while preserving the strong zero-shot synthesis capabilities of the original TTS model. Tianrui Wang, Meng Ge, Chunyu Qiang, Ziyang Ma 0001, Zikang Huang, Guanrou Yang, Xiaobao Wang, Chng Eng Siong, Xie Chen 0001, Longbiao Wang, Jianwu Dang 0001 |
NeurIPS | 6 |
| 2024 | VoiceFlow: Efficient Text-To-Speech with Rectified Flow MatchingabstractAlthough diffusion models in text-to-speech have become a popular choice due to their strong generative ability, the intrinsic complexity of sampling from diffusion models harms their efficiency. Alternatively, we propose VoiceFlow, an acoustic model that utilizes a rectified flow matching algorithm to achieve high synthesis quality with a limited number of sampling steps. VoiceFlow formulates the process of generating mel-spectrograms into an ordinary differential equation conditional on text inputs, whose vector field is then estimated. The rectified flow technique then effectively straightens its sampling trajectory for efficient synthesis. Subjective and objective evaluations on both single and multi-speaker corpora showed the superior synthesis quality of VoiceFlow compared to the diffusion counterpart. Ablation studies further verified the validity of the rectified flow technique in VoiceFlow. Chenpeng Du, Ziyang Ma 0001, Xie Chen 0001, Kai Yu 0004 |
ICASSP | 3 |
| 2024 | Leveraging Speech PTM, Text LLM, And Emotional TTS For Speech Emotion RecognitionabstractIn this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and speech synthesis technique, Azure TTS. First, we investigated the representation ability of different speech self-supervised pre-trained models, and we found that data2vec has a good representation ability on the SER task. Second, we employed a powerful large language model (LLM), GPT-4, and emotional text-to-speech (TTS) model, Azure TTS, to generate emotionally congruent text and speech. We carefully designed the text prompt and dataset construction, to obtain the synthetic emotional speech data with high quality. Third, we studied different ways of data augmentation to promote the SER task with synthetic speech, including random mixing, adversarial training, transfer learning, and curriculum learning. Experiments and ablation studies on the IEMOCAP dataset demonstrate the effectiveness of our method, compared with other data augmentation methods, and data augmentation with other synthetic data. Ziyang Ma 0001, Wen Wu 0007, Zhisheng Zheng, Qian Chen 0003, Shiliang Zhang, Xie Chen 0001 |
ICASSP | 1 |
| 2024 | Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTSabstractSelf-supervised learning (SSL) proficiency in speech-related tasks has driven research into utilizing discrete tokens for speech tasks like recognition and translation, which offer lower storage requirements and great potential to employ natural language processing techniques. However, these studies, mainly single-task focused, faced challenges like overfitting and performance degradation in speech recognition tasks, often at the cost of sacrificing performance in multi-task scenarios. This study presents a comprehensive comparison and optimization of discrete tokens generated by various leading SSL models in speech recognition and synthesis tasks. We aim to explore the universality of speech discrete tokens across multiple speech tasks. Experimental results demonstrate that discrete tokens achieve comparable results against systems trained on FBank features in speech recognition tasks and outperform mel-spectrogram features in speech synthesis in subjective and objective metrics. These findings suggest that universal discrete tokens have enormous potential in various speech-related tasks. Our work is open-source and publicly available at https://github.com/k2-fsa/icefall. Yifan Yang 0005, Feiyu Shen, Chenpeng Du, Ziyang Ma 0001, Kai Yu 0004, Daniel Povey, Xie Chen 0001 |
ICASSP | 4 |
| 2024 | Hourglass-AVSR: Down-Up Sampling-Based Computational Efficiency Model for Audio-Visual Speech RecognitionabstractRecently audio-visual speech recognition (AVSR), which better leverages video modality as additional information to extend automatic speech recognition (ASR), has shown promising results in complex acoustic environments. However, there is still substantial space to improve as complex computation of visual modules and ineffective fusion of audio-visual modalities. To eliminate these drawbacks, we propose a down-up sampling-based AVSR model (Hourglass-AVSR) to enjoy high efficiency and performance, whose time length is scaled during the intermediate processing, resembling an hourglass. Firstly, we propose a context and residual aware video upsampling approach to improve the recognition performance, which utilizes contextual information from visual representations and captures residual information between adjacent video frames. Secondly, we introduce a visual-audio alignment approach during the upsampling by explicitly incorporating boundary constraint loss. Besides, we propose a cross-layer attention fusion to capture the modality dependencies within each visual encoder layer. Experiments conducted on the MISP-AVSR dataset reveal that our proposed Hourglass-AVSR model outperforms ASR model by 12.9% and 20.8% relative concatenated minimum permutation character error rate (cpCER) reduction on far-field and middle-field test sets, respectively. Moreover, compared to other state-of-the-art AVSR models, our model exhibits the highest improvement in cpCER for the visual module. Furthermore, on the benefit of our down-up sampling approach, Hourglass-AVSR model reduces 54.2% overall computation costs with minor performance degradation. Fan Yu 0002, Haoxu Wang, Ziyang Ma 0001, Shiliang Zhang |
ICASSP | 3 |
| 2024 | BAT: Learning to Reason about Spatial Sounds with Large Language ModelsabstractSpatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analysis model with the natural language reasoning capabilities of a large language model (LLM) to replicate this innate ability. To address the lack of existing datasets of in-the-wild spatial sounds, we synthesized a binaural audio dataset using AudioSet and SoundSpaces 2.0. Next, we developed SpatialSoundQA, a spatial sound-based question-answering dataset, offering a range of QA tasks that train BAT in various aspects of spatial sound perception and reasoning. The acoustic front end encoder of BAT is a novel spatial audio encoder named Spatial Audio Spectrogram Transformer, or Spatial-AST, which by itself achieves strong performance across sound event detection, spatial localization, and distance estimation. By integrating Spatial-AST with LLaMA-2 7B model, BAT transcends standard Sound Event Localization and Detection (SELD) tasks, enabling the model to reason about the relationships between the sounds in its environment. Our experiments demonstrate BAT’s superior performance on both spatial sound perception and reasoning, showcasing the immense potential of LLMs in navigating and interpreting complex spatial audio environments. Zhisheng Zheng, Puyuan Peng, Ziyang Ma 0001, Xie Chen 0001, Eunsol Choi, David F. Harwath |
ICML | 3 |
| 2024 | EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
Yuzhe Liang, Ziyang Ma 0001, Zhisheng Zheng, Xie Chen 0001 |
IJCAI | 3 |
| 2024 | EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
Ziyang Ma 0001, Hezhao Zhang, Zhisheng Zheng, Xiquan Li, Jiaxin Ye, Xie Chen 0001, Thomas Hain |
INTERSPEECH | 1 |
| 2024 | TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers
Yakun Song, Zhuo Chen 0006, Ziyang Ma 0001, Guanrou Yang, Xie Chen 0001 |
INTERSPEECH | 4 |
| 2024 | LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR
Zheshu Song, Jianheng Zhuo, Yifan Yang 0005, Ziyang Ma 0001, Shixiong Zhang 0001, Xie Chen 0001 |
INTERSPEECH | 4 |
| 2024 | MaLa-ASR: Multimedia-Assisted LLM-Based ASR
Guanrou Yang, Ziyang Ma 0001, Fan Yu 0002, Zhifu Gao, Shiliang Zhang, Xie Chen 0001 |
INTERSPEECH | 2 |
| 2024 | NDVQ: Robust Neural Audio Codec With Normal Distribution-Based Vector QuantizationabstractBuilt upon vector quantization (VQ), discrete audio codec models have achieved great success in audio compression and autoregressive audio generation. However, existing models face substantial challenges in perceptual quality and signal distortion, especially when operating in extremely low bandwidth, rooted in the sensitivity of the VQ codebook to noise. This degradation poses significant challenges for several downstream tasks, such as codec-based speech synthesis. To address this issue, we propose a novel VQ method, Normal Distribution-based Vector Quantization (NDVQ), by introducing an explicit margin between the VQ codes via learning a variance. Specifically, our approach involves mapping the waveform to a latent space and quantizing it by selecting the most likely normal distribution, with each codebook entry representing a unique normal distribution defined by its mean and variance. Using these distribution-based VQ codec codes, a decoder reconstructs the input waveform. NDVQ is trained with additional distribution-related losses, alongside reconstruction and discrimination losses. Experiments demonstrate that NDVQ outperforms existing audio compression baselines, such as EnCodec, in terms of audio quality and zero-shot TTS, particularly in very low bandwidth scenarios. Zhikang Niu, Sanyuan Chen, Ziyang Ma 0001, Xie Chen 0001, Shujie Liu 0001 |
SLT | 4 |
| 2024 | CTC-Assisted LLM-Based Contextual ASRabstractContextual ASR or hotword customization holds substantial practical value. Despite the impressive performance of current end-to-end (E2E) automatic speech recognition (ASR) systems, they often face challenges in accurately recognizing rare words. Typical E2E contextual ASR models commonly feature complex architectures and decoding mechanisms, limited in performance and susceptible to interference from distractor words. With large language model (LLM)-based ASR models emerging as the new mainstream, we propose a CTC-Assisted LLM-Based Contextual ASR model with an efficient filtering algorithm. By using coarse CTC decoding results to filter potential relevant hotwords and incorporating them into LLM prompt input, our model attains WER/B-WER of $1.27 \% / 3.67 \%$ and $2.72 \% / 8.02 \%$ on the Librispeech test-clean and test-other sets targeting on recognizing rare long-tail words, demonstrating significant improvements compared to the baseline LLM-based ASR model, and substantially surpassing other related work. More remarkably, with the help of the large language model and proposed filtering algorithm, our contextual ASR model still performs well with 2000 biasing words.1 Guanrou Yang, Ziyang Ma 0001, Zhifu Gao, Shiliang Zhang, Xie Chen 0001 |
SLT | 2 |
| 2024 | E$^{3}$TTS: End-to-End Text-Based Speech Editing TTS System and Its ApplicationsabstractText-based speech editing aims at manipulating part of real audio by modifying the corresponding transcribed text, without being discernible by human auditory system. With the enhanced capability of neural Text-to-speech (TTS), researchers try to tackle speech editing problems with TTS methods. In this paper, we propose E$^{3}$TTS, a.k.a. end-to-end text-based speech editing TTS system, which combines a text encoder, a speech encoder, and a joint net for speech synthesis and speech editing. E$^{3}$TTS can insert, replace, and delete speech content at will, by manipulating the given text. Experiments show that our speech editing outperforms strong baselines on HiFiTTS and LibriTTS datasets, speakers of which are seen or unseen, respectively. Further, we introduce E$^{3}$TTS into data augmentation for automatic speech recognition (ASR) to mitigate the data insufficiency problem in code-switching and named entity recognition scenarios1. E$^{3}$TTS retains the coherence and reality of the recorded audio compared to past data augmentation methods. The experimental results show significant performance improvements over baseline systems with traditional TTS-based data augmentation. The code and samples of the proposed speech editing model are available at this repository.2 Ziyang Ma 0001, Chenpeng Du, Kai Yu 0004, Xie Chen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | Towards Weakly Supervised Text-to-Audio GroundingabstractText-to-audio grounding (TAG) task aims to predict the onsets and offsets of sound events described by natural language. This task can facilitate applications such as multimodal information retrieval. This paper focuses on weakly-supervised text-to-audio grounding (WSTAG), where frame-level annotations of sound events are unavailable, and only the caption of a whole audio clip can be utilized for training. WSTAG is superior to strongly-supervised approaches in its scalability to large audio-text datasets. Two WSTAG frameworks are studied in this paper: sentence-level and phrase-level. First, we analyze the limitations of mean pooling used in the previous WSTAG approach and investigate the effects of different pooling strategies. We then propose phrase-level WSTAG to use matching labels between audio clips and phrases for training. Advanced negative sampling strategies and self-supervision are proposed to enhance the accuracy of the weak labels and provide pseudo strong labels. Experimental results show that our system significantly outperforms previous WSTAG methods. Finally, we conduct extensive experiments to analyze the effects of several factors on phrase-level WSTAG. Xuenan Xu, Ziyang Ma 0001, Mengyue Wu, Kai Yu 0004 |
IEEE Trans. Multim. | 2 |
| 2023 | Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech RecognitionabstractSelf-supervised learning (SSL) has achieved great success in speech processing, but always with a large model size to increase the modeling capacity. This may limit its potential applications due to the expensive computation and memory costs introduced by the oversize model. Compression for SSL models has become an important research direction of practical value. To this end, we explore the effective distillation of HuBERT-based SSL models for automatic speech recognition. First, a comprehensive study of different student model structures is conducted. On top of this, as a supplement to the regression loss widely adopted in previous works, a discriminative loss is introduced for HuBERT to enhance the distillation performance, especially in low-resource scenarios. In addition, we design a simple and effective algorithm to distill the front-end input from waveform to Fbank feature, resulting in 17% parameter reduction and doubling inference speed, at marginal performance degradation. Changli Tang, Ziyang Ma 0001, Zhisheng Zheng, Xie Chen 0001, Weiqiang Zhang 0001 |
ASRU | 3 |
| 2023 | Fast-Hubert: an Efficient Training Framework for Self-Supervised Speech Representation LearningabstractRecent years have witnessed significant advancements in self-supervised learning (SSL) methods for speech-processing tasks. Various speech-based SSL models have been developed and present promising performance on a range of downstream tasks including speech recognition. However, existing speech-based SSL models face a common dilemma in terms of computational cost, which might hinder their potential application and in-depth academic research. To address this issue, we first analyze the computational cost of different modules during HuBERT pre-training and then introduce a stack of efficiency optimizations, which is named Fast-HuBERT in this paper. The proposed Fast-HuBERT can be trained in 1.1 days with 8 V100 GPUs on the Librispeech 960 h benchmark, without performance degradation, resulting in a 5.2x speedup, compared to the original implementation. Moreover, we explore two well-studied techniques in the Fast-HuBERT and demonstrate consistent improvements as reported in previous work.11The code for Fast-HuBERT training is available at https://github.com/yanghaha0908/FastHuBERT Guanrou Yang, Ziyang Ma 0001, Zhisheng Zheng, Yakun Song, Zhikang Niu, Xie Chen 0001 |
ASRU | 2 |
| 2023 | Improving Few-Shot Learning for Talking Face System with TTS Data AugmentationabstractAudio-driven talking face has attracted broad interest from academia and industry recently. However, data acquisition and labeling in audio-driven talking face are labor-intensive and costly. The lack of data resource results in poor synthesis effect. To alleviate this issue, we propose to use TTS (Text-To-Speech) for data augmentation to improve few-shot ability of the talking face system. The misalignment problem brought by the TTS audio is solved with the introduction of soft-DTW, which is first adopted in the talking face task. Moreover, features extracted by HuBERT are explored to utilize underlying information of audio, and found to be superior over other features. The proposed method achieves 17%, 14%, 38% dominance on MSE score, DTW score and user study preference repectively over the baseline model, which shows the effectiveness of improving few-shot learning for talking face system with TTS augmentation. Qi Chen 0009, Ziyang Ma 0001, Tao Liu 0068, Xu Tan 0003, Qu Lu, Kai Yu 0004, Xie Chen 0001 |
ICASSP | 2 |
| 2023 | Front-End Adapter: Adapting Front-End Input of Speech Based Self-Supervised Learning for Speech RecognitionabstractRecent years have witnessed a boom in self-supervised learning (SSL) in various areas including speech processing. Speech based SSL models present promising performance in a range of speech related tasks. However, the training of SSL models is computationally expensive and a common practice is to fine-tune a released SSL model on the specific task. It is essential to use consistent front-end input during pre-training and fine-tuning. This consistency may introduce potential issues when the optimal front-end is not the same as that used in pre-training. In this paper, we propose a simple but effective front-end adapter to address this front-end discrepancy. By minimizing the distance between the outputs of different front-ends, the filterbank feature (Fbank) can be compatible with SSL models which are pre-trained with waveform. The experiment results demonstrate the effectiveness of our proposed front-end adapter on several popular SSL models for the speech recognition task. Xie Chen 0001, Ziyang Ma 0001, Changli Tang, Zhisheng Zheng |
ICASSP | 2 |
| 2023 | Improving Code-Switching and Name Entity Recognition in ASR with Speech Editing based Data Augmentation
Zheshu Song, Ziyang Ma 0001, Chenpeng Du, Kai Yu 0004, Xie Chen 0001 |
INTERSPEECH | 3 |
| 2023 | MT4SSL: Boosting Self-Supervised Speech Representation Learning by Integrating Multiple Targets
Ziyang Ma 0001, Zhisheng Zheng, Changli Tang, Xie Chen 0001 |
INTERSPEECH | 1 |
| 2023 | Pushing the Limits of Unsupervised Unit Discovery for SSL Speech Representation
Ziyang Ma 0001, Zhisheng Zheng, Guanrou Yang, Yu Wang 0027, Chao Zhang 0031, Xie Chen 0001 |
INTERSPEECH | 1 |
| 2023 | Unsupervised Active Learning: Optimizing Labeling Cost-Effectiveness for Automatic Speech Recognition
Zhisheng Zheng, Ziyang Ma 0001, Yu Wang 0027, Xie Chen 0001 |
INTERSPEECH | 2 |
| 2021 | Hierarchical Deep Residual Reasoning for Temporal Moment LocalizationabstractTemporal Moment Localization (TML) in untrimmed videos is a challenging task in the field of multimedia, which aims at localizing the start and end points of the activity in the video, described by a sentence query. Existing methods mainly focus on mining the correlation between video and sentence representations or investigating the fusion manner of the two modalities. These works mainly understand the video and sentence coarsely, ignoring the fact that a sentence can be understood from various semantics, and the dominant words affecting the moment localization in the semantics are the action and object reference. Toward this end, we propose a Hierarchical Deep Residual Reasoning (HDRR) model, which decomposes the video and sentence into multi-level representations with different semantics to achieve a finer-grained localization. Furthermore, considering that videos with different resolution and sentences with different length have different difficulty in understanding, we design the simple yet effective Res-BiGRUs for feature fusion, which is able to grasp the useful information in a self-adapting manner. Extensive experiments conducted on Charades-STA and ActivityNet-Captions datasets demonstrate the superiority of our HDRR model compared with other state-of-the-art methods. Ziyang Ma 0001, Xianjing Han, Xuemeng Song, Yiran Cui, Liqiang Nie |
MMAsia | 1 |