Songjun Cao

dblp:264/4871 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021
YearPublicationVenuePosition
2026 Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception
abstract
The rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia security and trust. While existing countermeasures (CMs) perform well in single-type audio deepfake detection (ADD), their performance declines in cross-type scenarios. This paper is dedicated to studying the all-type ADD task. We are the first to comprehensively establish an all-type ADD benchmark to evaluate current CMs, incorporating cross-type deepfake detection across speech, sound, singing voice, and music. Then, we introduce the prompt tuning self-supervised learning (PT-SSL) training paradigm, which optimizes SSL front-end by learning specialized prompt tokens for ADD, requiring 458× fewer trainable parameters than fine-tuning (FT). Considering the auditory perception of different audio types, we propose the wavelet prompt tuning (WPT)-SSL method to capture type-invariant auditory deepfake information from the frequency domain without requiring additional training parameters, thereby enhancing performance over FT in the all-type ADD task. To achieve an universally CM, we utilize all types of deepfake audio for co-training. Experimental results demonstrate that WPT-XLSR-AASIST achieved the best performance, with an average EER of 3.58% across all evaluation sets.
Yuankun Xie, Ruibo Fu, Songjun Cao, Haonan Cheng, Long Ye
AAAI5
2026 OpenST: Toward open-set source tracing for neural codec deepfake audio
Yuankun Xie, Ruibo Fu, Zhengqi Wen, Songjun Cao, Chenxing Li, Haonan Cheng, Long Ye
Neurocomputing6
2025 M-MoE: Mixture of Mixture-of-Expert Model for CTC-based Streaming Multilingual ASR
abstract
The Mixture-of-Expert (MoE) structure has been effectively utilized in multilingual ASR tasks. However, the potential of external language information remains underutilized. In this paper, we introduce the Mixture of MoE (M-MoE) structure, featuring multiple language-specific MoEs and a language-unknown MoE. The language-unknown MoE reuses experts from language-specific MoEs. Inputs with language IDs are directed to language-specific MoEs, while those without IDs go to the language-unknown MoE. We propose a two-stage training method for the M-MoE-based model. Our unified model structure is suitable for streaming ASR tasks in both language-known and language-unknown scenarios. Experiments on a three-language dataset show that compared to the Conformer baseline, our model achieves an average of 12% and 9% relative improvement in language-known and language-unknown scenarios. Compared to the strong MoE baseline, there is an average 5% relative improvement in the language-known scenario.
Songjun Cao
ICASSP1
2025 DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models
abstract
Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding of conversational context. However, existing CSS systems are limited to deterministic prediction, overlooking the diversity of potential responses. Moreover, they rarely employ language model (LM)-based TTS backbones, limiting the naturalness and quality of synthesized speech. To address these issues, in this paper, we propose DiffCSS, an innovative CSS framework that leverages diffusion models and an LM-based TTS backbone to generate diverse, expressive, and contextually coherent speech. A diffusion-based context-aware prosody predictor is proposed to sample diverse prosody embeddings conditioned on multimodal conversational context. Then a prosody-controllable LM-based TTS backbone is developed to synthesize high-quality speech with sampled prosody embeddings. Experimental results demonstrate that the synthesized speech from DiffCSS is more diverse, contextually coherent, and expressive than existing CSS systems1.
Weihao Wu 0001, Yixuan Zhou 0002, Jingbei Li, Rui Niu, Songjun Cao, Zhiyong Wu 0001
ICASSP7
2025 MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
Yueteng Kang, Songjun Cao, Qiulin Li
INTERSPEECH3
2025 Monotonic Attention for Robust Text-to-Speech Synthesis in Large Language Model Frameworks
Songjun Cao
INTERSPEECH5
2025 SonarGuard2: Ultrasonic Face Liveness Detection Based on Adaptive Doppler Effect Feature Extraction
Ke-Yue Zhang, Taiping Yao, Songjun Cao, Shouhong Ding
INTERSPEECH4
2024 A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
Yangze Li, Songjun Cao, Lei Xie 0001
INTERSPEECH3
2022 Improving CTC-Based Speech Recognition Via Knowledge Transferring from Pre-Trained Language Models
abstract
Recently, end-to-end automatic speech recognition models based on connectionist temporal classification (CTC) have achieved impressive results, especially when fine-tuned from wav2vec2.0 models. Due to the conditional independence assumption, CTC-based models are always weaker than attention-based encoder-decoder models and require the assistance of external language models (LMs). To solve this issue, we propose two knowledge transferring methods that leverage pre-trained LMs, such as BERT and GPT2, to improve CTC-based models. The first method is based on representation learning, in which the CTC-based models use the representation produced by BERT as an auxiliary learning target. The second method is based on joint classification learning, which combines GPT2 for text modeling with a hybrid CTC/attention architecture. Experiment on AISHELL-1 corpus yields a character error rate (CER) of 4.2% on the test set. When compared to the vanilla CTC-based models fine-tuned from the wav2vec2.0 models, our knowledge transferring method reduces CER by 16.1% relatively without external LMs.
Keqi Deng, Songjun Cao, Gaofeng Cheng, Pengyuan Zhang
ICASSP2
2022 Censer: Curriculum Semi-supervised Learning for Speech Recognition Based on Self-supervised Pre-training
abstract
Recent studies have shown that the benefits provided by selfsupervised pre-training and self-training (pseudo-labeling) are complementary.Semi-supervised fine-tuning strategies under the pre-training framework, however, remain insufficiently studied.Besides, modern semi-supervised speech recognition algorithms either treat unlabeled data indiscriminately or filter out noisy samples with a confidence threshold.The dissimilarities among different unlabeled data are often ignored.In this paper, we propose Censer, a semi-supervised speech recognition algorithm based on self-supervised pre-training to maximize the utilization of unlabeled data.The pre-training stage of Censer adopts wav2vec2.0and the fine-tuning stage employs an improved semisupervised learning algorithm from slimIPL, which leverages unlabeled data progressively according to their pseudo labels' qualities.We also incorporate a temporal pseudo label pool and an exponential moving average to control the pseudo labels' update frequency and to avoid model divergence.Experimental results on Libri-Light and LibriSpeech datasets manifest our proposed method achieves better performance compared to existing approaches while being more unified.
Songjun Cao, Takahiro Shinozaki
INTERSPEECH2
2021 Improving Hybrid CTC/Attention End-to-End Speech Recognition with Pretrained Acoustic and Language Models
abstract
Recently, self-supervised pretraining has achieved impressive results in end-to-end (E2E) automatic speech recognition (ASR). However, the dominant sequence-to-sequence (S2S) E2E model is still hard to fully utilize the self-supervised pretraining methods because its decoder is conditioned on acoustic representation thus cannot be pretrained separately. In this paper, we propose a pretrained Transformer (Preformer) S2S ASR architecture based on hybrid CTC/attention E2E models to fully utilize the pretrained acoustic models (AMs) and language models (LMs). In our framework, the encoder is initialized with a pretrained AM (wav2vec2.0). The Preformer leverages CTC as an auxiliary task during training and inference. Furthermore, we design a one-cross decoder (OCD), which relaxes the dependence on acoustic representations so that it can be initialized with pretrained LM (DistilGPT2). Experiments are conducted on the AISHELL-1 corpus and achieve a 4.6% character error rate (CER) on the test set. Compared with our vanilla hybrid CTC/attention Transformer baseline, our proposed CTC/attention-based Preformer yields 27% relative CER reduction. To the best of our knowledge, this is the first work to utilize both pretrained AM and LM in a S2S ASR system.
Keqi Deng, Songjun Cao
ASRU2
2021 Improving Streaming Transformer Based ASR Under a Framework of Self-Supervised Learning
abstract
Recently self-supervised learning has emerged as an effective approach to improve the performance of automatic speech recognition (ASR).Under such a framework, the neural network is usually pre-trained with massive unlabeled data and then fine-tuned with limited labeled data.However, the nonstreaming architecture like bidirectional transformer is usually adopted by the neural network to achieve competitive results, which can not be used in streaming scenarios.In this paper, we mainly focus on improving the performance of streaming transformer under the self-supervised learning framework.Specifically, we propose a novel two-stage training method during finetuning, which combines knowledge distilling and self-training.The proposed training method achieves 16.3% relative word error rate (WER) reduction on Librispeech noisy test set.Finally, by only using the 100h clean subset of Librispeech as the labeled data and the rest (860h) as the unlabeled data, our streaming transformer based model obtains competitive WERs 3.5/8.7 on Librispeech clean/noisy test sets.
Songjun Cao, Yueteng Kang, Yanzhe Fu, Xiaoshuo Xu, Sining Sun
Interspeech1
2021 Improving Accent Identification and Accented Speech Recognition Under a Framework of Self-Supervised Learning
abstract
Recently, self-supervised pre-training has gained success in automatic speech recognition (ASR). However, considering the difference between speech accents in real scenarios, how to identify accents and use accent features to improve ASR is still challenging. In this paper, we employ the self-supervised pre-training method for both accent identification and accented speech recognition tasks. For the former task, a standard deviation constraint loss (SDC-loss) based end-to-end (E2E) architecture is proposed to identify accents under the same language. As for accented speech recognition task, we design an accent-dependent ASR system, which can utilize additional accent input features. Furthermore, we propose a frame-level accent feature, which is extracted based on the proposed accent identification model and can be dynamically adjusted. We pre-train our models using 960 hours unlabeled LibriSpeech dataset and fine-tune them on AESRC2020 speech dataset. The experimental results show that our proposed accent-dependent ASR system is significantly ahead of the AESRC2020 baseline and achieves $6.5\%$ relative word error rate (WER) reduction compared with our accent-independent ASR system.
Keqi Deng, Songjun Cao
Interspeech2
2021 Explore wav2vec 2.0 for Mispronunciation Detection
Xiaoshuo Xu, Yueteng Kang, Songjun Cao, Binghuai Lin
Interspeech3
2021 Improving Speech Recognition Accuracy of Local POI Using Geographical Models
abstract
Nowadays voice search for points of interest (POI) is becoming increasingly popular. However, speech recognition for local POI names still remains a challenge due to multi-dialect and long-tailed distribution of POI names. This paper improves speech recognition accuracy for local POI from two aspects. Firstly, a geographic acoustic model (Geo-AM) is proposed. The proposed Geo-AM deals with multi-dialect problem using dialect-specific input feature and dialect-specific top layers. Secondly, a group of geo-specific language models (Geo-LMs) are integrated into our speech recognition system to improve recognition accuracy of long-tailed and homophone POI names. During decoding, a specific Geo-LM is selected on-demand according to the user's geographic location. Experiments show that the proposed Geo-AM achieves 6.5%~10.1% relative character error rate (CER) reduction on an accent test set and the proposed Geo-AM and Geo-LMs totally achieve over 18.7% relative CER reduction on a voice search task for Tencent Map.
Songjun Cao
SLT1