Zexin Cai

dblp:218/6151 · DBLP profile ↗
← Back
17ranked-venue papers
9as first author
12since 2021 · last 2025
0009-0003-4495-1676ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 5 since 2021
YearPublicationVenuePosition
2025 GenVC: Self-Supervised Zero-Shot Voice Conversion
abstract
Most current zero-shot voice conversion methods rely on externally supervised components, particularly speaker encoders, for training. To explore alternatives that eliminate this dependency, this paper introduces GenVC, a novel framework that disentangles speaker identity and linguistic content from speech signals in a self-supervised manner. GenVC leverages speech tokenizers and an autoregressive, Transformer-based language model as its backbone for speech generation. This design supports large-scale training while enhancing both source speaker privacy protection and target speaker cloning fidelity. Experimental results demonstrate that GenVC achieves notably higher speaker similarity, with naturalness on par with leading zero-shot approaches. Moreover, due to its autoregressive formulation, GenVC introduces flexibility in temporal alignment, reducing the preservation of source prosody and speaker-specific traits, and making it highly effective for voice anonymization.11Audio samples, code, and model checkpoints are available at https://caizexin.github.io/GenVC/index.html
Zexin Cai, Henry Li Xinyuan, Ashi Garg, L. Paola García-Perera, Kevin Duh, Sanjeev Khudanpur, Matthew Wiesner, Nicholas Andrews
ASRU1
2025 Rapidly Adapting to New Voice Spoofing: Few-Shot Detection of Synthesized Speech Under Distribution Shifts
abstract
We address the challenge of detecting synthesized speech under distribution shifts—arising from unseen synthesis methods, speakers, languages, or audio conditions—relative to the training data. Fewshot learning methods are a promising way to tackle distribution shifts by rapidly adapting on the basis of a few in-distribution samples. We propose a self-attentive prototypical network to enable more robust fewshot adaptation. To evaluate our approach, we systematically compare the performance of traditional zero-shot detectors and the proposed fewshot detectors, carefully controlling training conditions to introduce distribution shifts at evaluation time. In conditions where distribution shifts hamper the zero-shot performance, our proposed few-shot adaptation technique can quickly adapt using as few as $\mathbf{1 0}$ in-distribution samples—achieving upto 32% relative EER reduction on deepfakes in Japanese language and 20% relative reduction on ASVspoof 2021 Deepfake dataset.
Ashi Garg, Zexin Cai, Henry Li Xinyuan, L. Paola García-Perera, Sanjeev Khudanpur, Matthew Wiesner, Nicholas Andrews
ASRU2
2025 Scalable Controllable Accented TTS
abstract
We propose a method to scale accented TTS training to large, accent-diverse datasets that often lack consistent, high-quality accent labels. Our approach relies on a speech geolocation model to infer accent labels directly from audio. To improve speaker generalization and encourage disentangling speaker from accent we explore timbre augmentation through kNN voice conversion. We validate our approach on CommonVoice by fine-tuning XTTS-v2 with accent labels inferred or improved via geolocation. According to various automated metrics based on embeddings extracted from an accent identification model, the resulting accented TTS model produces speech with better accent fidelity compared to XTTS-v2 fine-tuned on self-reported accent labels in CommonVoice, or other existing accented TTS models. According to human evaluation, it was clear that the geolocation model based data discovery and enhancement improved the naturalness and accent fidelity of generated speech. However, the effect of different data augmentation strategies was less clear.
Henry Li Xinyuan, Zexin Cai, Ashi Garg, Kevin Duh, L. Paola García-Perera, Sanjeev Khudanpur, Nicholas Andrews, Matthew Wiesner
ASRU2
2025 HLTCOE Submission to the VoicePrivacy Attacker Challenge
abstract
We describe our submission to the 2024 VoicePrivacy Attacker Challenge. We propose three main categories of methods to improve ASV performance against anonymized speech: improvements to the underlying classifier, alternative distance metrics when computing ASV scores, and kNN-VC normalization. By simultaneously employing one or more of these methods, we were able to achieve a significant reduction in EER against all of the submitted anonymization systems in the VoicePrivacy Challenge.
Henry Li Xinyuan, Ashi Garg, Zexin Cai, Kevin Duh, L. Paola García-Perera, Sanjeev Khudanpur, Nicholas Andrews, Matthew Wiesner
ICASSP3
2024 Invertible Voice Conversion with Parallel Data
abstract
This paper introduces an innovative deep learning framework for parallel voice conversion to mitigate inherent risks associated with such systems. Our approach focuses on developing an invertible model capable of countering potential spoofing threats. Specifically, we present a conversion model that allows for the retrieval of source voices, thereby facilitating the identification of the source speaker. This framework is constructed using a series of invertible modules composed of affine coupling layers to ensure the reversibility of the conversion process. We conduct comprehensive training and evaluation of the proposed framework using parallel training data. Our experimental results reveal that this approach achieves comparable performance to non-invertible systems in voice conversion tasks. Notably, the converted outputs can be seamlessly reverted to the original source inputs using the same parameters employed during the forwarding process. This advancement holds considerable promise for elevating the security and reliability of voice conversion.
Zexin Cai, Ming Li 0026
ICASSP1
2024 Privacy Versus Emotion Preservation Trade-Offs in Emotion-Preserving Speaker Anonymization
abstract
Advances in speech technology now allow unprecedented access to personally identifiable information through speech. To protect such information, the differential privacy field has explored ways to anonymize speech while preserving its utility, including linguistic and paralinguistic aspects. However, anonymizing speech while maintaining emotional state remains challenging. We explore this problem in the context of the VoicePrivacy 2024 challenge. Specifically, we developed various speaker anonymization pipelines and find that approaches either excel at anonymization or preserving emotion state, but not both simultaneously. Achieving both would require an in-domain emotion recognizer. Additionally, we found that it is feasible to train a semi-effective speaker verification system using only emotion representations, demonstrating the challenge of separating these two modalities.
Zexin Cai, Henry Li Xinyuan, Ashi Garg, L. Paola García-Perera, Kevin Duh, Sanjeev Khudanpur, Nicholas Andrews, Matthew Wiesner
SLT1
2024 The Database and Benchmark For the Source Speaker Tracing Challenge 2024
abstract
Voice conversion (VC) systems can transform audio to mimic another speaker’s voice, thereby attacking speaker verification (SV) systems. However, ongoing studies on source speaker verification (SSV) are hindered by limited data availability and methodological constraints. This paper presents the Source Speaker Tracking Challenge (SSTC) on STL 2024, which aims to fill the gap in the database and benchmark for the SSV task. In this study, we generate a large-scale converted speech database with 16 common VC methods and train a batch of baseline systems based on the MFA-Conformer architecture. In addition, we introduced a related task called conversion method recognition, with the aim of assisting the SSV task. We expect SSTC to be a platform for advancing the development of the SSV task and provide further insights into the performance and limitations of current SV systems against VC attacks. Further details about SSTC can be found here1.1https://sstc-challenge.github.io/
Ze Li 0003, Yuke Lin, Hongbin Suo, Pengyuan Zhang, Yanzhen Ren, Zexin Cai, Hiromitsu Nishizaki, Ming Li 0026
SLT7
2024 Integrating frame-level boundary detection and deepfake detection for locating manipulated regions in partially spoofed audio forgery attacks
Zexin Cai, Ming Li 0026
Comput. Speech Lang.1
2023 Identifying Source Speakers for Voice Conversion Based Spoofing Attacks on Speaker Verification Systems
abstract
An automatic speaker verification system aims to verify the speaker identity of a speech signal. However, a voice conversion system could manipulate a person’s speech signal to make it sound like another speaker’s voice and deceive the speaker verification system. Most countermeasures for voice conversion-based spoofing attacks are designed to discriminate bona fide speech from spoofed speech for speaker verification systems. In this paper, we investigate the problem of source speaker identification – inferring the identity of the source speaker given the voice converted speech. To perform source speaker identification, we simply add voice-converted speech data with the label of source speaker identity to the genuine speech dataset during speaker embedding network training. Experimental results show the feasibility of source speaker identification when training and testing with converted speeches from the same voice conversion model(s). In addition, our results demonstrate that having more converted utterances from various voice conversion model for training helps improve the source speaker identification performance on converted utterances from unseen voice conversion models.
Danwei Cai, Zexin Cai, Ming Li 0026
ICASSP2
2023 Waveform Boundary Detection for Partially Spoofed Audio
abstract
The present paper proposes a waveform boundary detection system for audio spoofing attacks containing partially manipulated segments. Partially spoofed/fake audio, where part of the utterance is replaced, either with synthetic or natural audio clips, has recently been reported as one scenario of audio deepfakes. As deepfakes can be a threat to social security, the detection of such spoofing audio is essential. Accordingly, we propose to address the problem with a deep learning-based frame-level detection system that can detect partially spoofed audio and locate the manipulated pieces. Our proposed method is trained and evaluated on data provided by the ADD2022 Challenge. We evaluate our detection model concerning various acoustic features and network configurations. As a result, our detection system achieves an equal error rate (EER) of 6.58% on the ADD2022 challenge test set, which is the best performance in partially spoofed audio detection systems that can locate manipulated clips.
Zexin Cai, Weiqing Wang 0004, Ming Li 0026
ICASSP1
2023 Cross-lingual multi-speaker speech synthesis with limited bilingual training data
Zexin Cai, Yaogen Yang, Ming Li 0026
Comput. Speech Lang.1
2022 SIG-VC: A Speaker Information Guided Zero-Shot Voice Conversion System for Both Human Beings and Machines
abstract
Nowadays, as more and more systems achieve good performance in traditional voice conversion (VC) tasks, people’s attention gradually turns to VC tasks under extreme conditions. In this paper, we propose a novel method for zero-shot voice conversion. We aim to obtain intermediate representations for speaker-content disentanglement of speech to better remove speaker information and get pure content information. Accordingly, our proposed framework contains a module that removes the speaker information from the acoustic feature of the source speaker. Moreover, speaker information control is added to our system to maintain the voice cloning performance. The proposed system is evaluated by subjective and objective metrics. Results show that our proposed system significantly reduces the trade-off problem in zero-shot voice conversion, while it also manages to have high spoofing power to the speaker verification system.
Zexin Cai, Xiaoyi Qin, Ming Li 0026
ICASSP2
2020 From Speaker Verification to Multispeaker Speech Synthesis, Deep Transfer with Feedback Constraint
abstract
High-fidelity speech can be synthesized by end-to-end text-tospeech models in recent years.However, accessing and controlling speech attributes such as speaker identity, prosody, and emotion in a text-to-speech system remains a challenge.This paper presents a system involving feedback constraints for multispeaker speech synthesis.We manage to enhance the knowledge transfer from the speaker verification to the speech synthesis by engaging the speaker verification network.The constraint is taken by an added loss related to the speaker identity, which is centralized to improve the speaker similarity between the synthesized speech and its natural reference audio.The model is trained and evaluated on publicly available datasets.Experimental results, including visualization on speaker embedding space, show significant improvement in terms of speaker identity cloning in the spectrogram level.In addition, synthesized samples are available online for listening. 1
Zexin Cai, Chuxiong Zhang, Ming Li 0026
INTERSPEECH1
2019 F0 Contour Estimation Using Phonetic Feature in Electrolaryngeal Speech Enhancement
abstract
Pitch plays a significant role in understanding a tone based language like Mandarin. In this paper, we present a new method that estimates F0 contour for electrolaryngeal (EL) speech enhancement in Mandarin. Our system explores the usage of phonetic feature to improve the quality of EL speech. First, we train an acoustic model for EL speech and generate the phoneme posterior probabilities feature sequence for each input EL speech utterance. Then we employ the phonetic feature for F0 contour generation rather than the acoustic feature. The experimental results indicate that the EL speech is significantly enhanced under the adoption of the phonetic feature. Experimental results demonstrate that the proposed method achieves notable improvement regarding the intelligibility and the similarity with normal speech.
Zexin Cai, Ming Li 0026
ICASSP1
2019 Polyphone Disambiguation for Mandarin Chinese Using Conditional Neural Network with Multi-Level Embedding Features
abstract
This paper describes a conditional neural network architecture for Mandarin Chinese polyphone disambiguation.The system is composed of a bidirectional recurrent neural network component acting as a sentence encoder to accumulate the context correlations, followed by a prediction network that maps the polyphonic character embeddings along with the conditions to corresponding pronunciations.We obtain the word-level condition from a pre-trained word-to-vector lookup table.One goal of polyphone disambiguation is to address the homograph problem existing in the front-end processing of Mandarin Chinese textto-speech system.Our system achieves an accuracy of 94.69% on a publicly available polyphonic character dataset.To further validate our choices on the conditional feature, we investigate polyphone disambiguation systems with multi-level conditions respectively.The experimental results show that both the sentence-level and the word-level conditional embedding features are able to attain good performance for Mandarin Chinese polyphone disambiguation.
Zexin Cai, Yaogen Yang, Chuxiong Zhang, Xiaoyi Qin, Ming Li 0026
INTERSPEECH1
2018 Insights in-to-End Learning Scheme for Language Identification
abstract
A novel interpretable end-to-end learning scheme for language identification is proposed. It is in line with the classical GMM i-vector methods both theoretically and practically. In the end-to-end pipeline, a general encoding layer is employed on top of the frontend CNN, so that it can encode the variable-length input sequence into an utterance level vector automatically. After comparing with the state-of-the-art GMM i-vector methods, we give insights into CNN, and reveal its role and effect in the whole pipeline. We further introduce a general encoding layer, illustrating the reason why they might be appropriate for language identification. We elaborate on several typical encoding layers, including a temporal average pooling layer, a recurrent encoding layer and a novel learnable dictionary encoding layer. We conducted experiment on NIST LRE07 closed-set task, and the results show that our proposed end-to-end systems achieve state-of-the-art performance.
Weicheng Cai, Zexin Cai, Wenbo Liu 0002, Ming Li 0026
ICASSP2
2018 A Novel Learnable Dictionary Encoding Layer for End-to-End Language Identification
abstract
A novel learnable dictionary encoding layer is proposed in this paper for end-to-end language identification. It is inline with the conventional GMM i-vector approach both theoretically and practically. We imitate the mechanism of traditional GMM training and Supervector encoding procedure on the top of CNN. The proposed layer can accumulate high-order statistics from variable-length input sequence and generate an utterance level fixed-dimensional vector representation. Unlike the conventional methods, our new approach provides an end-to-end learning framework, where the inherent dictionary are learned directly from the loss function. The dictionaries and the encoding representation for the classifier are learned jointly. The representation is orderless and therefore appropriate for language identification. We conducted a preliminary experiment on NIST LRE07 closed-set task, and the results reveal that our proposed dictionary encoding layer achieves significant error reduction comparing with the simple average pooling.
Weicheng Cai, Zexin Cai, Xiang Zhang 0014, Ming Li 0026
ICASSP2