VLDB 2026 Research / reviewers in the wild / expert
Xiaoxiao Miao
dblp:235/7067
· DBLP profile ↗
25ranked-venue papers
8as first author
23since 2021 · last 2027
0000-0002-6645-6524ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 15 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Privacy attacks on voice anonymization systems: Overview and key findings from the First VoicePrivacy Attacker Challenge
Natalia A. Tomashenko, Xiaoxiao Miao, Emmanuel Vincent 0001, Junichi Yamagishi |
Comput. Speech Lang. | 2 |
| 2026 | The third VoicePrivacy challenge: Preserving emotional expressiveness and linguistic content in voice anonymization
Natalia A. Tomashenko, Xiaoxiao Miao, Pierre Champion, Sarina Meyer, Michele Panariello, Xin Wang 0037, Nicholas W. D. Evans, Emmanuel Vincent 0001, Junichi Yamagishi, Massimiliano Todisco |
Comput. Speech Lang. | 2 |
| 2025 | SEF-MK: Speaker-Embedding-Free Voice Anonymization through Multi-k-means QuantizationabstractVoice anonymization protects speaker privacy by concealing identity while preserving linguistic and paralinguistic content. Self-supervised learning (SSL) representations encode linguistic features but preserve speaker traits. We propose a novel speaker-embedding-free framework called SEF-MK. Instead of using a single k-means model trained on the entire dataset, SEF-MK anonymizes SSL representations for each utterance by randomly selecting one of multiple k-means models, each trained on a different subset of speakers. We explore this approach from both attacker and user perspectives. Extensive experiments show that, compared to a single k -means model, SEF-MK with multiple $\mathbf{k}$-means models better preserves linguistic and emotional content from the user’s viewpoint. However, from the attacker’s perspective, utilizing multiple $\mathbf{k}$-means models boosts the effectiveness of privacy attacks. These insights can aid users in designing voice anonymization systems to mitigate attacker threats.11Code and audio samples can be found at https://github.com/Beilong-Tang/sef-mk Beilong Tang, Xiaoxiao Miao, Xin Wang 0037, Ming Li 0026 |
ASRU | 2 |
| 2025 | The First VoicePrivacy Attacker ChallengeabstractThe First VoicePrivacy Attacker Challenge is an ICASSP 2025 SP Grand Challenge which focuses on evaluating attacker systems against a set of voice anonymization systems submitted to the VoicePrivacy 2024 Challenge. Training, development, and evaluation datasets were provided along with a baseline attacker. Participants developed their attacker systems in the form of automatic speaker verification systems and submitted their scores on the development and evaluation data. The best attacker systems reduced the equal error rate (EER) by 25–44% relative w.r.t. the baseline. Natalia A. Tomashenko, Xiaoxiao Miao, Emmanuel Vincent 0001, Junichi Yamagishi |
ICASSP | 2 |
| 2025 | SecureSpeech: Prompt-based Speaker and Content ProtectionabstractGiven the increasing privacy concerns from identity theft and the re-identification of speakers through content in the speech field, this paper proposes a prompt-based speech generation pipeline that ensures dual anonymization of both speaker identity and spoken content. This is addressed through 1) generating a speaker identity un-linkable to the source speaker, controlled by descriptors, and 2) replacing sensitive content within the original text using a name entity recognition model and a large language model. The pipeline utilizes the anonymized speaker identity and text to generate high-fidelity, privacy-friendly speech via a text-to-speech synthesis model. Experimental results demonstrate an achievement of significant privacy protection while maintaining a decent level of content retention and audio quality. This paper also investigates the impact of varying speaker descriptions on the utility and privacy of generated speech to determine potential biases. Belinda Soh Hui Hui, Xiaoxiao Miao, Xin Wang 0037 |
IJCB | 2 |
| 2025 | Mitigating Language Mismatch in SSL-Based Speaker Anonymization
Wen-Chin Huang, Xin Wang 0037, Xiaoxiao Miao, Junichi Yamagishi |
INTERSPEECH | 4 |
| 2025 | Automated evaluation of children's speech fluency for low-resource languages
Nur Afiqah Abdul Latiff, Justin Kan, Rong Tong, Donny Soh, Xiaoxiao Miao, Ian McLoughlin 0001 |
INTERSPEECH | 6 |
| 2025 | LSPnet: an ultra-low bitrate hybrid neural codec
Ian McLoughlin 0001, Xiaoxiao Miao, A. S. Madhukumar |
INTERSPEECH | 3 |
| 2025 | Adapting general disentanglement-based speaker anonymization for enhanced emotion preservationabstractA general disentanglement-based speaker anonymization system typically separates speech into content, speaker, and prosody features using individual encoders. This paper explores how to adapt such a system when a new speech attribute, for example, emotion, needs to be preserved to a greater extent. Two strategies for this are examined. First, we show that integrating emotion embeddings from a pre-trained emotion encoder can help preserve emotional cues, even though this approach slightly compromises privacy protection. Alternatively, we propose an emotion compensation strategy as a post-processing step applied to anonymized speaker embeddings. This conceals the original speaker’s identity and reintroduces the emotional traits lost during speaker embedding anonymization. Specifically, we model the emotion attribute using support vector machines to learn separate boundaries for each emotion. During inference, the original speaker embedding is processed in two ways: one, by an emotion indicator to predict emotion and select the emotion-matched SVM accurately; and two, by a speaker anonymizer to conceal speaker characteristics. The anonymized speaker embedding is then modified along the corresponding SVM boundary towards an enhanced emotional direction to save the emotional cues. The proposed strategies are also expected to be useful for adapting a general disentanglement-based speaker anonymization system to preserve other target paralinguistic attributes, with potential for a range of downstream tasks 2 . Xiaoxiao Miao, Xin Wang 0037, Natalia A. Tomashenko, Cheng Lock Donny Soh, Ian McLoughlin 0001 |
Comput. Speech Lang. | 1 |
| 2025 | A Benchmark for Multi-Speaker AnonymizationabstractPrivacy-preserving voice protection approaches primarily suppress privacy-related information derived from paralinguistic attributes while preserving the linguistic content. Existing solutions focus particularly on single-speaker scenarios. However, they lack practicality for real-world applications, i.e., multi-speaker scenarios. In this paper, we present an initial attempt to provide a multi-speaker anonymization benchmark by defining the task and evaluation protocol, proposing benchmarking solutions, and discussing the privacy leakage of overlapping conversations. The proposed benchmark solutions are based on a cascaded system that integrates spectral-clustering-based speaker diarization and disentanglement-based speaker anonymization using a selection-based anonymizer. To improve utility, the benchmark solutions are further enhanced by two conversation-level speaker vector anonymization methods. The first method minimizes the differential similarity across speaker pairs in the original and anonymized conversations, which maintains original speaker relationships in the anonymized version. The other minimizes the aggregated similarity across anonymized speakers, which achieves better differentiation between speakers. Experiments conducted on both non-overlap simulated and real-world datasets demonstrate the effectiveness of the multi-speaker anonymization system with the proposed speaker anonymizers. Additionally, we analyzed overlapping speech regarding privacy leakage and provided potential solutions (Code and audio samples are available athttps://github.com/xiaoxiaomiao323/MSA), evaluation datasets can be download fromhttps://zenodo.org/records/14249171 Xiaoxiao Miao, Ruijie Tao, Chang Zeng, Xin Wang 0037 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Synvox2: Towards A Privacy-Friendly Voxceleb2 DatasetabstractThe success of deep learning in speaker recognition relies heavily on the use of large datasets. However, the data-hungry nature of deep learning methods has already being questioned on account the ethical, privacy, and legal concerns that arise when using large-scale datasets of natural speech collected from real human speakers. For example, the widely-used VoxCeleb2 dataset for speaker recognition is no longer accessible from the official website. To mitigate these concerns, this work presents an initiative to generate a privacyfriendly synthetic VoxCeleb2 dataset that ensures the quality of the generated speech in terms of privacy, utility, and fairness. We also discuss the challenges of using synthetic data for the downstream task of speaker verification. Xiaoxiao Miao, Xin Wang 0037, Erica Cooper, Junichi Yamagishi, Nicholas W. D. Evans, Massimiliano Todisco, Jean-François Bonastre, Mickael Rouvier |
ICASSP | 1 |
| 2024 | Target Speaker Extraction with Curriculum LearningabstractThis paper presents a novel approach to target speaker extraction (TSE) using Curriculum Learning (CL) techniques, addressing the challenge of distinguishing a target speaker's voice from a mixture containing interfering speakers.For efficient training, we propose designing a curriculum that selects subsets of increasing complexity, such as increasing similarity between target and interfering speakers, and that selects training data strategically.Our CL strategies include both variants using predefined difficulty measures (e.g.gender, speaker similarity, and signal-to-distortion ratio) and ones using the TSE's standard objective function, each designed to expose the model gradually to more challenging scenarios.Comprehensive testing on the Libri2talker dataset demonstrated that our CL strategies for TSE improved the performance, and the results markedly exceeded baseline models without CL about 1 dB. Xuechen Liu 0001, Xiaoxiao Miao, Junichi Yamagishi |
INTERSPEECH | 3 |
| 2024 | Spoofing-Aware Speaker Verification Robust Against Domain and Channel MismatchesabstractIn real-world applications, it is challenging to build a speaker verification system that is simultaneously robust against common threats, including spoofing attacks, channel mismatch, and domain mismatch. Traditional automatic speaker verification (ASV) systems often tackle these issues separately, leading to suboptimal performance when faced with simultaneous challenges. In this paper, we propose an integrated framework that incorporates pair-wise learning and spoofing attack simulation into the meta-learning paradigm to enhance robustness against these multifaceted threats. This novel approach employs an asymmetric dual-path model and a multi-task learning strategy to handle ASV, anti-spoofing, and spoofing-aware ASV tasks concurrently. A new testing dataset, CNComplex, is introduced to evaluate system performance under these combined threats. Experimental results demonstrate that our integrated model significantly improves performance over traditional ASV systems across various scenarios, showcasing its potential for real-world deployment. Additionally, the proposed framework’s ability to generalize across different conditions highlights its robustness and reliability, making it a promising solution for practical ASV applications. Chang Zeng, Xiaoxiao Miao, Xin Wang 0037, Erica Cooper, Junichi Yamagishi |
SLT | 2 |
| 2024 | Instructsing: High-Fidelity Singing Voice Generation Via Instructing YourselfabstractIt is challenging to accelerate the training process while ensuring both high-quality generated voices and acceptable inference speed. In this paper, we propose a novel neural vocoder called InstructSing, which can converge much faster compared with other neural vocoders while maintaining good performance by integrating differentiable digital signal processing and adversarial training. It includes one generator and two discriminators. Specifically, the generator incorporates a harmonic-plus-noise (HN) module to produce 8 kHz audio as an instructive signal. Subsequently, the HN module is connected with an extended WaveNet by an UNet-based module, which transforms the output of the HN module to a latent variable sequence containing essential periodic and aperiodic information. In addition to the latent sequence, the extended WaveNet also takes the melspectrogram as input to generate 48 kHz high-fidelity singing voices. In terms of discriminators, we combine a multi-period discriminator, as originally proposed in HiFiGAN, with a multi-resolution multiband STFT discriminator. Notably, InstructSing achieves comparable voice quality to other neural vocoders but with only one-tenth of the training steps on a 4 NVIDIA V100 GPU machine1. We plan to open-source our code and pretrained model once the paper get accepted.1Demo page: https://wavelandspeech.github.io/instructsing/ Chang Zeng, Xiaoxiao Miao, Zhonglin Jiang |
SLT | 3 |
| 2024 | Joint speaker encoder and neural back-end model for fully end-to-end automatic speaker verification with multiple enrollment utterancesabstractConventional automatic speaker verification systems can usually be decomposed into a front-end model such as time delay neural network (TDNN) for extracting speaker embeddings and a back-end model such as statistics-based probabilistic linear discriminant analysis (PLDA) or neural network-based neural PLDA (NPLDA) for similarity scoring. However, the sequential optimization of the front-end and back-end models may lead to a local minimum, which theoretically prevents the whole system from achieving the best optimization. Although some methods have been proposed for jointly optimizing the two models, such as the generalized end-to-end (GE2E) model and NPLDA E2E model, most of these methods have not fully investigated how to model the intra-relationship between multiple enrollment utterances. In this paper, we propose a new E2E joint method for speaker verification especially designed for the practical scenario of multiple enrollment utterances. To leverage the intra-relationship among multiple enrollment utterances, our model comes equipped with frame-level and utterance-level attention mechanisms. Additionally, focal loss is utilized to balance the importance of positive and negative samples within a mini-batch and focus on the difficult samples during the training process. We also utilize several data augmentation techniques, including conventional noise augmentation using MUSAN and RIRs datasets and a unique speaker embedding-level mixup strategy for better optimization. Chang Zeng, Xiaoxiao Miao, Xin Wang 0037, Erica Cooper, Junichi Yamagishi |
Comput. Speech Lang. | 2 |
| 2024 | The VoicePrivacy 2022 Challenge: Progress and Perspectives in Voice AnonymisationabstractThe VoicePrivacy Challenge promotes the development of voice anonymisation solutions for speech technology. In this paper we present a systematic overview and analysis of the second edition held in 2022. We describe the voice anonymisation task and datasets used for system development and evaluation, present the different attack models used for evaluation, and the associated objective and subjective metrics. We describe three anonymisation baselines, provide a summary description of the anonymisation systems developed by challenge participants, and report objective and subjective evaluation results for all. In addition, we describe post-evaluation analyses and a summary of related work reported in the open literature. Results show that solutions based on voice conversion better preserve utility, that an alternative which combines automatic speech recognition with synthesis achieves greater privacy, and that a privacy-utility trade-off remains inherent to current anonymisation solutions. Finally, we present our ideas and priorities for future VoicePrivacy Challenge editions. Michele Panariello, Natalia A. Tomashenko, Xin Wang 0037, Xiaoxiao Miao, Pierre Champion, Hubert Nourtel, Massimiliano Todisco, Nicholas W. D. Evans, Emmanuel Vincent 0001, Junichi Yamagishi |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Hiding Speaker's Sex in Speech Using Zero-Evidence Speaker Representation in an Analysis/Synthesis PipelineabstractThe use of modern vocoders in an analysis/synthesis pipeline allows us to investigate high-quality voice conversion that can be used for privacy purposes. Here, we propose to transform the speaker embedding and the pitch in order to hide the sex of the speaker. ECAPA-TDNN-based speaker representation fed into a HiFiGAN vocoder is protected using a neural-discriminant analysis approach, which is consistent with the zero-evidence concept of privacy. This approach significantly reduces the information in speech related to the speaker’s sex while preserving speech content and some consistency in the resulting protected voices. Paul-Gauthier Noé, Xiaoxiao Miao, Xin Wang 0037, Junichi Yamagishi, Jean-François Bonastre, Driss Matrouf |
ICASSP | 2 |
| 2023 | Improving Generalization Ability of Countermeasures for New Mismatch Scenario by Combining Multiple Advanced Regularization Terms
Chang Zeng, Xin Wang 0037, Xiaoxiao Miao, Erica Cooper, Junichi Yamagishi |
INTERSPEECH | 3 |
| 2023 | Speaker Anonymization Using Orthogonal Householder Neural NetworkabstractSpeaker anonymization aims to conceal a speaker's identity while preserving content information in speech. Current mainstream neural-network speaker anonymization systems disentangle speech into prosody-related, content, and speaker representations. The speaker representation is then anonymized by a selection-based speaker anonymizer that uses a mean vector over a set of randomly selected speaker vectors from an external pool of English speakers. However, the resulting anonymized vectors are subject to severe privacy leakage against powerful attackers, reduction in speaker diversity, and language mismatch problems for unseen-language speaker anonymization. To generate diverse, language-neutral speaker vectors, this paper proposes an anonymizer based on an orthogonal Householder neural network (OHNN). Specifically, the OHNN acts like a rotation to transform the original speaker vectors into anonymized speaker vectors, which are constrained to follow the distribution over the original speaker vector space. A basic classification loss is introduced to ensure that anonymized speaker vectors from different speakers have unique speaker identities. To further protect speaker identities, an improved classification loss and similarity loss are used to push original-anonymized sample pairs away from each other. Experiments on VoicePrivacy Challenge datasets in English and theAISHELL-3dataset in Mandarin demonstrate the proposed anonymizer's effectiveness. Xiaoxiao Miao, Xin Wang 0037, Erica Cooper, Junichi Yamagishi, Natalia A. Tomashenko |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Attention Back-End for Automatic Speaker Verification with Multiple Enrollment UtterancesabstractProbabilistic linear discriminant analysis (PLDA) or cosine similarity have been widely used in traditional speaker verification systems as back-end techniques to measure pairwise similarities. To make better use of multiple enrollment utterances, we propose a novel attention back-end model that can be used for both textindependent (TI) and text-dependent (TD) speaker verification, and we use scaled-dot self-attention and feed-forward self-attention networks as architectures that learn the intra-relationships of enrollment utterances. To verify the proposed model, we conduct a series of experiments on the CNCeleb and VoxCeleb datasets by combining it with several state-of-the-art speaker encoders including TDNN and ResNet. Experimental results obtained using multiple enrollment utterances on CNCeleb show that the proposed attention back-end model leads to lower EER and minDCF scores than its PLDA and cosine similarity counterparts for each speaker encoder, and an experiment on VoxCeleb demonstrates that our model can be used even for a single enrollment case. Chang Zeng, Xin Wang 0037, Erica Cooper, Xiaoxiao Miao, Junichi Yamagishi |
ICASSP | 4 |
| 2022 | Analyzing Language-Independent Speaker Anonymization Framework under Unseen ConditionsabstractIn our previous work, we proposed a language-independent speaker anonymization system based on self-supervised learning models.Although the system can anonymize speech data of any language, the anonymization was imperfect, and the speech content of the anonymized speech was distorted.This limitation is more severe when the input speech is from a domain unseen in the training data.This study analyzed the bottleneck of the anonymization system under unseen conditions.It was found that the domain (e.g., language and channel) mismatch between the training and test data affected the neural waveform vocoder and anonymized speaker vectors, which limited the performance of the whole system.Increasing the training data diversity for the vocoder was found to be helpful to reduce its implicit language and channel dependency.Furthermore, a simple correlation-alignment-based domain adaption strategy was found to be significantly effective to alleviate the mismatch on the anonymized speaker vectors.Audio samples 1 and source code 2 are available online. Xiaoxiao Miao, Xin Wang 0037, Erica Cooper, Junichi Yamagishi, Natalia A. Tomashenko |
INTERSPEECH | 1 |
| 2021 | Adaptive Margin Circle Loss for Speaker VerificationabstractDeep-Neural-Network (DNN) based speaker verification systems use the angular softmax loss with margin penalties to enhance the intra-class compactness of speaker embeddings, which achieved remarkable performance.In this paper, we propose a novel angular loss function called adaptive margin circle loss for speaker verification.The stage-based margin and chunk-based margin are applied to improve the angular discrimination of circle loss on the training set.The analysis on gradients shows that, compared with the previous angular loss like Additive Margin Softmax(Am-Softmax), circle loss has flexible optimization and definite convergence status.Experiments are carried out on the Voxceleb and SITW.By applying adaptive margin circle loss, our best system achieves 1.31%EER on Voxceleb1 and 2.13% on SITW core-core. Runqiu Xiao, Xiaoxiao Miao, Pengyuan Zhang, Liuping Luo |
Interspeech | 2 |
| 2021 | D-MONA: A dilated mixed-order non-local attention network for speaker and language recognition
Xiaoxiao Miao, Ian McLoughlin 0001, Pengyuan Zhang |
Neural Networks | 1 |
| 2019 | A New Time-Frequency Attention Mechanism for TDNN and CNN-LSTM-TDNN, with Application to Language Identification
Xiaoxiao Miao, Ian McLoughlin 0001, Yonghong Yan 0002 |
INTERSPEECH | 1 |
| 2018 | Improved Conditional Generative Adversarial Net Classification For Spoken Language RecognitionabstractRecent research on generative adversarial nets (GAN) for language identification (LID) has shown promising results. In this paper, we further exploit the latent abilities of GAN networks to firstly combine them with deep neural network (DNN)-based i-vector approaches and then to improve the LID model using conditional generative adversarial net (cGAN) classification. First, phoneme dependent deep bottleneck features (DBF) combined with output posteriors of a pre-trained DNN for automatic speech recognition (ASR) are used to extract i-vectors in the normal way. These i-vectors are then classified using cGAN, and we show an effective method within the cGAN to optimize parameters by combining both language identification and verification signals as supervision. Results show firstly that cGAN methods can significantly outperform DBF DNN i-vector methods where 49-dimensional i-vectors are used, but not where 600-dimensional vectors are used. Secondly, training a cGAN discriminator network for direct classification has further benefit for low dimensional i-vectors as well as short utterances with high dimensional i-vectors. However, incorporating a dedicated discriminator network output layer for classification and optimizing both classification and verification loss brings benefits in all test cases. Xiaoxiao Miao, Ian McLoughlin 0001, Shengyu Yao, Yonghong Yan 0002 |
SLT | 1 |