Qiongqiong Wang

dblp:81/10013 · DBLP profile ↗
← Back
21ranked-venue papers
10as first author
13since 2021 · last 2025
0000-0002-9903-0618ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 9 first-author · 10 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 7 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
abstract
Current large speech language models (SpeechLLMs) often exhibit limitations in empathetic reasoning, primarily due to the absence of training datasets that integrate both contextual content and paralinguistic cues. In this work, we propose two approaches to incorporate contextual paralinguistic information into model training: (1) an explicit method that provides paralinguistic metadata (e.g., emotion annotations) directly to the LLM, and (2) an implicit method that automatically generates novel training question-answer (QA) pairs using both categorical and dimensional emotion annotations alongside speech transcriptions. Our implicit method boosts performance (LLM-judged) by 38.41% on a human-annotated QA benchmark, reaching 46.02% when combined with the explicit approach, showing effectiveness in contextual paralinguistic understanding. We also validate the LLM judge by demonstrating its correlation with classification metrics, providing support for its reliability.
Qiongqiong Wang, Hardik Bhupendra Sailor, Jeremy H. M. Wong, Tianchi Liu 0004, Muhammad Huzaifah 0001, Nancy F. Chen, AiTi Aw
ASRU1
2025 Diversity and complementarity of speech encoders across diverse tasks in a multi-modal large language model
abstract
A Large Language Model (LLM) can be extended to understand speech inputs by using a speech encoder to compute embeddings from the speech, which are then used with a text prompt. Diverse information is expressed in speech and a wide variety of tasks can be performed. Different speech encoders may specialise toward different information types and tasks. This complementarity can be leveraged upon by using multiple speech encoders. This paper presents a comprehensive analysis of the diversity and complementarity between open-source speech encoders, when used in a multi-modal LLM framework. Experiments identify the encoders that excel in each type of downstream task, thereby guiding future system design. The diversity between encoders is measured, showing that Whisper tends to behave more differently. Diversity between encoders is compared across tasks, showing that semantic tasks tend to yield more diverse predictions. Early and late fusion show that complementarity can yield improvements.
Jeremy H. M. Wong, Muhammad Huzaifah 0001, Hardik B. Sailor, Kye Min Tan, Bin Wang 0040, Qiongqiong Wang, Xunlong Zou, Nancy F. Chen, AiTi Aw
ASRU7
2025 Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation
Qiongqiong Wang, Hardik B. Sailor, Tianchi Liu 0004, AiTi Aw
INTERSPEECH1
2024 Attentive Merging of Hidden Embeddings from Pre-trained Speech Model for Anti-spoofing Detection
Zihan Pan, Tianchi Liu 0004, Hardik B. Sailor, Qiongqiong Wang
INTERSPEECH4
2024 Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CTRSVDD) Challenge 2024
abstract
This work details our approach to achieving a leading system with a 1.79% pooled equal error rate (EER) on the evaluation set of the Controlled Singing Voice Deepfake Detection (CtrSVDD). The rapid advancement of generative AI models presents significant challenges for detecting AI-generated deepfake singing voices, attracting increased research attention. The Singing Voice Deepfake Detection (SVDD) Challenge 2024 aims to address this complex task. In this work, we explore the ensemble methods, utilizing speech foundation models to develop robust singing voice anti-spoofing systems. We also introduce a novel Squeeze-and-Excitation Aggregation (SEA) method, which efficiently and effectively integrates representation features from the speech foundation models, surpassing the performance of our other individual systems. Evaluation results confirm the efficacy of our approach in detecting deepfake singing voices. The codes can be accessed at https://github.com/Anmol2059/SVDD2024.
Anmol Guragain, Tianchi Liu 0004, Zihan Pan, Hardik B. Sailor, Qiongqiong Wang
SLT5
2024 Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-Spoofing
abstract
The effects of language mismatch impact speech anti-spoofing systems, while investigations and quantification of these effects remain limited. Existing anti-spoofing datasets are mainly in English, and the high cost of acquiring multilingual datasets hinders training language-independent models. We initiate this work by evaluating top-performing speech anti-spoofing systems that are trained on English data but tested on other languages, observing notable performance declines. We propose an innovative approach - Accent-based data expansion via TTS (ACCENT), which introduces diverse linguistic knowledge to monolingual-trained models, improving their cross-lingual capabilities. We conduct experiments on a large-scale dataset consisting of over 3 million samples, including 1.8 million training samples and nearly 1.2 million testing samples across 12 languages. The language mismatch effects are preliminarily quantified and remarkably reduced over 15% by applying the proposed ACCENT. This easily implementable method shows promise for multilingual and low-resource language scenarios.
Tianchi Liu 0004, Ivan Kukanov, Zihan Pan, Qiongqiong Wang, Hardik B. Sailor, Kong-Aik Lee
SLT4
2024 Cosine Scoring With Uncertainty for Neural Speaker Embedding
abstract
Uncertainty modeling in speaker representation aims to learn the variability present in speech utterances. While the conventional cosine-scoring is computationally efficient and prevalent in speaker recognition, it lacks the capability to handle uncertainty. To address this challenge, this paper proposes an approach for estimating uncertainty at the speaker embedding front-end and propagating it to the cosine scoring back-end. Experiments conducted on the VoxCeleb and SITW datasets confirmed the efficacy of the proposed method in handling uncertainty arising from embedding estimation. It achieved improvement with 8.5% and 9.8% average reductions in EER and minDCF compared to the conventional cosine similarity. It is also computationally efficient in practice.
Qiongqiong Wang, Kong-Aik Lee
IEEE Signal Process. Lett.1
2024 Golden Gemini is All You Need: Finding the Sweet Spots for Speaker Verification
abstract
The residual neural networks (ResNet) demonstrate the impressive performance in automatic speaker verification (ASV). They treat the time and frequency dimensions equally, following the default stride configuration designed for image recognition, where the horizontal and vertical axes exhibit similarities. This approach ignores the fact that time and frequency are asymmetric in speech representation. We address this issue and postulateGolden-Gemini Hypothesis,which posits the prioritization of temporal resolution over frequency resolution for ASV. The hypothesis is verified by conducting a systematic study on the impact of temporal and frequency resolutions on the performance, using a trellis diagram to represent the stride space. We further identify two optimal points, namelyGolden Gemini, which serves as a guiding principle for designing 2D ResNet-based ASV models. By following the principle, a state-of-the-art ResNet baseline model gains a significant performance improvement on VoxCeleb, SITW, and CNCeleb datasets with 7.70%/11.76% average EER/minDCF reductions, respectively, across different network depths (ResNet18, 34, 50, and 101), while reducing the number of parameters by 16.5% and FLOPs by 4.1%. We refer to it asGeminiResNet. Further investigation reveals the efficacy of the proposedGolden Geminioperating points across various training conditions and architectures. Furthermore, we present a new benchmark, namely theGeminiDF-ResNet, using a cutting-edge model.Codes and pre-trained models are available athttps://github.com/Tianchi-Liu9/Golden-Gemini-for-Speaker-Verification.
Tianchi Liu 0004, Kong-Aik Lee, Qiongqiong Wang, Haizhou Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Incorporating Uncertainty from Speaker Embedding Estimation to Speaker Verification
abstract
Speech utterances recorded under differing conditions exhibit varying degrees of confidence in their embedding estimates, i.e., uncertainty, even if they are extracted using the same neural network. This paper aims to incorporate the uncertainty estimate produced in the xi-vector network front-end with a probabilistic linear discriminant analysis (PLDA) back-end scoring for speaker verification. To achieve this we derive a posterior covariance matrix, which measures the uncertainty, from the frame-wise precisions to the embedding space. We propose a log-likelihood ratio function for the PLDA scoring with the uncertainty propagation. We also propose to replace the length normalization pre-processing technique with a length scaling technique for the application of uncertainty propagation in the back-end. Experimental results on the VoxCeleb-1, SITW test sets as well as a domain-mismatched CNCeleb1-E set show the effectiveness of the proposed techniques with 14.5%–41.3% EER reductions and 4.6%–25.3% minDCF reductions.
Qiongqiong Wang, Kong-Aik Lee, Tianchi Liu 0004
ICASSP1
2023 Disentangling Voice and Content with Self-Supervision for Speaker Recognition
abstract
For speaker recognition, it is difficult to extract an accurate speaker representation from speech because of its mixture of speaker traits and content. This paper proposes a disentanglement framework that simultaneously models speaker traits and content variability in speech. It is realized with the use of three Gaussian inference layers, each consisting of a learnable transition model that extracts distinct speech components. Notably, a strengthened transition model is specifically designed to model complex speech dynamics. We also propose a self-supervision method to dynamically disentangle content without the use of labels other than speaker identities. The efficacy of the proposed framework is validated via experiments conducted on the VoxCeleb and SITW datasets with 9.56\% and 8.24\% average reductions in EER and minDCF, respectively. Since neither additional model training nor data is specifically needed, it is easily applicable in practical use.
Tianchi Liu 0004, Kong-Aik Lee, Qiongqiong Wang, Haizhou Li 0001
NeurIPS3
2023 Generalized Domain Adaptation Framework for Parametric Back-End in Speaker Recognition
abstract
State-of-the-art speaker recognition systems comprise a speaker embedding front-end followed by a probabilistic linear discriminant analysis (PLDA) back-end. The effectiveness of these components relies on the availability of a large amount of labeled training data. In practice, it is common for domains (e.g., language, channel, demographic) in which a system is deployed to differ from that in which a system has been trained. To close the resulting gap, domain adaptation is often essential for PLDA models. Among two of its variants are Heavy-tailed PLDA (HT-PLDA) and Gaussian PLDA (G-PLDA). Though the former better fits real feature spaces than does the latter, its popularity has been severely limited by its computational complexity and, especially, by the difficulty, it presents in domain adaptation, which results from its non-Gaussian property. Various domain adaptation methods have been proposed for G-PLDA. This paper proposes a generalized framework for domain adaptation that can be applied to both of the above variants of PLDA for speaker recognition. It not only includes several existing supervised and unsupervised domain adaptation methods but also makes possible more flexible usage of available data in different domains. In particular, we introduce here two new techniques: (1) correlation-alignment in the model level, and (2) covariance regularization. To the best of our knowledge, this is the first proposed application of such techniques for domain adaptation w.r.t. HT-PLDA. The efficacy of the proposed techniques has been experimentally validated on NIST 2016, 2018, and 2019 Speaker Recognition Evaluation (SRE’16, SRE’18 and SRE’19) datasets.
Qiongqiong Wang, Koji Okabe, Kong-Aik Lee, Takafumi Koshinaka
IEEE Trans. Inf. Forensics Secur.1
2022 Scoring of Large-Margin Embeddings for Speaker Verification: Cosine or PLDA?
abstract
The emergence of large-margin softmax cross-entropy losses in training deep speaker embedding neural networks has triggered a gradual shift from parametric back-ends to a simpler cosine similarity measure for speaker verification. Popular parametric back-ends include the probabilistic linear discriminant analysis (PLDA) and its variants. This paper investigates the properties of margin-based cross-entropy losses leading to such a shift and aims to find scoring back-ends best suited for speaker verification. In addition, we revisit the pre-processing techniques which have been widely used in the past and assess their effectiveness on large-margin embeddings. Experiments on the state-of-the-art ECAPA-TDNN networks trained with various large-margin softmax cross-entropy losses show a substantial increment in intra-speaker compactness making the conventional PLDA superfluous. In this regard, we found that constraining the within-speaker covariance matrix could improve the performance of the PLDA. It is demonstrated through a series of experiments on the VoxCeleb-1 and SITW core-core test sets with 40.8% equal error rate (EER) reduction and 35.1% minimum detection cost (minDCF) reduction. It also outperforms cosine scoring consistently with reductions in EER and minDCF by 10.9% and 4.9%, respectively.
Qiongqiong Wang, Kong-Aik Lee, Tianchi Liu 0004
INTERSPEECH1
2021 Xi-Vector Embedding for Speaker Recognition
abstract
We present a Bayesian formulation for deep speaker embedding, wherein the xi-vector is the Bayesian counterpart of the x-vector, taking into account the uncertainty estimate. On the technology front, we offer a simple and straightforward extension to the now widely used x-vector. It consists of an auxiliary neural net predicting the frame-wise uncertainty of the input sequence. We show that the proposed extension leads to substantial improvement across all operating points, with a significant reduction in error rates and detection cost. On the theoretical front, our proposal integrates the Bayesian formulation of linear Gaussian model to speaker-embedding neural networks via the pooling layer. In one sense, our proposal integrates the Bayesian formulation of the i-vector to that of the x-vector. Hence, we refer to the embedding as the xi-vector, which is pronounced as /zai/ vector. Experimental results on the SITW evaluation set show a consistent improvement of over 17.5% in equal-error-rate and 10.9% in minimum detection cost.
Kong-Aik Lee, Qiongqiong Wang, Takafumi Koshinaka
IEEE Signal Process. Lett.2
2020 A Generalized Framework for Domain Adaptation of PLDA in Speaker Recognition
abstract
This paper proposes a generalized framework for domain adaptation of Probabilistic Linear Discriminant Analysis (PLDA) in speaker recognition. It not only includes several existing supervised and unsupervised domain adaptation methods but also makes possible more flexible usage of available data in different domains. In particular, we introduce here the two new techniques described below. (1) Correlation-alignment-based interpolation and (2) covariance regularization. The proposed correlation-alignment-based-interpolation method decreases minCprimaryup to 30.5% as compared with that from an out-of-domain PLDA model before adaptation, and minCprimaryis also 5.5% lower than with a conventional linear interpolation method with optimal interpolation weights. Further, the proposed regularization technique ensures robustness in interpolations w.r.t. varying interpolation weights, which in practice is essential.
Qiongqiong Wang, Koji Okabe, Kong-Aik Lee, Takafumi Koshinaka
ICASSP1
2020 NEC-TT Speaker Verification System for SRE'19 CTS Challenge
Kong-Aik Lee, Koji Okabe, Hitoshi Yamamoto, Qiongqiong Wang, Takafumi Koshinaka, Jiacen Zhang, Keisuke Ishikawa, Koichi Shinoda
INTERSPEECH4
2020 NEC-TT System for Mixed-Bandwidth and Multi-Domain Speaker Recognition
Kong-Aik Lee, Hitoshi Yamamoto, Koji Okabe, Qiongqiong Wang, Takafumi Koshinaka, Jiacen Zhang, Koichi Shinoda
Comput. Speech Lang.4
2019 The CORAL+ Algorithm for Unsupervised Domain Adaptation of PLDA
abstract
State-of-the-art speaker recognition systems comprise an x-vector (or i-vector) speaker embedding front-end followed by a probabilistic linear discriminant analysis (PLDA) backend. The effectiveness of these components relies on the availability of a large collection of labeled training data. In practice, it is common that the domains (e.g., language, demographic) in which the system is deployed differ from that we trained the system. To close the gap due to the domain mismatch, we propose an unsupervised PLDA adaptation algorithm to learn from a small amount of unlabeled in-domain data. The proposed method was inspired by a prior work on feature-based domain adaptation technique known as the correlation alignment (CORAL). We refer to the model-based adaptation technique proposed in this paper as CORAL+. The efficacy of the proposed technique is experimentally validated on the recent NIST 2016 and 2018 Speaker Recognition Evaluation (SRE'16, SRE'18) datasets.
Kong-Aik Lee, Qiongqiong Wang, Takafumi Koshinaka
ICASSP2
2019 The NEC-TT 2018 Speaker Verification System
Kong-Aik Lee, Hitoshi Yamamoto, Koji Okabe, Qiongqiong Wang, Takafumi Koshinaka, Jiacen Zhang, Koichi Shinoda
INTERSPEECH4
2018 Attention Mechanism in Speaker Recognition: What Does it Learn in Deep Speaker Embedding?
abstract
This paper presents an experimental study on deep speaker embedding with an attention mechanism that has been found to be a powerful representation learning technique in speaker recognition. In this framework, an attention model works as a frame selector that computes an attention weight for each frame-level feature vector, in accord with which an utterance-level representation is produced at the pooling layer in a speaker embedding network. In general, an attention model is trained together with the speaker embedding network on a single objective function, and thus those two components are tightly bound to one another. In this paper, we consider the possibility that the attention model might be decoupled from its parent network and assist other speaker embedding networks and even conventional i-vector extractors. This possibility is demonstrated through a series of experiments on a NIST Speaker Recognition Evaluation (SRE) task, with 9.0% EER reduction and 3.8% minCprimaryreduction when the attention weights are applied to i-vector extraction. Another experiment shows that DNN-based soft voice activity detection (VAD) can be effectively combined with the attention mechanism to yield further reduction of minCprimaryby 6.6% and 1.6% in deep speaker embedding and i-vector systems, respectively.
Qiongqiong Wang, Koji Okabe, Kong-Aik Lee, Hitoshi Yamamoto, Takafumi Koshinaka
SLT1
2017 Unsupervised Discriminative Training of PLDA for Domain Adaptation in Speaker Verification
Qiongqiong Wang, Takafumi Koshinaka
INTERSPEECH1
2016 Domain adaptation using maximum likelihood linear transformation for PLDA-based speaker verification
abstract
While i-vector-PLDA frameworks employing huge amounts of development data have achieved significant success in speaker recognition, it is infeasible to collect a sufficiently large amount of data for every real application. This paper proposes a method to perform supervised domain adaptation of PLDA in i-vector-based speaker recognition systems with available resource-rich mismatched data and small amounts of matched data, under two assumptions: (1) between-speaker and within-speaker covariances depend on domains; (2) features in one domain can be transformed into another domain by means of an affine transformation. Maximum likelihood linear transformation (MLLT) is used to infer the relationship between the datasets of two domains in training PLDA. The proposed method improves performance over that achieved without adaptation. Using a score fusion technique, it outperforms a conventional method based on linear combination.
Qiongqiong Wang, Hitoshi Yamamoto, Takafumi Koshinaka
ICASSP1