EDBT 2026 Demo / reviewers in the wild / expert
Qiongqiong Wang
dblp:81/10013
· DBLP profile ↗
21ranked-venue papers
10as first author
13since 2021 · last 2025
0000-0002-9903-0618ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 9 first-author · 10 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 7 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Incorporating Contextual Paralinguistic Understanding in Large Speech-Language ModelsabstractCurrent large speech language models (SpeechLLMs) often exhibit limitations in empathetic reasoning, primarily due to the absence of training datasets that integrate both contextual content and paralinguistic cues. In this work, we propose two approaches to incorporate contextual paralinguistic information into model training: (1) an explicit method that provides paralinguistic metadata (e.g., emotion annotations) directly to the LLM, and (2) an implicit method that automatically generates novel training question-answer (QA) pairs using both categorical and dimensional emotion annotations alongside speech transcriptions. Our implicit method boosts performance (LLM-judged) by 38.41% on a human-annotated QA benchmark, reaching 46.02% when combined with the explicit approach, showing effectiveness in contextual paralinguistic understanding. We also validate the LLM judge by demonstrating its correlation with classification metrics, providing support for its reliability. Qiongqiong Wang, Hardik Bhupendra Sailor, Jeremy H. M. Wong, Tianchi Liu 0004, Muhammad Huzaifah 0001, Nancy F. Chen, AiTi Aw |
ASRU | 1 |
| 2025 | Diversity and complementarity of speech encoders across diverse tasks in a multi-modal large language modelabstractA Large Language Model (LLM) can be extended to understand speech inputs by using a speech encoder to compute embeddings from the speech, which are then used with a text prompt. Diverse information is expressed in speech and a wide variety of tasks can be performed. Different speech encoders may specialise toward different information types and tasks. This complementarity can be leveraged upon by using multiple speech encoders. This paper presents a comprehensive analysis of the diversity and complementarity between open-source speech encoders, when used in a multi-modal LLM framework. Experiments identify the encoders that excel in each type of downstream task, thereby guiding future system design. The diversity between encoders is measured, showing that Whisper tends to behave more differently. Diversity between encoders is compared across tasks, showing that semantic tasks tend to yield more diverse predictions. Early and late fusion show that complementarity can yield improvements. Jeremy H. M. Wong, Muhammad Huzaifah 0001, Hardik B. Sailor, Kye Min Tan, Bin Wang 0040, Qiongqiong Wang, Xunlong Zou, Nancy F. Chen, AiTi Aw |
ASRU | 7 |
| 2025 | Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation
Qiongqiong Wang, Hardik B. Sailor, Tianchi Liu 0004, AiTi Aw |
INTERSPEECH | 1 |
| 2024 | Attentive Merging of Hidden Embeddings from Pre-trained Speech Model for Anti-spoofing Detection
Zihan Pan, Tianchi Liu 0004, Hardik B. Sailor, Qiongqiong Wang |
INTERSPEECH | 4 |
| 2024 | Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CTRSVDD) Challenge 2024abstractThis work details our approach to achieving a leading system with a 1.79% pooled equal error rate (EER) on the evaluation set of the Controlled Singing Voice Deepfake Detection (CtrSVDD). The rapid advancement of generative AI models presents significant challenges for detecting AI-generated deepfake singing voices, attracting increased research attention. The Singing Voice Deepfake Detection (SVDD) Challenge 2024 aims to address this complex task. In this work, we explore the ensemble methods, utilizing speech foundation models to develop robust singing voice anti-spoofing systems. We also introduce a novel Squeeze-and-Excitation Aggregation (SEA) method, which efficiently and effectively integrates representation features from the speech foundation models, surpassing the performance of our other individual systems. Evaluation results confirm the efficacy of our approach in detecting deepfake singing voices. The codes can be accessed at https://github.com/Anmol2059/SVDD2024. Anmol Guragain, Tianchi Liu 0004, Zihan Pan, Hardik B. Sailor, Qiongqiong Wang |
SLT | 5 |
| 2024 | Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-SpoofingabstractThe effects of language mismatch impact speech anti-spoofing systems, while investigations and quantification of these effects remain limited. Existing anti-spoofing datasets are mainly in English, and the high cost of acquiring multilingual datasets hinders training language-independent models. We initiate this work by evaluating top-performing speech anti-spoofing systems that are trained on English data but tested on other languages, observing notable performance declines. We propose an innovative approach - Accent-based data expansion via TTS (ACCENT), which introduces diverse linguistic knowledge to monolingual-trained models, improving their cross-lingual capabilities. We conduct experiments on a large-scale dataset consisting of over 3 million samples, including 1.8 million training samples and nearly 1.2 million testing samples across 12 languages. The language mismatch effects are preliminarily quantified and remarkably reduced over 15% by applying the proposed ACCENT. This easily implementable method shows promise for multilingual and low-resource language scenarios. Tianchi Liu 0004, Ivan Kukanov, Zihan Pan, Qiongqiong Wang, Hardik B. Sailor, Kong-Aik Lee |
SLT | 4 |
| 2024 | Cosine Scoring With Uncertainty for Neural Speaker EmbeddingabstractUncertainty modeling in speaker representation aims to learn the variability present in speech utterances. While the conventional cosine-scoring is computationally efficient and prevalent in speaker recognition, it lacks the capability to handle uncertainty. To address this challenge, this paper proposes an approach for estimating uncertainty at the speaker embedding front-end and propagating it to the cosine scoring back-end. Experiments conducted on the VoxCeleb and SITW datasets confirmed the efficacy of the proposed method in handling uncertainty arising from embedding estimation. It achieved improvement with 8.5% and 9.8% average reductions in EER and minDCF compared to the conventional cosine similarity. It is also computationally efficient in practice. Qiongqiong Wang, Kong-Aik Lee |
IEEE Signal Process. Lett. | 1 |
| 2024 | Golden Gemini is All You Need: Finding the Sweet Spots for Speaker VerificationabstractThe residual neural networks (ResNet) demonstrate the impressive performance in automatic speaker verification (ASV). They treat the time and frequency dimensions equally, following the default stride configuration designed for image recognition, where the horizontal and vertical axes exhibit similarities. This approach ignores the fact that time and frequency are asymmetric in speech representation. We address this issue and postulateGolden-Gemini Hypothesis,which posits the prioritization of temporal resolution over frequency resolution for ASV. The hypothesis is verified by conducting a systematic study on the impact of temporal and frequency resolutions on the performance, using a trellis diagram to represent the stride space. We further identify two optimal points, namelyGolden Gemini, which serves as a guiding principle for designing 2D ResNet-based ASV models. By following the principle, a state-of-the-art ResNet baseline model gains a significant performance improvement on VoxCeleb, SITW, and CNCeleb datasets with 7.70%/11.76% average EER/minDCF reductions, respectively, across different network depths (ResNet18, 34, 50, and 101), while reducing the number of parameters by 16.5% and FLOPs by 4.1%. We refer to it asGeminiResNet. Further investigation reveals the efficacy of the proposedGolden Geminioperating points across various training conditions and architectures. Furthermore, we present a new benchmark, namely theGeminiDF-ResNet, using a cutting-edge model.Codes and pre-trained models are available athttps://github.com/Tianchi-Liu9/Golden-Gemini-for-Speaker-Verification. Tianchi Liu 0004, Kong-Aik Lee, Qiongqiong Wang, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Incorporating Uncertainty from Speaker Embedding Estimation to Speaker VerificationabstractSpeech utterances recorded under differing conditions exhibit varying degrees of confidence in their embedding estimates, i.e., uncertainty, even if they are extracted using the same neural network. This paper aims to incorporate the uncertainty estimate produced in the xi-vector network front-end with a probabilistic linear discriminant analysis (PLDA) back-end scoring for speaker verification. To achieve this we derive a posterior covariance matrix, which measures the uncertainty, from the frame-wise precisions to the embedding space. We propose a log-likelihood ratio function for the PLDA scoring with the uncertainty propagation. We also propose to replace the length normalization pre-processing technique with a length scaling technique for the application of uncertainty propagation in the back-end. Experimental results on the VoxCeleb-1, SITW test sets as well as a domain-mismatched CNCeleb1-E set show the effectiveness of the proposed techniques with 14.5%–41.3% EER reductions and 4.6%–25.3% minDCF reductions. Qiongqiong Wang, Kong-Aik Lee, Tianchi Liu 0004 |
ICASSP | 1 |
| 2023 | Disentangling Voice and Content with Self-Supervision for Speaker RecognitionabstractFor speaker recognition, it is difficult to extract an accurate speaker representation from speech because of its mixture of speaker traits and content. This paper proposes a disentanglement framework that simultaneously models speaker traits and content variability in speech. It is realized with the use of three Gaussian inference layers, each consisting of a learnable transition model that extracts distinct speech components. Notably, a strengthened transition model is specifically designed to model complex speech dynamics. We also propose a self-supervision method to dynamically disentangle content without the use of labels other than speaker identities. The efficacy of the proposed framework is validated via experiments conducted on the VoxCeleb and SITW datasets with 9.56\% and 8.24\% average reductions in EER and minDCF, respectively. Since neither additional model training nor data is specifically needed, it is easily applicable in practical use. Tianchi Liu 0004, Kong-Aik Lee, Qiongqiong Wang, Haizhou Li 0001 |
NeurIPS | 3 |
| 2023 | Generalized Domain Adaptation Framework for Parametric Back-End in Speaker RecognitionabstractState-of-the-art speaker recognition systems comprise a speaker embedding front-end followed by a probabilistic linear discriminant analysis (PLDA) back-end. The effectiveness of these components relies on the availability of a large amount of labeled training data. In practice, it is common for domains (e.g., language, channel, demographic) in which a system is deployed to differ from that in which a system has been trained. To close the resulting gap, domain adaptation is often essential for PLDA models. Among two of its variants are Heavy-tailed PLDA (HT-PLDA) and Gaussian PLDA (G-PLDA). Though the former better fits real feature spaces than does the latter, its popularity has been severely limited by its computational complexity and, especially, by the difficulty, it presents in domain adaptation, which results from its non-Gaussian property. Various domain adaptation methods have been proposed for G-PLDA. This paper proposes a generalized framework for domain adaptation that can be applied to both of the above variants of PLDA for speaker recognition. It not only includes several existing supervised and unsupervised domain adaptation methods but also makes possible more flexible usage of available data in different domains. In particular, we introduce here two new techniques: (1) correlation-alignment in the model level, and (2) covariance regularization. To the best of our knowledge, this is the first proposed application of such techniques for domain adaptation w.r.t. HT-PLDA. The efficacy of the proposed techniques has been experimentally validated on NIST 2016, 2018, and 2019 Speaker Recognition Evaluation (SRE’16, SRE’18 and SRE’19) datasets. Qiongqiong Wang, Koji Okabe, Kong-Aik Lee, Takafumi Koshinaka |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | Scoring of Large-Margin Embeddings for Speaker Verification: Cosine or PLDA?abstractThe emergence of large-margin softmax cross-entropy losses in training deep speaker embedding neural networks has triggered a gradual shift from parametric back-ends to a simpler cosine similarity measure for speaker verification. Popular parametric back-ends include the probabilistic linear discriminant analysis (PLDA) and its variants. This paper investigates the properties of margin-based cross-entropy losses leading to such a shift and aims to find scoring back-ends best suited for speaker verification. In addition, we revisit the pre-processing techniques which have been widely used in the past and assess their effectiveness on large-margin embeddings. Experiments on the state-of-the-art ECAPA-TDNN networks trained with various large-margin softmax cross-entropy losses show a substantial increment in intra-speaker compactness making the conventional PLDA superfluous. In this regard, we found that constraining the within-speaker covariance matrix could improve the performance of the PLDA. It is demonstrated through a series of experiments on the VoxCeleb-1 and SITW core-core test sets with 40.8% equal error rate (EER) reduction and 35.1% minimum detection cost (minDCF) reduction. It also outperforms cosine scoring consistently with reductions in EER and minDCF by 10.9% and 4.9%, respectively. Qiongqiong Wang, Kong-Aik Lee, Tianchi Liu 0004 |
INTERSPEECH | 1 |
| 2021 | Xi-Vector Embedding for Speaker RecognitionabstractWe present a Bayesian formulation for deep speaker embedding, wherein the xi-vector is the Bayesian counterpart of the x-vector, taking into account the uncertainty estimate. On the technology front, we offer a simple and straightforward extension to the now widely used x-vector. It consists of an auxiliary neural net predicting the frame-wise uncertainty of the input sequence. We show that the proposed extension leads to substantial improvement across all operating points, with a significant reduction in error rates and detection cost. On the theoretical front, our proposal integrates the Bayesian formulation of linear Gaussian model to speaker-embedding neural networks via the pooling layer. In one sense, our proposal integrates the Bayesian formulation of the i-vector to that of the x-vector. Hence, we refer to the embedding as the xi-vector, which is pronounced as /zai/ vector. Experimental results on the SITW evaluation set show a consistent improvement of over 17.5% in equal-error-rate and 10.9% in minimum detection cost. Kong-Aik Lee, Qiongqiong Wang, Takafumi Koshinaka |
IEEE Signal Process. Lett. | 2 |
| 2020 | A Generalized Framework for Domain Adaptation of PLDA in Speaker RecognitionabstractThis paper proposes a generalized framework for domain adaptation of Probabilistic Linear Discriminant Analysis (PLDA) in speaker recognition. It not only includes several existing supervised and unsupervised domain adaptation methods but also makes possible more flexible usage of available data in different domains. In particular, we introduce here the two new techniques described below. (1) Correlation-alignment-based interpolation and (2) covariance regularization. The proposed correlation-alignment-based-interpolation method decreases minCprimaryup to 30.5% as compared with that from an out-of-domain PLDA model before adaptation, and minCprimaryis also 5.5% lower than with a conventional linear interpolation method with optimal interpolation weights. Further, the proposed regularization technique ensures robustness in interpolations w.r.t. varying interpolation weights, which in practice is essential. Qiongqiong Wang, Koji Okabe, Kong-Aik Lee, Takafumi Koshinaka |
ICASSP | 1 |
| 2020 | NEC-TT Speaker Verification System for SRE'19 CTS Challenge
Kong-Aik Lee, Koji Okabe, Hitoshi Yamamoto, Qiongqiong Wang, Takafumi Koshinaka, Jiacen Zhang, Keisuke Ishikawa, Koichi Shinoda |
INTERSPEECH | 4 |
| 2020 | NEC-TT System for Mixed-Bandwidth and Multi-Domain Speaker Recognition
Kong-Aik Lee, Hitoshi Yamamoto, Koji Okabe, Qiongqiong Wang, Takafumi Koshinaka, Jiacen Zhang, Koichi Shinoda |
Comput. Speech Lang. | 4 |
| 2019 | The CORAL+ Algorithm for Unsupervised Domain Adaptation of PLDAabstractState-of-the-art speaker recognition systems comprise an x-vector (or i-vector) speaker embedding front-end followed by a probabilistic linear discriminant analysis (PLDA) backend. The effectiveness of these components relies on the availability of a large collection of labeled training data. In practice, it is common that the domains (e.g., language, demographic) in which the system is deployed differ from that we trained the system. To close the gap due to the domain mismatch, we propose an unsupervised PLDA adaptation algorithm to learn from a small amount of unlabeled in-domain data. The proposed method was inspired by a prior work on feature-based domain adaptation technique known as the correlation alignment (CORAL). We refer to the model-based adaptation technique proposed in this paper as CORAL+. The efficacy of the proposed technique is experimentally validated on the recent NIST 2016 and 2018 Speaker Recognition Evaluation (SRE'16, SRE'18) datasets. Kong-Aik Lee, Qiongqiong Wang, Takafumi Koshinaka |
ICASSP | 2 |
| 2019 | The NEC-TT 2018 Speaker Verification System
Kong-Aik Lee, Hitoshi Yamamoto, Koji Okabe, Qiongqiong Wang, Takafumi Koshinaka, Jiacen Zhang, Koichi Shinoda |
INTERSPEECH | 4 |
| 2018 | Attention Mechanism in Speaker Recognition: What Does it Learn in Deep Speaker Embedding?abstractThis paper presents an experimental study on deep speaker embedding with an attention mechanism that has been found to be a powerful representation learning technique in speaker recognition. In this framework, an attention model works as a frame selector that computes an attention weight for each frame-level feature vector, in accord with which an utterance-level representation is produced at the pooling layer in a speaker embedding network. In general, an attention model is trained together with the speaker embedding network on a single objective function, and thus those two components are tightly bound to one another. In this paper, we consider the possibility that the attention model might be decoupled from its parent network and assist other speaker embedding networks and even conventional i-vector extractors. This possibility is demonstrated through a series of experiments on a NIST Speaker Recognition Evaluation (SRE) task, with 9.0% EER reduction and 3.8% minCprimaryreduction when the attention weights are applied to i-vector extraction. Another experiment shows that DNN-based soft voice activity detection (VAD) can be effectively combined with the attention mechanism to yield further reduction of minCprimaryby 6.6% and 1.6% in deep speaker embedding and i-vector systems, respectively. Qiongqiong Wang, Koji Okabe, Kong-Aik Lee, Hitoshi Yamamoto, Takafumi Koshinaka |
SLT | 1 |
| 2017 | Unsupervised Discriminative Training of PLDA for Domain Adaptation in Speaker Verification
Qiongqiong Wang, Takafumi Koshinaka |
INTERSPEECH | 1 |
| 2016 | Domain adaptation using maximum likelihood linear transformation for PLDA-based speaker verificationabstractWhile i-vector-PLDA frameworks employing huge amounts of development data have achieved significant success in speaker recognition, it is infeasible to collect a sufficiently large amount of data for every real application. This paper proposes a method to perform supervised domain adaptation of PLDA in i-vector-based speaker recognition systems with available resource-rich mismatched data and small amounts of matched data, under two assumptions: (1) between-speaker and within-speaker covariances depend on domains; (2) features in one domain can be transformed into another domain by means of an affine transformation. Maximum likelihood linear transformation (MLLT) is used to infer the relationship between the datasets of two domains in training PLDA. The proposed method improves performance over that achieved without adaptation. Using a score fusion technique, it outperforms a conventional method based on linear combination. Qiongqiong Wang, Hitoshi Yamamoto, Takafumi Koshinaka |
ICASSP | 1 |