Xinzhou Xu

dblp:173/6448 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0002-4017-5919ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Detecting Partially Spoofed Utterances Without Segment Annotation Through Fake Segment Mining-Based Graph Neural Networks
abstract
Determining whether an input utterance includes fake or bona fide segments without additional segment annotation is a frequently investigated topic in Partially Spoofed Speech Detection (PSSD). Nevertheless, most existing works on this topic usually fail to sufficiently consider inter-segment temporal dependence within an utterance, and further, these works usually overlook representative possibly-fake segments, which may be critical for detecting spoofed utterances. In response, we propose an approach of Fake segment Mining based Graph neural network (FMG) to address these issues. For the first issue, we employ a Graph Neural Network (GNN) based module for processing segment-level representations, through comprising Adjacent Temporal Dependence (ATD) and Across Temporal Correlation (ATC) branches for jointly modelling the GNNs’ adjacent segments’ dependence and different segments’ correlation. Then, regarding the second issue, we propose a Fake Segment Mining (FSM) module, which contains attentive pooling, fake segment prototype loss and entropy loss parts, in order to achieve utterance-level predictions and pinpoint representative fake segments. Afterwards, experimental evaluations on spoofed-speech datasets demonstrate that, the proposed approach outperforms compared models, showcasing its effectiveness for detecting partially spoofed utterances without segment annotation.
Zirui Ge, Xinzhou Xu, Zhen Yang 0001
IEEE Trans. Inf. Forensics Secur.2
2025 GNCL: A Graph Neural Network with Consistency Loss for Segment-Level Spoofed Speech Detection
abstract
Segment-level spoofed speech detection focuses on recognizing fake or synthetic segments within identifying partially spoofed speech. Nevertheless, existing models for this segment-level task usually overlook latent local relationships between fake and bona fide segments, and further, a lack of inter-branch consistency may lead to insufficient information sharing between different domains. In this regard, we propose an approach of a Graph Neural network with Consistency Loss (GNCL) for segment-level spoofed speech detection. The proposed approach contains a speech representation extraction module, a graph neural network module for modeling local differences, and a consistency-enhanced loss function. Experimental evaluations on the partial spoof dataset demonstrate that, the proposed approach outperforms compared approaches in spoofed-segment detection in terms of the equal error rate, showcasing its effectiveness for the segment-level spoofed speech detection.
Zirui Ge, Xinzhou Xu, Björn W. Schuller
ICASSP2
2025 GateM2Former: Gated Feature Selection and Expert Modeling in Multimodal Emotion Recognition
abstract
In recent years, multimodal emotion recognition (MER) has gained significant attention due to its potential to integrate information from diverse signals. However, existing methods often struggle to effectively capture complex interactions and contextual information both inter- and intra-modalities, and even to extract the salient representations from pre-trained models. To address these issues, we propose a novel model, gated Mixture of Multimodal Experts (MoME) and Mixtral of Experts (MixMoE) models, namely GateM2Former. The gate mechanism is used to select the most relevant representations from pre-trained models. The MoME and MixMoE expert modules respectively learn the individual characteristics of each modality and the intrinsic alignment and potential interactions between modalities. Besides, we design a hierarchical merge structure to better suit the long sequence scenario (i. e., speech in our case). To verify the effectiveness of the introduced model, we conducted extensive experiments on the IEMOCAP and MELD datasets. The results show that GateM2Former, with a universal multimodal structure, is able to achieve the best results on IEMOCAP and MELD compared with other latest approaches.
Zhongren Dong, Runming Wang, Xinzhou Xu, Zixing Zhang 0001
ICASSP4
2025 Guest Editorial Extremely Low-Resource Autonomous Affective Learning
Xinzhou Xu, Björn W. Schuller, Elisabeth André, Erik Cambria
IEEE Trans. Affect. Comput.1
2025 Modality Imbalance? Dynamic Multi-Modal Knowledge Distillation in Automatic Alzheimer's Disease Recognition
abstract
Alzheimer's disease (AD), as the most prevalent form of dementia, necessitates early identification and treatment for the critical enhancement of patients' quality of life. Recent studies strive to explore advanced machine learning approaches with multiple information cues, such as speech and text, to automatically and precisely detect this disease from conversations. However, these multi-modality-based approaches often suffer from a modality-imbalance challenge that leads to performance degradation. That is, the multi-modal model performs worse than the best mono-modal model, although the former contains more information. To address this issue, we propose a Dynamic Multi-Modal Knowledge Distillation (DMMKD) approach, which dynamically identify the dominant modality and the weak modality, and opt to conduct an inter(cross)-modal or intra-modal knowledge distillation. The core idea is to balance the individual learning speed in the multi-modal learning process by boosting the weak modality with the dominant modality. To evaluate the effectiveness of the introduced DMMKD algorithm, we conducted extensive experiments on two publicly available and widely used AD datasets, i. e., ADReSSo and ADReSS-M. Compared to the multi-modal approaches without dealing with the modality imbalance issue, the introduced DMMKD indicates substantial performance improvements by 15.4% and 10.9% in terms of relative accuracy on the ADReSSo and ADReSS-M datasets, respectively. Moreover, when compared to the state-of-the-art models for automatic AD detection, the DMMKD achieves the best performance of 91.5% and 87.0% accuracies on the two datasets, respectively.
Zhongren Dong, Xinzhou Xu, Zixing Zhang 0001
IEEE J. Biomed. Health Informatics3
2024 DGPN: A Dual Graph Prototypical Network for Few-Shot Speech Spoofing Algorithm Recognition
Zirui Ge, Xinzhou Xu, Björn W. Schuller
INTERSPEECH2
2023 Zero-Shot Speech Emotion Recognition Using Generative Learning with Reconstructed Prototypes
abstract
Zero-shot Speech Emotion Recognition (SER) enables machines to perceive unseen-emotional speech without knowing any samples from these emotional states, which is helpful in audio-based autonomous affective computing. However, existing works on zero-shot SER directly employ original prototypes and only consider inter-domain knowledge transfer through learning unseen-emotional classifiers. In this regard, we propose a zero-shot SER approach using generative learning with reconstructed prototypes in this paper. Within the proposed approach, we first reconstruct prototypes using the alignment from paralinguistic features to semantic prototypes. Then, generative learning is performed to build the connection from the reconstructed prototypes to the features. Afterwards, zero-shot experiments on emotional-speech data demonstrate that the proposed approach achieves better performance compared with the state-of-the-art approaches.
Xinzhou Xu, Zixing Zhang 0001, Björn W. Schuller
ICASSP1
2022 Rethinking Auditory Affective Descriptors Through Zero-Shot Emotion Recognition in Speech
abstract
Zero-shot speech emotion recognition (SER) endows machines with the ability of sensing unseen-emotional states in speech, compared with conventional SER endeavors on supervised cases. On addressing the zero-shot SER task, auditory affective descriptors (AADs) are typically employed to transfer affective knowledge from seen- to unseen-emotional states. However, it remains unknown which types of AADs can well describe emotional states in speech during the transfer. In this regard, we define and research on three types of AADs, namely, per-emotion semantic-embedding, per-emotion manually annotated, and per-sample manually annotated AADs, through zero-shot emotion recognition in speech. This leads to a systematic design including prototype- and annotation-based zero-shot SER modules, relying on the input from per-emotion and per-sample AADs, respectively. We then perform extensive experimental comparisons between human and machines’ AADs on the French emotional speech corpus CINEMO for positive-negative (PN) and within-negative (WN) tasks. The experimental results indicate that semantic-embedding prototypes from pretrained models can outperform manually annotated emotional dimensions in zero-shot SER. The results further demonstrate that it is possible for machines to understand and describe affective information in speech better than human beings, with the help of sufficient pretrained models.
Xinzhou Xu, Zixing Zhang 0001, Xijian Fan, Li Zhao 0003, Laurence Devillers, Björn W. Schuller
IEEE Trans. Comput. Soc. Syst.1
2022 Exploring Zero-Shot Emotion Recognition in Speech Using Semantic-Embedding Prototypes
abstract
Speech Emotion Recognition (SER) makes it possible for machines to perceive affective information. Our previous research differed from conventional SER endeavours in that it focused on recognising unseen emotions in speech autonomously through machine learning. Such a step would enable the automatic leaning of unknown emerging emotional states. This type of learning framework, however, still relied on manual annotations to obtain multiple samples of each emotion. In order to reduce this additional workload, herein, we propose a zero-shot SER framework employing a per-emotion semantic-embedding paradigm to describe emotions in zero-shot SER, instead of using the sample-wise descriptors. Aiming to optimise the relationship between emotions, prototypes, and speech samples, this framework includes two types of learning strategies: Sample-wise learning and emotion-wise learning. These strategies apply a novel learning process to speech samples and emotions, respectively, via specifically designed semantic-embedding prototypes. We verify the utility of these approaches by performing an extensive experimental evaluation on two corpora on three aspects, namely the influence of different types of learning strategies, emotional-pair comparison, and the selections of semantic-embedding prototypes and paralinguistic features. The experimental results indicate that it is applicable to use semantic-embedding prototypes for zero-shot emotion recognition in speech, despite the influence of choosing optimal strategies and prototypes.
Xinzhou Xu, Nicholas Cummins, Zixing Zhang 0001, Li Zhao 0003, Björn W. Schuller
IEEE Trans. Multim.1
2019 Autonomous Emotion Learning in Speech: A View of Zero-Shot Speech Emotion Recognition
abstract
Conventionally, speech emotion recognition is achieved using passive learning approaches.Differing from such approaches, we herein propose and develop a dynamic method of autonomous emotion learning based on zero-shot learning.The proposed methodology employs emotional dimensions as the attributes in the zero-shot learning paradigm, resulting in two phases of learning, namely attribute learning and label learning.Attribute learning connects the paralinguistic features and attributes utilising speech with known emotional labels, while label learning aims at defining unseen emotions through the attributes.The experimental results achieved on the CINEMO corpus indicate that zero-shot learning is a useful technique for autonomous speech-based emotion learning, achieving accuracies considerably better than chance level and an attribute-based gold-standard setup.Furthermore, different emotion recognition tasks, emotional attributes, and employed approaches strongly influence system performance.
Xinzhou Xu, Nicholas Cummins, Zixing Zhang 0001, Li Zhao 0003, Björn W. Schuller
INTERSPEECH1
2019 Connecting Subspace Learning and Extreme Learning Machine in Speech Emotion Recognition
abstract
Speech emotion recognition (SER) is a powerful tool for endowing computers with the capacity to process information about the affective states of users in human-machine interactions. Recent research has shown the effectiveness of graph embedding-based subspace learning and extreme learning machine applied to SER, but there are still various drawbacks in these two techniques that limit their application. Regarding subspace learning, the change from linearity to nonlinearity is usually achieved through kernelization, whereas extreme learning machines only take label information into consideration at the output layer. In order to overcome these drawbacks, this paper leverages extreme learning machines for dimensionality reduction and proposes a novel framework to combine spectral regression-based subspace learning and extreme learning machines. The proposed framework contains three stages-data mapping, graph decomposition, and regression. At the data mapping stage, various mapping strategies provide different views of the samples. At the graph decomposition stage, specifically designed embedding graphs provide a possibility to better represent the structure of data through generating virtual coordinates. Finally, at the regression stage, dimension-reduced mappings are achieved by connecting the virtual coordinates and data mapping. Using this framework, we propose several novel dimensionality reduction algorithms, apply them to SER tasks, and compare their performance to relevant state-of-the-art methods. Our results on several paralinguistic corpora show that our proposed techniques lead to significant improvements.
Xinzhou Xu, Eduardo Coutinho, Li Zhao 0003, Björn W. Schuller
IEEE Trans. Multim.1
2018 Semisupervised Autoencoders for Speech Emotion Recognition
abstract
Despite the widespread use of supervised learning methods for speech emotion recognition, they are severely restricted due to the lack of sufficient amount of labelled speech data for the training. Considering the wide availability of unlabelled speech data, therefore, this paper proposes semisupervised autoencoders to improve speech emotion recognition. The aim is to reap the benefit from the combination of the labelled data and unlabelled data. The proposed model extends a popular unsupervised autoencoder by carefully adjoining a supervised learning objective. We extensively evaluate the proposed model on the INTERSPEECH 2009 Emotion Challenge database and other four public databases in different scenarios. Experimental results demonstrate that the proposed model achieves state-of-the-art performance with a very small number of labelled data on the challenge task and other tasks, and significantly outperforms other alternative methods.
Xinzhou Xu, Zixing Zhang 0001, Sascha Frühholz, Björn W. Schuller
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Universum Autoencoder-Based Domain Adaptation for Speech Emotion Recognition
abstract
One of the serious obstacles to the applications of speech emotion recognition systems in real-life settings is the lack of generalization of the emotion classifiers. Many recognition systems often present a dramatic drop in performance when tested on speech data obtained from different speakers, acoustic environments, linguistic content, and domain conditions. In this letter, we propose a novel unsupervised domain adaptation model, called Universum autoencoders, to improve the performance of the systems evaluated in mismatched training and test conditions. To address the mismatch, our proposed model not only learns discriminative information from labeled data, but also learns to incorporate the prior knowledge from unlabeled data into the learning. Experimental results on the labeled Geneva Whispered Emotion Corpus database plus other three unlabeled databases demonstrate the effectiveness of the proposed method when compared to other domain adaptation methods.
Xinzhou Xu, Zixing Zhang 0001, Sascha Frühholz, Björn W. Schuller
IEEE Signal Process. Lett.2
2017 A Two-Dimensional Framework of Multiple Kernel Subspace Learning for Recognizing Emotion in Speech
abstract
As a highly active topic in computational paralinguistics, speech emotion recognition (SER) aims to explore ideal representations for emotional factors in speech. In order to improve the performance of SER, multiple kernel learning (MKL) dimensionality reduction has been utilized to obtain effective information for recognizing emotions. However, the solution of MKL usually provides only one nonnegative mapping direction for multiple kernels; this may lead to loss of valuable information. To address this issue, we propose a two-dimensional framework for multiple kernel subspace learning. This framework provides more linear combinations on the basis of MKL without nonnegative constraints, which preserves more information in the learning procedures. It also leverages both of MKL and two-dimensional subspace learning, combining them into a unified structure. To apply the framework to SER, we also propose an algorithm, namely generalised multiple kernel discriminant analysis (GMKDA), by employing discriminant embedding graphs in this framework. GMKDA takes advantage of the additional mapping directions for multiple kernels in the proposed framework. In order to evaluate the performance of the proposed algorithm a wide range of experiments is carried out on several key emotional corpora. These experimental results demonstrate that the proposed methods can achieve better performance compared with some conventional and subspace learning methods in dealing with SER.
Xinzhou Xu, Nicholas Cummins, Zixing Zhang 0001, Li Zhao 0003, Björn W. Schuller
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 Multiscale kernel locally penalised discriminant analysis exemplified by emotion recognition in speech
abstract
We propose a novel method to learn multiscale kernels with locally penalised discriminant analysis, namely Multiscale-Kernel Locally Penalised Discriminant Analysis (MS-KLPDA). As an exemplary use-case, we apply it to recognise emotions in speech. Specifically, we employ the term of locally penalised discriminant analysis by controlling the weights of marginal sample pairs, while the method learns kernels with multiple scales. Evaluated in a series of experiments on emotional speech corpora, our proposed MS-KLPDA is able to outperform the previous research of Multiscale-Kernel Fisher Discriminant Analysis and some conventional methods in solving speech emotion recognition.
Xinzhou Xu, Maryna Gavryukova, Zixing Zhang 0001, Li Zhao 0003, Björn W. Schuller
ICMI1
2015 Dimensionality reduction for speech emotion features by multiscale kernels
abstract
To achieve efficient and compact low-dimensional features for speech emotion recognition, this paper proposes a novel feature reduction method using multiscale kernels in the framework of graph embedding.With Fisher discriminant embedding graph, multiscale Gaussian kernels are used in constructing optimal linear combination of Gram matrices for multiple kernel learning.To evaluate the proposed method, comprehensive experiments, using different public feature sets from the open-source toolbox openSMILE on various corpora, show that the proposed method achieves better performance compared with conventional linear dimensionality reduction methods and singlekernel methods.
Xinzhou Xu, Wenming Zheng, Li Zhao 0003, Björn W. Schuller
INTERSPEECH1