EDBT 2026 Demo / reviewers in the wild / expert
Mingxing Xu
dblp:67/7820
· DBLP profile ↗
47ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 29 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Spherical CNNs With Lifting-Based Adaptive Wavelets for Pooling and UnpoolingabstractPooling and unpooling are indispensable in constructing hierarchical spherical convolutional neural networks (HS-CNNs). Most existing models employ simple downsampling-based pooling, which ignores the sampling theorem and cannot adapt to different spherical signals (with different spectra) and tasks (dependent on different frequency components), thus suffering a significant information loss. Besides, signals reconstructed by the widely-adopted padding-based unpooling may also change unwantedly the spectra of original signals. To address these, we propose a novel framework of HS-CNNs with lifting structures to learn adaptive spherical wavelets for pooling and unpooling, named LiftHS-CNNs. Specifically, we learn spherical wavelets with a lifting structure to adaptively partition the input signal into low- and high-frequency sub-bands, with the down-scaled representations for pooling generated to preserve more information in the low-frequency sub-band. The lifting structure consists of learnable update and predict operators parameterized with graph attention to jointly consider the signal's characteristics and underlying geometries. We then propose an unpooling operation invertible to the lifting-based pooling for restoring the up-scaled representations, which can well preserve spectral characteristics of the original signal. Particular properties (i.e., spatial locality, vanishing moments, and stability) of the learned wavelets and the information preserving ability of the proposed pooling and unpooling are further studied. Experiments on benchmark spherical datasets for a wide range of tasks verify the superiority of our LiftHS-CNNs. Mingxing Xu, Wenrui Dai, Siheng Chen, Junni Zou, Pascal Frossard, Hongkai Xiong |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | ISL-MED: A General Iterative Self-Learning Framework for Speech Complex Emotion DetectionabstractSpeech complex emotion detection (SCED) aims to identify all emotion categories and their intensities in speech, which is crucial for understanding the speaker’s genuine intentions. A significant challenge inherent to SCED is the coarse nature of manual annotations, such as one-hot labels. These labels merely indicate the presence or absence of an emotion, thereby lacking the granularity required to guide models in capturing crucial intensity information. To overcome this limitation, we propose a novel framework: the Iterative Self-Learning-based Multiple Emotion Detector (ISL-MED). This framework utilizes an iterative self-learning approach to infer the latent complex emotion distribution directly from coarse one-hot labels. Specifically, ISL-MED employs multiple dedicated emotion detectors, each responsible for estimating the intensity component for a distinct emotion category. Notably, these detectors can be instantiated using various existing speech emotion recognition (SER) models, and their quantity can be flexibly configured based on task requirements. Furthermore, this paper proposes a data selection strategy based on Curriculum Learning and Human-Machine Consensus (HMCC). This strategy enhances model performance and accelerates convergence by systematically identifying and excluding highly ambiguous samples from the training set. We validated the effectiveness of ISL-MED on the complex emotions dataset CNSCED for speech complex emotion detection tasks and further evaluated its generalization capability on the IEMOCAP dataset for single emotion recognition tasks. Xinxin Luo, Chang Feng, Hankiz Yilahun, Mingxing Xu, Askar Hamdulla, Thomas Fang Zheng |
IJCNN | 5 |
| 2024 | Enhancing Quantised End-to-End ASR Models Via PersonalisationabstractRecent end-to-end automatic speech recognition (ASR) models have become increasingly larger, making them particularly challenging to be deployed on resource-constrained devices. Model quantisation is an effective solution that sometimes causes the word error rate (WER) to increase. In this paper, a novel strategy of personalisation for a quantised model (PQM) is proposed, which combines speaker adaptive training (SAT) with model quantisation to improve the performance of heavily compressed models. Specifically, PQM uses a 4-bit NormalFloat Quantisation (NF4) approach for model quantisation and low-rank adaptation (LoRA) for SAT. Experiments have been performed on the LibriSpeech and the TED-LIUM 3 corpora. Remarkably, with a 7x reduction in model size and 1% additional speaker-specific parameters, 15.1% and 23.3% relative WER reductions were achieved on quantised Whisper and Conformer-based attention-based encoder-decoder ASR models respectively, comparing to the original full precision models. Qiuming Zhao, Guangzhi Sun, Chao Zhang 0031, Mingxing Xu, Thomas Fang Zheng |
ICASSP | 4 |
| 2024 | Advancing Respiratory Sound Classification: Integration of Audio Spectrogram Transformer with ConnectMix and NEFTune Augmentation
Runze Huang, Mingxing Xu, Thomas Fang Zheng |
ICONIP (10) | 2 |
| 2024 | Emotional Atmosphere Soft Label for Emotion Recognition in Conversations
Chang Feng, Hankiz Yilahun, Mingxing Xu, Askar Hamdulla, Thomas Fang Zheng |
ICONIP (3) | 4 |
| 2024 | A Joint Noise Disentanglement and Adversarial Training Framework for Robust Speaker VerificationabstractAutomatic Speaker Verification (ASV) suffers from performance degradation in noisy conditions. To address this issue, we propose a novel adversarial learning framework that incorporates noise-disentanglement to establish a noise-independent speaker invariant embedding space. Specifically, the disentanglement module includes two encoders for separating speaker related and irrelevant information, respectively. The reconstruction module serves as a regularization term to constrain the noise. A feature-robust loss is also used to supervise the speaker encoder to learn noise-independent speaker embeddings without losing speaker information. In addition, adversarial training is introduced to discourage the speaker encoder from encoding acoustic condition information for achieving a speaker-invariant embedding space. Experiments on VoxCeleb1 indicate that the proposed method improves the performance of the speaker verification system under both clean and noisy conditions. Xujiang Xing, Mingxing Xu, Thomas Fang Zheng |
INTERSPEECH | 2 |
| 2024 | SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
Qiuming Zhao, Guangzhi Sun, Chao Zhang 0031, Mingxing Xu, Thomas Fang Zheng |
INTERSPEECH | 4 |
| 2024 | Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
Shuai Wang 0016, Guangzhi Sun, Chao Zhang 0031, Mingxing Xu, Thomas Fang Zheng |
INTERSPEECH | 6 |
| 2024 | Hierarchical Multi-Path and Multi-Model Selection For Fake Speech DetectionabstractThe variety of spoofing algorithms used in generating speech poses obstacles to fake speech detection. Earlier methods have demonstrated complementary effects for detection. This paper proposes a novel hierarchical multi-path multi-model selection method for fake speech detection. It is designed to dynamically select and utilise the most suitable model from a set of complementary models. In our method, four basic detection models are incorporated, each offering partial but complementary detection abilities, to enhance balanced performance on diverse fake speech. The models are trained through a multi-path schema and the selection mechanism is structured hierarchically to improve the generalisation ability. Our method achieves an Equal Error Rate (EER) of 0.37% on the ASVspoof 2019 LA dataset, and outperforms other state-of-the-art method on the cross-domain and cross-dataset scenarios. A statistical analysis of EERs against thirteen unknown attacks reveals our method’s superiority, evidenced by the lowest standard deviation of 0.24, further underscoring our method’s robustness against a range of attacks. Chang Feng, Guangzhi Sun, Shuai Wang 0016, Chao Zhang 0031, Mingxing Xu, Thomas Fang Zheng |
SLT | 7 |
| 2024 | Nonlinear control strategies for 3-DOF control moment gyroscope using deep reinforcement learning
Jianxiang Zhang, Mingxing Xu, Liang Guo 0011 |
Neural Comput. Appl. | 4 |
| 2023 | Robust Point Cloud Classification With Permutohedral Lattice-based RepresentationabstractDeep learning models have greatly improved the accuracy of point cloud classification. Nevertheless, existing deep learning models are vulnerable to data corruptions such as Gaussian noise and outliers which are inevitable in real-world point cloud collection. To address this challenge, we develop a novel point cloud classification model that is robust to data corruptions, with permutohedral lattice-based representation and density-improved hierarchical feature extraction. Specifically, raw noisy point clouds are firstly projected into a regular per-mutohedral lattice space to obtain the quantized representations. Subsequently, we propose to leverage the local density to improve the farthest point sampling (FPS) and spectral graph convolution for robust hierarchical feature extraction, inspired by that the local point density implicitly reveals the reliability and semantic information of each point. The local density can be efficiently calculated from the permutohedral lattice-based representation. Extensive experiments on the benchmark dataset (i.e., ModelNet-C) verify the robustness of the proposed model. In comparison to the state-of-the-art methods, the proposed model achieves superior performance on the corrupted dataset while maintaining competitive classification accuracy on the clean dataset. Mingxing Xu, Wenrui Dai, Cewu Lu, Weisheng Hu, Junfeng Du, Hongkai Xiong |
VCIP | 2 |
| 2023 | Automatic Representative Frame Selection and Intrathoracic Lymph Node Diagnosis With Endobronchial Ultrasound Elastography VideosabstractEndobronchial ultrasound (EBUS) elastography videos have shown great potential to supplement intrathoracic lymph node diagnosis. However, it is laborious and subjective for the specialists to select the representative frames from the tedious videos and make a diagnosis, and there lacks a framework for automatic representative frame selection and diagnosis. To this end, we propose a novel deep learning framework that achieves reliable diagnosis by explicitly selecting sparse representative frames and guaranteeing the invariance of diagnostic results to the permutations of video frames. Specifically, we develop a differentiable sparse graph attention mechanism that jointly considers frame-level features and the interactions across frames to select sparse representative frames and exclude disturbed frames. Furthermore, instead of adopting deep learning-based frame-level features, we introduce the normalized color histogram that considers the domain knowledge of EBUS elastography images and achieves superior performance. To our best knowledge, the proposed framework is the first to simultaneously achieve automatic representative frame selection and diagnosis with EBUS elastography videos. Experimental results demonstrate that it achieves an average accuracy of 81.29% and area under the receiver operating characteristic curve (AUC) of 0.8749 on the collected dataset of 727 EBUS elastography videos, which is comparable to the performance of the expert-based clinical methods based on manually-selected representative frames. Mingxing Xu, Junxiang Chen, Jin Li 0057, Xinxin Zhi, Wenrui Dai, Jiayuan Sun, Hongkai Xiong |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | Surrogate modeling for spacecraft thermophysical models using deep learning
Liang Guo 0011, Yang Zhang 0071, Mingxing Xu, Defu Tian |
Neural Comput. Appl. | 4 |
| 2021 | Learning from Multiple Noisy Augmented Data Sets for Better Cross-Lingual Spoken Language UnderstandingabstractLack of training data presents a grand challenge to scaling out spoken language understanding (SLU) to low-resource languages.Although various data augmentation approaches have been proposed to synthesize training data in low-resource target languages, the augmented data sets are often noisy, and thus impede the performance of SLU models.In this paper we focus on mitigating noise in augmented data.We develop a denoising training approach.Multiple models are trained with data produced by various augmented methods.Those models provide supervision signals to each other.The experimental results show that our method outperforms the existing state of the art by 3.05 and 4.24 percentage points on two benchmark datasets, respectively.The code will be made open sourced on github. Yingmei Guo, Linjun Shou, Jian Pei 0001, Ming Gong 0001, Mingxing Xu, Zhiyong Wu 0001, Daxin Jiang |
EMNLP (1) | 5 |
| 2021 | Cross-Database Replay Detection in Terminal-Dependent Speaker Verification
Xingliang Cheng, Mingxing Xu, Thomas Fang Zheng |
Interspeech | 2 |
| 2020 | Structure-Aware Graph Construction For Point Cloud Segmentation With Graph Convolutional NetworksabstractThe k-nearest neighbors (KNN) algorithm has been widely adopted to construct graph convolutional networks (GCNs) for point cloud segmentation. However, the ℓ2norm cannot discriminate the multi-dimensional structures within a point cloud. In this paper, we propose a novel structure-aware graph construction for point clouds that compensates the ℓ2norm with per-dimension differences of the signal. The proposed method dynamically calculates the similarity ratio to determine the dimension-based proximity of the pair of points. Consequently, it improves both the spatial and spectral GCNs with the capability of aggregating information from relevant neighbors for point cloud segmentation. As a model-agnostic method, it can be seamlessly embedded into arbitrary GCN architectures during the graph construction phase. Experimental results demonstrate that the proposed method can improve classification accuracy around the joint areas of objects. Shanghong Wang, Wenrui Dai, Mingxing Xu, Junni Zou, Hongkai Xiong |
ICME | 3 |
| 2020 | Depth Estimation From Light Field Using Graph-Based Structure-Aware AnalysisabstractExisting light field depth map estimation approaches only utilize partial angular views in occlusion areas and local spatial dependencies in the optimization. This paper proposes a novel two-stage light field depth estimation method via graph spectral analysis to exploit the complete correlations and dependencies within angular patches and spatial images. The initial depth map estimation leverages the undirected graph to jointly consider occluded and unoccluded views within each angular patch. The estimated depth minimizes the structural incoherence of its corresponding angular patch with the focused one by evaluating the highest graph frequency component. Subsequently, depth map refinement optimizes the initial depth map with the color consistency and smoothness formulated by weighted adjacency matrix. The structural constraints are efficiently employed using low-pass graph filtering with Chebyshev polynomial approximation. Experimental results demonstrate that the proposed method improves the depth map estimation, especially in the edge regions. Wenrui Dai, Mingxing Xu, Junni Zou, Xiaopeng Zhang 0008, Hongkai Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Emotion Recognition from Variable-Length Speech Segments Using Deep Learning on Spectrograms
Xi Ma, Zhiyong Wu 0001, Jia Jia 0001, Mingxing Xu, Helen M. Meng, Lianhong Cai |
INTERSPEECH | 4 |
| 2018 | Imbalance Learning-based Framework for Fear Recognition in the MediaEval Emotional Impact of Movies Task
Xingliang Cheng, Mingxing Xu, Thomas Fang Zheng |
INTERSPEECH | 3 |
| 2017 | Learning cross-lingual knowledge with multilingual BLSTM for emphasis detection with limited training dataabstractBidirectional long short-term memory (BLSTM) recurrent neural network (RNN) has achieved state-of-the-art performance in many sequence processing problems given its capability in capturing contextual information. However, for languages with limited amount of training data, it is still difficult to obtain a high quality BLSTM model for emphasis detection, the aim of which is to recognize the emphasized speech segments from natural speech. To address this problem, in this paper, we propose a multilingual BLSTM (MTL-BLSTM) model where the hidden layers are shared across different languages while the softmax output layer is language-dependent. The MTL-BLSTM can learn cross-lingual knowledge and transfer this knowledge to both languages to improve the emphasis detection performance. Experimental results demonstrate our method can outperform the comparison methods over 2-15.6% and 2.9-15.4% on the English corpus and Mandarin corpus in terms of relative F1-measure, respectively. Yishuang Ning, Zhiyong Wu 0001, Runnan Li, Jia Jia 0001, Mingxing Xu, Helen M. Meng, Lianhong Cai |
ICASSP | 5 |
| 2017 | Speaker segmentation using deep speaker vectors for fast speaker change scenariosabstractA novel speaker segmentation approach based on deep neural network is proposed and investigated. This approach uses deep speaker vectors (d-vectors) to represent speaker characteristics and to find speaker change points. The d-vector is a kind of frame-level speaker discriminative feature, whose discriminative training process corresponds to the goal of discriminating a speaker change point from a single speaker speech segment in a short time window. Following the traditional metric-based segmentation, each analysis window contains two sub-windows and is shifting along the audio stream to detect speaker change points, where the speaker characteristics are represented by the means of deep speaker vectors for all frames in each window. Experimental investigations conducted in fast speaker change scenarios show that the proposed method can detect speaker change points more quickly and more effectively than the commonly used segmentation methods. Renyu Wang, Mingliang Gu, Lantian Li, Mingxing Xu, Thomas Fang Zheng |
ICASSP | 4 |
| 2017 | Speech Emotion Recognition with Emotion-Pair Based Framework Considering Emotion Distribution Information in Dimensional Emotion Space
Xi Ma, Zhiyong Wu 0001, Jia Jia 0001, Mingxing Xu, Helen M. Meng, Lianhong Cai |
INTERSPEECH | 4 |
| 2017 | Multi-scale Context Based Attention for Dynamic Music Emotion PredictionabstractDynamic music emotion prediction is to recognize the continuous emotion information in music, which is necessary for music retrieval and recommendation. In this paper, we adopt the dimensional valence-arousal (V-A) emotion model to represent the dynamic emotion in music. In our opinion, music and V-A emotion label do not have the one-to-one correspondence in the time domain, while the expression of music emotion at one moment is the accumulation of previous music content for a period of time, so we propose Long Short-Term Memory (LSTM) based sequence-to-one mapping for dynamic music emotion prediction. Based on this sequence-to-one music emotion mapping, it is proved that different time scales' preceding content has an influence on the LSTM model's performance, so we further propose the Multi-scale Context based Attention (MCA) for dynamic music emotion prediction. We evaluate our proposed method on the database of Emotion in Music task at MediaEval 2015, and the results show that our proposed method outperforms most of the models using the same features and achieves a competitive performance with the state-of-the-art methods. Xinxing Li, Mingxing Xu, Jia Jia 0001, Lianhong Cai |
ACM Multimedia | 3 |
| 2016 | A deep bidirectional long short-term memory based multi-scale approach for music dynamic emotion predictionabstractMusic Dynamic Emotion Prediction is a challenging and significant task. In this paper, We adopt the dimensional valence-arousal (V-A) emotion model to represent the dynamic emotion in music. Considering the high context correlation among the music feature sequence and the advantage of Bidirectional Long Short-Term Memory (BLSTM) in capturing sequence information, we propose a multi-scale approach, Deep BLSTM (DBLSTM) based multi-scale regression and fusion with Extreme Learning Machine (ELM), to predict the V-A values in music. We achieved the best performance on the database of Emotion in Music task in MediaEval 2015 compared with other submitted results. The experimental results demonstrated the effectiveness of our novel proposed multi-scale DBLSTM-ELM model. Xinxing Li, Haishu Xianyu, Jiashen Tian, Wenxiao Chen, Fanhang Meng, Mingxing Xu, Lianhong Cai |
ICASSP | 6 |
| 2016 | Question detection from acoustic features using recurrent neural network with gated recurrent unitabstractQuestion detection is of importance for many speech applications. Only parts of the speech utterances can provide useful clues for question detection. Previous work of question detection using acoustic features in Mandarin conversation is weak in capturing such proper time context information, which could be modeled essentially in recurrent neural network (RNN) structure. In this paper, we conduct an investigation on recurrent approaches to cope with this problem. Based on gated recurrent unit (GRU), we build different RNN and bidirectional RNN (BRNN) models to extract efficient features at segment and utterance level. The particular advantage of GRU is it can determine a proper time scale to extract high-level contextual features. Experimental results show that the features extracted within proper time scale make the classifier perform better than the baseline method with pre-designed lexical and acoustic feature set. Yaodong Tang, Zhiyong Wu 0001, Helen M. Meng, Mingxing Xu, Lianhong Cai |
ICASSP | 5 |
| 2016 | SVR based double-scale regression for dynamic emotion prediction in musicabstractDynamic music emotion prediction is to recognize the continuous emotion contained in music, and has various applications. In recent years, dynamic music emotion recognition is widely studied, while the inside structure of the emotion in music remains unclear. We conduct a data observation based on the database provided by Free Music Archive (FMA), and find that emotion dynamic shows different properties under different scales. According to the data observation, we propose a new method, Double-scale Support Vector Regression (DS-SVR), to dynamically recognize the music emotion. The new method decouples two scales of emotion dynamics apart, and recognizes them separately. We apply the DS-SVR to MediaEval 2015, Emotion in Music database, and achieve an outstanding performance, significantly better than the baseline provided by organizer. Haishu Xianyu, Xinxing Li, Wenxiao Chen, Fanhang Meng, Jiashen Tian, Mingxing Xu, Lianhong Cai |
ICASSP | 6 |
| 2016 | DBLSTM-based multi-scale fusion for dynamic emotion prediction in musicabstractDynamic Music Emotion Prediction is crucial to the emerging applications of music retrieval and recommendation. Considering the influence of temporal context and hierarchical structure on emotion in music, we propose a Deep Bidirectional Long Short-Term Memory (DBLSTM) based multi-scale regression method. In this method, a post-processing component is utilised for individual DBSLTM output to further enhance the ability of temporal context processing and a fusion component is to integrate the output of all DBLSTM models with different scales. In addition, we investigate how the difference of sequence length between the training and predicting phase affects the performance of DBLSTM. We conduct our experiments on a public database of Emotion in Music task at MediaEval 2015, and the result shows that our method achieves significant improvement when compared with the state-of-art methods. Xinxing Li, Jiashen Tian, Mingxing Xu, Yishuang Ning, Lianhong Cai |
ICME | 3 |
| 2016 | Heterogeneity-entropy based unsupervised feature learning for personality prediction with cross-media dataabstractPersonality prediction has broad prospects of application in real life. It can be accomplished by analyzing massive and variant data in social networks, which conveys one's personal traits through user generated contents, user's social relationships and behaviors. However, it is difficult to design an effective feature representation from such complex data to predict user's personality as well as high-level and abstract psychological concepts. In this paper, we propose a novel unsupervised cross-modal feature learning algorithm, named Heterogeneity Entropy Neural Network (HENN), to extract the common information between modalities and map it to the user's personality. HENN is constructed hierarchically on Deep Belief Networks (DBNs) and Auto-encoder (AE) with a modified loss function, in which an additional term named Heterogeneity Entropy (HE) is added to measure common information among different modalities. Experiments on a cross-media dataset collected from two famous Chinese social network platforms, i.e., Renren and SinaMicroblog, demonstrate the superiority of our method over several existing algorithms. Haishu Xianyu, Mingxing Xu, Zhiyong Wu 0001, Lianhong Cai |
ICME | 2 |
| 2016 | Combining CNN and BLSTM to Extract Textual and Acoustic Features for Recognizing Stances in Mandarin Ideological Debate Competition
Linchuan Li, Zhiyong Wu 0001, Mingxing Xu, Helen M. Meng, Lianhong Cai |
INTERSPEECH | 3 |
| 2016 | Analysis on Gated Recurrent Unit Based Question Detection Approach
Yaodong Tang, Zhiyong Wu 0001, Helen M. Meng, Mingxing Xu, Lianhong Cai |
INTERSPEECH | 4 |
| 2014 | Improved keyword spotting system by optimizing posterior confidence measure vector using feed-forward neural networkabstractIn this paper, a novel method based on feedforward neural network is proposed to optimize the confidence measure for improving a mandarine keyword spotting system. Keyword spotting is to detect the occurrences of a pre-defined list of keywords in the input speech, and confidence measure is an critical part in the verification stage of keyword spotting. Posterior confidence has been widely used and was verified to be effective. In some previous works, the optimization of posterior confidence has been proposed, which linearly transforms the phone-level confidence into the word-level confidence. On this basis, we propose a neural network based method that make a non-linear transformation. In addition, a sparse activation and back-propagation strategy is proposed to make this method feasible and work fast. In the experiments, the proposed method is compared to other two previous methods. To evaluate performance, two most commonly used measures are considered: AUC and EER. The experimental result shows that the proposed method is effective and achieved the best performance among three methods. Mingxing Xu, Lianhong Cai |
IJCNN | 2 |
| 2012 | Compensation of Intrinsic Variability with Factor Analysis Modeling for Robust Speaker Verification
Mingxing Xu |
INTERSPEECH | 2 |
| 2012 | Comparison of adaptation methods for GMM-SVM based speech emotion recognitionabstractThe required length of the utterance is one of the key factors affecting the performance of automatic emotion recognition. To gain the accuracy rate of emotion distinction, adaptation algorithms that can be manipulated on short utterances are highly essential. Regarding this, this paper compares two classical model adaptation methods, maximum a posteriori (MAP) and maximum likelihood linear regression (MLLR), in GMM-SVM based emotion recognition, and tries to find which method can perform better on different length of the enrollment of the utterances. Experiment results show that MLLR adaptation performs better for very short enrollment utterances (with the length shorter than 2s) while MAP adaptation is more effective for longer utterances. Jianbo Jiang, Zhiyong Wu 0001, Mingxing Xu, Jia Jia 0001, Lianhong Cai |
SLT | 3 |
| 2009 | ANN based decision fusion for speech emotion recognitionabstractAbstract As a hot research field, speech emotion recognition has attracted increasing attentions from both academic and business. In this paper, we proposed a method to recognize speech emotions adopting ANNs and to fuse two kinds of recognitions using different features at the decision level. Each emotional utterance is recognized by some individual recognizers firstly. Then the outputs of these recognizers were fused adopting the voting strategy. Furthermore, the dimensionality of supervectors constructed from spectral features is reduced through PCA. Experimental results demonstrated that the proposed decision fusion is effective and the dimensionality reduction is feasible. Index Terms : speech emotion recognition, ANN, decision fusion 1. Introduction Speech is a dominant tool for communication, and it is also an important and effective approach for transmitting information and human emotions. With the increasing role of speech interfaces in human-machine interaction applications, speech emotion recognition becomes more and more important recently. Speech emotion recognition is an interesting and challenging speech technology, which can be applied to broad areas, such as environment of call center [1], treatment of mental and psychological diseases [2], development of education and entertainment software [3], and so on. Speech emotion recognition deals with how to make the computer automatically recognize various emotions in speech signal by extracting and analyzing some acoustic features. A key problem of speech emotion recognition is that which kinds of speech features can be used to represent human emotions. Some researchers have investigated the relations between features and emotions. With their efforts, many speech features were found to be used for emotion recognition. Statistical features based on prosody and voice quality have been widely used in speech emotion recognition and demonstrated considerable recognition success [4, 5]. Besides statistical features, spectral or cepstral features are another effective group for describing emotional states [6, 7]. Since these features all have played significant roles in speech emotion recognition, it is necessary to explore an effective way to complementarily fuse both two kinds of features to further enhance the performance of emotion recognition.Another key issue of speech emotion recognition is how to choose an effective method to classify speech emotions. So far, many pattern classification methods have been used for speech emotion recognition [6-9], such as Support Vector Machine (SVM), k-Nearest Neighbors (k-NN), Gaussian Mixture Model (GMM), Hidden Markov Model (HMM), Artificial Neural Networks (ANN), and so on. These methods are all feasible, but their performances are different with each other seriously. The SVMs based method has been shown to be robust and performs well. But some fresh researches have indicated that ideal performance may be obtained using ANNs as well. However, it is difficult to determine which kind of ANNs is suitable for emotion recognition and it is necessary to compare its performance with the SVMs. In this paper, the ANN based decision fusion for speech emotion recognition was presented. Firstly four different ANNs were used to recognize various emotions. Then the voting scheme was adopted to fuse recognitions using two kinds of features at the decision level. Experimental results demonstrated that the proposed approach improved the performance of ANN based recognition and its accuracy was comparable with SVM based method.The remainder of this paper is organized as follows. The features used for speech emotion recognition are introduced in Section 2. The principles of PCA and ANN are briefly described in Section 3 and Section 4 respectively. The proposed decision fusion is depicted in Section 5. In Section 6, experiments and discussions of experimental results are presented. In Section 7, conclusions are drawn and future works are suggested. Mingxing Xu, Dali Yang |
INTERSPEECH | 2 |
| 2007 | A Cohort-Based Speaker Model Synthesis for Mismatched Channels in Speaker VerificationabstractMismatch between enrollment and test data is one of the top performance degrading factors in speaker recognition applications. This mismatch is particularly true over public telephone networks, where input speech data is collected over different handsets and transmitted over different channels from one trial to the next. In this paper, a cohort-based speaker model synthesis (SMS) algorithm, designed for synthesizing robust speaker models without requiring channel-specific enrollment data, is proposed. This algorithm utilizes a priori knowledge of channels extracted from speaker-specific cohort sets to synthesize such speaker models. The cohort selection in the proposed new SMS can be either speaker-specific or Gaussian component based. Results on the China Criminal Police College (CCPC) speaker recognition corpus, which contains utterances from both landline and mobile channel, show the new algorithms yield significant speaker verification performance improvement over Htnorm and universal background model (UBM)-based speaker model synthesis. Thomas Fang Zheng, Mingxing Xu, Frank K. Soong |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | Cohort-Based Speaker Model Synthesis for Channel Robust Speaker RecognitionabstractSpeaker recognition over a public telephone network involves various types of transmission channels and handsets, which leads to mismatched channels (between the enrolled models and the test utterances), and hence to a significant decline in the speaker recognition performance. In this paper a cohort-based speaker model synthesis algorithm, which aims at synthesizing speaker models for channels where no enrollment data is available is proposed. This algorithm applies a priori knowledge of channels extracted from speaker-specific cohort sets to synthesize speaker models. Results for the China Criminal Police College (CCPC) speaker recognition corpus, which contains utterances from both a landline and a mobile channel, show significant improvements over the HT-norm and UBM-based speaker model synthesis algorithms Thomas Fang Zheng, Mingxing Xu |
ICASSP (1) | 3 |
| 2003 | Using word confidence measure for OOV words detection in a spontaneous spoken dialog systemabstractDeveloping a real-life spoken dialogue system must face with many practical issues, where the out-of-vocabulary (OOV) words problem is one of the key difficulties. This paper presents the OOV detection mechanism based on the word confidence scoring developed for the d-Ear Attendant system, a spontaneous spoken dialogue system. In the d-Ear Attendant system, an explicit filler model is originally used to detect the presence of OOV words [1]. Although this approach has a satisfactory OOV detection rate, it badly degrades the accuracy of in-vocabulary (IV) detection by 4.4 % absolutely (from 97 % to 92.6%). Such the degradation will not be acceptable in a practical system. By using a few commonly used acoustic confidence features and some new context confidence features, our confidence measure method not only is able to detect the word level speech recognition errors, but also has a good ability for OOV words detection with an acceptable false alarm rate. For example, with a false rejection rate of 2.5%, the false acceptance rate of 26 % is achieved. 1. Thomas Fang Zheng, Mingxing Xu |
INTERSPEECH | 4 |
| 2002 | Improved katz smoothing for language modeling in speech recognitonabstractIn this paper, a new method is proposed to improve the canonical Katz back-off smoothing technique in language modeling. The process of Katz smoothing is detailedly analyzed and the global discounting parameters are selected for discounting. Further more, a modified version of the formula for discounting parameters is proposed, in which the discounting parameters are determined by not only the occurring counts of the n-gram units but also the low-order history frequencies. This modification makes the smoothing more reasonable for those n-gram units that have homophonic (same in pronunciation) histories. The new method is tested on a Chinese Pinyin-to-character (where Pinyin is the pronunciation string) conversion system and the results show that the improved method can achieve a surprising reduction both in perplexity and Chinese character error rate. 1. Genqing Wu, Thomas Fang Zheng, Wenhu Wu, Mingxing Xu |
INTERSPEECH | 4 |
| 2001 | Topic Forest: a plan-based dialog management structureabstractThere are many task-oriented dialog systems, but few of them can cope with the issues such as the multiple-topic issue, the topic changing issue, information sharing among different topics, and the difference in importance for different information items. To provide efficient solutions, a plan-based dialog management structure named Topic Forest is proposed, which makes the mixed-initiative dialog control easier. The Topic Forest based reasoning engine with a certain strategy for both remembering and forgetting is also described. The reasoning strategy is designed to be domain-independent; therefore it makes the dialog management model easy to be ported to other different domains. Thomas Fang Zheng, Mingxing Xu |
ICASSP | 3 |
| 2001 | Robust parsing in spoken dialogue systemsabstractThe rule-based parsing is a prevalent method for the natural language understanding (NLU) and has been introduced in dialogue systems for spoken language processing (SLP). However, additional measures must be taken to cope with the severe spoken linguistic phenomena, such as garbage, repetition, ellipsis, word disordering, fragment and ill form, which frequently occur in the spoken language. We propose in this paper a robust parsing scheme, which integrates the following methods. Keywords are used as terminal symbols; hence the symbol set of the grammar is purely within the semantical category. The definition of the grammar is extended to accommodate four types of rules, called up-tying, by-passing, up-messing, and over-crossing respectively. An improved chart parser, named marionette, is designed to parse the semantic grammar instance. The robust parsing scheme has been adopted in an air traveling information service system, called EasyFlight, and has achieved a high performance when dealing with the spontaneous speech. 1. Pengju Yan, Thomas Fang Zheng, Mingxing Xu |
INTERSPEECH | 3 |
| 2000 | Language understanding component for Chinese dialogue systemabstractIn this paper we present the design and the implement of the language understanding component of a Chinese spoken language dialogue system EasyNav. In pursuing the coherence with the goal of understanding, we design the structure of system with speech decoding and language understanding integrated closely. Thus the language understanding component need to be restrictive besides portable. Actually we implement a general syntactic parser and domain-specific semantic parser for the purpose. The grammar rules written for understanding are suitable for spoken language. The feature of spoken language also exists throughout the system. 1. Yinfei Huang, Thomas Fang Zheng, Mingxing Xu, Pengju Yan, Wenhu Wu |
INTERSPEECH | 3 |
| 2000 | An equivalent-class based MMI learning method for MGCPMabstractIn this paper, we present an Equivalent-Class Based Maximum Mutual Information (ECB-MMI) learning method for our previously proposed Mixed Gaussian Continuous Probability Model (MGCPM). Similar to HMMs, the defined object function for MGCPM training considers the mutual information among different models so as to maximally separate the Speech Recognition Units (SRUs) in model space. Experimental result shows that for MGCPM the MMI training method can improve the recognition rate by 5 % compared to the traditional training method MLE (Maximum Likelihood Estimation). Because the computation amount of MMI algorithm is very large, we propose an N-Best strategy to find the corresponding equivalent class (EC) in order to reduce complexity. Our experimental result shows that this criterion works very well. Chunhua Luo, Thomas Fang Zheng, Mingxing Xu |
INTERSPEECH | 3 |
| 2000 | Semi-continuous segmental probability modeling for continuous speech recognitionabstractIn this paper the design of semi-continuous segmental probability models (SCSPMs) in large vocabulary continuous speech recognition is presented. The tied Gaussian densities are trained using data from all states of all utterances while the mixture weights are estimated using data from the state being trained individually. The SCSPMs tie all the densities of all states from all Speech Recognition Units (SRUs) to form a shared pdf codebook, thus the number of Gaussian densities is greatly reduced. Several pruning methods are reviewed and then a new pruning criterion is proposed in order to reduce the number of tied mixture Gaussian densities while there is only a small subset of mixture Gaussian densities with larger tying weights. Our preliminary experiments show that the SCSPM incorporated with the pruning techniques can lessen the size of model storage and speed up the system with little degradation in the accuracy compared to the prior continuous model. Thomas Fang Zheng, Mingxing Xu, Ditang Fang |
INTERSPEECH | 3 |
| 1999 | An effective scoring method for speaking skill evaluation systemabstractThe Speaking Skill Evaluation (SSE) technologies are derived from speech recognition technologies and are used for language learning and instructing. In this paper, an effective automatic pronunciation scoring method for SSE systems is proposed. The Center-Distance Continuous Probability Model (CDCPM) is incorporated to model the speech. The Merging-Based Syllable Detection Automaton (MBSDA) and the Non-Linear Partition (NLP) method are used to perform the time alignment. And the Critical Area Percentage (CAP) based scoring method is used to score the learner’s pronunciations or reject invalid utterances. Subjective assessments show that this method is concise, fast, and effective. The SSE system based on it has achieved a satisfying performance. Zhanjiang Song, Thomas Fang Zheng, Mingxing Xu, Wenhu Wu |
EUROSPEECH | 3 |
| 1999 | A fast and effective state decoding algorithmabstract\n Contains fulltext :\n 75022.pdf (author's version ) (Open Access)\n Mingxing Xu, Thomas Fang Zheng, Wenhu Wu |
EUROSPEECH | 1 |
| 1999 | Easytalk: a large-vocabulary speaker-independent Chinese dictation machine
Thomas Fang Zheng, Zhanjiang Song, Mingxing Xu, Jian Wu 0034, Yinfei Huang, Wenhu Wu |
EUROSPEECH | 3 |
| 1999 | HarkMan - A vocabulary-independent keyword spotter for spontaneous Chinese speech
Thomas Fang Zheng, Mingxing Xu, Xiaolong Mou, Jian Wu 0034, Wenhu Wu, Ditang Fang |
J. Comput. Sci. Technol. | 2 |