Li Zhao 0003

dblp:97/4708-3 · DBLP profile ↗
← Back
58ranked-venue papers
1as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 11 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Exploiting spatial information and target speaker phoneme loss for multichannel directional speech enhancement and recognition
Cong Pang, Ye Ni, Lin Zhou 0001, Li Zhao 0003, Feifei Xiong
Comput. Speech Lang.4
2025 Lightweight Attentive ConvNeXt-TCN for Causal Target Sound Extraction
abstract
Target sound extraction (TSE) aims to isolate specific sounds from complex acoustic mixtures. While various causal TSE models have been developed for real-time processing, most existing causal models operate in the time domain and do not effectively leverage frequency-domain information. This paper proposes a novel time-frequency domain model, ACN-TCN, which integrates the ConvNeXt design paradigm into the temporal convolutional network (TCN) to jointly model temporal and spectral information. Additionally, a convolutional self-attention mechanism is introduced to improve feature selection. Since shallow information is easily lost in a deep network, we incorporate a feature enhancement module to effectively integrate shallow features. Experimental results demonstrate that the proposed model improves signal-to-noise ratio (SNR) by 2.5 dB and scale-invariant signal-to-noise ratio (SI-SNR) by 3.4 dB compared to the state-of-the-art causal TSE method in single-target extraction task. Furthermore, ACN-TCN reduces the number of parameters by approximately 40% compared to previous models. We provide code and audio samples:https://github.com/Xiang-M-J/ACN-TCN
MinJie Xiang, Ruiyu Liang, Ye Ni, Li Zhao 0003, Björn W. Schuller
IEEE Signal Process. Lett.4
2024 An improved TF-GSC for dual-microphone interference suppression in the specific direction
Cong Pang, Jingjie Fan, Ruiyu Liang, Li Zhao 0003, Jiaming Cheng 0005
Multim. Tools Appl.4
2024 Residual Fusion Probabilistic Knowledge Distillation for Speech Enhancement
abstract
In recent years, a great deal of research has focused on in developing neural network (NN)-based speech enhancement (SE) models, which have achieved promising results. However, NN-based models typically require expensive computations to achieve remarkable performance, constraining their deployment in real-world scenarios, especially when hardware resources are limited or when latency requirements are strict. To reduce this computational burden, we propose a unified residual fusion probabilistic knowledge distillation (KD) method for the SE task, in which knowledge is transferred from a deep teacher to a shallower student model. Previous KD approaches commonly focused on narrowing the output distances between teachers and students, but research on the intermediate representation of these models is lacking. In this paper, we first study the cross-layer residual feature fusion strategy, which enables the student model to distill knowledge contained in multiple teacher layers from shallow to deep. Second, a frame weighting probabilistic distillation loss is proposed to assign more emphasis to frames containing essential information and preserve pairwise probabilistic similarities in the representation space. The proposed distillation framework is applied to the dual-path dilated convolutional recurrent network (DPDCRN), which won the championship of the SE track in the L3DAS23 challenge. Extensive experiments are conducted on single-channel and multichannel SE datasets. Objective evaluations show that the proposed KD strategy outperforms other distillation methods and considerably improves the enhancement effect of the low-complexity student model (with only 17% of the teacher's parameters).
Jiaming Cheng 0005, Ruiyu Liang, Lin Zhou 0001, Li Zhao 0003, Chengwei Huang, Björn W. Schuller
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 Layer-Adapted Implicit Distribution Alignment Networks for Cross-Corpus Speech Emotion Recognition
abstract
In this article, we propose a new unsupervised domain adaptation (DA) method called layer-adapted implicit distribution alignment networks (LIDANs) to address the challenge of cross-corpus speech emotion recognition (SER). LIDAN extends our previous ICASSP work, deep implicit distribution alignment networks (DIDANs), whose key contribution lies in the introduction of a novel regularization term called implicit distribution alignment (IDA). This term allows DIDAN trained on source (training) speech samples to remain applicable to predicting emotion labels for target (testing) speech samples, regardless of corpus variance in cross-corpus SER. To further enhance this method, we extend IDA to layer-adapted IDA (LIDA), resulting in LIDAN. This layer-adapted extension consists of three modified IDA terms that consider emotion labels at different levels of granularity. These terms are strategically arranged within different fully connected layers in LIDAN, aligning with the increasing emotion-discriminative abilities with respect to the layer depth. This arrangement enables LIDAN to more effectively learn emotion-discriminative and corpus-invariant features for SER across various corpora compared to DIDAN. It is also worthy to mention that unlike most existing methods that rely on estimating statistical moments to describe preassumed explicit distributions, both IDA and LIDA take a different approach. They utilize an idea of target sample reconstruction to directly bridge the feature distribution gap without making assumptions about their distribution type. As a result, DIDAN and LIDAN can be viewed as implicit cross-corpus SER methods. To evaluate LIDAN, we conducted extensive cross-corpus SER experiments on EmoDB, eNTERFACE, and CASIA corpora. The experimental results demonstrate that LIDAN surpasses recent state-of-theart explicit unsupervised DA methods in tackling cross-corpus SER tasks.
Yan Zhao 0037, Yuan Zong, Jincen Wang, Hailun Lian, Cheng Lu 0005, Li Zhao 0003, Wenming Zheng
IEEE Trans. Comput. Soc. Syst.6
2023 Dual-Path Dilated Convolutional Recurrent Network with Group Attention for Multi-Channel Speech Enhancement
abstract
This paper proposes a dual-path convolutional recurrent network with group attention for ICASSP Signal Processing Grand Challenge: L3DAS23 Challenge. We design a structure based on convolutional encoder-decoder, and frequency-time blocks based on group attention are introduced in the middle. The encoder is used to extract the local representation from the complex spectrum, the correlation along the frequency axis and the time axis are captured through groups of time-frequency processing modules and the key information in the feature flow is extracted by the group attention. As a result, our system ranks the 1st place of the 3D speech enhancement task in L3DAS23 Challenge, and significantly outperforms the baseline, while achieving 0.101 WER and 0.902 STOI on the blind test-set.
Jiaming Cheng 0005, Cong Pang, Ruiyu Liang, Jingjie Fan, Li Zhao 0003
ICASSP5
2023 Deep Implicit Distribution Alignment Networks for cross-Corpus Speech Emotion Recognition
abstract
In this paper, we propose a novel deep transfer learning method called deep implicit distribution alignment networks (DIDAN) to deal with cross-corpus speech emotion recognition (SER) problem, in which the labeled training (source) and unlabeled testing (target) speech signals come from different corpora. Specifically, DIDAN first adopts a simple deep regression network consisting of a set of convolutional and fully connected layers to directly regress the source speech spectrums into the emotional labels such that the proposed DIDAN can own the emotion discriminative ability. Then, such ability is transferred to be also applicable to the target speech samples regardless of corpus variance by resorting to a well-designed regularization term called implicit distribution alignment (IDA). Unlike widely-used maximum mean discrepancy (MMD) and its variants, the proposed IDA absorbs the idea of sample reconstruction to implicitly align the distribution gap, which enables DIDAN to learn both emotion discriminative and corpus invariant features from speech spectrums. To evaluate the proposed DIDAN, extensive cross-corpus SER experiments on widely-used speech emotion corpora are carried out. Experimental results show that the proposed DIDAN can outperform lots of recent state-of-the-art methods in coping with the cross-corpus SER tasks.
Yan Zhao 0037, Jincen Wang, Yuan Zong, Wenming Zheng, Hailun Lian, Li Zhao 0003
ICASSP6
2023 Multimodal emotion recognition from facial expression and speech based on feature fusion
Guichen Tang, Ke Li 0049, Ruiyu Liang, Li Zhao 0003
Multim. Tools Appl.5
2023 Speech Denoising and Compensation for Hearing Aids Using an FTCRN-Based Metric GAN
abstract
Hearing aids aims to improve speech intelligibility for hearing impaired patients to levels comparable to those for normal hearing listeners. However, the interference of environmental noises greatly increase the difficulty of hearing loss compensation. Most related research only focuses on one aspect of noise reduction and hearing loss compensation. In this letter, we propose a metric generative adversarial framework based on a frequency-time convolution recurrent network for joint noise reduction and hearing loss compensation. The audiogram is extended along the frequency axis to form embedded features. A metric discriminator is introduced and the optimization of the generator is guided by an evaluation score related to hearing loss compensation. Additional perceptual-based losses are set to stabilize optimization. Experimental results show that the proposed method can better reduce noise and compensate for hearing loss compared with other algorithms.
Jiaming Cheng 0005, Ruiyu Liang, Li Zhao 0003, Chengwei Huang, Björn W. Schuller
IEEE Signal Process. Lett.3
2022 Cross-Layer Similarity Knowledge Distillation for Speech Enhancement
abstract
Speech enhancement (SE) algorithms based on deep neural networks (DNNs) often encounter challenges of limited hardware resources or strict latency requirements when deployed in real-world scenarios. However, a strong enhancement effect typically requires a large DNN. In this paper, a knowledge distillation framework for SE is proposed to compress the DNN model. We study the strategy of cross-layer connection paths, which fuses multi-level information from the teacher and transfers it to the student. To adapt to the SE task, we propose a frame-level similarity distillation loss. We apply this method to the deep complex convolution recurrent network (DCCRN) and make targeted adjustments. Experimental results show that the proposed method considerably improves the enhancement effect of the compressed DNN and outperforms other distillation methods.
Jiaming Cheng 0005, Ruiyu Liang, Li Zhao 0003, Björn W. Schuller, Yiyuan Peng
INTERSPEECH4
2022 Deep Transductive Transfer Regression Network for Cross-Corpus Speech Emotion Recognition
Yan Zhao 0037, Jincen Wang, Ru Ye, Yuan Zong, Wenming Zheng, Li Zhao 0003
INTERSPEECH6
2022 A Sparse-Based Transformer Network With Associated Spatiotemporal Feature for Micro-Expression Recognition
abstract
Despite a lot of work in excavating the emotion descriptor from the hidden information, learning an effective spatiotemporal feature is a challenging issue for micro-expression recognition due to the fact that the micro-expression has a small difference in dynamic change and occurs in localized facial regions. Therefore, these properties of micro-expression suggest that the representation is sparse in the spatiotemporal domain. In this letter, a high-performance spatiotemporal feature learning based on sparse transformer is presented to solve the above issue. We extract the strong associated spatiotemporal feature by distinguishing the spatial attention map and attentively fusing the temporal feature. Thus, the feature map extracted from the critical relation will be fully utilized, while the superfluous relation will be masked. Our proposed method achieves remarkable results compared to state-of-the-art methods, proving that the sparse representation can be successfully integrated into the self-attention mechanism for micro-expression recognition.
Yuan Zong, Hongli Chang, Yushun Xiao, Li Zhao 0003
IEEE Signal Process. Lett.5
2022 Rethinking Auditory Affective Descriptors Through Zero-Shot Emotion Recognition in Speech
abstract
Zero-shot speech emotion recognition (SER) endows machines with the ability of sensing unseen-emotional states in speech, compared with conventional SER endeavors on supervised cases. On addressing the zero-shot SER task, auditory affective descriptors (AADs) are typically employed to transfer affective knowledge from seen- to unseen-emotional states. However, it remains unknown which types of AADs can well describe emotional states in speech during the transfer. In this regard, we define and research on three types of AADs, namely, per-emotion semantic-embedding, per-emotion manually annotated, and per-sample manually annotated AADs, through zero-shot emotion recognition in speech. This leads to a systematic design including prototype- and annotation-based zero-shot SER modules, relying on the input from per-emotion and per-sample AADs, respectively. We then perform extensive experimental comparisons between human and machines’ AADs on the French emotional speech corpus CINEMO for positive-negative (PN) and within-negative (WN) tasks. The experimental results indicate that semantic-embedding prototypes from pretrained models can outperform manually annotated emotional dimensions in zero-shot SER. The results further demonstrate that it is possible for machines to understand and describe affective information in speech better than human beings, with the help of sufficient pretrained models.
Xinzhou Xu, Zixing Zhang 0001, Xijian Fan, Li Zhao 0003, Laurence Devillers, Björn W. Schuller
IEEE Trans. Comput. Soc. Syst.5
2022 Exploring Zero-Shot Emotion Recognition in Speech Using Semantic-Embedding Prototypes
abstract
Speech Emotion Recognition (SER) makes it possible for machines to perceive affective information. Our previous research differed from conventional SER endeavours in that it focused on recognising unseen emotions in speech autonomously through machine learning. Such a step would enable the automatic leaning of unknown emerging emotional states. This type of learning framework, however, still relied on manual annotations to obtain multiple samples of each emotion. In order to reduce this additional workload, herein, we propose a zero-shot SER framework employing a per-emotion semantic-embedding paradigm to describe emotions in zero-shot SER, instead of using the sample-wise descriptors. Aiming to optimise the relationship between emotions, prototypes, and speech samples, this framework includes two types of learning strategies: Sample-wise learning and emotion-wise learning. These strategies apply a novel learning process to speech samples and emotions, respectively, via specifically designed semantic-embedding prototypes. We verify the utility of these approaches by performing an extensive experimental evaluation on two corpora on three aspects, namely the influence of different types of learning strategies, emotional-pair comparison, and the selections of semantic-embedding prototypes and paralinguistic features. The experimental results indicate that it is applicable to use semantic-embedding prototypes for zero-shot emotion recognition in speech, despite the influence of choosing optimal strategies and prototypes.
Xinzhou Xu, Nicholas Cummins, Zixing Zhang 0001, Li Zhao 0003, Björn W. Schuller
IEEE Trans. Multim.5
2021 Cross-Corpus Speech Emotion Recognition Using Joint Distribution Adaptive Regression
abstract
In this paper, we focus on the research of cross-corpus speech emotion recognition (SER), in which the training and testing speech signals in cross-corpus SER belong to dierent speech corpus. Due to this fact, mismatched feature distributions may exist between the training and testing speech feature sets degrading the performance of most originally well-performing SER methods. To deal with cross-corpus SER, we propose a novel domain adaptation (DA) method called joint distribution adaptive regression (JDAR). The basic idea of JDAR is to learn a regression matrix by jointly considering the marginal and conditional probability distribution between the training and testing speech signals and hence their feature distribution dierence can be alleviated in the subspace spanned by the learned regression matrix. To evaluate the proposed JDAR, we conduct extensive cross-corpus SER experiments on EmoDB, eNTERFACE, and CASIA speech databases. Experimental results show that the proposed JDAR achieves satisfactory performance and outperforms most of state-of-the-art subspace learning based DA methods.
Lin Jiang 0007, Yuan Zong, Wenming Zheng, Li Zhao 0003
ICASSP5
2021 A Deep Adaptation Network for Speech Enhancement: Combining a Relativistic Discriminator With Multi-Kernel Maximum Mean Discrepancy
abstract
In deep-learning-based speech enhancement (SE) systems, trained models are often used to handle unseen noise types and language environments in real-life scenarios. However, since production environments differ from training conditions, mismatch problems arise that may cause a serious decrease in the performance of an SE system. In this study, a domain adaptive method combining two adaptation strategies is proposed to improve the generalization of unlabeled noisy speech. In the proposed encoder-decoder-based SE framework, a domain discriminator and a domain confusion adaptation layer are introduced to conduct adversarial training. The model has two main innovations. First, the algorithm optimizes adversarial training by introducing a relativistic discriminator that relies on relative values by applying the difference, thus avoiding possible bias and better reflecting domain differences. Second, the multi-kernel maximum mean discrepancy (MK-MMD) between domains is taken as the regularization term of the domain adversarial loss, thereby further decreasing the edge distribution distance between domains. The proposed model improves the adaptability to unseen noises by encouraging the feature encoder to generate domain-invariant features. The model was evaluated using cross-noise and cross-language-and-noise experiments, and the results show that the proposed method provides considerable improvements over the baseline without an adaptation in the perceptual evaluation of speech quality (PESQ), the short time objective intelligibility (STOI) and the frequency-weighted signal-to-noise ratio (FWSNR).
Jiaming Cheng 0005, Ruiyu Liang, Zhenlin Liang, Li Zhao 0003, Chengwei Huang, Björn W. Schuller
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Transfer Learning From Speech Synthesis to Voice Conversion With Non-Parallel Training Data
abstract
We present a novel voice conversion (VC) framework by learning from a text-to-speech (TTS) synthesis system, that is called TTS-VC transfer learning or TTL-VC for short. We first develop a multi-speaker speech synthesis system with sequence-to-sequence encoder-decoder architecture, where the encoder extracts the linguistic representations of input text, while the decoder, conditioned on target speaker embedding, takes the context vectors and the attention recurrent network cell output to generate target acoustic features. We take advantage of the fact that TTS system maps input text to speaker independent context vectors, thus re-purpose such a mapping to supervise the training of the latent representations of an encoder-decoder voice conversion system. In the voice conversion system, the encoder takes speech instead of text as the input, while the decoder is functionally similar to the TTS decoder. As we condition the decoder on a speaker embedding, the system can be trained on non-parallel data for any-to-any voice conversion. During voice conversion training, we present both text and speech to speech synthesis and voice conversion networks respectively. At run-time, the voice conversion network uses its own encoder-decoder architecture without the need of text input. Experiments show that the proposed TTL-VC system outperforms two competitive voice conversion baselines consistently, namely phonetic posteriorgram and AutoVC methods, in terms of speech quality, naturalness, and speaker similarity.
Mingyang Zhang 0003, Yi Zhou 0020, Li Zhao 0003, Haizhou Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Matrix Decomposition Based Low-Complexity FIR Filter: Further Results
abstract
The matrix decomposition (MD) based FIR filter design technique can synthesize a traditional FIR filter with much less implementation complexity. A scheme for obtaining more sparse coefficients of a MD-FIR filter is proposed, which consists of two parts. The first part is the previous procedure of designing a MD-FIR filter (i.e., design an initial MD-FIR filter using a certain MD method and optimize its coefficients). The second part is the proposed procedure of obtaining more sparse coefficients, where the MD-FIR filters' coefficients also need to be optimized. The performances of the various initial MD-FIR filters, which are obtained based on the various MD methods, in the implementation of this scheme are experimentally compared. The MD-FIR filter's coefficients can be effectively optimized by the trust-region iterative-gradient-searching (TR-IGS) algorithm. We present further results on the TR-IGS. The error bound for each iteration of TR-IGS is analyzed. A convergent online implementation scheme of the TR-IGS is presented and analyzed theoretically and experimentally. The proof of the convergence is provided. A sufficient condition for determining whether a theoretical termination point of the scheme is a strict local optimum point is provided. The step-size optimization problem for the TR-IGS is analyzed theoretically and experimentally.
Hao Wang 0004, Zhijin Zhao, Li Zhao 0003
ISCAS3
2020 ADHD classification by dual subspace learning using resting-state functional connectivity
Ying Chen 0013, Yibin Tang, Xiaofeng Liu 0006, Li Zhao 0003, Zhishun Wang
Artif. Intell. Medicine5
2020 DNN-based speech enhancement with self-attention on feature dimension
Jiaming Cheng 0005, Ruiyu Liang, Li Zhao 0003
Multim. Tools Appl.3
2020 DeepConversion: Voice conversion with limited parallel training data
abstract
A deep neural network approach to voice conversion usually depends on a large amount of parallel training data from source and target speakers. In this paper, we propose a novel conversion pipeline, DeepConversion, that leverages a large amount of non-parallel, multi-speaker data, but requires only a small amount of parallel training data. It is believed that we can represent the shared characteristics of speakers by training a speaker independent general model on a large amount of publicly available, non-parallel, multi-speaker speech data. Such general model can then be used to learn the mapping between source and target speaker more effectively from a limited amount of parallel training data. We also propose a strategy to make full use of the parallel data in all models along the pipeline. In particular, the parallel data is used to adapt the general model towards the source-target speaker pair to achieve a coarse grained conversion, and to develop a compact Error Reduction Network (ERN) for a fine-grained conversion. The parallel data is also used to adapt the WaveNet vocoder towards the source-target pair. The experiments show that DeepConversion that only uses a limited amount of parallel training data, consistently outperforms the traditional approaches that use a large amount of parallel training data, in both objective and subjective evaluations.
Mingyang Zhang 0003, Berrak Sisman, Li Zhao 0003, Haizhou Li 0001
Speech Commun.3
2019 Autonomous Emotion Learning in Speech: A View of Zero-Shot Speech Emotion Recognition
abstract
Conventionally, speech emotion recognition is achieved using passive learning approaches.Differing from such approaches, we herein propose and develop a dynamic method of autonomous emotion learning based on zero-shot learning.The proposed methodology employs emotional dimensions as the attributes in the zero-shot learning paradigm, resulting in two phases of learning, namely attribute learning and label learning.Attribute learning connects the paralinguistic features and attributes utilising speech with known emotional labels, while label learning aims at defining unseen emotions through the attributes.The experimental results achieved on the CINEMO corpus indicate that zero-shot learning is a useful technique for autonomous speech-based emotion learning, achieving accuracies considerably better than chance level and an attribute-based gold-standard setup.Furthermore, different emotion recognition tasks, emotional attributes, and employed approaches strongly influence system performance.
Xinzhou Xu, Nicholas Cummins, Zixing Zhang 0001, Li Zhao 0003, Björn W. Schuller
INTERSPEECH5
2019 Connecting Subspace Learning and Extreme Learning Machine in Speech Emotion Recognition
abstract
Speech emotion recognition (SER) is a powerful tool for endowing computers with the capacity to process information about the affective states of users in human-machine interactions. Recent research has shown the effectiveness of graph embedding-based subspace learning and extreme learning machine applied to SER, but there are still various drawbacks in these two techniques that limit their application. Regarding subspace learning, the change from linearity to nonlinearity is usually achieved through kernelization, whereas extreme learning machines only take label information into consideration at the output layer. In order to overcome these drawbacks, this paper leverages extreme learning machines for dimensionality reduction and proposes a novel framework to combine spectral regression-based subspace learning and extreme learning machines. The proposed framework contains three stages-data mapping, graph decomposition, and regression. At the data mapping stage, various mapping strategies provide different views of the samples. At the graph decomposition stage, specifically designed embedding graphs provide a possibility to better represent the structure of data through generating virtual coordinates. Finally, at the regression stage, dimension-reduced mappings are achieved by connecting the virtual coordinates and data mapping. Using this framework, we propose several novel dimensionality reduction algorithms, apply them to SER tasks, and compare their performance to relevant state-of-the-art methods. Our results on several paralinguistic corpora show that our proposed techniques lead to significant improvements.
Xinzhou Xu, Eduardo Coutinho, Li Zhao 0003, Björn W. Schuller
IEEE Trans. Multim.5
2017 A Two-Dimensional Framework of Multiple Kernel Subspace Learning for Recognizing Emotion in Speech
abstract
As a highly active topic in computational paralinguistics, speech emotion recognition (SER) aims to explore ideal representations for emotional factors in speech. In order to improve the performance of SER, multiple kernel learning (MKL) dimensionality reduction has been utilized to obtain effective information for recognizing emotions. However, the solution of MKL usually provides only one nonnegative mapping direction for multiple kernels; this may lead to loss of valuable information. To address this issue, we propose a two-dimensional framework for multiple kernel subspace learning. This framework provides more linear combinations on the basis of MKL without nonnegative constraints, which preserves more information in the learning procedures. It also leverages both of MKL and two-dimensional subspace learning, combining them into a unified structure. To apply the framework to SER, we also propose an algorithm, namely generalised multiple kernel discriminant analysis (GMKDA), by employing discriminant embedding graphs in this framework. GMKDA takes advantage of the additional mapping directions for multiple kernels in the proposed framework. In order to evaluate the performance of the proposed algorithm a wide range of experiments is carried out on several key emotional corpora. These experimental results demonstrate that the proposed methods can achieve better performance compared with some conventional and subspace learning methods in dealing with SER.
Xinzhou Xu, Nicholas Cummins, Zixing Zhang 0001, Li Zhao 0003, Björn W. Schuller
IEEE ACM Trans. Audio Speech Lang. Process.6
2016 Speech emotion recognition using transfer non-negative matrix factorization
abstract
In practical situations, the emotional speech utterances are often collected from different devices and conditions, which will obviously affect the recognition performance. To address this issue, in this paper, a novel transfer non-negative matrix factorization (TNMF) method is presented for cross-corpus speech emotion recognition. First, the NMF algorithm is adopted to learn a latent common feature space for the source and target datasets. Then, the discrepancies between the feature distributions of different corpora are considered, and the maximum mean discrepancy (MMD) algorithm is used for the similarity measurement. Finally, the TNMF approach, which integrates the NMF and MMD algorithms, is proposed. Experiments are carried out on two popular datasets, and the results verify that the TNMF method can significantly outperform the automatic and competitive methods for cross-corpus speech emotion recognition.
Peng Song 0002, Shifeng Ou, Wenming Zheng, Yun Jin, Li Zhao 0003
ICASSP5
2016 Multiscale kernel locally penalised discriminant analysis exemplified by emotion recognition in speech
abstract
We propose a novel method to learn multiscale kernels with locally penalised discriminant analysis, namely Multiscale-Kernel Locally Penalised Discriminant Analysis (MS-KLPDA). As an exemplary use-case, we apply it to recognise emotions in speech. Specifically, we employ the term of locally penalised discriminant analysis by controlling the weights of marginal sample pairs, while the method learns kernels with multiple scales. Evaluated in a series of experiments on emotional speech corpora, our proposed MS-KLPDA is able to outperform the previous research of Multiscale-Kernel Fisher Discriminant Analysis and some conventional methods in solving speech emotion recognition.
Xinzhou Xu, Maryna Gavryukova, Zixing Zhang 0001, Li Zhao 0003, Björn W. Schuller
ICMI5
2015 Dimensionality reduction for speech emotion features by multiscale kernels
abstract
To achieve efficient and compact low-dimensional features for speech emotion recognition, this paper proposes a novel feature reduction method using multiscale kernels in the framework of graph embedding.With Fisher discriminant embedding graph, multiscale Gaussian kernels are used in constructing optimal linear combination of Gram matrices for multiple kernel learning.To evaluate the proposed method, comprehensive experiments, using different public feature sets from the open-source toolbox openSMILE on various corpora, show that the proposed method achieves better performance compared with conventional linear dimensionality reduction methods and singlekernel methods.
Xinzhou Xu, Wenming Zheng, Li Zhao 0003, Björn W. Schuller
INTERSPEECH4
2014 A feature selection and feature fusion combination method for speaker-independent speech emotion recognition
abstract
To enhance the recognition rate of speaker independent speech emotion recognition, a feature selection and feature fusion combination method based on multiple kernel learning is presented. Firstly, multiple kernel learning is used to obtain sparse feature subsets. The features selected at least n times are recombined into another subset named n-subset. The optimal n is determined by 10 cross-validation experiments. Secondly, feature fusion is made at the kernel level. Not only each kind of feature is associated with a kernel, but also the full feature set is associated with a kernel which is not considered in the previous studies. All of the kernels are added together to obtain a combination kernel. The final recognition rate for 7 kinds of emotions on Berlin Database is 83.10%, which outperforms state-of-the-art results and shows the effectiveness of our method. It is also proved that MFCCs play a crucial role in speech emotion recognition.
Yun Jin, Peng Song 0002, Wenming Zheng, Li Zhao 0003
ICASSP4
2014 Text-independent voice conversion using speaker model alignment method from non-parallel speech
Peng Song 0002, Yun Jin, Wenming Zheng, Li Zhao 0003
INTERSPEECH4
2014 Unsupervised learning of phonemes of whispered speech in a noisy environment based on convolutive non-negative matrix factorization
Jian Zhou 0006, Ruiyu Liang, Li Zhao 0003, Cairong Zou
Inf. Sci.3
2013 Non-parallel training for voice conversion based on adaptation method
abstract
In this paper, we propose a simple and efficient non-parallel training scheme for voice conversion (VC). First, the speaker models are adapted from the background model using maximum a posteriori (MAP) technique. Then, by utilizing the parameters of adapted speaker models, the Gaussian normalization and mean transformation methods are proposed for VC, respectively. In addition, to improve the conversion performance of the proposed methods, a combination approach is further presented. Finally, objective and subjective experiments are carried out to evaluate the performance of the proposed scheme, the results demonstrate that our scheme can obtain comparable performance with the traditional GMM method based on parallel corpus.
Peng Song 0002, Wenming Zheng, Li Zhao 0003
ICASSP3
2012 An optimal design of the extrapolated impulse response filter with analytical solutions
Hao Wang 0004, Li Zhao 0003, Haixian Wang, Shiliang Fang, Pingxiang Liu, Xuanmin Du
Signal Process.2
2011 Detecting Practical Speech Emotion in a Cognitive Task
abstract
In this paper we analysis the speech emotions related to cognitive process. An automatic system is established for detecting speech emotions including anxiety, hesitation, confidence and joy. In order to obtain a naturalistic database we use noise to induce negative emotions, sleep deprivation is also used for this purpose. The lack of sleep is an important cause for anxiety. Annotation of emotional speech is then done manually with a self evaluation for the felt emotions in each utterance. Acoustic features are extracted both for valence dimension and arousal dimension including voice quality features. For the recognition algorithm Gaussian Mixture Model is adopted for detecting each type of emotions from neutral speech. Based on the detection of each emotion in the continuous recognition of emotion states an error correcting method is proposed. With the previous emotion states and cognitive performance the detection errors in current emotion state is corrected with empirical probability. Experimental results show that our system can detect "practical" speech emotions related to cognitive process. With the proposed error correcting method the recognition performance is improved compared to the baseline system based on Gaussian Mixture Model. We believe the detection of these "practical" emotions is important for real world applications, especially for helping people cope with negative emotions in cognitive activities.
Cairong Zou, Chengwei Huang, Li Zhao 0003
ICCCN4
2010 Two-dimensional canonical correlation analysis and its application in small sample size face recognition
Ning Sun 0001, Zhen-hai Ji, Cairong Zou, Li Zhao 0003
Neural Comput. Appl.4
2010 Acoustic feedback cancellation based on weighted adaptive projection subgradient method in hearing aids
Qingyun Wang 0004, Li Zhao 0003, Jie Qiao, Cairong Zou
Signal Process.2
2009 Speaker Recognition Based on GMM with an Embedded TDNN
Cunbao Chen, Li Zhao 0003
ICONIP (2)2
2009 The Application of Wavelet Neural Network Optimized by Particle Swarm in Localization of Acoustic Emission Source
Aidong Deng, Li Zhao 0003, Xin Wei 0001
ICONIP (2)2
2009 Image Denoising Using Trivariate Shrinkage Filter in the Wavelet Domain and Joint Bilateral Filter in the Spatial Domain
abstract
This correspondence proposes an efficient algorithm for removing Gaussian noise from corrupted image by incorporating a wavelet-based trivariate shrinkage filter with a spatial-based joint bilateral filter. In the wavelet domain, the wavelet coefficients are modeled as trivariate Gaussian distribution, taking into account the statistical dependencies among intrascale wavelet coefficients, and then a trivariate shrinkage filter is derived by using the maximum a posteriori (MAP) estimator. Although wavelet-based methods are efficient in image denoising, they are prone to producing salient artifacts such as low-frequency noise and edge ringing which relate to the structure of the underlying wavelet. On the other hand, most spatial-based algorithms output much higher quality denoising image with less artifacts. However, they are usually too computationally demanding. In order to reduce the computational cost, we develop an efficient joint bilateral filter by using the wavelet denoising result rather than directly processing the noisy image in the spatial domain. This filter could suppress the noise while preserve image details with small computational cost. Extension to color image denoising is also presented. We compare our denoising algorithm with other denoising techniques in terms of PSNR and visual quality. The experimental results indicate that our algorithm is competitive with other denoising techniques.
Hancheng Yu, Li Zhao 0003, Haixian Wang
IEEE Trans. Image Process.2
2008 An efficient algorithm for Kernel two-dimensional principal component analysis
Ning Sun 0001, Haixian Wang, Zhen-hai Ji, Cairong Zou, Li Zhao 0003
Neural Comput. Appl.5
2008 An Efficient Procedure for Removing Random-Valued Impulse Noise in Images
abstract
In this letter, we present an efficient algorithm for the removal of random-valued impulse noise from a corrupted image by using a reference image. The proposed method uses a statistic of rank-ordered relative differences to identify pixels which are likely to be corrupted by impulse noise. Once a noisy pixel is identified, its value is restored by a simple weighted mean filter. Simulation results indicate that our algorithm provides a significant improvement over many other existing techniques.
Hancheng Yu, Li Zhao 0003, Haixian Wang
IEEE Signal Process. Lett.2
2007 2DCCA: A Novel Method for Small Sample Size Face Recognition
abstract
In the traditional canonical correlation analysis (CCA) based face recognition methods, the size of sample is always smaller than the dimension of sample. This problem is so called the small sample size (SSS) problem. In order to solve this problem, a new supervised learning method called two-dimensional CCA (2DCCA) is developed in this paper. Different from traditional CCA method, 2DCCA directly extracts the features from image matrix rather than matrix-to-vector transformation. In practice, the covariance matrix extracted by 2DCCA is always full rank. Hence the small sample size (SSS) problem can be effectively dealt with by this new developed method. The theory foundation of 2DCCA method is firstly developed, and the construction method for the class-membership matrix Y which is used to precisely represent the relationship between samples and classes in the 2DCCA framework is then clarified. Simultaneously, the analytic form of the generalized inverse of such class-membership matrix is derived. From our experiment results on face recognition, we clearly find that not only the SSS problem can be effectively solved, but also better recognition performance than several other CCA based methods has been achieved
Cairong Zou, Ning Sun 0001, Zhen-hai Ji, Li Zhao 0003
WACV4
2006 A Fast Feature Extraction Method for Kernel 2DPCA
Ning Sun 0001, Haixian Wang, Zhen-hai Ji, Cairong Zou, Li Zhao 0003
ICIC (1)5
2006 Facial Expression Recognition Based on BoostingTree
Ning Sun 0001, Wenming Zheng, Changyin Sun 0001, Cairong Zou, Li Zhao 0003
ISNN (2)5
2006 Gender Classification Based on Boosting Local Binary Pattern
Ning Sun 0001, Wenming Zheng, Changyin Sun 0001, Cairong Zou, Li Zhao 0003
ISNN (2)5
2006 Face recognition using common faces method
Yunhui He, Li Zhao 0003, Cairong Zou
Pattern Recognit.2
2006 Facial expression recognition using kernel canonical correlation analysis (KCCA)
abstract
In this correspondence, we address the facial expression recognition problem using kernel canonical correlation analysis (KCCA). Following the method proposed by Lyons et al. and Zhang et al., we manually locate 34 landmark points from each facial image and then convert these geometric points into a labeled graph (LG) vector using the Gabor wavelet transformation method to represent the facial features. On the other hand, for each training facial image, the semantic ratings describing the basic expressions are combined into a six-dimensional semantic expression vector. Learning the correlation between the LG vector and the semantic expression vector is performed by KCCA. According to this correlation, we estimate the associated semantic expression vector of a given test image and then perform the expression classification according to this estimated semantic expression vector. Moreover, we also propose an improved KCCA algorithm to tackle the singularity problem of the Gram matrix. The experimental results on the Japanese female facial expression database and the Ekman's "Pictures of Facial Affect" database illustrate the effectiveness of the proposed method.
Wenming Zheng, Cairong Zou, Li Zhao 0003
IEEE Trans. Neural Networks4
2005 Expression Recognition Using Elastic Graph Matching
Yujia Cao, Wenming Zheng, Li Zhao 0003, Cairong Zou
ACII3
2005 Speech Emotional Recognition Using Global and Time Sequence Structure Features with MMD
Li Zhao 0003, Yujia Cao, Cairong Zou
ACII1
2005 Discriminative Features Extraction in Minor Component Subspace
Wenming Zheng, Cairong Zou, Li Zhao 0003
ACII3
2005 Boosted Independent Features for Face Expression Recognition
Lianghua He, Die Hu 0002, Cairong Zou, Li Zhao 0003
ISNN (2)5
2005 Weighted maximum margin discriminant analysis with kernels
Wenming Zheng, Cairong Zou, Li Zhao 0003
Neurocomputing3
2005 An Improved Algorithm for Kernel Principal Component Analysis
Wenming Zheng, Cairong Zou, Li Zhao 0003
Neural Process. Lett.3
2005 Foley-Sammon optimal discriminant vectors using kernel approach
abstract
A new nonlinear feature extraction method called kernel Foley-Sammon optimal discriminant vectors (KFSODVs) is presented in this paper. This new method extends the well-known Foley-Sammon optimal discriminant vectors (FSODVs) from linear domain to a nonlinear domain via the kernel trick that has been used in support vector machine (SVM) and other commonly used kernel-based learning algorithms. The proposed method also provides an effective technique to solve the so-called small sample size (SSS) problem which exists in many classification problems such as face recognition. We give the derivation of KFSODV and conduct experiments on both simulated and real data sets to confirm that the KFSODV method is superior to the previous commonly used kernel-based learning algorithms in terms of the performance of discrimination.
Wenming Zheng, Li Zhao 0003, Cairong Zou
IEEE Trans. Neural Networks2
2004 Face recognition using two novel nearest neighbor classifiers
abstract
In this paper, two novel classifiers, based on locally nearest neighborhood rules, called nearest neighbor line (NNL) and nearest neighbor plane (NNP), are presented for face recognition. The underlying idea of both classifiers is the local linear combination technique that has been previously used in locally linear embedding (LLE) for nonlinear dimension reduction. In comparison to other linear combination based classifiers, such as the nearest feature line (NFL) and the nearest feature plane (NFP), the proposed method has a much lower computation cost. Furthermore, the experimental results on the ORL face database have shown that the performance of both proposed methods are competitive to the NFL and NFP in face classification.
Wenming Zheng, Cairong Zou, Li Zhao 0003
ICASSP (5)3
2004 Facial Expression Recognition Using Kernel Discriminant Plane
Wenming Zheng, Cairong Zou, Li Zhao 0003
ISNN (1)4
2004 A Modified Algorithm for Generalized Discriminant Analysis
abstract
Generalized discriminant analysis (GDA) is an extension of the classical linear discriminant analysis (LDA) from linear domain to a nonlinear domain via the kernel trick. However, in the previous algorithm of GDA, the solutions may suffer from the degenerate eigenvalue problem (i.e., several eigenvectors with the same eigenvalue), which makes them not optimal in terms of the discriminant ability. In this letter, we propose a modified algorithm for GDA (MGDA) to solve this problem. The MGDA method aims to remove the degeneracy of GDA and find the optimal discriminant solutions, which maximize the between-class scatter in the subspace spanned by the degenerate eigenvectors of GDA. Theoretical analysis and experimental results on the ORL face database show that the MGDA method achieves better performance than the GDA method.
Wenming Zheng, Li Zhao 0003, Cairong Zou
Neural Comput.2
2004 An efficient algorithm to solve the small sample size problem for LDA
Wenming Zheng, Li Zhao 0003, Cairong Zou
Pattern Recognit.2
2004 Locally nearest neighbor classifiers for pattern classification
Wenming Zheng, Li Zhao 0003, Cairong Zou
Pattern Recognit.2