Peng Song 0002

dblp:58/3960-2 · DBLP profile ↗
← Back
55ranked-venue papers
8as first author
40since 2021 · last 2026
0000-0002-6567-663XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 4 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 ATGFB-MFF: Adaptive Text-Guided Fiber Bundle Feature Fusion with LLMs for Multimodal Sentiment Analysis and Emotion Recognition in Conversations
abstract
Multimodal Sentiment Analysis (MSA) and Emotion Recognition in Conversations (ERC) have rapidly developed into pivotal tasks in artificial intelligence. Large Language Models (LLMs) offer powerful semantic reasoning and computational capabilities, showing great potential for understanding emotional content. However, when applied to multimodal sentiment data, LLMs face significant challenges, including the inability to directly process heterogeneous data, difficulties in coping with feature misalignment and suboptimal cross-modal fusion. To address these challenges, we propose a novel multimodal sentiment inference framework named ATGFB-MFF which grounded in fiber bundle theory. This method decomposes multimodal features into an adaptive text-guided shared semantic space and fiber offset spaces to achieve structured alignment and fusion. Then the fused features are converted into structured pseudo-token sequences for effective inference via frozen LLMs. We also introduce two loss functions respectively called shared space consistency loss and fiber offset regularization loss which are used to improve representation stability. Extensive experiments on four benchmark datasets demonstrate that ATGFB-MFF consistently outperforms state-of-the-art baselines. These results highlight the efficacy of geometric structural modeling in unlocking the potential of LLMs for multimodal sentiment inference.
Zhaowei Liu 0001, Weiqing Yan, Peng Song 0002, Yongchao Song, Rufei Gao
WWW4
2026 Fast multi-view unsupervised feature selection via structure correlation guidance
Changjia Wang, Peng Song 0002
Neurocomputing2
2026 COALN-MvC: A continuous optimized anchor learning network for multi-view clustering
Beihua Yang, Peng Song 0002, Yunpeng Zeng
Knowl. Based Syst.2
2026 Robust structure-preservation tensorized representation for multi-view unsupervised feature selection
Peng Song 0002, Changjia Wang, Beihua Yang, Zhaowei Liu 0001
Neural Networks2
2026 Enhanced anchor contrastive multi-view representations learning network for clustering
Beihua Yang, Peng Song 0002
Neural Networks2
2026 Hypergraph regularization-based anchor learning for multi-view clustering
Yunpeng Zeng, Peng Song 0002, Beihua Yang, Changjia Wang, Guanghao Du, Yanwei Yu, Wenming Zheng
Pattern Recognit.2
2026 Dual-Hypergraph Based Symmetrical Self-Representation Learning for Cross-Domain Facial Expression Recognition
Yuhan Cheng, Peng Song 0002, Siqi Fu, Xingxin Wan, Changjia Wang, Wenming Zheng
IEEE Trans. Comput. Soc. Syst.2
2026 Coupled Sparse Subspace Alignment-Based Domain Adaptation for Speech Emotion Recognition
abstract
Speech emotion recognition (SER) is crucial for human–computer interaction (HCI), yet remains a challenge in cross-domain scenarios. Emotional expressions vary significantly across speakers, languages, and recording conditions, leading to serious domain shifts. Coupled subspace learning has recently attracted considerable attention in domain adaptation (DA) as an effective approach to mitigating domain discrepancies by capturing common and domain-specific information. However, existing algorithms suffer from two limitations: 1) most methods directly adopt classifiers [e.g., support vector machine (SVM)] to incorporate discriminative information of the target domain, but such strategies lack flexibility and adaptability; and 2) the features learned from coupled projection matrices are redundant and poorly discriminative. To address these issues, we propose a novel DA approach named coupled sparse subspace alignment (CSSA) for cross-domain SER. Specifically, CSSA first performs latent representation learning on the unlabeled target domain data, in which the latent representation matrix is then optimized into a pseudolabel matrix to provide emotional guidance. Meanwhile, it models the source and target domains separately by sparse regression, thereby learning both discriminative and domain-specific information. Subsequently, CSSA performs coupled subspace alignment to reduce the domain discrepancy, where the dual projection matrices are progressively aligned to enhance their similarity. Additionally, a graph Laplacian regularization is applied to the cross-domain data to capture the local geometric structure. Extensive experiments on five public SER datasets demonstrate the superiority of CSSA over state-of-the-art DA methods.
Siqi Fu, Peng Song 0002, Wenming Zheng
IEEE Trans. Comput. Soc. Syst.2
2026 Dynamic Graph Consistent Weighted Subspace Learning for Cross-Domain Speech Emotion Recognition
abstract
In recent years, cross-domain speech emotion recognition (SER) has attracted considerable interest. Most transfer subspace learning based SER methods lack unified adaptive constraints, making it difficult to balance discriminative capability and domain alignment, which limits their cross-domain generalization. To address these problems, we propose a novel domain adaptation (DA) approach called dynamic graph consistent weighted subspace learning (DGCWSL). Specifically, DGCWSL first projects samples from the source and target domains into a shared low-dimensional discriminative subspace, then performs cross-domain instance reconstruction, representing each target as a weighted combination of source instances. In parallel, a dynamic graph is constructed to capture local structural information between domains while preserving the data manifold. Subsequently, label supervision and discriminative learning between domains are achieved through linear regression. Furthermore, we introduce an adaptive weighted matrix that enforces consistent feature contributions across the distance metric, instance alignment, and discriminative regression, thereby mitigating both overfitting and underfitting. Finally, extensive experiments are conducted on four public datasets. The results confirm the superiority of DGCWSL over several state-of-the-art DA methods.
Peng Song 0002, Siqi Fu, Zhaowei Liu 0001, Changjia Wang, Wenming Zheng
IEEE Trans. Comput. Soc. Syst.2
2025 High-order correlation preserved multi-view unsupervised feature selection
Meng Duan, Peng Song 0002, Shixuan Zhou, Yuanbo Cheng, Jinshuai Mu, Wenming Zheng
Eng. Appl. Artif. Intell.2
2025 Essential anchor graph learning for incomplete multi-view clustering
Peng Song 0002, Jinshuai Mu, Yuanbo Cheng, Zhaohu Liu, Wenming Zheng
Eng. Appl. Artif. Intell.1
2025 RM-BGNN: A weakly informative Bayesian graph neural network based on residual mechanism
Jihao Dong, Zhaowei Liu 0001, Peng Song 0002, Jinglei Liu, Anzuo Jiang
Neurocomputing4
2025 Weighted tensor-based consistent anchor graph learning for multi-view clustering
Guanghao Du, Peng Song 0002, Yuanbo Cheng, Zhaowei Liu 0001, Yanwei Yu, Wenming Zheng
Neurocomputing2
2025 Robust multi-view unsupervised feature selection via latent consensus structure learning
Changjia Wang, Peng Song 0002
Inf. Sci.2
2025 Enhanced tensor based embedding anchor learning for multi-view clustering
Beihua Yang, Peng Song 0002, Yuanbo Cheng, Shixuan Zhou, Zhaowei Liu 0001
Inf. Sci.2
2025 Deep concept subspace embedding for multi-view clustering
Guanghao Du, Peng Song 0002, Beihua Yang
Knowl. Based Syst.2
2025 Low-rank tensor based smooth representation learning for multi-view unsupervised feature selection
Changjia Wang, Peng Song 0002, Meng Duan, Shixuan Zhou, Yuanbo Cheng
Knowl. Based Syst.2
2025 Label completion based concept factorization for incomplete multi-view clustering
Beihua Yang, Peng Song 0002, Yuanbo Cheng, Zhaowei Liu 0001, Yanwei Yu
Knowl. Based Syst.2
2025 Common Discriminative Latent Space Learning for Cross-Domain Speech Emotion Recognition
abstract
Cross-domain speech emotion recognition (SER) has received increasing attention in recent years. Existing transfer subspace learning and regression-based SER methods have the following drawbacks. The features in the subspace are still insufficiently representative and discriminative, and direct regression would lead to information loss. To address these problems, we present a novel common discriminative latent space learning (CDLSL) method for cross-domain SER. To be specific, we first obtain a common latent space by imposing a projection matrix on the cross-domain data. Meanwhile, we impose an uncorrelated constraint on the projection matrix to ensure that the features are representative and discriminative after dimension reduction. Then, we implement a graph regularization term on the latent representations of the samples to capture the local similarity information. Furthermore, to obtain a more discriminative common latent space, we introduce the label information by aligning the latent space with the relaxed label space, while mitigating the information loss for regression. Extensive experimental results validate the superiority of the proposed method over the state-of-the-art competitors.
Siqi Fu, Peng Song 0002, Hao Wang 0269, Zhaowei Liu 0001, Wenming Zheng
IEEE Trans. Comput. Soc. Syst.2
2024 Multi-Source Unsupervised Transfer Components Learning for Cross-Domain Speech Emotion Recognition
abstract
As an important research direction in the field of speech signal processing, cross-domain speech emotion recognition (SER) has attracted extensive attention. In practice, it is challenging to collect enough labeled samples from single source domain to train robust classifiers. To this end, this paper presents a novel method named multi-source unsupervised transfer components learning (MUTCL) for cross-domain SER. In MUTCL, we first adopt a PCA-like strategy and apply it to multi-source domains, aiming to preserve both intra-domain individuality and inter-domain commonality principal components within each domain. Simultaneously, a simple alignment strategy is developed to guide cross-domain samples to have similar structures, thus preserving more transfer components. Moreover, an adaptive weight strategy is utilized to determine the contribution of each source domain. We conduct experiments on five benchmark datasets, and the results show that MUTCL achieves excellent performance compared with some state-of-the-art methods.
Shenjie Jiang, Peng Song 0002, Shaokai Li, Wenming Zheng
ICASSP2
2024 Incomplete multi-view clustering via confidence graph completion based tensor decomposition
Yuanbo Cheng, Peng Song 0002
Expert Syst. Appl.2
2024 Comprehensive multi-view self-representations for clustering
abstract
Subspace learning-based methods have shown excellent performance for multi-view clustering, yet have the following problems: (1) most existing methods obtain the subspace representation from the original space, which might contain noises and cannot guarantee a clean enough subspace representation; (2) existing methods mainly focus on the consistency of the subspace representation, while the unique information of each view is not sufficiently exploited. To solve these two problems, we propose a novel multi-view subspace clustering method called comprehensive multi-view self-representations (CMSR). Specifically, we learn the original coefficient matrix of each view through the self-representation, which can reduce the noise of the original space to some extent. Then, we learn the subspace representation of the original coefficient matrix and decompose it into a consistent coefficient matrix and multiple diverse coefficient matrices, which can exploit the consistent and complementary information of multi-view data. Further, we impose the Schatten p -norm constraint on the consistent coefficient matrix to capture robust consistent information. Finally, the comprehensive results on eight real datasets demonstrate the versatility and effectiveness of the proposed method.
Yuanbo Cheng, Peng Song 0002, Jinshuai Mu, Yanwei Yu, Wenming Zheng
Expert Syst. Appl.2
2024 Deep low-rank tensor embedding for multi-view subspace clustering
Zhaohu Liu, Peng Song 0002
Expert Syst. Appl.2
2024 Consistency-exclusivity guided unsupervised multi-view feature selection
Shixuan Zhou, Peng Song 0002
Neurocomputing2
2024 Clean affinity matrix induced hyper-Laplacian regularization for unsupervised multi-view feature selection
Peng Song 0002, Shixuan Zhou, Jinshuai Mu, Meng Duan, Yanwei Yu, Wenming Zheng
Inf. Sci.1
2024 Common Latent Embedding Space for Cross-Domain Facial Expression Recognition
abstract
In practical facial expression recognition (FER), the training data and test data are often obtained from different domains. It is obvious that the domain disparity could significantly degrade the recognition performance. To tackle this challenging cross-domain FER problem, we put forward a novel method termed common latent embedding space (CLES). To be specific, first, we obtain a common embedding space for cross-domain samples by matrix factorization (MF). Then, the dual-graph Laplacian is applied to this common embedding space to narrow the gap across distinct domains and, meanwhile, explores the inherent geometric information. Furthermore, to characterize the global relationship of the cross-domain samples, the self-representation strategy is used to guide the learning of the common embedding space. Finally, comprehensive experiments on four benchmark databases indicate that the proposed method can achieve better performance in comparison with the state-of-the-art methods on cross-domain FER tasks.
Peng Song 0002, Shaokai Li, Wenming Zheng
IEEE Trans. Comput. Soc. Syst.2
2024 Graph-Diffusion-Based Domain-Invariant Representation Learning for Cross-Domain Facial Expression Recognition
abstract
The precondition that most of the existing facial expression recognition (FER) algorithms have succeeded lies in that the training (source) and test (target) samples are independent of each other and identically distributed. However, it is too strict to satisfy this precondition in the real-world. To this end, we propose a novel graph-diffusion-based domain-invariant representation learning (GDRL) model for the cross-domain FER scenario where there exist distribution shifts between various domains. Specifically, a low-dimensional space mapping strategy is first adopted to diminish the domain mismatch. Then, by skillfully combining the local graph embedding and affinity graph diffusion, the local geometric structures can be effectively modeled and the deeper higher-order relationships of samples from various domains can be captured. In addition, in order to better guide the transfer process and learn a more discriminative and invariant representation, we take into account the label consistency. Experimental results on four laboratory-controlled databases and two in-the-wild databases demonstrate that our proposed model can yield better recognition performance compared with state-of-the-art domain adaptation methods.
Peng Song 0002, Wenming Zheng
IEEE Trans. Comput. Soc. Syst.2
2023 A Generalized Subspace Distribution Adaptation Framework for Cross-Corpus Speech Emotion Recognition
abstract
In this paper, we propose a novel transfer learning framework, named generalized subspace distribution adaptation (GSDA), to tackle the challenging cross-corpus speech emotion recognition problem. First, we learn a common low-dimensional feature subspace by utilizing a generalized subspace learning method. Second, we develop a novel distance metric to reduce the divergence between the source and target corpora, which can efficiently explore the similarity and dissimilarity information in the process of knowledge transfer. Third, to demonstrate the effectiveness of our framework, we apply GSDA to the traditional subspace learning algorithms. Finally, we conduct extensive experiments by using the low-level features and deep features on three popular emotional databases, i.e., Berlin, IEMOCAP, and CVE. The results demonstrate that the proposed framework can achieve better performance than several state-of-the-art transfer learning approaches.
Shaokai Li, Peng Song 0002, Yun Jin, Wenming Zheng
ICASSP2
2023 Unsupervised Transfer Components Learning for Cross-Domain Speech Emotion Recognition
Shenjie Jiang, Peng Song 0002, Shaokai Li, Keke Zhao, Wenming Zheng
INTERSPEECH2
2023 Joint Instance Reconstruction and Feature Subspace Alignment for Cross-Domain Speech Emotion Recognition
Keke Zhao, Peng Song 0002, Shaokai Li, Wenming Zheng
INTERSPEECH2
2023 DeepFormer: a hybrid network based on convolutional neural network and flow-attention mechanism for identifying the function of DNA sequences
abstract
Identifying the function of DNA sequences accurately is an essential and challenging task in the genomic field. Until now, deep learning has been widely used in the functional analysis of DNA sequences, including DeepSEA, DanQ, DeepATT and TBiNet. However, these methods have the problems of high computational complexity and not fully considering the distant interactions among chromatin features, thus affecting the prediction accuracy. In this work, we propose a hybrid deep neural network model, called DeepFormer, based on convolutional neural network (CNN) and flow-attention mechanism for DNA sequence function prediction. In DeepFormer, the CNN is used to capture the local features of DNA sequences as well as important motifs. Based on the conservation law of flow network, the flow-attention mechanism can capture more distal interactions among sequence features with linear time complexity. We compare DeepFormer with the above four kinds of classical methods using the commonly used dataset of 919 chromatin features of nearly 4.9 million noncoding DNA sequences. Experimental results show that DeepFormer significantly outperforms four kinds of methods, with an average recall rate at least 7.058% higher than other methods. Furthermore, we confirmed the effectiveness of DeepFormer in capturing functional variation using Alzheimer's disease, pathogenic mutations in alpha-thalassemia and modification in CCCTC-binding factor (CTCF) activity. We further predicted the maize chromatin accessibility of five tissues and validated the generalization of DeepFormer. The average recall rate of DeepFormer exceeds the classical methods by at least 1.54%, demonstrating strong robustness.
Zhou Yao, Wenjing Zhang 0003, Peng Song 0002, Yuxue Hu, Jianxiao Liu
Briefings Bioinform.3
2023 Dual-graph regularized concept factorization for multi-view clustering
Jinshuai Mu, Peng Song 0002, Shaokai Li
Expert Syst. Appl.2
2023 Tensor-based consensus learning for incomplete multi-view clustering
Jinshuai Mu, Peng Song 0002, Yanwei Yu, Wenming Zheng
Expert Syst. Appl.2
2023 Soft-label guided non-negative matrix factorization for unsupervised feature selection
Shixuan Zhou, Peng Song 0002
Expert Syst. Appl.2
2023 Structural regularization based discriminative multi-view unsupervised feature selection
Shixuan Zhou, Peng Song 0002, Yanwei Yu, Wenming Zheng
Knowl. Based Syst.2
2023 Learning Transferable Sparse Representations for Cross-Corpus Facial Expression Recognition
abstract
An assumption widely used in traditional facial expression recognition algorithms is that the training and testing are conducted on the same dataset. However, this assumption does not hold in practice, in which the training data and testing data are often from different datasets. In this scenario, directly deploying these algorithms would lead to severe information loss and performance degradation due to the domain shift. To address this challenging problem, in this article, we propose a novel transferable sparse subspace representation method (TSSR) for cross-corpus facial expression recognition. Specifically, in order to reduce the cross-corpus mismatch, inspired by sparse subspace clustering, we advocate reconstructing the source and target samples using the source data points based on$\ell _1-$norm sparse representation. Each data point in source and target corpora can be ideally represented as a combination of a few other source points from its own subspace. Moreover, we take into account the local geometrical information within the cross-corpus data by adopting a graph Laplacian regularizer, which can efficiently preserve the local manifold structure and better transfer knowledge between two corpora. Finally, extensive experiments on several facial expression datasets are conducted to evaluate the recognition performance of TSSR. Experimental results demonstrate the superiority of the proposed method over some state-of-the-art methods.
Peng Song 0002, Wenming Zheng
IEEE Trans. Affect. Comput.2
2023 Joint Local-Global Discriminative Subspace Transfer Learning for Facial Expression Recognition
abstract
Traditional facial expression recognition (FER) has achieved satisfactory results to some extent, and most of the current methods are trained and evaluated on a single database. However, in real applications, the training and testing images are often collected in different scenarios, which will lead to performance degeneration. To tackle this problem, in this paper, we propose a novel transfer learning approach, named joint local-global discriminative subspace transfer learning (LGDSTL), for cross-database FER. In LGDSTL, first, we develop a joint local-global graph as the distance metric, in which we not only consider the local discriminative geometric structure for each database, but also consider a global graph to transfer knowledge. In this way, the discrepancy between the two databases will be significantly reduced. Then, we present a pairwise regression function to guide the discriminative subspace transfer learning. Additionally, a data reconstruction constraint is introduced to preserve the main discriminative information. Finally, comparative studies on six popular benchmarks demonstrate the effectiveness of the proposed approach.
Wenjing Zhang 0003, Peng Song 0002, Wenming Zheng
IEEE Trans. Affect. Comput.2
2023 Multi-Source Discriminant Subspace Alignment for Cross-Domain Speech Emotion Recognition
abstract
Cross-domain speech emotion recognition (SER) is an effective strategy to improve the generalization ability of emotion classification models, which is an important research direction in speech signal processing. However, since the speech signals are non-stationary, it is difficult to train a robust classifier from single-source emotional corpus. To solve this shortcoming, we propose a novel method named multi-source discriminant subspace alignment (MDSA) for cross-domain SER. In MDSA, we first conduct linear discriminant analysis (LDA) in the multi-source domain. Then, the instances in the multi-source discriminant subspace are used to linearly reconstruct the instances in the target subspace. At the same time, the reconstruction contribution of each source discriminant subspace is determined by adaptive weights. Furthermore, the multi-source discriminant subspace is aligned by reducing the loss between projections, which can make our model more robust. In this way, MDSA considers both the alignment of cross-domain data distribution and the structural information of cross-domain instances. Finally, extensive experiments are conducted on five standard emotional corpora, i.e., Berlin, IEMOCAP, CVE, EMOVO, and TESS, and the results demonstrate the proposed MDSA is superior to several state-of-the-art transfer learning algorithms in terms of performance. The codes are available athttps://github.com/shaokai1209/MDSA.
Shaokai Li, Peng Song 0002, Wenming Zheng
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Coupled Discriminant Subspace Alignment for Cross-database Speech Emotion Recognition
Shaokai Li, Peng Song 0002, Keke Zhao, Wenjing Zhang 0003, Wenming Zheng
INTERSPEECH2
2022 Incomplete multi-view clustering via virtual-label guided matrix factorization
Peng Song 0002
Expert Syst. Appl.2
2020 Dynamic Representation Learning for Large-Scale Attributed Networks
abstract
Network embedding, which aims at learning low-dimensional representations of nodes in a network, has drawn much attention for various network mining tasks, ranging from link prediction to node classification. In addition to network topological information, there also exist rich attributes associated with network structure, which exerts large effects on the network formation. Hence, many efforts have been devoted to tackling attributed network embedding tasks. However, they are also limited in their assumption of static network data as they do not account for evolving network structure as well as changes in the associated attributes. Furthermore, scalability is a key factor when performing representation learning on large-scale networks with huge number of nodes and edges. In this work, we address these challenges by developing the DRLAN-Dynamic Representation Learning framework for large-scale Attributed Networks. The DRLAN model generalizes the dynamic attributed network embedding from two perspectives: First, we develop an integrative learning framework with an offline batch embedding module to preserve both the node and attribute proximities, and online network embedding model that recursively updates learned representation vectors. Second, we design a recursive pre-projection mechanism to efficiently model the attribute correlations based on the associative property of matrices. Finally, we perform extensive experiments on three real-world network datasets to show the superiority of DRLAN against state-of-the-art network embedding techniques in terms of both effectiveness and efficiency. The source code is available at: https://github.com/ZhijunLiu95/DRLAN.
Chao Huang 0001, Yanwei Yu, Peng Song 0002, Baode Fan, Junyu Dong
CIKM4
2020 Student Performance Prediction Based on Multi-view Network Embedding
Jianian Li, Yanwei Yu, Yunhong Lu, Peng Song 0002
PRCV (3)4
2020 Find you if you drive: Inferring home locations for vehicles with surveillance camera data
Yanwei Yu, Peng Song 0002, Xianfeng Tang, Lei Cao 0004, Xiangrong Tong
Knowl. Based Syst.3
2020 Scalable KDE-based top-n local outlier detection over large-scale data streams
Yanwei Yu, Peng Song 0002, Yangyang Fan, Xiangrong Tong
Knowl. Based Syst.3
2020 Feature Selection Based Transfer Subspace Learning for Speech Emotion Recognition
abstract
Cross-corpus speech emotion recognition has recently received considerable attention due to the widespread existence of various emotional speech. It takes one corpus as the training data aiming to recognize emotions of another corpus, and generally involves two basic problems, i.e., feature matching and feature selection. Many previous works study these two problems independently, or just focus on solving the first problem. In this paper, we propose a novel algorithm, called feature selection based transfer subspace learning (FSTSL), to address these two problems. To deal with the first problem, a latent common subspace is learnt by reducing the difference of different corpora and preserving the important properties. Meanwhile, we adopt the l2,1-norm on the projection matrix to deal with the second problem. Besides, to guarantee the subspace to be robust and discriminative, the geometric information of data is exploited simultaneously in the proposed FSTSL framework. Empirical experiments on cross-corpus speech emotion recognition tasks demonstrate that our proposed method can achieve encouraging results in comparison with state-of-the-art algorithms.
Peng Song 0002, Wenming Zheng
IEEE Trans. Affect. Comput.1
2020 EEG Emotion Recognition Using Dynamical Graph Convolutional Neural Networks
abstract
In this paper, a multichannel EEG emotion recognition method based on a novel dynamical graph convolutional neural networks (DGCNN) is proposed. The basic idea of the proposed EEG emotion recognition method is to use a graph to model the multichannel EEG features and then perform EEG emotion classification based on this model. Different from the traditional graph convolutional neural networks (GCNN) methods, the proposed DGCNN method can dynamically learn the intrinsic relationship between different electroencephalogram (EEG) channels, represented by an adjacency matrix, via training a neural network so as to benefit for more discriminative EEG feature extraction. Then, the learned adjacency matrix is used to learn more discriminative features for improving the EEG emotion recognition. We conduct extensive experiments on the SJTU emotion EEG dataset (SEED) and DREAMER dataset. The experimental results demonstrate that the proposed method achieves better recognition performance than the state-of-the-art methods, in which the average recognition accuracy of 90.4 percent is achieved for subject dependent experiment while 79.95 percent for subject independent cross-validation one on the SEED database, and the average accuracies of 86.23, 84.54 and 85.02 percent are respectively obtained for valence, arousal and dominance classifications on the DREAMER database.
Tengfei Song, Wenming Zheng, Peng Song 0002, Zhen Cui 0001
IEEE Trans. Affect. Comput.3
2020 Transfer Sparse Discriminant Subspace Learning for Cross-Corpus Speech Emotion Recognition
abstract
Cross-corpus speech emotion recognition has attracted much attention due to the widespread existence of various emotional speech in life. It takes one corpus for training and another corpus for testing, and generally involves the following two basic problems: the corpus-invariant feature representation and relevance across different corpora. To deal with these two problems, we propose a novel transfer learning method called transfer sparse discriminant subspace learning (TSDSL) in this article. Specifically, to solve the first problem, we learn a common feature subspace of different corpora by introducing the discriminative learning and ℓ2,1-norm penalty, which can learn the most discriminative features across different corpora. To address the second problem, we construct a novel nearest neighbor graph as the distance metric, in which the similarity between different corpora can be measured simultaneously. Extensive experiments are carried out on cross-corpus speech emotion recognition tasks, and the results show that our method can achieve competitive performance compared with state-of-the-art algorithms.
Peng Song 0002
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Transfer Linear Subspace Learning for Cross-Corpus Speech Emotion Recognition
abstract
Speech emotion recognition has received an increasing interest in recent years, which is often conducted on the assumption that speech utterances in training and testing datasets are obtained under the same conditions. However, in reality, this assumption does not hold as the speech data are often collected from different devices or environments. Hence, there exists discrepancy between the training and testing data, which will have an adverse effect on recognition performance. In this paper, we examine the problem of cross-corpus speech emotion recognition. To address it, we present a novel transfer linear subspace learning (TLSL) framework to learn a common feature subspace for source and target datasets. In TLSL, a nearest neighbor graph algorithm is used to measure the similarity between different corpora, and a feature grouping strategy is introduced to divide the emotional features into two categories, i.e., high transferable part (HTP) versus low transferable part (LTP). To explore the proposed TLSL with different scenarios, we propose two kinds of TLSL approaches, called transfer unsupervised linear subspace learning (TULSL) and transfer supervised linear subspace learning (TSLSL), and provide the corresponding solutions for the optimization problems. Extensive experiments on several benchmark datasets validate the effectiveness of TLSL for cross-corpus speech emotion recognition.
Peng Song 0002
IEEE Trans. Affect. Comput.1
2017 Discovering Evolving Moving Object Groups from Massive-Scale Trajectory Streams
abstract
The increasing pervasiveness of object tracking technologies leads to huge volumes of spatio-temporal data collected in the form of trajectory streams. The discovery of useful group patterns from moving objects' movement behaviors in trajectory streams is critical for real time applications ranging from transportation management to military surveillance. In this work we propose a novel type of group pattern, called evolving group, which models the unusual group events of moving objects that travel together within density connected clusters in evolving streaming trajectories. Our theoretical analysis and empirical study on the Beijing Taxi data demonstrate its effectiveness in capturing development, evolution and trend of group events of moving objects in streaming context. Furthermore, we propose a discovery framework that efficiently supports online detection of evolving groups over massive-scale trajectory streams using sliding window. It contains three phases along with a set of novel optimization techniques designed to minimize the computation costs. Our comprehensive empirical study demonstrates that our discovery framework is effective and efficient on real-world high volume trajectory streams.
Ruoshan Lan, Yanwei Yu, Lei Cao 0004, Peng Song 0002
MDM4
2017 Phase-Sensitive Decision-Directed SNR Estimator for Single-Channel Speech Enhancement
abstract
The a priori signal-to-noise ratio (SNR) plays an essential role in many speech enhancement systems. Most of the existing approaches to estimate the a priori SNR only exploit the amplitude spectra while making the phase neglected. Considering the fact that incorporating phase information into a speech processing system can significantly improve the speech quality, this paper proposes a phase-sensitive decision-directed (DD) approach for the a priori SNR estimate. By representing the short-time discrete Fourier transform (STFT) signal spectra geometrically in a complex plane, the proposed approach estimates the a priori SNR using both the magnitude and phase information while making no assumptions about the phase difference between clean speech and noise spectra. Objective evaluations in terms of the spectrograms, segmental SNR, log-spectral distance (LSD) and short-time objective intelligibility (STOI) measures are presented to demonstrate the superiority of the proposed approach compared to several competitive methods at different noise conditions and input SNR levels.
Shifeng Ou, Peng Song 0002, Ying Gao 0007
Int. J. Pattern Recognit. Artif. Intell.2
2016 Speech emotion recognition using transfer non-negative matrix factorization
abstract
In practical situations, the emotional speech utterances are often collected from different devices and conditions, which will obviously affect the recognition performance. To address this issue, in this paper, a novel transfer non-negative matrix factorization (TNMF) method is presented for cross-corpus speech emotion recognition. First, the NMF algorithm is adopted to learn a latent common feature space for the source and target datasets. Then, the discrepancies between the feature distributions of different corpora are considered, and the maximum mean discrepancy (MMD) algorithm is used for the similarity measurement. Finally, the TNMF approach, which integrates the NMF and MMD algorithms, is proposed. Experiments are carried out on two popular datasets, and the results verify that the TNMF method can significantly outperform the automatic and competitive methods for cross-corpus speech emotion recognition.
Peng Song 0002, Shifeng Ou, Wenming Zheng, Yun Jin, Li Zhao 0003
ICASSP1
2016 Cross-corpus speech emotion recognition based on transfer non-negative matrix factorization
Peng Song 0002, Wenming Zheng, Shifeng Ou, Yun Jin, Jinglei Liu, Yanwei Yu
Speech Commun.1
2014 A feature selection and feature fusion combination method for speaker-independent speech emotion recognition
abstract
To enhance the recognition rate of speaker independent speech emotion recognition, a feature selection and feature fusion combination method based on multiple kernel learning is presented. Firstly, multiple kernel learning is used to obtain sparse feature subsets. The features selected at least n times are recombined into another subset named n-subset. The optimal n is determined by 10 cross-validation experiments. Secondly, feature fusion is made at the kernel level. Not only each kind of feature is associated with a kernel, but also the full feature set is associated with a kernel which is not considered in the previous studies. All of the kernels are added together to obtain a combination kernel. The final recognition rate for 7 kinds of emotions on Berlin Database is 83.10%, which outperforms state-of-the-art results and shows the effectiveness of our method. It is also proved that MFCCs play a crucial role in speech emotion recognition.
Yun Jin, Peng Song 0002, Wenming Zheng, Li Zhao 0003
ICASSP2
2014 Text-independent voice conversion using speaker model alignment method from non-parallel speech
Peng Song 0002, Yun Jin, Wenming Zheng, Li Zhao 0003
INTERSPEECH1
2013 Non-parallel training for voice conversion based on adaptation method
abstract
In this paper, we propose a simple and efficient non-parallel training scheme for voice conversion (VC). First, the speaker models are adapted from the background model using maximum a posteriori (MAP) technique. Then, by utilizing the parameters of adapted speaker models, the Gaussian normalization and mean transformation methods are proposed for VC, respectively. In addition, to improve the conversion performance of the proposed methods, a combination approach is further presented. Finally, objective and subjective experiments are carried out to evaluate the performance of the proposed scheme, the results demonstrate that our scheme can obtain comparable performance with the traditional GMM method based on parallel corpus.
Peng Song 0002, Wenming Zheng, Li Zhao 0003
ICASSP1