Heping Song

dblp:06/3008 · DBLP profile ↗
← Back
4ranked-venue papers in the field
0as first author
4since 2021 · last 2023
0000-0002-8583-2804ORCID · corroborated

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 3Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2023 Large-scale non-negative subspace clustering based on Nyström approximation
Hongjie Jia, Qize Ren, Longxia Huang, Qirong Mao, Liangjun Wang, Heping Song
Inf. Sci.6
2022 Extended variational inference for Dirichlet process mixture of Beta-Liouville distributions for proportional data modeling
abstract
Bayesian estimation of parameters in the Dirichlet mixture process of the Beta-Liouville distribution (i.e., the infinite Beta-Liouville mixture model) has recently gained considerable attention due to its modeling capability for proportional data. However, applying the conventional variational inference (VI) framework cannot derive an analytically tractable solution since the variational objective function cannot be explicitly calculated. In this paper, we adopt the recently proposed extended VI framework to derive the closed-form solution by further lower bounding the original variational objective function in the VI framework. This method is capable of simultaneously determining the model's complexity and estimating the model's parameters. Moreover, due to the nature of Bayesian nonparametric approaches, it can also avoid the problems of underfitting and overfitting. Extensive experiments were conducted on both synthetic and real data, generated from two real-world challenging applications, namely, object detection and text categorization, and its superior performance and effectiveness of the proposed method have been demonstrated.
Yuping Lai, Wenbo Guan, Lijuan Luo, Qiang Ruan, Yuan Ping 0003, Heping Song, Hongying Meng
Int. J. Intell. Syst.6
2021 Cross lingual speech emotion recognition via triple attentive asymmetric convolutional neural network
abstract
The application of cross-corpus for speech emotion recognition (SER) via domain adaptation methods have gain high acknowledgment for developing good robust emotion recognition systems using different corpora or datasets. However, the issue of cross-lingual still remains a challenge in SER and needs more attention to resolve the scenario of applying different language types in both training and testing. In this paper, we propose a triple attentive asymmetric convolutional neural network to address the recognition of emotions for cross-lingual and cross-corpus speech in an unsupervised approach. The proposed method adopts the joint supervision of softmax loss and center loss to learn high power discriminative feature representations for target domain via the use of high quality pseudo-labels. The proposed model uses three attentive convolutional neural networks asymmetrically, where two of the networks are used to artificially label unlabeled target samples as a result of their predictions from training on source labeled samples and the other network is used to obtain salient target discriminative features from the pseudo-labeled target samples. We evaluate our proposed method on three different language types (i.e., English, German, and Italian) data sets. The experimental results indicate that, our proposed method achieves higher prediction accuracy over other state-of-the-art methods.
Ocquaye Elias Nii Noi, Qirong Mao, Yanfei Xue, Heping Song
Int. J. Intell. Syst.4
2021 Learning to disentangle emotion factors for facial expression recognition in the wild
abstract
Facial expression recognition (FER) in the wild is a very challenging problem due to different expressions under complex scenario (e.g., large head pose, illumination variation, occlusions, etc.), leading to suboptimal FER performance. Accuracy in FER heavily relies on discovering superior discriminative, emotion-related features. In this paper, we propose an end-to-end module to disentangle latent emotion discriminative factors from the complex factors variables for FER to obtain salient emotion features. The training of proposed method contains two stages. First of all, emotion samples are used to obtain the latent representation using a variational auto-encoder with reconstruction penalization. Furthermore, the latent representation as the input is thrown into a disentangling layer to learn a set of discriminative emotion factors through the attention mechanism (e.g., a Squeeze-and-Excitation block) that encourages to separate emotion-related factors and nonaffective factors. Experimental results on public benchmark databases (RAF-DB and FER2013) show that our approach has remarkable performance in complex scenes than current state-of-the-art methods.
Qing Zhu 0002, Lijian Gao, Heping Song, Qirong Mao
Int. J. Intell. Syst.3