VLDB 2026 Research / reviewers in the wild / expert
Ercheng Pei
dblp:160/0008
· DBLP profile ↗
14ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0003-3582-6809ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mambaer: Mamba with knowledge-learning hierarchical attention for facial expression recognition
Ercheng Pei, Xiaofeng Wei, Zhanxuan Hu, Hailong Ning, Xiaochun An, Xiaoge Li |
Multim. Syst. | 1 |
| 2026 | LLM-driven fine-grained emotion parsing and parameterized mapping for conversational TTS
Xiaochun An, Xiaoge Li, Ercheng Pei, Qingli Yan |
Pattern Recognit. | 4 |
| 2025 | A Hidden Stumbling Block in Generalized Category Discovery: Distracted AttentionabstractGeneralized Category Discovery (GCD) aims to classify unlabeled data from both known and unknown categories by leveraging knowledge from labeled known categories. While existing methods have made notable progress, they often overlook a hidden stumbling block in GCD: distracted attention. Specifically, when processing unlabeled data, models tend to focus not only on key objects in the image but also on task-irrelevant background regions, leading to suboptimal feature extraction. To remove this stumbling block, we propose Attention Focusing (AF), an adaptive mechanism designed to sharpen the model's focus by pruning non-informative tokens. AF consists of two simple yet effective components: Token Importance Measurement (TIME) and Token Adaptive Pruning (TAP), working in a cascade. TIME quantifies token importance across multiple scales, while TAP prunes non-informative tokens by utilizing the multi-scale importance scores provided by TIME. AF is a lightweight, plug-and-play module that integrates seamlessly into existing GCD methods with minimal computational overhead. When incorporated into one prominent GCD method, SimGCD, AF achieves up to 15.4% performance improvement over the baseline with minimal computational overhead. The implementation code is provided in https://github.com/Afleve/AFGCD. Qiyu Xu, Zhanxuan Hu, Yu Duan 0001, Ercheng Pei, Yonghang Tai |
ICCV | 4 |
| 2025 | Classifier ensemble based source-free domain adaptation for time series classification
Ercheng Pei, Wangdong Zhao, Zhanxuan Hu, Hailong Ning |
Knowl. Based Syst. | 1 |
| 2024 | An ensemble learning-enhanced multitask learning method for continuous affect recognition from facial imagesabstractContinuous affect recognition from facial images aims to estimate the values of multiple affective dimensions from a facial image sequence. To leverage relevant information between multiple affective dimensions, multitask learning has been used in the estimation of continuous affective states . Most of the existing multitask continuous affect recognition methods focus on designing elaborate multitask networks. Meanwhile, a few research works consider using multitask training strategies for continuous affect recognition. In general, existing multitask continuous affect recognition methods face the problem of unstable training effects. In this work, to improve the stability of multitask learning , we propose an ensemble learning-enhanced multitask network architecture for continuous affect recognition. In addition, we introduce a novel adaptive weighted loss-based multitask learning strategy to effectively train the proposed multitask continuous affect recognition model. Experimental results, on the RECOLA, SEMAINE and AFEW-VA datasets for continuous affect recognition, demonstrate the potential of the proposed method compared to state-of-the-art methods. Ercheng Pei, Zhanxuan Hu, Hailong Ning, Abel Díaz Berenguer |
Expert Syst. Appl. | 1 |
| 2023 | Relate Auditory Speech To Eeg By Shallow-Deep Attention-Based NetworkabstractElectroencephalography (EEG) plays a vital role in detecting how brain responses to different stimulus. In this paper, we propose a novel Shallow-Deep Attention-based Network (SDANet) to classify the correct auditory stimulus evoking the EEG signal. It adopts the Attention-based Correlation Module (ACM) to discover the connection between auditory speech and EEG from global aspect, and the Shallow-Deep Similarity Classification Module (SDSCM) to decide the classification result via the embeddings learned from the shallow and deep layers. Moreover, various training strategies and data augmentation are used to boost the model robustness. Experiments are conducted on the dataset provided by Auditory EEG challenge (ICASSP Signal Processing Grand Challenge 2023). Results show that the proposed model has a significant gain over the baseline on the match-mismatch track. Fan Cui, Liyong Guo, Jiyao Liu, Ercheng Pei, Dongmei Jiang |
ICASSP | 5 |
| 2023 | A Bayesian Filtering Framework for Continuous Affect Recognition From Facial ImagesabstractContinuous affective state estimation from facial information is a task which requires the prediction of time series of emotional state outputs from a facial image sequence. Modeling the spatial-temporal evolution of facial information plays an important role in affective state estimation. One of the most widely used methods is Recurrent Neural Networks (RNN). RNNs provide an attractive framework for propagating information over a sequence using a continuous-valued hidden layer representation. In this work, we propose to instead learn rich affective state dynamics. We model human affect as a dynamical system and define the affective state in terms of valence, arousal and their higher-order derivatives. We then pose the affective state estimation problem as a jointly trained state estimator for high-dimensional input images, combining an RNN and a Bayesian Filter, i.e. Kalman filters (KF) and Extended Kalman filters (EKF), so that all weights in the resulting network can be trained using backpropagation. We use a recently proposed general framework for designing and learning discriminative state estimators framed as computational graphs. Such approach can handle high dimensional observations and efficiently optimize, in an end-to-end fashion, the state estimator. In addition, to deal with the asynchrony between emotion labels and input images, caused by the inherent reaction lag of the annotators, we introduce a convolutional layer that aligns features with emotion labels. Experimental results, on the RECOLA and SEMAINE datasets for continuous emotion prediction, illustrate the potential of the proposed framework compared to recent state-of-the-art models. Ercheng Pei, Meshia Cédric Oveneke, Dongmei Jiang, Hichem Sahli |
IEEE Trans. Multim. | 1 |
| 2022 | Audio-visual collaborative representation learning for Dynamic Saliency Prediction
Hailong Ning, Bin Zhao 0001, Zhanxuan Hu, Ercheng Pei |
Knowl. Based Syst. | 5 |
| 2022 | Leveraging the Deep Learning Paradigm for Continuous Affect Estimation from Facial ExpressionsabstractContinuous affect estimation from facial expressions has attracted increased attention in the affective computing research community. This paper presents a principled framework for estimating continuous affect from video sequences. Based on recent developments, we address the problem of continuous affect estimation by leveraging the Bayesian filtering paradigm, i.e., considering affect as a latent dynamical system corresponding to a general feeling of pleasure with a degree of arousal, and recursively estimating its state using a sequence of visual observations. To this end, we advance the state-of-the-art as follows: (i) Canonical face representation (CFR): a novel algorithm for two-dimensional face frontalization, (ii) Convex unsupervised representation learning (CURL): a novel frequency-domain convex optimization algorithm for unsupervised training of deep convolutional neural networks (CNN)s, and (iii) Deep extended Kalman filtering (DEKF): an extended Kalman filtering-based algorithm for affect estimation from a sequence of CNN observations. The performance of the resulting CFR-CURL-DEKF algorithmic framework is empirically evaluated on publicly available benchmark datasets for facial expression recognition (CK+) and continuous affect estimation (AVEC 2012 and 2014). Meshia Cédric Oveneke, Ercheng Pei, Abel Díaz Berenguer, Dongmei Jiang, Hichem Sahli |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | Action Unit Driven Facial Expression Synthesis from a Single Image with Patch Attentive GANabstractAbstract Recent advances in generative adversarial networks (GANs) have shown tremendous success for facial expression generation tasks. However, generating vivid and expressive facial expressions at Action Units (AUs) level is still challenging, due to the fact that automatic facial expression analysis for AU intensity itself is an unsolved difficult task. In this paper, we propose a novel synthesis‐by‐analysis approach by leveraging the power of GAN framework and state‐of‐the‐art AU detection model to achieve better results for AU‐driven facial expression generation. Specifically, we design a novel discriminator architecture by modifying the patch‐attentive AU detection network for AU intensity estimation and combine it with a global image encoder for adversarial learning to force the generator to produce more expressive and realistic facial images. We also introduce a balanced sampling approach to alleviate the imbalanced learning problem for AU synthesis. Extensive experimental results on DISFA and DISFA+ show that our approach outperforms the state‐of‐the‐art in terms of photo‐realism and expressiveness of the facial expression quantitatively and qualitatively. Le Yang 0009, Ercheng Pei, Meshia Cédric Oveneke, Mitchel Alioscha-Pérez, Dongmei Jiang, Hichem Sahli |
Comput. Graph. Forum | 3 |
| 2021 | Monocular 3D Facial Expression Features for Continuous Affect RecognitionabstractAutomated facial expression analysis from image sequences for continuous emotion recognition is a very challenging task due to the loss of the three-dimensional information during the image formation process. State-of-the-art relied on estimating dynamic textures features and convolutional neural network features to derive spatio-temporal features. Despite their great success, such features are insensitive to micro facial muscle deformations and are affected by identity, face pose, illumination variation, and self-occlusion. In this work, we argue that retrieving, from image sequences, 3D facial spatio-temporal information, which describes the natural facial muscle deformation, provides a semantical and efficient way of representation and is useful for emotion recognition. In this paper, we propose a framework for extracting three-dimensional facial spatio-temporal features from monocular image sequences using an extended 3D Morphable Model (3DMM) which disentangles the identity factor from the facial expressions of a specific person. An LSTM model is used to evaluate the effectiveness of the proposed spatio-temporal features on video-based facial expression recognition task and continuous affect recognition task. Experimental results, on the AFEW6.0 datasets for facial expression recognition, and the RECOLA and SEMAINE datasets for continuous emotion prediction, illustrate the potential of the proposed 3D spatio-temporal features for facial expressions analysis and continuous affect recognition, as well as their efficiency compared to recent state-of-the-art features. Ercheng Pei, Meshia Cédric Oveneke, Dongmei Jiang, Hichem Sahli |
IEEE Trans. Multim. | 1 |
| 2020 | An efficient model-level fusion approach for continuous affect recognition from audiovisual signals
Ercheng Pei, Dongmei Jiang, Hichem Sahli |
Neurocomputing | 1 |
| 2019 | Continuous affect recognition with weakly supervised learning
Ercheng Pei, Dongmei Jiang, Mitchel Alioscha-Pérez, Hichem Sahli |
Multim. Tools Appl. | 1 |
| 2015 | Multimodal dimensional affect recognition using deep bidirectional long short-term memory recurrent neural networksabstractIn this paper we propose the deep bidirectional long short-term memory recurrent neural network (DBLSTM-RNN) based single modal and multi-modal affect recognition frameworks. In the single modal framework DBLSTM with moving average (MA), audio or visual features are input into the DBLSTM-RNN model, whose output estimations of a dimension are smoothed by the moving average filter. After the smoothed estimations are expanded to the frame rate of the ground truth labels, another MA is adopted for smoothing the final results. In the multi-modal framework DBLSTM-DBLSTM-MA, the initial estimations from the audio and visual modalities via the first layer of DBLSTM-RNNs are input into a second layer of DBLSTM-RNN, whose outputs are smoothed by MA. The smoothed estimations are then expanded to the frame rate of the ground truth labels and smoothed again by another MA. Affect recognition experiments are carried out on the training set and development set of the AVEC2014 database, results show that the proposed DBLSTM-MA framework outperforms linear regression, support vector regression (SVR), and BLSTM for single modal dimension estimation. For audio visual multi-modal affect recognition, DBLSTM-DBLSTM-MA obtains better or comparable performance than the state of the art results in the competition of AVEC2014, with the average correlation coefficient (COR) reaches 0.599 on the Freeform database, 0.630 on the Northwind database, and 0.615 on the Freeform-Northwind database. Ercheng Pei, Le Yang 0009, Dongmei Jiang, Hichem Sahli |
ACII | 1 |