Shaoqi Yan

dblp:353/9273 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Hi-EF: Benchmarking Emotion Forecasting in Human-interaction
abstract
Affective Forecasting is an psychology task that involves predicting an individual's future emotional responses, often hampered by reliance on external factors leading to inaccuracies, and typically remains at a qualitative analysis stage. To address these challenges, we narrows the scope of Affective Forecasting by introducing the concept of Human-interaction-based Emotion Forecasting (EF). This task is set within the context of a two-party interaction, positing that an individual's emotions are significantly influenced by their interaction partner's emotional expressions and informational cues. This dynamic provides a structured perspective for exploring the patterns of emotional change, thereby enhancing the feasibility of emotion forecasting.
Haoran Wang 0006, Xinji Mai, Zeng Tao, Junxiong Lin, Xuan Tong, Ivy Pan, Shaoqi Yan, Yan Wang 0068, Shuyong Gao
AAAI7
2025 OUS: Bridging Scene Context and Facial Features to Overcome the Rigid Cognitive Problem
abstract
Dynamic Facial Expression Recognition (DFER) is crucial for affective computing but often overlooks the impact of scene context. We have identified a significant issue in current DFER tasks: human annotators typically integrate emotions from various angles, including environmental cues and body language, whereas existing DFER methods tend to consider the scene as noise that needs to be filtered out, focusing solely on facial information. We refer to this as the Rigid Cognitive Problem. The Rigid Cognitive Problem can lead to discrepancies between the cognition of annotators and models in some samples. To align more closely with the human cognitive paradigm of emotions, we propose an Overall Understanding of the Scene DFER method (OUS). OUS effectively integrates scene and facial features, combining scene-specific emotional knowledge for DFER. Extensive experiments on the two largest datasets in the DFER field, DFEW and FERV39k, demonstrate that OUS significantly outperforms existing methods. By analyzing the Rigid Cognitive Problem, OUS successfully understands the complex relationship between scene context and emotional expression, closely aligning with human emotional understanding in real-world scenarios.
Xinji Mai, Haoran Wang 0006, Zeng Tao, Junxiong Lin, Shaoqi Yan, Yan Wang 0068, Jiawen Yu, Xuan Tong
AAAI5
2025 D2SP: Dynamic Dual-Stage Purification Framework for Dual Noise Mitigation in Vision-based Affective Recognition
abstract
The current advancements in Dynamic Facial Expression Recognition (DFER) methods mainly focus on better capturing the spatial and temporal features of facial expressions. However, DFER datasets contain a substantial amount of noisy samples, and few have addressed the issue of handling this noise. We identified two types of noise: one is caused by low-quality data resulting from factors such as occlusion, dim lighting, and blurriness; the other arises from mislabeled data due to annotation bias by annotators. Addressing the two types of noise, we have meticulously crafted a Dynamic Dual-Stage Purification (D2SP) Framework. This initiative aims to dynamically purify the DFER datasets of these two types of noise, ensuring that only high-quality and correctly labeled data is used in the training process. To mitigate low-quality samples, we introduce the Coarse-Grained Pruning (CGP) stage, which computes sample weights and prunes those low-weight samples. After CGP, the Fine-Grained Correction (FGC) stage evaluates prediction stability to correct mislabeled data. Moreover, D2SP is conceived as a general, plug-and-play framework, tailored to integrate seamlessly with prevailing DFER methods. Extensive experiments covering prevalent DFER datasets and deploying multiple benchmark methods have substantiated D2SP’s ability to enhance performance metrics.
Haoran Wang 0006, Xinji Mai, Zeng Tao, Xuan Tong, Junxiong Lin, Yan Wang 0068, Jiawen Yu, Shaoqi Yan, Ziheng Zhou 0005
CVPR8
2025 Observe finer to select better: Learning key frame extraction via semantic coherence for dynamic facial expression recognition in the wild
Shaoqi Yan, Yan Wang 0068, Xinji Mai, Zeng Tao, Wei Song 0007, Qing Zhao 0007, Boyang Wang 0003, Haoran Wang 0006, Shuyong Gao
Inf. Sci.1
2024 Adaptive Multi-modal Fusion of Spatially Variant Kernel Refinement with Diffusion Model for Blind Image Super-Resolution
Junxiong Lin, Yan Wang 0068, Zeng Tao, Boyang Wang 0003, Qing Zhao 0007, Haorang Wang, Xuan Tong, Xinji Mai, Yuxuan Lin 0001, Wei Song 0007, Jiawen Yu, Shaoqi Yan
ECCV (52)12
2024 Suppressing Uncertainties in Degradation Estimation for Blind Super-Resolution
Junxiong Lin, Zen Tao, Xuan Tong, Xinji Mai, Haoran Wang 0006, Boyang Wang 0003, Yan Wang 0068, Qing Zhao 0007, Jiawen Yu, Yuxuan Lin 0001, Shaoqi Yan, Shuyong Gao
ACM Multimedia11
2024 All rivers run into the sea: Unified Modality Brain-Inspired Emotional Central Mechanism
abstract
In the field of affective computing, fully leveraging information from a variety of sensory modalities is essential for the comprehensive understanding and processing of human emotions. Inspired by the process through which the human brain handles emotions and the theory of cross-modal plasticity, we propose UMBEnet, a brain-like unified modal affective processing network. The primary design of UMBEnet includes a Dual-Stream (DS) structure that fuses inherent prompts with a Prompt Pool and a Sparse Feature Fusion (SFF) module. The design of the Prompt Pool is aimed at integrating information from different modalities, while inherent prompts are intended to enhance the system's predictive guidance capabilities and effectively manage knowledge related to emotion classification. Moreover, considering the sparsity of effective information across different modalities, the SSF module aims to make full use of all available sensory data through the sparse integration of modality fusion prompts and inherent prompts, maintaining high adaptability and sensitivity to complex emotional states. Extensive experiments on the largest benchmark datasets in the Dynamic Facial Expression Recognition (DFER) field, including DFEW, FERV39k, and MAFW, have proven that UMBEnet consistently outperforms the current state-of-the-art methods. Notably, in scenarios of Modality Missingness and multimodal contexts, UMBEnet significantly surpasses the leading current methods, demonstrating outstanding performance and adaptability in tasks that involve complex emotional understanding with rich multimodal information. Code can be obtained at https://github.com/Xinji-Mai/UMBEnet.
Xinji Mai, Junxiong Lin, Haoran Wang 0006, Zeng Tao, Yan Wang 0068, Shaoqi Yan, Xuan Tong, Jiawen Yu, Boyang Wang 0003, Ziheng Zhou 0005, Qing Zhao 0007, Shuyong Gao
ACM Multimedia6
2024 Empower smart cities with sampling-wise dynamic facial expression recognition via frame-sequence contrastive learning
Shaoqi Yan, Yan Wang 0068, Xinji Mai, Qing Zhao 0007, Wei Song 0007, Zeng Tao, Haoran Wang 0006, Shuyong Gao
Comput. Commun.1
2024 Mixed noise-guided mutual constraint framework for unsupervised anomaly detection in smart industries
Qing Zhao 0007, Yan Wang 0068, Yuxuan Lin 0001, Shaoqi Yan, Wei Song 0007, Boyang Wang 0003, Yang Chang, Lizhe Qi
Comput. Commun.4
2024 MGR3Net: Multigranularity Region Relation Representation Network for Facial Expression Recognition in Affective Robots
abstract
Automatic facial expression recognition (FER) based on face images is essential for affective robots, which are designed for interactive companions and intelligent healthcare. Although existing DL-based FERs have made significant progress, an accurate FER model in robots is challenging due to the subtle differences in facial expressions across various scenarios. To address this issue, we propose a multigranularity region relation representation network (MGR3Net) to improve the robustness and generalization of FER via attention-guided global-local fusion. The MGR3Net is composed of three modules: multigranularity attention (MGA), holistic-regional feature extractor (HRFE), and hybrid feature fusion. In the MGA module, we first process each holistic cropped face image into three granularity of face regions from coarse to fine, which are four region-cropped faces,$2^{2}$face partitions, and$4^{2}$face partitions. Then, we propose the region attention relation cell to model the relationship between each region and the aggregated representation while preserving the spatial information of the local features. In the HRFE module, we align multigranularity features from the coarse space to the finer space and extract one holistic embedding and multiple region embeddings for each granularity. Finally, we use a hybrid-level fusion strategy to combine global-local features from the three granularities for final classification. Extensive experiments demonstrate that the MGR3Net outperforms the state-of-the-art methods evaluated on the in-the-lab datasets, in-the-wild datasets, and occlusion/pose-based sets.
Yan Wang 0068, Shaoqi Yan, Wei Song 0007, Antonio Liotta, Jing Liu 0050, Dingkang Yang, Shuyong Gao
IEEE Trans. Ind. Informatics2
2024 MSC-AD: A Multiscene Unsupervised Anomaly Detection Dataset for Small Defect Detection of Casting Surface
abstract
Intelligent detection of product surface defects in the industrial scene is the key to ensuring product quality. On general benchmarks, current unsupervised anomaly detection techniques have achieved significant success. When used in complex industrial environments (e.g., large industrial components with small defects), the model needs to be able to adapt to different imaging scenarios (e.g., illumination and resolution) and accurately detect and localize anomalies, but its performance is still far from satisfactory. Besides, the complex and unstable optical lighting environment for collecting such data poses major challenges in establishing unified benchmarks for optical lighting and imaging resolution in defect detection. To fill this gap, we build a standard imaging system-based multiscene unsupervised anomaly detection dataset, coined as MSC-AD. In particular, it provides 12 imaging scenes, i.e., a cross combination of low-to-high three illuminations and 150 × 150 to 600 × 600 four resolutions, in which six types of large casting surfaces with different structures include five kinds of small defects with sample-level and pixel-level precise ground truth. We systematically investigate representative baseline methods and empirical analysis on this dataset to obtain a number of interesting findings, e.g., how to detach from distinctly different imaging scenes, and how to distinguish between subtly normal–anomaly classes. To the best of our knowledge, MSC-AD is the first multi-illumination, multiresolution, multisurface, and multidefect dataset built in a standard imaging system.
Qing Zhao 0007, Yan Wang 0068, Boyang Wang 0003, Junxiong Lin, Shaoqi Yan, Wei Song 0007, Antonio Liotta, Jiawen Yu, Shuyong Gao
IEEE Trans. Ind. Informatics5
2023 IntrNet: Weakly Supervised Segmentation of Thyroid Nodules Based on Intra-image and Inter-image Semantic Information
Jie Gao 0008, Shaoqi Yan, Xuzhou Fu, Zhiqiang Liu 0002, Mei Yu 0004
ICIC (2)2
2023 Freq-HD: An Interpretable Frequency-based High-Dynamics Affective Clip Selection Method for in-the-Wild Facial Expression Recognition in Videos
abstract
The in-the-wild dynamic facial expression recognition (DFER) has been challenging due to several high-dynamics factors such as limited dynamic expression-related frames and variable non-expression noise in facial expression sequences. To provide more expression-related clips for DFER models, we propose a novel and interpretable frequency-based method (Freq-HD) for high-dynamics affective clip selection. It can select clips containing pure expression changes from sequences and aid different DFER network structures in recognizing in-the-wild dynamic facial expressions more accurately and efficiently. We first design a novel spatial-temporal frequency analysis (STFA) module to compute the dynamics values of each clip by using sliding windows and spatial-temporal frequency analysis. Moreover, we propose a multi-band complementary selection (MBC) module to amend the inappropriate reaction of the dynamics values of different spatial frequency bands in STFA when expression-irrelevant noise occurs. Specifically, the MBC uses an ingenious mapping method to generate the inhibitory factors to complement and separate the dynamics of expressions and non-expressions in different frequency bands. The Freq-HD can select the most expression-correlated clips and the consisting frames, which could be incorporated into any existing DFER models. We extensively evaluate the Freq-HD on two in-the-wild datasets and four DFER baselines, showing that our method significantly improves the subsequent network performance while using fewer input frames and reducing computation cost. More ablation studies and visualization analysis provide further empirical evidence of the effectiveness of our method.
Zeng Tao, Yan Wang 0068, Zhaoyu Chen 0001, Boyang Wang 0003, Shaoqi Yan, Kaixun Jiang, Shuyong Gao
ACM Multimedia5
2023 A Capture to Registration Framework for Realistic Image Super-Resolution in the Industry Environment
abstract
The acquisition and processing of visual data in industrial environments are of paramount importance. High-resolution (HR) images offer superior clarity and richer textural detail compared to low-resolution (LR) images. On the one hand, owing to the incorporation of richer information, HR images demonstrate substantially enhanced performance compared to LR images in downstream applications, such as anomaly detection. On the other hand, they provide valuable insights to designers and quality inspectors who require a detailed understanding of the images. Currently, the majority of research on super-resolution focuses on natural scenes such as cities and fields, however, the development of datasets for industrial scenes is still in its infancy. To address the image distortion in building realistic LR-HR image pairs in the industry environment, we design a capture to registration framework. It consists of the standard imaging system, physical calibration of the imaging system, as well as the rigid to elastic registration of the LR-HR image pairs. Thus, we build the first realistic industrial sence super-resolution dataset (IndSR), comprises of 50 sets of calibrated images with three scale factors and five typical defects. To benchmark IndSR, we employ quantitative, qualitative, and task-oriented studies to evaluate the representative super-resolution and anomaly detection methods. Besides, we systematically investigate and discuss the performances and results of the existing SISR methods to advance research in the field of super-resolution in industry environment. The IndSR dataset can be available from https://byw4ng.github.io/IndSR/.
Boyang Wang 0003, Yan Wang 0068, Qing Zhao 0007, Junxiong Lin, Zeng Tao, Pinxue Guo, Zhaoyu Chen 0001, Kaixun Jiang, Shaoqi Yan, Shuyong Gao
ACM Multimedia9