Yunxi Wang

dblp:314/7394 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0004-1432-8338ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Image recognition and object detection · 32% Graph learning · 32% Face, body and person analysis · 28%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning › graph neural network › attention-based graph neural network
graph attention network
1.012026
Topology-Aware Vision Transformers for Enhanced Scene Recognition · AAAI 2026
Computer vision › Image recognition and object detection
scene recognition
1.012026
Topology-Aware Vision Transformers for Enhanced Scene Recognition · AAAI 2026
Computer vision › Face, body and person analysis › facial expression analysis
facial expression recognition
0.912025
Fusion of Granular-Ball Visual Spatial Representations for Enhanced Facial Expression Recognition · IJCAI 2025
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.312026
Topology-Aware Vision Transformers for Enhanced Scene Recognition · AAAI 2026

Methods — techniques the papers use, named apart from their topics

topology attention guidance · 1.0multimodal fusion · 1.0graph attention mask · 1.0granular-ball computing · 0.9attention mechanism · 0.9
YearPublicationVenuePosition
2026 Topology-Aware Vision Transformers for Enhanced Scene Recognition
abstract
Scene recognition (SR) is a fundamental task in computer vision (CV). In recent years, Transformer-based methods have achieved remarkable success in scene recognition tasks. Most existing approaches primarily rely on visual features, while failing to effectively model the structural relationships within scenes, which are crucial for accurate scene recognition. To this end, we propose Topology Attention Network for Scene Recognition (TANSR), an innovative method that leverages topological relationships from graphs to guide scene recognition. Specifically, Graph Attention Mask Generation Network (GAMGN) generates topology-aware masks from graph representations constructed by Graph Generation Module (GGM) and integrates them with patch embeddings by Topology Attention Guidance (TAG), enabling the transformer's attention mechanism to incorporate topological information. Furthermore, we introduce an innovative attention-driven multimodal fusion strategy that integrates graph-derived topological cues with visual patch embeddings, substantially enhancing the transformer’s capability to capture topological information and improving performance in complex scene recognition tasks. We evaluate TANSR on the benchmarks MIT-67, Scene-15 and SUN397, where it achieves consistent state-of-the-art (SOTA) performance, including 98.58% accuracy on MIT-67.
Yunxi Wang, Shuaiyu Liu, Qiling Li, Yazhou Ren 0001, Xiaorong Pu
AAAI1
2025 Fusion of Granular-Ball Visual Spatial Representations for Enhanced Facial Expression Recognition
abstract
Facial Expression Recognition (FER) is a fundamental problem in computer vision. Despite recent advances, significant challenges remain. Current methods primarily focus on extracting visual representations while overlooking other valuable information. To address this limitation, we propose a novel method called Component Separation and Granular-ball Space Bootstrap Fusion (CS-GBSBF), which leverages granular balls to transform visual images to spatial graphs, thereby enlarging the spatial information embedded in images. Our method separates the face into different components and utilizes the spatial information to bootstrap the fusion. More specifically, CS-GBSBF mainly consists of three crucial networks: Represent Extraction Network (REN), Represent Separation Network (RSN) and Represent Fusion Network (RFN). First, granular balls are used to represent expression images as graphs, which are fed into REN along with images. Then, RSN separates basic visual/spatial representations extracted from REN into a set of component visual/spatial representations. Next, RFN utilizes spatial representations to bootstrap component visual integration. A significant challenge in two-stream models is feature alignment, for which we have developed Attention Guidance Module (AGM) and Bootstrap Alignment Loss (L_BA) in REN and RFN, respectively. Results of experiment on eight databases show that CS-GBSBF consistently achieves higher recognition accuracy than several state-of-the-art methods. The code is available at https://github.com/Lsy235/CS-GBSBF.
Shuaiyu Liu, Qiyao Shen, Yunxi Wang, Yazhou Ren 0001, Guoyin Wang 0001
IJCAI3