Xiang Shen 0002

dblp:71/8452-2 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-6894-2550ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021
YearPublicationVenuePosition
2026 Multi-Granularity Modal Interaction and Fusion framework for vision-language tasks
Yangshuyi Xu, Guangzhong Liu, Xiang Shen 0002, Xiuying Wang 0001, Huiyu Zhou 0001
Eng. Appl. Artif. Intell.3
2026 Multimodal context-aware consistency alignment for vision-language tasks
Xiang Shen 0002, Dezhi Han, Chin-Chen Chang 0001, Yangshuyi Xu, Chongqing Chen
Expert Syst. Appl.1
2026 Enhancing image-text matching through contextual fine-grained alignment
FanRong Meng, Dezhi Han, Xiang Shen 0002, Chongqing Chen
Vis. Comput.3
2025 A triple-branch hybrid dynamic-static alignment strategy for vision-language tasks
Xiang Shen 0002, Chongqing Chen, Dezhi Han, Yangshuyi Xu, Xiuying Wang 0001, Huiyu Zhou 0001
Neural Networks1
2025 SAFFNet: self-attention based on Fourier frequency domain filter network for visual question answering
Jingya Shi, Dezhi Han, Chongqing Chen, Xiang Shen 0002
Vis. Comput.4
2025 Vman: visual-modified attention network for multimodal paradigms
Dezhi Han, Chongqing Chen, Xiang Shen 0002, Huafeng Wu
Vis. Comput.4
2025 Enhanced small-target detection in SAR images via SIE-YOLO11: a deep learning approach
Jihang Wang, Dezhi Han, Xiang Shen 0002, Bing Han 0009, Zhongdai Wu
Vis. Comput.3
2025 Enhancing image-text matching through multi-level semantic consistency alignment
Liqi Zhu, Dezhi Han, Xiang Shen 0002, Chongqing Chen, Kuanching Li
Vis. Comput.3
2024 Relational reasoning and adaptive fusion for visual question answering
Xiang Shen 0002, Dezhi Han, Liang Zong, Jie Hua 0001
Appl. Intell.1
2024 KTMN: Knowledge-driven Two-stage Modulation Network for visual question answering
abstract
Existing visual question answering (VQA) methods introduce the Transformer as the backbone architecture for intra- and inter-modal interactions, demonstrating its effectiveness in dependency relationship modeling and information alignment. However, the Transformer’s inherent attention mechanisms tend to be affected by irrelevant information and do not utilize the positional information of objects in the image during the modelling process, which hampers its ability to adequately focus on key question words and crucial image regions during answer inference. Considering this issue is particularly pronounced on the visual side, this paper designs a Knowledge-driven Two-stage Modulation self-attention mechanism to optimize the internal interaction modeling of image sequences. In the first stage, we integrate textual context knowledge and the geometric knowledge of visual objects to modulate and optimize the query and key matrices. This effectively guides the model to focus on visual information relevant to the context and geometric knowledge during the information selection process. In the second stage, we design an information comprehensive representation to apply a secondary modulation to the interaction results from the first modulation. This further guides the model to fully consider the overall context of the image during inference, enhancing its global understanding of the image content. On this basis, we propose a Knowledge-driven Two-stage Modulation Network (KTMN) for VQA, which enables fine-grained filtering of redundant image information while more precisely focusing on key regions. Finally, extensive experiments conducted on the datasets VQA v2 and CLEVR yielded Overall accuracies of 71.36% and 99.20%, respectively, providing ample validation of the proposed method’s effectiveness and rationality. Source code is available at https://github.com/shijingya/KTMN .
Jingya Shi, Dezhi Han, Chongqing Chen, Xiang Shen 0002
Multim. Syst.4
2023 Local self-attention in transformer for visual question answering
Xiang Shen 0002, Dezhi Han, Chongqing Chen, Jie Hua 0001, GaoFeng Luo
Appl. Intell.1
2023 CLVIN: Complete language-vision interaction network for visual question answering
Chongqing Chen, Dezhi Han, Xiang Shen 0002
Knowl. Based Syst.3