VLDB 2026 Research / reviewers in the wild / expert
Chongqing Chen
dblp:260/1419
· DBLP profile ↗
15ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-0521-0184ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Geometry-guided explicit dual-stream alignment network for visual question answering
Chongqing Chen, Dezhi Han, Huafeng Wu, Kuanching Li |
Expert Syst. Appl. | 1 |
| 2026 | Multimodal context-aware consistency alignment for vision-language tasks
Xiang Shen 0002, Dezhi Han, Chin-Chen Chang 0001, Yangshuyi Xu, Chongqing Chen |
Expert Syst. Appl. | 5 |
| 2026 | Enhancing image-text matching through contextual fine-grained alignment
FanRong Meng, Dezhi Han, Xiang Shen 0002, Chongqing Chen |
Vis. Comput. | 4 |
| 2025 | Towards bias-aware visual question answering: Rectifying and mitigating comprehension biases
Chongqing Chen, Dezhi Han, Chin-Chen Chang 0001 |
Expert Syst. Appl. | 1 |
| 2025 | A triple-branch hybrid dynamic-static alignment strategy for vision-language tasks
Xiang Shen 0002, Chongqing Chen, Dezhi Han, Yangshuyi Xu, Xiuying Wang 0001, Huiyu Zhou 0001 |
Neural Networks | 2 |
| 2025 | SAFFNet: self-attention based on Fourier frequency domain filter network for visual question answering
Jingya Shi, Dezhi Han, Chongqing Chen, Xiang Shen 0002 |
Vis. Comput. | 3 |
| 2025 | Vman: visual-modified attention network for multimodal paradigms
Dezhi Han, Chongqing Chen, Xiang Shen 0002, Huafeng Wu |
Vis. Comput. | 3 |
| 2025 | Enhancing image-text matching through multi-level semantic consistency alignment
Liqi Zhu, Dezhi Han, Xiang Shen 0002, Chongqing Chen, Kuanching Li |
Vis. Comput. | 4 |
| 2024 | FastPFM: a multi-scale ship detection algorithm for complex scenes based on SAR imagesabstractSynthetic Aperture Radar (SAR) is renowned for its all-weather capabilities, exceptional penetration, and high-resolution imaging, making SAR-based ship detection crucial for maritime surveillance and sea rescue operations. However, various challenges, such as blurred ship contours, complex backgrounds, and uneven scale distribution, can impede detection performance improvement. In this study, we propose FastPFM, a novel ship detection model developed to address these challenges. Firstly, we utilize FasterNet as the backbone network to reduce computational redundancy, enhancing feature extraction efficiency and overall computational performance. Additionally, we employ the Feature Bi-level Routing Transformation model (FBM) to obtain global feature information and enhance focus on target regions. Secondly, the PFM module is engineered to collect multi-scale target information effectively by establishing connections across stages, thereby improving fusion of target features. Thirdly, an extra target feature fusion layer is introduced to enhance small ship detection precision and accommodate multi-scale targets. Finally, comprehensive tests on SSDD and HRSID datasets validate FastPFM's efficacy. Compared to the baseline model YOLOX, FastPFM achieves a 5.5% and 4.4% improvement in detection accuracy, respectively. Furthermore, FastPFM demonstrates comparable or superior performance to other detection algorithms, achieving 92.1% and 83.1% accuracy on AP50, respectively. Dezhi Han, Chongqing Chen, Zhongdai Wu |
Connect. Sci. | 3 |
| 2024 | KTMN: Knowledge-driven Two-stage Modulation Network for visual question answeringabstractExisting visual question answering (VQA) methods introduce the Transformer as the backbone architecture for intra- and inter-modal interactions, demonstrating its effectiveness in dependency relationship modeling and information alignment. However, the Transformer’s inherent attention mechanisms tend to be affected by irrelevant information and do not utilize the positional information of objects in the image during the modelling process, which hampers its ability to adequately focus on key question words and crucial image regions during answer inference. Considering this issue is particularly pronounced on the visual side, this paper designs a Knowledge-driven Two-stage Modulation self-attention mechanism to optimize the internal interaction modeling of image sequences. In the first stage, we integrate textual context knowledge and the geometric knowledge of visual objects to modulate and optimize the query and key matrices. This effectively guides the model to focus on visual information relevant to the context and geometric knowledge during the information selection process. In the second stage, we design an information comprehensive representation to apply a secondary modulation to the interaction results from the first modulation. This further guides the model to fully consider the overall context of the image during inference, enhancing its global understanding of the image content. On this basis, we propose a Knowledge-driven Two-stage Modulation Network (KTMN) for VQA, which enables fine-grained filtering of redundant image information while more precisely focusing on key regions. Finally, extensive experiments conducted on the datasets VQA v2 and CLEVR yielded Overall accuracies of 71.36% and 99.20%, respectively, providing ample validation of the proposed method’s effectiveness and rationality. Source code is available at https://github.com/shijingya/KTMN . Jingya Shi, Dezhi Han, Chongqing Chen, Xiang Shen 0002 |
Multim. Syst. | 3 |
| 2024 | MPCCT: Multimodal vision-language learning paradigm with context-based compact Transformer
Chongqing Chen, Dezhi Han, Chin-Chen Chang 0001 |
Pattern Recognit. | 1 |
| 2023 | Local self-attention in transformer for visual question answering
Xiang Shen 0002, Dezhi Han, Chongqing Chen, Jie Hua 0001, GaoFeng Luo |
Appl. Intell. | 4 |
| 2023 | NAS-YOLOX: a SAR ship detection using neural architecture search and multi-scale attentionabstractDue to the advantages of all-weather capability and high resolution, synthetic aperture radar (SAR) image ship detection has been widely applied in the military, civilian, and other domains.However, SAR-based ship detection suffers from limitations such as strong scattering of targets, multiple scales, and background interference, leading to low detection accuracy.To address these limitations, this paper presents a novel SAR ship detection method, NAS-YOLOX, which leverages the efficient feature fusion of the neural architecture search feature pyramid network (NAS-FPN) and the effective feature extraction of the multi-scale attention mechanism.Specifically, NAS-FPN replaces the PAFPN in the baseline YOLOX, greatly enhances the fusion performance of the model's multi-scale feature information, and a dilated convolution feature enhancement module (DFEM) is designed and integrated into the backbone network to improve the network's receptive field and target information extraction capabilities.Furthermore, a multi-scale channel-spatial attention (MCSA) mechanism is conceptualised to enhance focus on target regions, improve small-scale target detection, and adapt to multi-scale targets.Additionally, extensive experiments conducted on benchmark datasets, HRSID and SSDD, demonstrate that NAS-YOLOX achieves comparable or superior performance compared to other state-ofthe-art ship detection models and reaches best accuracies of 91.1% and 97.2% on AP 0.5 , respectively. Dezhi Han, Mingming Cui, Chongqing Chen |
Connect. Sci. | 4 |
| 2023 | CLVIN: Complete language-vision interaction network for visual question answering
Chongqing Chen, Dezhi Han, Xiang Shen 0002 |
Knowl. Based Syst. | 1 |
| 2022 | CAAN: Context-Aware attention network for visual question answering
Chongqing Chen, Dezhi Han, Chin-Chen Chang 0001 |
Pattern Recognit. | 1 |