Chengyang Fang

dblp:317/1115 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-3583-9119ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 IC-Bench: Benchmarking robustness of large multimodal models to common corruptions on image captioning
Xuelin Liu, Xinpeng Fang, Jiebin Yan, Chengyang Fang, Yuming Fang 0001
Pattern Recognit.4
2026 DSRAS: Dual-Stage Reasoning and Answer Selection for Video-Text Visual Question Answering
abstract
Video text-based visual question answering (Video TextVQA) aims to answer questions by spatio-temporal joint reasoning over textual and visual information in a video. Existing methods have achieved remarkable progress using uniform sampling strategies. However, uniform frame sampling may introduce noisy frames and miss keyframes. Meanwhile, current methods treat all video frames equally, which is suboptimal, especially since video text question answering tasks primarily target questions that include text within the video. To address the aforementioned issues, we introduce a novel Dual-Stage Reasoning and Answer Selection (DSRAS) model, which not only adaptively focuses on and extracts keyframes, but also significantly enhances attention to video frames containing text through an answer selection mechanism. Specifically, we propose a Dual-Stage Reasoning (DSR) module to achieve adaptive frame selection. Then, we introduce an Answer Selection Module (ASM) to guide our model to focus on keyframes containing textual information. Extensive experiments demonstrate that our model outperforms existing approaches on the RoadTextVQA and M4-ViteVQA datasets.
Chengyang Fang, Xiankun Wan, Wenhui Jiang 0001, Yuming Fang 0001
IEEE Signal Process. Lett.1
2025 Separate, Locate, and Align: Determine Context Relation of Scene Text From Multiple Perspectives in TextVQA
abstract
Text-based Visual Question Answering (TextVQA) focuses on answering questions about the scene text in images. Most works in this field uses transformer based models to modeling the interaction of question and scene texts which means the scene texts will be treated as a natural language sentence and concatenated in reading order as a part of input. However, they ignore the fact that different from words in natural language sentence which have inherent context relation, the context relation of scene texts in images need to be determined. To tackle this problem, we propose a novel method named Separate, Locate and Align (SLA) that discriminate the context relation of scene texts from semantic, visual and spatial aspects. Specifically, based on scene texts with similar visual information (e.g. background color, font color, font style, etc.) having semantic contextual relations, we propose a Text Semantic Separate (TSS) module to discriminate the semantic relation between different scene texts according to their visual contextual information. Then, we introduce a Spatial Circle Position (SCP) module that helps the model discriminate the spatial relation between different scene texts. Last, we design a Visual Alignment (VA) module to help the model distinguish the visual relationships between different scene texts according to the color distribution differences. Extensive experiments show that our method outperforms existing alternatives on TextVQA and ST-VQA datasets without pre-training tasks.
Chengyang Fang, Wenhui Jiang 0001, Yuming Fang 0001, Yuxin Peng 0001, Yang Liu 0293
IEEE Trans. Circuits Syst. Video Technol.1
2024 Segment then Match: Find the Carrier before Reasoning in Scene-Text VQA
abstract
Text-based Visual Question Answering (TextVQA) requires models to answer questions about the scene text in images by reasoning the context between the scene text and the question. Previous works demonstrated that clustering the scene text could help the model understand the context between different scene texts in the image. However, these methods cluster scene text solely based on spatial information, resulting in close scene text being grouped together even in the absence of semantic contextual relationships. In order to solve the above problem, we proposed a Segment then Match method. Specifically, we propose an OCR-carrier Segmentation and Matching module to segment texts and carriers in the scene image and help all OCR texts find the carrier that belongs to them. Then, we propose a Hierarchical Visual Feature Fusion module to facilitate semantic relevance judgment of OCR text from multiple visual perspectives, thereby aiding the answer reasoning process. Our proposed method outperforms state-of-the-art methods by 3.65% and 3.31% on TextVQA and ST-VQA datasets, respectively. Extensive experiments validate the effectiveness of our method.
Chengyang Fang, Jiapeng Liu 0006, Dayong Hu, Can Ma
ICASSP1
2024 Prompting Large Language Models with Fine-Grained Visual Relations from Scene Graph for Visual Question Answering
abstract
Visual Question Answering (VQA) is a task that requires models to comprehend both questions and images. An increasing number of works are leveraging the strong reasoning capabilities of Large Language Models (LLMs) to address VQA. These methods typically utilize image captions as visual text description to aid LLMs in comprehending images. However, these captions often overlooking the relations of fine-grained objects, which will limit the reasoning capability of LLMs. In this paper, we present PFVR, a modular framework that Prompts LLMs with Fine-grained Visual Relationships for VQA. PFVR primarily consists of an answer-guided generation module (AGG) and a question-guided filtering module (QGF). The two modules can combine to extract the fine-grained visual relations from scene graph, which will finally serve as crucial context for LLMs to comprehend the image. Extensive experiments conducted on the popular VQA dataset, GQA, confirm PFVR achieves state-of-the-art results compared to other strong VQA competitors, demonstrating its exceptional effectiveness.
Jiapeng Liu 0006, Chengyang Fang, Dayong Hu, Can Ma
ICASSP2
2023 CATS: A Pragmatic Chinese Answer-to-Sequence Dataset with Large Scale and High Quality
abstract
Liang Li, Ruiying Geng, Chengyang Fang, Bing Li, Can Ma, Rongyu Cao, Binhua Li, Fei Huang, Yongbin Li. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Liang Li 0006, Ruiying Geng, Chengyang Fang, Bing Li 0001, Can Ma, Rongyu Cao, Binhua Li, Fei Huang 0002, Yongbin Li 0001
ACL (1)3
2023 Separate and Locate: Rethink the Text in Text-based Visual Question Answering
abstract
Text-based Visual Question Answering (TextVQA) aims at answering questions about the text in images. Most works in this field focus on designing network structures or pre-training tasks. All these methods list the OCR texts in reading order (from left to right and top to bottom) to form a sequence, which is treated as a natural language ''sentence''. However, they ignore the fact that most OCR words in the TextVQA task do not have a semantical contextual relationship. In addition, these approaches use 1-D position embedding to construct the spatial relation between OCR tokens sequentially, which is not reasonable. The 1-D position embedding can only represent the left-right sequence relationship between words in a sentence, but not the complex spatial position relationship. To tackle these problems, we propose a novel method named Separate and Locate (SaL) that explores text contextual cues and designs spatial position embedding to construct spatial relations between OCR texts. Specifically, we propose a Text Semantic Separate (TSS) module that helps the model recognize whether words have semantic contextual relations. Then, we introduce a Spatial Circle Position (SCP) module that helps the model better construct and reason the spatial position relationships between OCR texts. Our SaL model outperforms the baseline model by 4.44% and 3.96% accuracy on TextVQA and ST-VQA datasets. Compared with the pre-training state-of-the-art method pre-trained on 64 million pre-training samples, our method, without any pre-training tasks, still achieves 2.68% and 2.52% accuracy improvement on TextVQA and ST-VQA. Our code and models will be released at https://github.com/fangbufang/SaL.
Chengyang Fang, Can Ma, Dayong Hu
ACM Multimedia1
2022 Towards Escaping from Language Bias and OCR Error: Semantics-Centered Text Visual Question Answering
abstract
Texts in scene images convey critical information for scene understanding and reasoning. The abilities of reading and rea-soning matter for the model in the text-based visual question answering (TextVQA) process. However, current TextVQA models do not center on the text and suffer from several limitations. The model is easily dominated by language biases and optical character recognition (OCR) errors due to the ab-sence of semantic guidance in the answer prediction process. In this paper, we propose a novel Semantics-Centered Net-work (SC-Net) that consists of an instance-level contrastive semantic prediction module (ICSP) and a semantics-centered transformer module (SCT). Equipped with the two modules, the semantics-centered model can resist the language biases and the accumulated errors from OCR. Extensive experiments on TextVQA and ST-VQA datasets show the effectiveness of our model. SC- Net surpasses previous works with a notice-able margin and is more reasonable for the TextVQA task.
Chengyang Fang, Gangyan Zeng, Yu Zhou 0015, Daiqing Wu, Can Ma, Dayong Hu, Weiping Wang 0005
ICME1