VLDB 2026 Research / reviewers in the wild / expert
Gangyan Zeng
dblp:261/0517
· DBLP profile ↗
20ranked-venue papers
5as first author
20since 2021 · last 2025
0000-0003-2696-8549ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal CluesabstractVideo text-based visual question answering (TextVQA) is a practical task that aims to answer questions by jointly reasoning textual and visual information in a given video. Inspired by the development of TextVQA in image domain, existing Video TextVQA approaches leverage a language model (e.g. T5) to process text-rich multiple frames and generate answers auto-regressively. Nevertheless, the spatio-temporal relationships among visual entities (including scene text and objects) will be disrupted and models are susceptible to interference from unrelated information, resulting in irrational reasoning and inaccurate answering. To tackle these challenges, we propose the TEA (stands for "Track the Answer'') method that better extends the generative TextVQA framework from image to video. TEA recovers the spatio-temporal relationships in a complementary way and incorporates OCR-aware clues to enhance the quality of reasoning questions. Extensive experiments on several public Video TextVQA datasets validate the effectiveness and generalization of our framework. TEA outperforms existing TextVQA methods, video-language pretraining methods and video large language models by great margins. The code will be publicly released. Gangyan Zeng, Huawen Shen, Daiqing Wu, Yu Zhou 0015, Can Ma |
AAAI | 2 |
| 2025 | Towards Natural Language-Based Document Image Retrieval: New Dataset and BenchmarkabstractDocument image retrieval (DIR) aims to retrieve document images from a gallery according to a given query. Existing DIR methods are primarily based on image queries that retrieve documents within the same coarse semantic category, e.g., newspapers or receipts. However, these methods struggle to effectively retrieve document images in real-world scenarios where textual queries with fine-grained semantics are usually provided. To bridge this gap, we introduce a new Natural Language-based Document Image Retrieval (NL-DIR) benchmark with corresponding evaluation metrics. In this work, natural language descriptions serve as semantically rich queries for the DIR task. The NL-DIR dataset contains 41K authentic document images, each paired with five high-quality, fine-grained semantic queries generated and evaluated through large language models in conjunction with manual verification. We perform zero-shot and fine-tuning evaluations of existing mainstream contrastive vision-language models and OCR-free visual document understanding (VDU) models. A two-stage retrieval method is further investigated for performance improvement while achieving both time and space efficiency. We hope the proposed NL-DIR benchmark can bring new opportunities and facilitate research for the VDU community. Datasets and codes will be publicly available at huggingface.co/datasets/nianbing/NL-DIR. Xugong Qin, Jun Jie Ou Yang, Peng Zhang 0044, Gangyan Zeng, Hailun Lin |
CVPR | 5 |
| 2025 | CLIP is Almost All You Need: Towards Parameter-Efficient Scene Text Retrieval without OCRabstractScene Text Retrieval (STR) seeks to identify all images containing a given query string. Existing methods typically rely on an explicit Optical Character Recognition (OCR) process of text spotting or localization, which is susceptible to complex pipelines and accumulated errors. To settle this, we resort to the Contrastive Language-Image Pre-training (CLIP) models, which have demonstrated the capacity to perceive and understand scene text, making it possible to achieve strictly OCR-free STR. From the perspective of parameter-efficient transfer learning, a lightweight visual position adapter is proposed to provide a positional information complement for the visual encoder. Besides, we introduce a visual context dropout technique to improve the alignment of local visual features. A novel, parameter-free cross-attention mechanism transfers the contrastive relationship between images and text to that between visual tokens and text, producing a rich cross-modal representation, which can be utilized for efficient reranking with a linear classifier. The resulting model, CAYN, which proves that CLIP is Almost all You Need for STR with no more than 0.50M additional parameters required, achieves new state-of-the-art performance on the STR task, with 92.46%/89.49%/85.98% mAP on the SVT/IIIT-STR/TTR datasets. Our findings demonstrate that CLIP can serve as a reliable and efficient solution for OCR-free STR. Xugong Qin, Peng Zhang 0044, Jun Jie Ou Yang, Gangyan Zeng, Wanqian Zhang, Pengwen Dai |
CVPR | 4 |
| 2025 | PerturbCTC: Improving Alignment in Scene Text Recognition with Feature Perturbation Based CTC
Zhijie Shen, Yaqiang Wu, Gangyan Zeng, Dongbao Yang, Yu Zhou 0015 |
ICDAR (4) | 5 |
| 2025 | Gather and Trace: Rethinking Video TextVQA from an Instance-oriented PerspectiveabstractVideo text-based visual question answering (Video TextVQA) aims to answer questions by explicitly reading and reasoning about the text involved in a video. Most works in this field follow a frame-level framework which suffers from redundant text entities and implicit relation modeling, resulting in limitations in both accuracy and efficiency. In this paper, we rethink the Video TextVQA task from an instance-oriented perspective and propose a novel model termed GAT (Gather and Trace). First, to obtain accurate reading result for each video text instance, a context-aggregated instance gathering module is designed to integrate the visual appearance, layout characteristics, and textual contents of the related entities into a unified textual representation. Then, to capture dynamic evolution of text in the video flow, an instance-focused trajectory tracing module is utilized to establish spatio-temporal relationships between instances and infer the final answer. Extensive experiments on several public Video TextVQA datasets validate the effectiveness and generalization of our framework. GAT outperforms existing Video TextVQA methods, video-language pretraining methods, and video large language models in both accuracy and inference speed. Notably, GAT surpasses the previous state-of-the-art Video TextVQA methods by 3.86% in accuracy and achieves ten times of faster inference speed than video large language models. The source code is available at https://github.com/zhangyan-ucas/GAT. Gangyan Zeng, Daiqing Wu, Huawen Shen, Binbin Li 0003, Yu Zhou 0015, Can Ma, Xiaojun Bi 0002 |
ACM Multimedia | 2 |
| 2025 | When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and UnderstandingabstractLarge Multimodal Models (LMMs) have achieved impressive progress in visual perception and reasoning. However, when confronted with visually ambiguous or non-semantic scene text, they often struggle to accurately spot and understand the content, frequently generating semantically plausible yet visually incorrect answers, which we refer to as semantic hallucination.
In this work, we investigate the underlying causes of semantic hallucination and identify a key finding: Transformer layers in LLM with stronger attention focus on scene text regions are less prone to producing semantic hallucinations.
Thus, we propose a training-free semantic hallucination mitigation framework comprising two key components: (1) ZoomText, a coarse-to-fine strategy that identifies potential text regions without external detectors; and (2) Grounded Layer Correction, which adaptively leverages the internal representations from layers less prone to hallucination to guide decoding, correcting hallucinated outputs for non-semantic samples while preserving the semantics of meaningful ones. To enable rigorous evaluation, we introduce TextHalu-Bench, a benchmark of 1,740 samples spanning both semantic and non-semantic cases, with manually curated question–answer pairs designed to probe model hallucinations.
Extensive experiments demonstrate that our method not only effectively mitigates semantic hallucination but also achieves strong performance on public benchmarks for scene text spotting and understanding. Hangui Lin, Yexin Liu, Gangyan Zeng, Yu Zhou 0015, Ser-Nam Lim, Harry Yang, Nicu Sebe |
NeurIPS | 5 |
| 2025 | Towards Fine-Grained Document Tampering Detection: New Dataset and Benchmark
Xugong Qin, Jiuqiang Tian, Jiayi Sheng, Tiantian Xia, Yuyi Wang 0008, Gangyan Zeng |
PRCV (7) | 7 |
| 2025 | TrapLLM: An LLM-powered Interactive Log-based Honeypot for Real-world Network AttacksabstractWith the continual escalation of cyberattack tactics, zero-day exploits and advanced persistent threats (APTs) characterized by high stealth and dynamic evolution have posed significant challenges to traditional honeypot systems. Existing approaches are limited in service fidelity, interactive intelligence, and awareness of attacker intent, making it difficult to effectively lure advanced attackers or reconstruct threat chains from massive volumes of log data. To address these challenges, this paper introduces large language models (LLMs) as the core driving force to construct an architecture that integrates log-driven data governance with dynamic response generation. Leveraging the semantic understanding and generative capabilities of LLMs, the proposed method enables fine-grained identification of attacker intent, reconstruction of event sequences, and adaptive responses driven by retrieval-augmented generation (RAG), thereby realizing an intelligent closed-loop defense. This innovative integration overcomes the constraints of traditional static rules and low interaction emulation, achieving attack intent analysis and adaptive deception response, and providing an efficient pathway for proactive threat hunting. During a 25day deployment in a real production network environment, the system utilized 25 diversion nodes and 8 types of emulated services to capture over 1.39 million raw attack logs. Analysis revealed multiple attack attempts targeting 9 known Common Vulnerabilities and Exposures (CVE) vulnerabilities, along with a substantial number of high-severity attacks for which specific CVE identifiers could not be determined. Experimental results demonstrate that the proposed approach significantly improves both the accuracy of threat awareness and the timeliness of response, providing reliable support for the evolution of intelligent defense mechanisms. Yunjun Ma, Gangyan Zeng, Peng Zhang 0044, Fuyuan Zhang, Ran Lin, Huan Qian |
TrustCom | 2 |
| 2025 | TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language ModelabstractExisting scene text spotters are designed to locate and transcribe texts from images. However, it is challenging for a spotter to achieve precise detection and recognition of scene texts simultaneously. Inspired by the glimpse-focus spotting pipeline of human beings and impressive performances of Pre-trained Language Models (PLMs) on visual tasks, we ask: (1) “Can machines spot texts without precise detection just like human beings?”, and if yes, (2) “Is text block another alternative for scene text spotting other than word or character?” To this end, our proposed scene text spotter leverages advanced PLMs to enhance performance without fine-grained detection. Specifically, we first use a simple detector for block-level text detection to obtain rough positional information. Then, we fine-tune a PLM using a large-scale OCR dataset to achieve accurate recognition. Benefiting from the comprehensive language knowledge gained during the pre-training phase, the PLM-based recognition module effectively handles complex scenarios, including multi-line, reversed, occluded, and incomplete-detection texts. Taking advantage of the fine-tuned language model on scene recognition benchmarks and the paradigm of text block detection, extensive experiments demonstrate the superior performance of our scene text spotter across multiple public benchmarks. Additionally, we attempt to spot texts directly from an entire scene image to demonstrate the potential of PLMs, even Large Language Models (LLMs). Jiahao Lyu 0002, Gangyan Zeng, Enze Xie, Wei Wang 0315, Can Ma, Yu Zhou 0015 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Improving Multimodal Rumor Detection via Dynamic Graph Modeling
Xiaoxu Hu, Xugong Qin, Peng Zhang 0044, Gangyan Zeng, Runbo Zhao, Xinjian Huang |
ICPR (18) | 5 |
| 2024 | Perception-Enhanced Generative Transformer for Key Information Extraction from Documents
Runbo Zhao, Jun Jie Ou Yang, Xugong Qin, Gangyan Zeng, Xiaoxu Hu, Peng Zhang 0044 |
ICPR (31) | 5 |
| 2024 | Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text RetrievalabstractScene text retrieval aims to find all images containing the query text from an image gallery. Current efforts tend to adopt an Optical Character Recognition (OCR) pipeline, which requires complicated text detection and/or recognition processes, resulting in inefficient and inflexible retrieval. Different from them, in this work we propose to explore the intrinsic potential of Contrastive Language-Image Pre-training (CLIP) for OCR-free scene text retrieval. Through empirical analysis, we observe that the main challenges of CLIP as a text retriever are: 1) limited text perceptual scale, and 2) entangled visual-semantic concepts. To this end, a novel model termed FDP (Focus, Distinguish, and Prompt) is developed. FDP first focuses on scene text via shifting the attention to the text area and probing the hidden text knowledge, and then divides the query text into content word and function word for processing, in which a semantic-aware prompting scheme and a distracted queries assistance module are utilized. Extensive experiments show that FDP significantly enhances the inference speed while achieving better or competitive retrieval accuracy compared to existing methods. Notably, on the IIIT-STR benchmark, FDP surpasses the state-of-the-art model by 4.37% with a 4 times faster speed. Furthermore, additional experiments under phrase-level and attribute-aware scene text retrieval settings validate FDP's particular advantages in handling diverse forms of query text. The source code will be available at https://github.com/Gyann-z/FDP. Gangyan Zeng, Yuan Zhang 0013, Dongbao Yang, Peng Zhang 0044, Yiwen Gao 0001, Xugong Qin, Yu Zhou 0015 |
ACM Multimedia | 1 |
| 2024 | Show Exemplars and Tell Me What You See: In-Context Learning with Frozen Large Language Models for TextVQA
Gangyan Zeng, Huawen Shen, Can Ma, Yu Zhou 0015 |
PRCV (7) | 2 |
| 2023 | Filling in the Blank: Rationale-Augmented Prompt Tuning for TextVQAabstractRecently, generative Text-based visual question answering (TextVQA) methods, which are often based on language models, have exhibited impressive results and drawn increasing attention. However, due to the inconsistencies in both input forms and optimization objectives, the power of pretrained language models is not fully explored, resulting in the need for large amounts of training data. In this work, we rethink the characteristics of the TextVQA task and find that scene text is indeed a special kind of language embedded in images. To this end, we propose a text-centered generative framework FITB (stands for Filling In The Blank), in which multimodal information is mainly represented in textual form and rationale-augmented prompting is involved. Specifically, an infilling-based prompt strategy is utilized to formulate TextVQA as a novel problem of filling in the blank with proper scene text according to the language context. Furthermore, aiming to prevent the model from language bias overfitting, we design a rough answer grounding module to provide visual rationales for promoting multimodal reasoning. Extensive experiments verify the superiority of FITB in both fully-supervised and zero-shot/few-shot settings. Notably, even with a saving of about 64M data, FITB surpasses the state-of-the-art method by 3.00% and 1.99% on TextVQA and ST-VQA datasets, respectively. Gangyan Zeng, Yuan Zhang 0013, Yu Zhou 0015, Bo Fang 0003, Weiping Wang 0005 |
ACM Multimedia | 1 |
| 2023 | Feature Enhancement with Text-Specific Region Contrast for Scene Text Detection
Xurui Sun, Jiahao Lyu 0002, Yifei Zhang 0005, Gangyan Zeng, Bo Fang 0003, Yu Zhou 0015, Enze Xie, Can Ma |
PRCV (7) | 4 |
| 2023 | Beyond OCR + VQA: Towards end-to-end reading and reasoning for robust and accurate textvqaabstractText-based visual question answering (TextVQA), which answers a visual question by considering both visual contents and scene texts, has attracted increasing attention recently. Most existing methods employ an optical character recognition (OCR) module as a pre-processor to read texts, then combine it with a visual question answering (VQA) framework. However, inaccurate OCR results may lead to cumulative error propagation , and the correlation between text reading and text-based reasoning is not fully exploited. In this work, we integrate OCR into the flow of TextVQA, targeting the mutual reinforcement of OCR and VQA tasks. Specifically, a visually enhanced text embedding module is proposed to predict semantic features from the visual information of texts, by which texts can be reasonably understood even without accurate recognition. Further, two elaborate schemes are developed to leverage contextual information in VQA to modify OCR results. The first scheme is a reading modification module that adaptively selects the answer results according to the contexts. Second, we propose an efficient end-to-end text reading and reasoning network, where the downstream VQA signal contributes to the optimization of text reading. Extensive experiments show that our method outperforms existing alternatives in terms of accuracy and robustness, whether ground truth OCR annotations are used or not. Gangyan Zeng, Yuan Zhang 0013, Yu Zhou 0015, Weiping Wang 0005, Xu-Cheng Yin |
Pattern Recognit. | 1 |
| 2022 | Towards Escaping from Language Bias and OCR Error: Semantics-Centered Text Visual Question AnsweringabstractTexts in scene images convey critical information for scene understanding and reasoning. The abilities of reading and rea-soning matter for the model in the text-based visual question answering (TextVQA) process. However, current TextVQA models do not center on the text and suffer from several limitations. The model is easily dominated by language biases and optical character recognition (OCR) errors due to the ab-sence of semantic guidance in the answer prediction process. In this paper, we propose a novel Semantics-Centered Net-work (SC-Net) that consists of an instance-level contrastive semantic prediction module (ICSP) and a semantics-centered transformer module (SCT). Equipped with the two modules, the semantics-centered model can resist the language biases and the accumulated errors from OCR. Extensive experiments on TextVQA and ST-VQA datasets show the effectiveness of our model. SC- Net surpasses previous works with a notice-able margin and is more reasonable for the TextVQA task. Chengyang Fang, Gangyan Zeng, Yu Zhou 0015, Daiqing Wu, Can Ma, Dayong Hu, Weiping Wang 0005 |
ICME | 2 |
| 2022 | TextBlock: Towards Scene Text Spotting without Fine-grained DetectionabstractScene text spotting systems which integrate text detection and recognition modules have witnessed a lot of success in recent years. Existing works mostly follow the framework of word/character-level fine-grained detection and isolated-instance recognition, which overemphasize the role of detector and ignore the rich context information in recognition. After rethinking the conventional framework, and inspired by the glimpse-focus spotting pipeline of human beings, we ask:1) "can machine spot text without accurate detection just like human beings?", and if yes, 2) "is text block another alternative for scene text spotting other than word or character?". Based on these questions, we propose a new perspective of coarse-grained detection with multi-instance recognition for text spotting. Specifically, a pioneering network termed TextBlock is developed, and a heuristic text block generation method as well as a multi-instance block-level recognition module are proposed. In this way, the burden of detection is relieved, and the contextual semantic information is well explored for recognition. To train the block-level recognizer, a synthetic dataset including about 800K images is formed. As a by-product of attention, fine-grained detection can be recovered with the recognizer. Equipped with a detector without many bells and whistles (e.g., Faster R-CNN), TextBlock achieves competitive or even better performance compared with previous sophisticated text spotters on several public benchmarks. As a primary attempt, we expect this framework will have a potential impact on scene text spotting research in the future. Yuan Zhang 0013, Yu Zhou 0015, Gangyan Zeng, Youhui Guo, Haiying Wu, Weiping Wang 0005 |
ACM Multimedia | 4 |
| 2021 | Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQAabstractText-based visual question answering (TextVQA) requires analyzing both the visual contents and texts in an image to answer a question, which is more practical than general visual question answering (VQA). Existing efforts tend to regard optical character recognition (OCR) as a pre-processing and then combine it with a VQA framework. It makes the performance of multimodal reasoning and question answering highly depend on the accuracy of OCR. In this work, we address this issue with two perspectives. First, we take advantages of multimodal cues to complete the semantic information of texts. A visually enhanced text embedding is proposed to enable understanding of texts without accurately recognizing them. Second, we further leverage rich contextual information to modify the answer texts even if the OCR module does not correctly recognize them. In addition, the visual objects are endued with semantic representations to enable objects in the same semantic space as OCR tokens. Equipped with these techniques, the cumulative error propagation caused by poor OCR performance is effectively suppressed. Extensive experiments on TextVQA and ST-VQA datasets demonstrate that our approach achieves the state-of-the-art performance in terms of accuracy and robustness. Gangyan Zeng, Yuan Zhang 0013, Yu Zhou 0015 |
ACM Multimedia | 1 |
| 2021 | A Cost-Efficient Framework for Scene Text Detection in the Wild
Gangyan Zeng, Yuan Zhang 0013, Yu Zhou 0015 |
PRICAI (1) | 1 |