Xugong Qin

dblp:247/9503 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
15since 2021 · last 2026
0009-0004-3130-3220ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 One2Seq: One-Token Wise Decoder for Efficient Scene Text Recognition
abstract
Auto-regressive (AR)-based decoders, owing to their flexibility in handling variable-length outputs and their strong capability in modeling character-level dependencies, have emerged as the predominant decoding paradigm in the field of scene text recognition (STR). However, AR-based decoders suffer from attention drift, slow decoding speed, and difficulty capturing global dependencies, restricting their performance in various scenarios. In this paper, we propose a novel paradigm for AR-based decoding, called One-Token to Sequence (One2Seq), to address the above issues. Unlike existing methods, we encode the semantic features into a single context token and design a One-Token Wise Decoder to perform the decoding, which alleviates the attention drift caused by the accumulation of semantic information. Moreover, we proposed Positioal-aware Hash Embedding to embed the decoded characters, ensuring the order information is obtained in the context token. By continuously updating this token, One2Seq fully leverages the decoded semantic information while avoiding the computational overhead associated with the growing query sequence. Furthermore, to leverage global information for decoding, we propose Dynamic Global Infusion to dynamically integrates global visual features into the context token. Equipped with the enriched context token, the model has an enhanced ability to extract discriminative local features under the guidance of global context, thereby enhancing recognition accuracy. Extensive experiments reveal that, with its ingenious design, One2Seq exhibits marked superiority on both accuracy and decoding speed compared to existing STR models.
Zhibin Ma, Pengwen Dai, Xugong Qin
AAAI4
2025 Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark
abstract
Document image retrieval (DIR) aims to retrieve document images from a gallery according to a given query. Existing DIR methods are primarily based on image queries that retrieve documents within the same coarse semantic category, e.g., newspapers or receipts. However, these methods struggle to effectively retrieve document images in real-world scenarios where textual queries with fine-grained semantics are usually provided. To bridge this gap, we introduce a new Natural Language-based Document Image Retrieval (NL-DIR) benchmark with corresponding evaluation metrics. In this work, natural language descriptions serve as semantically rich queries for the DIR task. The NL-DIR dataset contains 41K authentic document images, each paired with five high-quality, fine-grained semantic queries generated and evaluated through large language models in conjunction with manual verification. We perform zero-shot and fine-tuning evaluations of existing mainstream contrastive vision-language models and OCR-free visual document understanding (VDU) models. A two-stage retrieval method is further investigated for performance improvement while achieving both time and space efficiency. We hope the proposed NL-DIR benchmark can bring new opportunities and facilitate research for the VDU community. Datasets and codes will be publicly available at huggingface.co/datasets/nianbing/NL-DIR.
Xugong Qin, Jun Jie Ou Yang, Peng Zhang 0044, Gangyan Zeng, Hailun Lin
CVPR2
2025 CLIP is Almost All You Need: Towards Parameter-Efficient Scene Text Retrieval without OCR
abstract
Scene Text Retrieval (STR) seeks to identify all images containing a given query string. Existing methods typically rely on an explicit Optical Character Recognition (OCR) process of text spotting or localization, which is susceptible to complex pipelines and accumulated errors. To settle this, we resort to the Contrastive Language-Image Pre-training (CLIP) models, which have demonstrated the capacity to perceive and understand scene text, making it possible to achieve strictly OCR-free STR. From the perspective of parameter-efficient transfer learning, a lightweight visual position adapter is proposed to provide a positional information complement for the visual encoder. Besides, we introduce a visual context dropout technique to improve the alignment of local visual features. A novel, parameter-free cross-attention mechanism transfers the contrastive relationship between images and text to that between visual tokens and text, producing a rich cross-modal representation, which can be utilized for efficient reranking with a linear classifier. The resulting model, CAYN, which proves that CLIP is Almost all You Need for STR with no more than 0.50M additional parameters required, achieves new state-of-the-art performance on the STR task, with 92.46%/89.49%/85.98% mAP on the SVT/IIIT-STR/TTR datasets. Our findings demonstrate that CLIP can serve as a reliable and efficient solution for OCR-free STR.
Xugong Qin, Peng Zhang 0044, Jun Jie Ou Yang, Gangyan Zeng, Wanqian Zhang, Pengwen Dai
CVPR1
2025 Towards Fine-Grained Document Tampering Detection: New Dataset and Benchmark
Xugong Qin, Jiuqiang Tian, Jiayi Sheng, Tiantian Xia, Yuyi Wang 0008, Gangyan Zeng
PRCV (7)1
2024 MHPS: Multimodality-Guided Hierarchical Policy Search for Knowledge Graph Reasoning
abstract
Recently, path inference-based knowledge graph reasoning (KGR) methods have attracted great attention due to their good performance and interpretability. However, as the number of hops increases, the search space grows exponentially, making the reward sparse and the process of reasoning difficult. To alleviate this problem, we propose the Multimodality-guided Hierarchical Policy Search (MHPS) for KGR, which introduces multimodal hierarchical guidance to each layer of policies during policy search. On the one hand, multimodal guidance reserves rich information on different dimensions, providing more opportunities to find better paths. On the other hand, this leads to better interaction between the two agents, resulting in more concise guidance for policy stepping. Experimental results on two public datasets demonstrate that the proposed approach outperforms state-of-the-art methods on multi-hop KGR.
Xugong Qin, Peng Zhang 0044, Yongquan He, Xinjian Huang, Ming Zhou 0010, Liehuang Zhu, Qingfeng Tan
ICASSP2
2024 Enhancing VPN Traffic Recognition Through CatBoost Feature Extraction and Stacking Ensemble Learning
abstract
A virtual private network (VPN) often serves as an accessory to conceal the online identities of malicious activities. The identification of VPN tunnels has become a prevalent method for detecting potential security threats or abnormalities. Nevertheless, current deep packet inspection and deep learning approaches encounter challenges such as limited scalability or low accuracy. We introduce a novel approach to address the problems by proposing a supervised protocol-wide flow representation learning approach. Our approach leverages the semantic information inherent in the protocol to generate optimal feature embeddings automatically. Additionally, we propose a stacking ensemble machine algorithm to enhance the accuracy of VPN tunnel identification using the generated feature embeddings. We have implemented a prototype named SA-VPN and conducted a comprehensive evaluation of its effectiveness and efficiency using a significant volume of VPN traffic flows. The results demonstrate that our tool surpasses the performance of current state-of-the-art VPN tunnel identification tools.
Ming Zhou 0010, Peng Zhang 0044, Xugong Qin, Xiaoxu Hu
ICC4
2024 Improving Multimodal Rumor Detection via Dynamic Graph Modeling
Xiaoxu Hu, Xugong Qin, Peng Zhang 0044, Gangyan Zeng, Runbo Zhao, Xinjian Huang
ICPR (18)3
2024 Perception-Enhanced Generative Transformer for Key Information Extraction from Documents
Runbo Zhao, Jun Jie Ou Yang, Xugong Qin, Gangyan Zeng, Xiaoxu Hu, Peng Zhang 0044
ICPR (31)4
2024 Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
abstract
Scene text retrieval aims to find all images containing the query text from an image gallery. Current efforts tend to adopt an Optical Character Recognition (OCR) pipeline, which requires complicated text detection and/or recognition processes, resulting in inefficient and inflexible retrieval. Different from them, in this work we propose to explore the intrinsic potential of Contrastive Language-Image Pre-training (CLIP) for OCR-free scene text retrieval. Through empirical analysis, we observe that the main challenges of CLIP as a text retriever are: 1) limited text perceptual scale, and 2) entangled visual-semantic concepts. To this end, a novel model termed FDP (Focus, Distinguish, and Prompt) is developed. FDP first focuses on scene text via shifting the attention to the text area and probing the hidden text knowledge, and then divides the query text into content word and function word for processing, in which a semantic-aware prompting scheme and a distracted queries assistance module are utilized. Extensive experiments show that FDP significantly enhances the inference speed while achieving better or competitive retrieval accuracy compared to existing methods. Notably, on the IIIT-STR benchmark, FDP surpasses the state-of-the-art model by 4.37% with a 4 times faster speed. Furthermore, additional experiments under phrase-level and attribute-aware scene text retrieval settings validate FDP's particular advantages in handling diverse forms of query text. The source code will be available at https://github.com/Gyann-z/FDP.
Gangyan Zeng, Yuan Zhang 0013, Dongbao Yang, Peng Zhang 0044, Yiwen Gao 0001, Xugong Qin, Yu Zhou 0015
ACM Multimedia7
2024 Granularity-Aware Single-Point Scene Text Spotting With Sequential Recurrence Self-Attention
abstract
Scene text spotting, a unified framework between text detection and text recognition, has made great progress in recent years. Existing methods usually adopt the fully-supervised learning strategy, which relies on time-consuming location annotations, particularly for scene texts with arbitrary shapes. In this paper, we propose a weakly-supervised scene text spotting method via the location labels of single points with the corresponding text transcriptions. Due to the weak location annotations for challenging scene texts, previous weakly-supervised methods adopting the convolution neural network structure make it hard to model the different-scale text feature representations under blurring or nosing scenarios. In addition, as the single-point location can only cover part of the text instance, it will burden the confusion of sequential-like scene text recognition. To address these issues, we present a novel sequential recurrence self-attention for granularity-aware single-point scene text spotting. Specifically, we first enhance the scene text feature representations with different scales by integrating the global intra-interaction of high-level features with the low-level local features. Then, based on the granularity-aware text features, we decode them into text transcriptions in the sequential recurrence self-attention manner to capture the sequence-dependent relation in character-level semantics and locations. Extensive experiments show that our proposed method outperforms existing state-of-the-art weakly-supervised scene text spotters by a large margin.
Xunquan Tong, Pengwen Dai, Xugong Qin, Rui Wang 0032, Wenqi Ren
IEEE Trans. Circuits Syst. Video Technol.3
2023 Towards Robust Real-Time Scene Text Detection: From Semantic to Instance Representation Learning
abstract
Due to the flexible representation of arbitrary-shaped scene text and simple pipeline, bottom-up segmentation-based methods begin to be mainstream in real-time scene text detection. Despite great progress, these methods show deficiencies in robustness and still suffer from false positives and instance adhesion. Different from existing methods which integrate multiple-granularity features or multiple outputs, we resort to the perspective of representation learning in which auxiliary tasks are utilized to enable the encoder to jointly learn robust features with the main task of per-pixel classification during optimization. For semantic representation learning, we propose global-dense semantic contrast (GDSC), in which a vector is extracted for global semantic representation, then used to perform element-wise contrast with the dense grid features. To learn instance-aware representation, we propose to combine top-down modeling (TDM) with the bottom-up framework to provide implicit instance-level clues for the encoder. With the proposed GDSC and TDM, the encoder network learns stronger representation without introducing any parameters and computations during inference. Equipped with a very light decoder, the detector can achieve more robust real-time scene text detection. Experimental results on four public datasets show that the proposed method can outperform or be comparable to the state-of-the-art on both accuracy and speed. Specifically, the proposed method achieves 87.2% F-measure with 48.2 FPS on Total-Text and 89.6% F-measure with 36.9 FPS on MSRA-TD500 on a single GeForce RTX 2080 Ti GPU.
Xugong Qin, Pengyuan Lv, Chengquan Zhang, Yu Zhou 0015, Peng Zhang 0044, Hailun Lin, Weiping Wang 0005
ACM Multimedia1
2022 UNITS: Unsupervised Intermediate Training Stage for Scene Text Detection
abstract
Recent scene text detection methods are almost based on deep learning and data-driven. Synthetic data is commonly adopted for pre-training due to expensive annotation cost. However, there are obvious domain discrepancies between synthetic data and real-world data. It may lead to suboptimal performance to directly adopt the model initialized by synthetic data in the fine-tuning stage. In this paper, we propose a new training paradigm for scene text detection, which introduces an UNsupervised Intermediate Training Stage (UNITS) that builds a buffer path to real-world data and can alleviate the gap between the pre-training stage and fine-tuning stage. Three training strategies are further explored to perceive information from real-world data in an unsupervised way. With UNITS, scene text detectors are improved without introducing any parameters and computations during inference. Extensive experimental results show consistent performance improvements on three public datasets.
Youhui Guo, Yu Zhou 0015, Xugong Qin, Enze Xie, Weiping Wang 0005
ICME3
2021 Which and Where to Focus: A Simple yet Accurate Framework for Arbitrary-Shaped Nearby Text Detection in Scene Images
Youhui Guo, Yu Zhou 0015, Xugong Qin, Weiping Wang 0005
ICANN (5)3
2021 FC2RN: A Fully Convolutional Corner Refinement Network for Accurate Multi-Oriented Scene Text Detection
abstract
Accurate detection of multi-oriented text that accounts for a large proportion in real practice is of great significance. The performance has improved rapidly on common benchmarks in recent years. However, dense long text case and the quality of detection are easy to be overlooked. Direct regression may produce low-quality and incomplete detections due to the constrain of the receptive field; proposal-based methods could alleviate this but might introduce redundant context due to RoI operation, degrading the performance. To address the dilemma, a novel proposed corner-aware convolution in which the sampling positions tightly cover the text area is utilized to encode an initial corner prediction into the feature maps, which can be further used to produce a refined corner prediction. We embed the proposed module into an anchor-free baseline model, leading to a simple and effective fully convolutional corner refinement network (FC2RN). Experimental results on four public datasets including MSRATD500, ICDAR2015, RCTW-17, and COCO-Text demonstrate that FC2RN can outperform state-of-the-art methods.
Xugong Qin, Yu Zhou 0015, Youhui Guo, Dayan Wu, Weiping Wang 0005
ICASSP1
2021 Mask is All You Need: Rethinking Mask R-CNN for Dense and Arbitrary-Shaped Scene Text Detection
abstract
Due to the large success in object detection and instance segmentation, Mask R-CNN attracts great attention and is widely adopted as a strong baseline for arbitrary-shaped scene text detection and spotting. However, two issues remain to be settled. The first is dense text case, which is easy to be neglected but quite practical. There may exist multiple instances in one proposal, which makes it difficult for the mask head to distinguish different instances and degrades the performance. In this work, we argue that the performance degradation results from the learning confusion issue in the mask head. We propose to use an MLP decoder instead of the "deconv-conv" decoder in the mask head, which alleviates the issue and promotes robustness significantly. And we propose instance-aware mask learning in which the mask head learns to predict the shape of the whole instance rather than classify each pixel to text or non-text. With instance-aware mask learning, the mask branch can learn separated and compact masks. The second is that due to large variations in scale and aspect ratio, RPN needs complicated anchor settings, making it hard to maintain and transfer across different datasets. To settle this issue, we propose an adaptive label assignment in which all instances especially those with extreme aspect ratios are guaranteed to be associated with enough anchors. Equipped with these components, the proposed method named MAYOR achieves state-of-the-art performance on five benchmarks including DAST1500, MSRA-TD500, ICDAR2015, CTW1500, and Total-Text.
Xugong Qin, Yu Zhou 0015, Youhui Guo, Dayan Wu, Zhihong Tian 0001, Weiping Wang 0005
ACM Multimedia1
2020 Gaussian Constrained Attention Network for Scene Text Recognition
abstract
Scene text recognition has been a hot topic in computer vision. Recent methods adopt the attention mechanism for sequence prediction which achieve convincing results. However, we argue that the existing attention mechanism faces the problem of attention diffusion, in which the model may not focus on a certain character area. In this paper, we propose Gaussian Constrained Attention Network to deal with this problem. It is a 2D attention-based method integrated with a novel Gaussian Constrained Refinement Module, which predicts an additional Gaussian mask to refine the attention weights. Different from adopting an additional supervision on the attention weights simply, our proposed method introduces an explicit refinement. In this way, the attention weights will be more concentrated and the attention-based recognition network achieves better performance. The proposed Gaussian Constrained Refinement Module is flexible and can be applied to existing attention-based methods directly. The experiments on several benchmark datasets demonstrate the effectiveness of our proposed method. Our code has been available at https://github.com/Pay20Y/GCAN.
Xugong Qin, Yu Zhou 0015, Weiping Wang 0005
ICPR2
2019 Curved Text Detection in Natural Scene Images with Semi- and Weakly-Supervised Learning
abstract
Detecting curved text in the wild is very challenging. Recently, most state-of-the-art methods are segmentation based and require pixel-level annotations. We propose a novel scheme to train an accurate text detector using only a small amount of pixel-level annotated data and a large amount of data annotated with rectangles or even unlabeled data. A light model is first obtained by training with the pixel-level annotated data and then used to annotate unlabeled or weakly labeled data. A novel strategy which utilizes ground-truth bounding boxes to generate pseudo mask annotations is proposed in weakly-supervised learning. Experimental results on CTW1500 and Total-Text demonstrate that our method can substantially reduce the requirement of pixel-level annotated data. Our method can also generalize well across the two datasets. The performance of the proposed method is comparable with the state-of-the-art methods with only 10% pixel-level annotated data and 90% rectangle-level weakly annotated data.
Xugong Qin, Yu Zhou 0015, Dongbao Yang, Weiping Wang 0005
ICDAR1