Pei Fu

dblp:202/1707 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Databases, data management, data science and information retrieval · 3
YearPublicationVenuePosition
2026 AutoLink: Autonomous Schema Exploration and Expansion for Scalable Schema Linking in Text-to-SQL at Scale
abstract
For industrial-scale text-to-SQL, supplying the entire database schema to Large Language Models (LLMs) is impractical due to context window limits and irrelevant noise. Schema linking, which filters the schema to a relevant subset, is therefore critical. However, existing methods incur prohibitive costs, struggle to trade off recall and noise, and scale poorly to large databases. We present AutoLink, an autonomous agent framework that reformulates schema linking as an iterative, agent-driven process. Guided by an LLM, AutoLink dynamically explores and expands the linked schema subset, progressively identifying necessary schema components without inputting the full database schema. Our experiments demonstrate AutoLink's superior performance, achieving state-of-the-art strict schema linking recall of 97.4% on Bird-Dev and 91.2% on Spider 2.0-Lite, with competitive execution accuracy, i.e., 68.7% EX on Bird-Dev (better than CHESS) and 34.9% EX on Spider 2.0-Lite (ranking 2nd on the official leaderboard). Crucially, AutoLink exhibits exceptional scalability, maintaining high recall, efficient token consumption, and robust execution accuracy on large schemas (e.g., over 3,000 columns) where existing methods severely degrade—making it a highly scalable, high-recall schema-linking solution for industrial text-to-SQL systems.
Yuanlei Zheng, Zhenbiao Cao, Xiaojin Zhang 0002, Zhongyu Wei, Pei Fu, Zhenbo Luo, Wei Chen 0088, Xiang Bai
AAAI6
2026 Doc-V^*: Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA
abstract
Yuanlei Zheng, Pei Fu, Hang Li, Ziyang Wang, Yuyi Zhang, Wenyu Ruan, Xiaojin Zhang, Zhongyu Wei, Zhenbo Luo, Jian Luan, Wei Chen, Xiang Bai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuanlei Zheng, Pei Fu, Hang Li 0001, Wenyu Ruan, Xiaojin Zhang 0002, Zhongyu Wei, Zhenbo Luo, Jian Luan 0001, Wei Chen 0088, Xiang Bai
ACL (1)2
2025 InstructOCR: Instruction Boosting Scene Text Spotting
abstract
In the field of scene text spotting, previous OCR methods primarily relied on image encoders and pre-trained text information, but they often overlooked the advantages of incorporating human language instructions. To address this gap, we propose InstructOCR, an innovative instruction-based scene text spotting model that leverages human language instructions to enhance the understanding of text within images. Our framework employs both text and image encoders during training and inference, along with instructions meticulously designed based on text attributes. This approach enables the model to interpret text more accurately and flexibly. Extensive experiments demonstrate the effectiveness of our model and we achieve state-of-the-art results on widely used benchmarks. Furthermore, the proposed framework can be seamlessly applied to scene text VQA tasks. By leveraging instruction strategies during pre-training, the performance on downstream VQA tasks can be significantly improved, with a 2.6% increase on the TextVQA dataset and a 2.1% increase on the ST-VQA dataset. These experimental results provide insights into the benefits of incorporating human language instructions for OCR-related tasks.
Chen Duan, Qianyi Jiang, Pei Fu, Shengxi Li, Shan Guo, Junfeng Luo
AAAI3
2025 Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
abstract
Multi-modal Large Language Models (MLLMs) have introduced a novel dimension to document understanding, i.e., they endow large language models with visual comprehension capabilities; however, how to design a suitable image-text pre-training task for bridging the visual and language modality in document-level MLLMs remains underexplored. In this study, we introduce a novel visuallanguage alignment method that casts the key issue as a Visual Question Answering with Mask generation (VQA-Mask) task, optimizing two tasks simultaneously: VQA-based text parsing and mask generation. The former allows the model to implicitly align images and text at the semantic level. The latter introduces an additional mask generator (discarded during inference) to explicitly ensure alignment between visual texts within images and their corresponding image regions at a spatially-aware level. Together, they can prevent model hallucinations when parsing visual text and effectively promote spatially-aware feature representation learning. To support the proposed VQAMask task, we construct a comprehensive image-mask generation pipeline and provide a large-scale dataset with 6M data (MTMask6M). Subsequently, we demonstrate that introducing the proposed mask generation task yields competitive document-level understanding performance. Leveraging the proposed VQAMask, we introduce Marten, a trainingefficient MLLM tailored for document-level understanding. Extensive experiments show that our Marten consistently achieves significant improvements among 8B-MLLMs in document-centric tasks. Code and datasets are available at https://github.com/PriNing/Marten.
Tongkun Guan, Pei Fu, Chen Duan, Qianyi Jiang, Zhentao Guo, Shan Guo, Junfeng Luo, Wei Shen 0002, Xiaokang Yang 0001
CVPR3
2025 A Token-Level Text Image Foundation Model for Document Understanding
Tongkun Guan, Pei Fu, Zhengtao Guo, Wei Shen 0002, Tiezhu Yue, Chen Duan, Qianyi Jiang, Junfeng Luo, Xiaokang Yang 0001
ICCV3
2025 BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
abstract
In the field of AI-driven human-GUI interaction automation, while rapid advances in multimodal large language models and reinforcement fine-tuning techniques have yielded remarkable progress, a fundamental challenge persists: their interaction logic significantly deviates from natural human-GUI communication patterns. To address this gap, we propose Blink–Think–Link (BTL), a brain-inspired framework for human-GUI interaction that mimics the human cognitive process between users and graphical interfaces. The system decomposes interactions into three biologically plausible phases: (1) \textbf{Blink} - rapid detection and attention to relevant screen areas, analogous to saccadic eye movements; (2) \textbf{Think} - higher-level reasoning and decision-making, mirroring cognitive planning; and (3) \textbf{Link} - generation of executable commands for precise motor control, emulating human action selection mechanisms. Additionally, we introduce two key technical innovations for BTL framework: (1) Blink Data Generation - an automated annotation pipeline specifically optimized for blink data, and (2) {BTL Reward – the first rule-based reward mechanism that enables reinforcement learning driven by both process and outcome.} Building upon this framework, we develop a GUI agent model named BTL-UI, which demonstrates competitive performance across both static GUI understanding and dynamic interaction tasks in comprehensive benchmarks. These results provide conclusive empirical validation of the framework's efficacy in developing advanced GUI agents.
Shaojie Zhang 0004, Ruoceng Zhang, Pei Fu, Shiqi Cui, Zhenbo Luo, Jian Luan 0001
NeurIPS3
2024 ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and Spotting
abstract
In recent years, text-image joint pre-training techniques have shown promising results in various tasks. However, in Optical Character Recognition (OCR) tasks, aligning text instances with their corresponding text regions in images poses a challenge, as it requires effective alignment between text and OCR-Text (referring to the text in images as OCR-Text to distinguish from the text in natural language) rather than a holistic understanding of the overall image content. In this paper, we propose a new pre-training method called OCR-Text Destylization Modeling (ODM) that transfers Diverse styles of text found in images to a uniform style based on the text prompt. With ODM, we achieve better alignment between text and OCR-Text and enable pre-trained models to adapt to the complex and diverse styles of scene text detection and spotting tasks. Additionally, we have designed a new labeling generation method specifically for ODM and combined it with our proposed Text-Controller module to address the challenge of annotation costs in OCR tasks, allowing a larger amount of unlabeled data to participate in pre-training. Extensive experiments on multiple public datasets demonstrate that our method significantly improves performance and outperforms current pre-training methods in scene text detection and spotting tasks. Code is available at ODM.
Chen Duan, Pei Fu, Shan Guo, Qianyi Jiang, Xiaoming Wei
CVPR2
2018 R2 CNN: Rotational Region CNN for Arbitrarily-Oriented Scene Text Detection
abstract
Scene text detection is challenging as the input may have different orientations, sizes, font styles, lighting conditions, perspective distortions and languages. This paper addresses the problem by designing a Rotational Region CNN (R2CNN). R2CNN includes a Text Region Proposal Network (Text-RPN) to estimate approximate text regions and a multitask refinement network to get the precise inclined box. Our work has the following features. First, we use a novel multi-task regression method to support arbitrarily-oriented scene text detection. Second, we introduce multiple ROIPoolings to address the scene text detection problem for the first time. Third, we use an inclined Non-Maximum Suppression (NMS) to post-process the detection candidates. Experiments show that our method outperforms the state-of-the-art on standard benchmarks: ICDAR 2013, ICDAR 2015, COCO-Text and MSRA-TD500.
Yingying Jiang 0001, Shuli Yang, Pei Fu, Zhenbo Luo
ICPR7
2017 End-to-End Scene Text Recognition in Videos Based on Multi Frame Tracking
abstract
Text detection and recognition in scene images and videos attract much attention in computer vision recently. However, most existing text detection and recognition methods only focus on static images. In this paper an end-to-end scene text recognition method based on multi frame tracking is proposed for text in videos, in which temporal information is employed to improve performance. First, an end-to-end text recognition method based on a unified deep neural network is used to detect and recognize text in each frame of the input video. Then, multi frame text tracking is employed through associations of texts in current frame and several previous frames to obtain final results. Experiments on ICDAR datasets demonstrate that the proposed method outperforms the state-of-the-art methods in end-to-end video text recognition.
Yingying Jiang 0001, Shuli Yang, Pei Fu, Zhenbo Luo
ICDAR6
2017 Selecting Fine-Tuned Features for Layout Analysis of Historical Documents
abstract
In this paper, we investigate fine-tuned features learned by deep neural networks in the context of layout analysis. Pre-training and fine-tuning are techniques used in deep neural networks to learn representations (features) of input. However, it is not clear if the fine-tuned features are all useful for a following classification task. We investigate this problem using feature selection. Firstly, features are learned by a deep neural network, where stacked autoencoders are used for pre-training and then the whole network is fine-tuned. Then, a feature selection method is used to select relevant features for classification. We observe that despite fine-tuning, a significant number of the features are still redundant or irrelevant for layout classification. Furthermore, features from the top layer of the stacked autoencoders are generally more relevant for classification than those from lower layers.
Hao Wei 0001, Mathias Seuret, Marcus Liwicki, Rolf Ingold, Pei Fu
ICDAR5
2017 Deep Residual Text Detection Network for Scene Text
abstract
Scene text detection is a challenging problem in computer vision. In this paper, we propose a novel text detection network based on prevalent object detection frameworks. In order to obtain stronger semantic feature, we adopt ResNet as feature extraction layers and exploit multi-level feature by combining hierarchical convolutional networks. A vertical proposal mechanism is utilized to avoid proposal classification, while regression layer remains working to improve localization accuracy. Our approach evaluated on ICDAR2013 dataset achieves 0.91 F-measure, which outperforms previous state-of-the-art results in scene text detection.
Yingying Jiang 0001, Shuli Yang, Pei Fu, Zhenbo Luo
ICDAR6