Qi Zheng 0002

dblp:52/1078-2 · DBLP profile ↗
← Back
16ranked-venue papers
0as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 8 since 2021
YearPublicationVenuePosition
2025 ProcTag: Process Tagging for Assessing the Efficacy of Document Instruction Data
abstract
Recently, large language models (LLMs) and multimodal large language models (MLLMs) have demonstrated promising results on document visual question answering (VQA) task, particularly after training on document instruction datasets. An effective evaluation method for document instruction data is crucial in constructing instruction data with high efficacy, which, in turn, facilitates the training of LLMs and MLLMs for document VQA. However, most existing evaluation methods for instruction data are limited to the textual content of the instructions themselves, thereby hindering the effective assessment of document instruction datasets and constraining their construction. In this paper, we propose ProcTag, a data-oriented method that assesses the efficacy of document instruction data. ProcTag innovatively performs tagging on the execution process of instructions rather than the instruction text itself. By leveraging the diversity and complexity of these tags to assess the efficacy of the given dataset, ProcTag enables selective sampling or filtering of document instructions. Furthermore, DocLayPrompt, a novel semi-structured layout-aware document prompting strategy, is proposed for effectively representing documents. Experiments demonstrate that sampling existing open-sourced and generated document VQA/instruction datasets with ProcTag significantly outperforms current methods for evaluating instruction data. Impressively, with ProcTag-based sampling in the generated document datasets, only 30.5 percent of the document instructions are required to achieve 100 percent efficacy compared to the complete dataset.
Yufan Shen, Chuwei Luo, Zhaoqing Zhu, Qi Zheng 0002, Jiajun Bu, Cong Yao
AAAI5
2025 A Simple yet Effective Layout Token in Large Language Models for Document Understanding
abstract
Recent methods that integrate spatial layouts with text for document understanding in large language models (LLMs) have shown promising results. A commonly used method is to represent layout information as text tokens and interleave them with text content as inputs to the LLMs. However, such a method still demonstrates limitations, as it requires additional position IDs for tokens that are used to represent layout information. Due to the constraint on max position IDs, assigning them to layout information reduces those available for text content, reducing the capacity for the model to learn from the text during training, while also introducing a large number of potentially untrained position IDs during long-context inference, which can hinder performance on document understanding tasks. To address these issues, we propose LayTokenLLM, a simple yet effective method for document understanding. LayTokenLLM represents layout information as a single token per text segment and uses a specialized positional encoding scheme. It shares position IDs between text and layout tokens, eliminating the need for additional position IDs. This design maintains the model’s capacity to learn from text while mitigating long-context issues during inference. Furthermore, a novel pre-training objective called Next Interleaved Text and Layout Token Prediction (NTLP) is devised to enhance cross-modality learning between text and layout tokens. Extensive experiments show that LayTokenLLM outperforms existing layout-integrated LLMs and MLLMs of similar scales on multi-page document understanding tasks, as well as most single-page tasks.
Zhaoqing Zhu, Chuwei Luo, Zirui Shao, Feiyu Gao, Hangdi Xing, Qi Zheng 0002, Ji Zhang 0011
CVPR6
2025 Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding
abstract
Multimodal large language models (MLLMs) have shown impressive capabilities in document understanding, a rapidly growing research area with significant industrial demand.As a multimodal task, document understanding requires models to possess both perceptual and cognitive abilities.However, due to different types of annotation noise in training, current MLLMs often face conflicts between perception and cognition.Taking a document VQA task (cognition) as an example, an MLLM might generate answers that do not match the corresponding visual content identified by its OCR (perception).This conflict suggests that the MLLM might struggle to establish an intrinsic connection between the information it "sees" and what it "understands".Such conflicts challenge the intuitive notion that cognition is consistent with perception, hindering the performance and explainability of MLLMs.In this paper, we define the conflicts between cognition and perception as Cognition and Perception (C&P) knowledge conflicts, a form of multimodal knowledge conflicts, and systematically assess them with a focus on document understanding.Our analysis reveals that even GPT-4o, a leading MLLM, achieves only 75.26% C&P consistency.To mitigate the C&P knowledge conflicts, we propose a novel method called Multimodal Knowledge Consistency Fine-tuning.Our method reduces C&P knowledge conflicts across all tested MLLMs and enhances their performance in both cognitive and perceptual tasks.
Zirui Shao, Feiyu Gao, Zhaoqing Zhu, Chuwei Luo, Hangdi Xing, Qi Zheng 0002, Ming Yan 0008, Jiajun Bu
EMNLP7
2025 Bi-VLDoc: bidirectional vision-language modeling for visually-rich document understanding
Chuwei Luo, Guozhi Tang, Qi Zheng 0002, Cong Yao, Yang Xue 0001, Luo Si
Int. J. Document Anal. Recognit.3
2025 LORE++: Logical location regression network for table structure recognition with pre-training
Rujiao Long, Hangdi Xing, Zhibo Yang 0003, Qi Zheng 0002, Fei Huang 0002, Cong Yao
Pattern Recognit.4
2024 LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
abstract
Recently, leveraging large language models (LLMs) or multimodal large language models (MLLMs) for document understanding has been proven very promising. However, previous works that employ LLMs/MLLMs for document understanding have not fully explored and utilized the document layout information, which is vital for precise document understanding. In this paper, we propose LayoutLLM, an LLM/MLLM based method for document understanding. The core of LayoutLLM is a layout instruction tuning strategy, which is specially designed to enhance the comprehension and utilization of document layouts. The proposed layout instruction tuning strategy consists of two components: Layout-aware Pre-training and Layout-aware Supervised Fine-tuning. To capture the characteristics of document layout in Layout-aware Pre-training, three groups of pretraining tasks, corresponding to document-level, region-level and segment-level information, are introduced. Furthermore, a novel module called layout chain-of-thought (LayoutCoT) is devised to enable LayoutLLM to focus on regions relevant to the question and generate accurate answers. LayoutCoT is effective for boosting the performance of document understanding. Meanwhile, it brings a certain degree of interpretability, which could facilitate manual inspection and correction. Experiments on standard benchmarks show that the proposed LayoutLLM significantly outperforms existing methods that adopt open-source 7B LLMs/MLLMs for document understanding.
Chuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng 0002, Cong Yao
CVPR4
2024 WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation
Zirui Shao, Feiyu Gao, Hangdi Xing, Zepeng Zhu, Jiajun Bu, Qi Zheng 0002, Cong Yao
ECCV (8)7
2024 DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing
abstract
Parsing documents from pixels, such as pictures and scanned PDFs, into hierarchical structures is extensively demanded in the daily routines of data storage, retrieval and understanding.However, previously the research on this topic has been largely hindered since most existing datasets are small-scale, or contain documents of only a single type, which are characterized by a lack of document diversity.Moreover, there is a significant discrepancy in the annotation standards across datasets.In this paper, we introduce a large and diverse document hierarchy parsing (DHP) dataset to compensate for the data scarcity and inconsistency problem.We aim to set a new standard as a more practical, long-standing benchmark.Meanwhile, we present a new DHP framework designed to grasp both fine-grained text content and coarsegrained pattern at layout element level, enhancing the capacity of pre-trained text-layout models in handling the multi-page and multi-level challenges in DHP.Through exhaustive experiments, we validate the effectiveness of our proposed dataset and method 1 .
Hangdi Xing, Changxu Cheng, Feiyu Gao, Zirui Shao, Jiajun Bu, Qi Zheng 0002, Cong Yao
EMNLP7
2023 LORE: Logical Location Regression Network for Table Structure Recognition
abstract
Table structure recognition (TSR) aims at extracting tables in images into machine-understandable formats. Recent methods solve this problem by predicting the adjacency relations of detected cell boxes, or learning to generate the corresponding markup sequences from the table images. However, they either count on additional heuristic rules to recover the table structures, or require a huge amount of training data and time-consuming sequential decoders. In this paper, we propose an alternative paradigm. We model TSR as a logical location regression problem and propose a new TSR framework called LORE, standing for LOgical location REgression network, which for the first time combines logical location regression together with spatial location regression of table cells. Our proposed LORE is conceptually simpler, easier to train and more accurate than previous TSR models of other paradigms. Experiments on standard benchmarks demonstrate that LORE consistently outperforms prior arts. Code is available at https:// github.com/AlibabaResearch/AdvancedLiterateMachinery/tree/main/DocumentUnderstanding/LORE-TSR.
Hangdi Xing, Feiyu Gao, Rujiao Long, Jiajun Bu, Qi Zheng 0002, Liangcheng Li, Cong Yao
AAAI5
2023 GeoLayoutLM: Geometric Pre-training for Visual Information Extraction
abstract
Visual information extraction (VIE) plays an important role in Document Intelligence. Generally, it is divided into two tasks: semantic entity recognition (SER) and relation extraction (RE). Recently, pre-trained models for documents have achieved substantial progress in VIE, particularly in SER. However, most of the existing models learn the geometric representation in an implicit way, which has been found insufficient for the RE task since geometric information is especially crucial for RE. Moreover, we reveal another factor that limits the performance of RE lies in the objective gap between the pre-training phase and the finetuning phase for RE. To tackle these issues, we propose in this paper a multi-modal framework, named GeoLayoutLM, for VIE. GeoLayoutLM explicitly models the geometric relations in pre-training, which we call geometric pre-training. Geometric pre-training is achieved by three specially designed geometry-related pre-training tasks. Additionally, novel relation heads, which are pre-trained by the geometric pre-training tasks and fine-tuned for RE, are elaborately designed to enrich and enhance the feature representation. According to extensive experiments on standard VIE benchmarks, GeoLayoutLM achieves highly competitive scores in the SER task and significantly outperforms the previous state-of-the-arts for RE (e.g., the F1 score of RE on FUNSD is boosted from 80.35% to 89.45%)11https://github.com/AlibabaResearch/AdvancedLiterateMachinery.
Chuwei Luo, Changxu Cheng, Qi Zheng 0002, Cong Yao
CVPR3
2023 GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render Tree
abstract
Inexhaustible web content carries abundant perceptible information beyond text.Unfortunately, most prior efforts in pre-trained Language Models (LMs) ignore such cyberrichness, while few of them only employ plain HTMLs, and crucial information in the rendered web, such as visual, layout, and style, are excluded.Intuitively, those perceptible web information can provide essential intelligence to facilitate content understanding tasks.This study presents an innovative Gestalt Enhanced Markup (GEM) Language Model inspired by Gestalt psychological theory for hosting heterogeneous visual information from the render tree into the language model without requiring additional visual input.Comprehensive experiments on multiple downstream tasks, i.e., web question answering and web information extraction, validate GEM superiority.
Zirui Shao, Feiyu Gao, Zhongda Qi, Hangdi Xing, Jiajun Bu, Qi Zheng 0002, Xiaozhong Liu 0001
EMNLP7
2023 LISTER: Neighbor Decoding for Length-Insensitive Scene Text Recognition
abstract
The diversity in length constitutes a significant characteristic of text. Due to the long-tail distribution of text lengths, most existing methods for scene text recognition (STR) only work well on short or seen-length text, lacking the capability of recognizing longer text or performing length extrapolation. This is a crucial issue, since the lengths of the text to be recognized are usually not given in advance in real-world applications, but it has not been adequately investigated in previous works. Therefore, we propose in this paper a method called Length-Insensitive Scene TExt Recognizer (LISTER), which remedies the limitation regarding the robustness to various text lengths. Specifically, a Neighbor Decoder is proposed to obtain accurate character attention maps with the assistance of a novel neighbor matrix regardless of the text lengths. Besides, a Feature Enhancement Module is devised to model the long-range dependency with low computation cost, which is able to perform iterations with the neighbor decoder to enhance the feature map progressively. To the best of our knowledge, we are the first to achieve effective length-insensitive scene text recognition. Extensive experiments demonstrate that the proposed LISTER algorithm exhibits obvious superiority on long text recognition and the ability for length extrapolation, while comparing favourably with the previous state-of-the-art methods on standard benchmarks for STR (mainly short text)1.
Changxu Cheng, Peng Wang 0103, Cheng Da, Qi Zheng 0002, Cong Yao
ICCV4
2023 Vision Grid Transformer for Document Layout Analysis
abstract
Document pre-trained models and grid-based models have proven to be very effective on various tasks in Document AI. However, for the document layout analysis (DLA) task, existing document pre-trained models, even those pretrained in a multi-modal fashion, usually rely on either textual features or visual features. Grid-based models for DLA are multi-modality but largely neglect the effect of pre-training. To fully leverage multi-modal information and exploit pre-training techniques to learn better representation for DLA, in this paper, we present VGT, a two-stream Vision Grid Transformer, in which Grid Transformer (GiT) is proposed and pre-trained for 2D token-level and segment-level semantic understanding. Furthermore, a new dataset named D4LA, which is so far the most diverse and detailed manually-annotated benchmark for document layout analysis, is curated and released. Experiment results have illustrated that the proposed VGT model achieves new state-of-the-art results on DLA tasks, e.g. PubLayNet (95.7%→96.2%), DocBank (79.6%→84.1%), and D4LA (67.7%→68.8%). The code and models as well as the D4LA dataset will be made publicly available1.
Cheng Da, Chuwei Luo, Qi Zheng 0002, Cong Yao
ICCV3
2020 An End-to-End OCR Text Re-organization Sequence Learning for Rich-Text Detail Image Comprehension
Liangcheng Li, Feiyu Gao, Jiajun Bu, Yongpan Wang, Qi Zheng 0002
ECCV (25)6
2019 SegLink++: Detecting Dense and Arbitrary-shaped Scene Text by Instance-aware Component Grouping
Jun Tang 0008, Zhibo Yang 0003, Yongpan Wang, Qi Zheng 0002, Yongchao Xu, Xiang Bai
Pattern Recognit.4
2018 ICPR2018 Contest on Robust Reading for Multi-Type Web Images
abstract
Electronic commerce has infiltrated every aspect of our daily lives, which offers great convenience for shopping, advertising, etc. Text in the web images is responsible to convey essential information for consumers. Algorithms that read text in these web images can facilitate applications of various types, such as goods surveillance, products classification, and intelligent retrieval or recommendation. Despite of various existing text reading tasks, this contest introduces a novel large-scale dataset named MTWI that contains 20,000 images, which is the first dataset that is mainly constructed by Chinese and English web text. Three tasks (web text recognition, web text detection, and end-to-end web text detection and recognition) were set up for encouraging more research on the web text reading problem. The contest was held from February 2, 2018 to May 26, 2018 with 289 valid submissions from 4,282 registered teams. Throughout this report, we describe the details of this new dataset, the purposes and definitions of the tasks, the evaluation protocols, and the summaries of the results.
Mengchao He, Zhibo Yang 0003, Sheng Zhang 0024, Canjie Luo, Feiyu Gao, Qi Zheng 0002, Yongpan Wang, Xin Zhang 0013
ICPR7