VLDB 2026 Research / reviewers in the wild / expert
Feiyu Gao
dblp:34/9056
· DBLP profile ↗
12ranked-venue papers
0as first author
9since 2021 · last 2026
0009-0009-3206-5347ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SceneVTG++: Controllable Multilingual Visual Text Generation in the WildabstractGenerating visual text in natural scene images is a challenging task with many unsolved problems. Different from generating text on artificially designed images (such as posters, covers, and cartoons), existing methods for natural scene visual text generation still have significant deficiencies: methods based on rendering engines rely on manually crafted rules, which struggle to adapt to diverse backgrounds and leave obvious artificial traces, while their text layouts may be placed in unreasonable areas (e.g., sky or ground) and text content is semantically disconnected from the scene; diffusion model-based methods, on the other hand, face difficulties in generating small characters, depend on manually designed prompts to ensure reasonable layout and content, fail to generate text at precise locations, and cannot effectively control text attributes (e.g., font and color). In this paper, we propose a two-stage method named SceneVTG++ to address these issues. SceneVTG++ comprises two core components: a Text Layout and Content Generator (TLCG) and a Controllable Local Text Diffusion (CLTD). The former leverages the world knowledge and visual reasoning capabilities of multimodal large language models to identify reasonable text areas and recommend scene-relevant text content based on natural scene background images; the latter generates controllable multilingual text using a diffusion model, ensuring alignment with the outputs of TLCG. Through extensive experiments, we verified the effectiveness of both TLCG and CLTD, and demonstrated that SceneVTG++ achieves state-of-the-art performance in natural scene visual text generation. Additionally, the images generated by SceneVTG++ exhibit superior utility for training natural scene optical character recognition (OCR) tasks, including text detection and text recognition. Codes and datasets will be made publicly available. Jiawei Liu 0006, Feiyu Gao, Zhibo Yang 0003, Peng Wang 0028, Junyang Lin, Xinggang Wang, Wenyu Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | A Simple yet Effective Layout Token in Large Language Models for Document UnderstandingabstractRecent methods that integrate spatial layouts with text for document understanding in large language models (LLMs) have shown promising results. A commonly used method is to represent layout information as text tokens and interleave them with text content as inputs to the LLMs. However, such a method still demonstrates limitations, as it requires additional position IDs for tokens that are used to represent layout information. Due to the constraint on max position IDs, assigning them to layout information reduces those available for text content, reducing the capacity for the model to learn from the text during training, while also introducing a large number of potentially untrained position IDs during long-context inference, which can hinder performance on document understanding tasks. To address these issues, we propose LayTokenLLM, a simple yet effective method for document understanding. LayTokenLLM represents layout information as a single token per text segment and uses a specialized positional encoding scheme. It shares position IDs between text and layout tokens, eliminating the need for additional position IDs. This design maintains the model’s capacity to learn from text while mitigating long-context issues during inference. Furthermore, a novel pre-training objective called Next Interleaved Text and Layout Token Prediction (NTLP) is devised to enhance cross-modality learning between text and layout tokens. Extensive experiments show that LayTokenLLM outperforms existing layout-integrated LLMs and MLLMs of similar scales on multi-page document understanding tasks, as well as most single-page tasks. Zhaoqing Zhu, Chuwei Luo, Zirui Shao, Feiyu Gao, Hangdi Xing, Qi Zheng 0002, Ji Zhang 0011 |
CVPR | 4 |
| 2025 | Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document UnderstandingabstractMultimodal large language models (MLLMs) have shown impressive capabilities in document understanding, a rapidly growing research area with significant industrial demand.As a multimodal task, document understanding requires models to possess both perceptual and cognitive abilities.However, due to different types of annotation noise in training, current MLLMs often face conflicts between perception and cognition.Taking a document VQA task (cognition) as an example, an MLLM might generate answers that do not match the corresponding visual content identified by its OCR (perception).This conflict suggests that the MLLM might struggle to establish an intrinsic connection between the information it "sees" and what it "understands".Such conflicts challenge the intuitive notion that cognition is consistent with perception, hindering the performance and explainability of MLLMs.In this paper, we define the conflicts between cognition and perception as Cognition and Perception (C&P) knowledge conflicts, a form of multimodal knowledge conflicts, and systematically assess them with a focus on document understanding.Our analysis reveals that even GPT-4o, a leading MLLM, achieves only 75.26% C&P consistency.To mitigate the C&P knowledge conflicts, we propose a novel method called Multimodal Knowledge Consistency Fine-tuning.Our method reduces C&P knowledge conflicts across all tested MLLMs and enhances their performance in both cognitive and perceptual tasks. Zirui Shao, Feiyu Gao, Zhaoqing Zhu, Chuwei Luo, Hangdi Xing, Qi Zheng 0002, Ming Yan 0008, Jiajun Bu |
EMNLP | 2 |
| 2024 | WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation
Zirui Shao, Feiyu Gao, Hangdi Xing, Zepeng Zhu, Jiajun Bu, Qi Zheng 0002, Cong Yao |
ECCV (8) | 2 |
| 2024 | Visual Text Generation in the Wild
Jiawei Liu 0006, Feiyu Gao, Wenyu Liu 0001, Xinggang Wang, Peng Wang 0028, Fei Huang 0002, Cong Yao, Zhibo Yang 0003 |
ECCV (53) | 3 |
| 2024 | DocHieNet: A Large and Diverse Dataset for Document Hierarchy ParsingabstractParsing documents from pixels, such as pictures and scanned PDFs, into hierarchical structures is extensively demanded in the daily routines of data storage, retrieval and understanding.However, previously the research on this topic has been largely hindered since most existing datasets are small-scale, or contain documents of only a single type, which are characterized by a lack of document diversity.Moreover, there is a significant discrepancy in the annotation standards across datasets.In this paper, we introduce a large and diverse document hierarchy parsing (DHP) dataset to compensate for the data scarcity and inconsistency problem.We aim to set a new standard as a more practical, long-standing benchmark.Meanwhile, we present a new DHP framework designed to grasp both fine-grained text content and coarsegrained pattern at layout element level, enhancing the capacity of pre-trained text-layout models in handling the multi-page and multi-level challenges in DHP.Through exhaustive experiments, we validate the effectiveness of our proposed dataset and method 1 . Hangdi Xing, Changxu Cheng, Feiyu Gao, Zirui Shao, Jiajun Bu, Qi Zheng 0002, Cong Yao |
EMNLP | 3 |
| 2023 | LORE: Logical Location Regression Network for Table Structure RecognitionabstractTable structure recognition (TSR) aims at extracting tables in images into machine-understandable formats. Recent methods solve this problem by predicting the adjacency relations of detected cell boxes, or learning to generate the corresponding markup sequences from the table images. However, they either count on additional heuristic rules to recover the table structures, or require a huge amount of training data and time-consuming sequential decoders. In this paper, we propose an alternative paradigm. We model TSR as a logical location regression problem and propose a new TSR framework called LORE, standing for LOgical location REgression network, which for the first time combines logical location regression together with spatial location regression of table cells. Our proposed LORE is conceptually simpler, easier to train and more accurate than previous TSR models of other paradigms. Experiments on standard benchmarks demonstrate that LORE consistently outperforms prior arts. Code is available at https:// github.com/AlibabaResearch/AdvancedLiterateMachinery/tree/main/DocumentUnderstanding/LORE-TSR. Hangdi Xing, Feiyu Gao, Rujiao Long, Jiajun Bu, Qi Zheng 0002, Liangcheng Li, Cong Yao |
AAAI | 2 |
| 2023 | GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render TreeabstractInexhaustible web content carries abundant perceptible information beyond text.Unfortunately, most prior efforts in pre-trained Language Models (LMs) ignore such cyberrichness, while few of them only employ plain HTMLs, and crucial information in the rendered web, such as visual, layout, and style, are excluded.Intuitively, those perceptible web information can provide essential intelligence to facilitate content understanding tasks.This study presents an innovative Gestalt Enhanced Markup (GEM) Language Model inspired by Gestalt psychological theory for hosting heterogeneous visual information from the render tree into the language model without requiring additional visual input.Comprehensive experiments on multiple downstream tasks, i.e., web question answering and web information extraction, validate GEM superiority. Zirui Shao, Feiyu Gao, Zhongda Qi, Hangdi Xing, Jiajun Bu, Qi Zheng 0002, Xiaozhong Liu 0001 |
EMNLP | 2 |
| 2021 | Parsing Table Structures in the WildabstractThis paper tackles the problem of table structure parsing (TSP) from images in the wild. In contrast to existing studies that mainly focus on parsing well-aligned tabular images with simple layouts from scanned PDF documents, we aim to establish a practical table structure parsing system for real-world scenarios where tabular input images are taken or scanned with severe deformation, bending or occlusions. For designing such a system, we propose an approach named Cycle-CenterNet on the top of CenterNet with a novel cycle-pairing module to simultaneously detect and group tabular cells into structured tables. In the cycle-pairing module, a new pairing loss function is proposed for the network training. Alongside with our Cycle-CenterNet, we also present a large-scale dataset, named Wired Table in the Wild (WTW), which includes well-annotated structure parsing of multiple style tables in several scenes like photo, scanning files, web pages, etc.. In experiments, we demonstrate that our Cycle-CenterNet consistently achieves the best accuracy of table structure parsing on the new WTW dataset by 24.6% absolute improvement evaluated by the TEDS metric. A more comprehensive experimental analysis also validates the advantages of our proposed methods for the TSP task. Rujiao Long, Nan Xue 0001, Feiyu Gao, Zhibo Yang 0003, Yongpan Wang, Gui-Song Xia |
ICCV | 4 |
| 2020 | An End-to-End OCR Text Re-organization Sequence Learning for Rich-Text Detail Image Comprehension
Liangcheng Li, Feiyu Gao, Jiajun Bu, Yongpan Wang, Qi Zheng 0002 |
ECCV (25) | 2 |
| 2018 | ICPR2018 Contest on Robust Reading for Multi-Type Web ImagesabstractElectronic commerce has infiltrated every aspect of our daily lives, which offers great convenience for shopping, advertising, etc. Text in the web images is responsible to convey essential information for consumers. Algorithms that read text in these web images can facilitate applications of various types, such as goods surveillance, products classification, and intelligent retrieval or recommendation. Despite of various existing text reading tasks, this contest introduces a novel large-scale dataset named MTWI that contains 20,000 images, which is the first dataset that is mainly constructed by Chinese and English web text. Three tasks (web text recognition, web text detection, and end-to-end web text detection and recognition) were set up for encouraging more research on the web text reading problem. The contest was held from February 2, 2018 to May 26, 2018 with 289 valid submissions from 4,282 registered teams. Throughout this report, we describe the details of this new dataset, the purposes and definitions of the tasks, the evaluation protocols, and the summaries of the results. Mengchao He, Zhibo Yang 0003, Sheng Zhang 0024, Canjie Luo, Feiyu Gao, Qi Zheng 0002, Yongpan Wang, Xin Zhang 0013 |
ICPR | 6 |
| 2010 | Accelerating Strategy for Functional Test of NoC Communication FabricabstractTesting of network-on-chip (NoC) communication fabric draws more and more concern recently. Previous research show that the functional test of NoC can achieve nearly 100% structural fault coverage for the network switches, and the test approach based on functional test can be applied at full operation speed which original scan-based test approaches are hard to reach. The main novel contribution of this paper is an accelerating strategy for the NoC functional testing, including the “throughway” design of switches and its corresponding test configurations. The main concept is to organize the test configurations so that after boundary switches are proved to be fault-free, they could be reused as TAM of the other inner switches. Experimental results shows that comparing to some prior work, the proposed recursive test scheme using the accelerating strategy could reduce the complexity of test application time from O(N2) to O(N). The test scheme also has lower ambiguity of detected faults with same high fault-coverage. Experimental results show the proposed structure has low area cost and power consumption. Feiyu Gao |
Asian Test Symposium | 5 |