VLDB 2026 Research / reviewers in the wild / expert
Zirui Shao
dblp:280/7728
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Vision and language · 46% Information extraction and text analysis · 23% Trustworthy machine learning · 17% | |
| Human-computer interaction and pervasive computing
2 papers |
Accessibility and assistive technology · 54% User interface design and tools · 46% |
Topics — the 12 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › multimodal understanding
multimodal document understanding |
1.7 | 2 | 2025 | Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding · EMNLP 2025 A Simple yet Effective Layout Token in Large Language Models for Document Understanding · CVPR 2025 |
Natural language and speech › Information extraction and text analysis
document understanding |
0.9 | 1 | 2025 | A Simple yet Effective Layout Token in Large Language Models for Document Understanding · CVPR 2025 |
Computer vision › Vision and language › multimodal understanding
GUI understanding |
0.9 | 1 | 2025 | MP-GUI: Modality Perception with MLLMs for GUI Understanding · CVPR 2025 |
Natural language and speech › Language models and text generation
instruction tuning |
0.9 | 1 | 2025 | BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks · EMNLP 2025 |
Machine learning › Trustworthy machine learning › multimodal trustworthiness
multimodal knowledge conflict |
0.9 | 1 | 2025 | Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding · EMNLP 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | MP-GUI: Modality Perception with MLLMs for GUI Understanding · CVPR 2025 |
Natural language and speech › Information extraction and text analysis › document analysis
document parsing |
0.8 | 1 | 2024 | DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing · EMNLP 2024 |
Computer vision › Vision and language › multimodal understanding
multimodal web understanding |
0.7 | 1 | 2023 | GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render Tree · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis
web information extraction |
0.7 | 1 | 2023 | GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render Tree · EMNLP 2023 |
Computer vision › Vision and language › visual question answering
document visual question answering |
0.3 | 1 | 2025 | Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding · EMNLP 2025 |
Machine learning › Deep learning architectures and training
positional encoding |
0.3 | 1 | 2025 | A Simple yet Effective Layout Token in Large Language Models for Document Understanding · CVPR 2025 |
Natural language and speech › Question answering and dialogue systems › open-domain question answering
web question answering |
0.2 | 1 | 2023 | GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render Tree · EMNLP 2023 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.6pretraining objective · 0.9positional encoding · 0.9multimodal large language model · 0.9modality perception · 0.9fusion gate · 0.9fine-tuning · 0.9automatic data collection · 0.9OCR · 0.9multi-page document parsing · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MP-GUI: Modality Perception with MLLMs for GUI UnderstandingabstractGraphical user interface (GUI) has become integral to modern society, making it crucial to be understood for human-centric systems. However, unlike natural images or documents, GUIs comprise artificially designed graphical elements arranged to convey specific semantic meanings. Current multi-modal large language models (MLLMs) already proficient in processing graphical and textual components suffer from hurdles in GUI understanding due to the lack of explicit spatial structure modeling. Moreover, obtaining high-quality spatial structure data is challenging due to privacy issues and noisy environments. To address these challenges, we present MP-GUI, a specially designed MLLM for GUI understanding. MP-GUI features three precisely specialized perceivers to extract graphical, textual, and spatial modalities from the screen as GUI-tailored visual clues, with spatial structure refinement strategy and adaptively combined via a fusion gate to meet the specific preferences of different GUI understanding tasks. To cope with the scarcity of training data, we also introduce a pipeline for automatically data collecting. Extensive experiments demonstrate that MP-GUI achieves impressive results on various GUI understanding tasks with limited data. Our codes and datasets are publicly available at https://github.com/BigTaige/MP-GUI. Weizhi Chen, Leyang Yang, Sheng Zhou 0004, Shengchu Zhao, Hanbei Zhan, Jiongchao Jin, Liangcheng Li, Zirui Shao, Jiajun Bu |
CVPR | 9 |
| 2025 | A Simple yet Effective Layout Token in Large Language Models for Document UnderstandingabstractRecent methods that integrate spatial layouts with text for document understanding in large language models (LLMs) have shown promising results. A commonly used method is to represent layout information as text tokens and interleave them with text content as inputs to the LLMs. However, such a method still demonstrates limitations, as it requires additional position IDs for tokens that are used to represent layout information. Due to the constraint on max position IDs, assigning them to layout information reduces those available for text content, reducing the capacity for the model to learn from the text during training, while also introducing a large number of potentially untrained position IDs during long-context inference, which can hinder performance on document understanding tasks. To address these issues, we propose LayTokenLLM, a simple yet effective method for document understanding. LayTokenLLM represents layout information as a single token per text segment and uses a specialized positional encoding scheme. It shares position IDs between text and layout tokens, eliminating the need for additional position IDs. This design maintains the model’s capacity to learn from text while mitigating long-context issues during inference. Furthermore, a novel pre-training objective called Next Interleaved Text and Layout Token Prediction (NTLP) is devised to enhance cross-modality learning between text and layout tokens. Extensive experiments show that LayTokenLLM outperforms existing layout-integrated LLMs and MLLMs of similar scales on multi-page document understanding tasks, as well as most single-page tasks. Zhaoqing Zhu, Chuwei Luo, Zirui Shao, Feiyu Gao, Hangdi Xing, Qi Zheng 0002, Ji Zhang 0011 |
CVPR | 3 |
| 2025 | BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain TasksabstractTianyuan Huang, Zepeng Zhu, Hangdi Xing, Zirui Shao, Zhi Yu, Chaoxiong Yang, Jiaxian He, Xiaozhong Liu, Jiajun Bu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zepeng Zhu, Hangdi Xing, Zirui Shao, Chaoxiong Yang, Jiaxian He, Xiaozhong Liu 0001, Jiajun Bu |
EMNLP | 4 |
| 2025 | Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document UnderstandingabstractMultimodal large language models (MLLMs) have shown impressive capabilities in document understanding, a rapidly growing research area with significant industrial demand.As a multimodal task, document understanding requires models to possess both perceptual and cognitive abilities.However, due to different types of annotation noise in training, current MLLMs often face conflicts between perception and cognition.Taking a document VQA task (cognition) as an example, an MLLM might generate answers that do not match the corresponding visual content identified by its OCR (perception).This conflict suggests that the MLLM might struggle to establish an intrinsic connection between the information it "sees" and what it "understands".Such conflicts challenge the intuitive notion that cognition is consistent with perception, hindering the performance and explainability of MLLMs.In this paper, we define the conflicts between cognition and perception as Cognition and Perception (C&P) knowledge conflicts, a form of multimodal knowledge conflicts, and systematically assess them with a focus on document understanding.Our analysis reveals that even GPT-4o, a leading MLLM, achieves only 75.26% C&P consistency.To mitigate the C&P knowledge conflicts, we propose a novel method called Multimodal Knowledge Consistency Fine-tuning.Our method reduces C&P knowledge conflicts across all tested MLLMs and enhances their performance in both cognitive and perceptual tasks. Zirui Shao, Feiyu Gao, Zhaoqing Zhu, Chuwei Luo, Hangdi Xing, Qi Zheng 0002, Ming Yan 0008, Jiajun Bu |
EMNLP | 1 |
| 2024 | WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation
Zirui Shao, Feiyu Gao, Hangdi Xing, Zepeng Zhu, Jiajun Bu, Qi Zheng 0002, Cong Yao |
ECCV (8) | 1 |
| 2024 | DocHieNet: A Large and Diverse Dataset for Document Hierarchy ParsingabstractParsing documents from pixels, such as pictures and scanned PDFs, into hierarchical structures is extensively demanded in the daily routines of data storage, retrieval and understanding.However, previously the research on this topic has been largely hindered since most existing datasets are small-scale, or contain documents of only a single type, which are characterized by a lack of document diversity.Moreover, there is a significant discrepancy in the annotation standards across datasets.In this paper, we introduce a large and diverse document hierarchy parsing (DHP) dataset to compensate for the data scarcity and inconsistency problem.We aim to set a new standard as a more practical, long-standing benchmark.Meanwhile, we present a new DHP framework designed to grasp both fine-grained text content and coarsegrained pattern at layout element level, enhancing the capacity of pre-trained text-layout models in handling the multi-page and multi-level challenges in DHP.Through exhaustive experiments, we validate the effectiveness of our proposed dataset and method 1 . Hangdi Xing, Changxu Cheng, Feiyu Gao, Zirui Shao, Jiajun Bu, Qi Zheng 0002, Cong Yao |
EMNLP | 4 |
| 2023 | GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render TreeabstractInexhaustible web content carries abundant perceptible information beyond text.Unfortunately, most prior efforts in pre-trained Language Models (LMs) ignore such cyberrichness, while few of them only employ plain HTMLs, and crucial information in the rendered web, such as visual, layout, and style, are excluded.Intuitively, those perceptible web information can provide essential intelligence to facilitate content understanding tasks.This study presents an innovative Gestalt Enhanced Markup (GEM) Language Model inspired by Gestalt psychological theory for hosting heterogeneous visual information from the render tree into the language model without requiring additional visual input.Comprehensive experiments on multiple downstream tasks, i.e., web question answering and web information extraction, validate GEM superiority. Zirui Shao, Feiyu Gao, Zhongda Qi, Hangdi Xing, Jiajun Bu, Qi Zheng 0002, Xiaozhong Liu 0001 |
EMNLP | 1 |