EDBT 2026 Demo / reviewers in the wild / expert
Hangdi Xing
dblp:342/2712
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-1770-005XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Information extraction and text analysis · 37% Vision and language · 30% Trustworthy machine learning · 18% | |
| Human-computer interaction and pervasive computing
2 papers |
Accessibility and assistive technology · 54% User interface design and tools · 46% |
Topics — the 11 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › multimodal understanding
multimodal document understanding |
1.7 | 2 | 2025 | Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding · EMNLP 2025 A Simple yet Effective Layout Token in Large Language Models for Document Understanding · CVPR 2025 |
Natural language and speech › Information extraction and text analysis
document understanding |
1.5 | 2 | 2025 | A Simple yet Effective Layout Token in Large Language Models for Document Understanding · CVPR 2025 LORE: Logical Location Regression Network for Table Structure Recognition · AAAI 2023 |
Natural language and speech › Language models and text generation
instruction tuning |
0.9 | 1 | 2025 | BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks · EMNLP 2025 |
Machine learning › Trustworthy machine learning › multimodal trustworthiness
multimodal knowledge conflict |
0.9 | 1 | 2025 | Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis › document analysis
document parsing |
0.8 | 1 | 2024 | DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing · EMNLP 2024 |
Computer vision › Vision and language › multimodal understanding
multimodal web understanding |
0.7 | 1 | 2023 | GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render Tree · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis › document understanding › table recognition
table structure recognition |
0.7 | 1 | 2023 | LORE: Logical Location Regression Network for Table Structure Recognition · AAAI 2023 |
Natural language and speech › Information extraction and text analysis
web information extraction |
0.7 | 1 | 2023 | GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render Tree · EMNLP 2023 |
Computer vision › Vision and language › visual question answering
document visual question answering |
0.3 | 1 | 2025 | Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding · EMNLP 2025 |
Machine learning › Deep learning architectures and training
positional encoding |
0.3 | 1 | 2025 | A Simple yet Effective Layout Token in Large Language Models for Document Understanding · CVPR 2025 |
Natural language and speech › Question answering and dialogue systems › open-domain question answering
web question answering |
0.2 | 1 | 2023 | GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render Tree · EMNLP 2023 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.6pretraining objective · 0.9positional encoding · 0.9multimodal large language model · 0.9fine-tuning · 0.9OCR · 0.9multi-page document parsing · 0.8spatial location regression · 0.7logical location regression · 0.7gestalt psychology-inspired modeling · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Simple yet Effective Layout Token in Large Language Models for Document UnderstandingabstractRecent methods that integrate spatial layouts with text for document understanding in large language models (LLMs) have shown promising results. A commonly used method is to represent layout information as text tokens and interleave them with text content as inputs to the LLMs. However, such a method still demonstrates limitations, as it requires additional position IDs for tokens that are used to represent layout information. Due to the constraint on max position IDs, assigning them to layout information reduces those available for text content, reducing the capacity for the model to learn from the text during training, while also introducing a large number of potentially untrained position IDs during long-context inference, which can hinder performance on document understanding tasks. To address these issues, we propose LayTokenLLM, a simple yet effective method for document understanding. LayTokenLLM represents layout information as a single token per text segment and uses a specialized positional encoding scheme. It shares position IDs between text and layout tokens, eliminating the need for additional position IDs. This design maintains the model’s capacity to learn from text while mitigating long-context issues during inference. Furthermore, a novel pre-training objective called Next Interleaved Text and Layout Token Prediction (NTLP) is devised to enhance cross-modality learning between text and layout tokens. Extensive experiments show that LayTokenLLM outperforms existing layout-integrated LLMs and MLLMs of similar scales on multi-page document understanding tasks, as well as most single-page tasks. Zhaoqing Zhu, Chuwei Luo, Zirui Shao, Feiyu Gao, Hangdi Xing, Qi Zheng 0002, Ji Zhang 0011 |
CVPR | 5 |
| 2025 | BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain TasksabstractTianyuan Huang, Zepeng Zhu, Hangdi Xing, Zirui Shao, Zhi Yu, Chaoxiong Yang, Jiaxian He, Xiaozhong Liu, Jiajun Bu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zepeng Zhu, Hangdi Xing, Zirui Shao, Chaoxiong Yang, Jiaxian He, Xiaozhong Liu 0001, Jiajun Bu |
EMNLP | 3 |
| 2025 | Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document UnderstandingabstractMultimodal large language models (MLLMs) have shown impressive capabilities in document understanding, a rapidly growing research area with significant industrial demand.As a multimodal task, document understanding requires models to possess both perceptual and cognitive abilities.However, due to different types of annotation noise in training, current MLLMs often face conflicts between perception and cognition.Taking a document VQA task (cognition) as an example, an MLLM might generate answers that do not match the corresponding visual content identified by its OCR (perception).This conflict suggests that the MLLM might struggle to establish an intrinsic connection between the information it "sees" and what it "understands".Such conflicts challenge the intuitive notion that cognition is consistent with perception, hindering the performance and explainability of MLLMs.In this paper, we define the conflicts between cognition and perception as Cognition and Perception (C&P) knowledge conflicts, a form of multimodal knowledge conflicts, and systematically assess them with a focus on document understanding.Our analysis reveals that even GPT-4o, a leading MLLM, achieves only 75.26% C&P consistency.To mitigate the C&P knowledge conflicts, we propose a novel method called Multimodal Knowledge Consistency Fine-tuning.Our method reduces C&P knowledge conflicts across all tested MLLMs and enhances their performance in both cognitive and perceptual tasks. Zirui Shao, Feiyu Gao, Zhaoqing Zhu, Chuwei Luo, Hangdi Xing, Qi Zheng 0002, Ming Yan 0008, Jiajun Bu |
EMNLP | 5 |
| 2025 | LORE++: Logical location regression network for table structure recognition with pre-training
Rujiao Long, Hangdi Xing, Zhibo Yang 0003, Qi Zheng 0002, Fei Huang 0002, Cong Yao |
Pattern Recognit. | 2 |
| 2024 | WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation
Zirui Shao, Feiyu Gao, Hangdi Xing, Zepeng Zhu, Jiajun Bu, Qi Zheng 0002, Cong Yao |
ECCV (8) | 3 |
| 2024 | DocHieNet: A Large and Diverse Dataset for Document Hierarchy ParsingabstractParsing documents from pixels, such as pictures and scanned PDFs, into hierarchical structures is extensively demanded in the daily routines of data storage, retrieval and understanding.However, previously the research on this topic has been largely hindered since most existing datasets are small-scale, or contain documents of only a single type, which are characterized by a lack of document diversity.Moreover, there is a significant discrepancy in the annotation standards across datasets.In this paper, we introduce a large and diverse document hierarchy parsing (DHP) dataset to compensate for the data scarcity and inconsistency problem.We aim to set a new standard as a more practical, long-standing benchmark.Meanwhile, we present a new DHP framework designed to grasp both fine-grained text content and coarsegrained pattern at layout element level, enhancing the capacity of pre-trained text-layout models in handling the multi-page and multi-level challenges in DHP.Through exhaustive experiments, we validate the effectiveness of our proposed dataset and method 1 . Hangdi Xing, Changxu Cheng, Feiyu Gao, Zirui Shao, Jiajun Bu, Qi Zheng 0002, Cong Yao |
EMNLP | 1 |
| 2023 | LORE: Logical Location Regression Network for Table Structure RecognitionabstractTable structure recognition (TSR) aims at extracting tables in images into machine-understandable formats. Recent methods solve this problem by predicting the adjacency relations of detected cell boxes, or learning to generate the corresponding markup sequences from the table images. However, they either count on additional heuristic rules to recover the table structures, or require a huge amount of training data and time-consuming sequential decoders. In this paper, we propose an alternative paradigm. We model TSR as a logical location regression problem and propose a new TSR framework called LORE, standing for LOgical location REgression network, which for the first time combines logical location regression together with spatial location regression of table cells. Our proposed LORE is conceptually simpler, easier to train and more accurate than previous TSR models of other paradigms. Experiments on standard benchmarks demonstrate that LORE consistently outperforms prior arts. Code is available at https:// github.com/AlibabaResearch/AdvancedLiterateMachinery/tree/main/DocumentUnderstanding/LORE-TSR. Hangdi Xing, Feiyu Gao, Rujiao Long, Jiajun Bu, Qi Zheng 0002, Liangcheng Li, Cong Yao |
AAAI | 1 |
| 2023 | GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render TreeabstractInexhaustible web content carries abundant perceptible information beyond text.Unfortunately, most prior efforts in pre-trained Language Models (LMs) ignore such cyberrichness, while few of them only employ plain HTMLs, and crucial information in the rendered web, such as visual, layout, and style, are excluded.Intuitively, those perceptible web information can provide essential intelligence to facilitate content understanding tasks.This study presents an innovative Gestalt Enhanced Markup (GEM) Language Model inspired by Gestalt psychological theory for hosting heterogeneous visual information from the render tree into the language model without requiring additional visual input.Comprehensive experiments on multiple downstream tasks, i.e., web question answering and web information extraction, validate GEM superiority. Zirui Shao, Feiyu Gao, Zhongda Qi, Hangdi Xing, Jiajun Bu, Qi Zheng 0002, Xiaozhong Liu 0001 |
EMNLP | 4 |