Hangdi Xing

dblp:342/2712 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-1770-005XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Information extraction and text analysis · 37% Vision and language · 30% Trustworthy machine learning · 18%
Human-computer interaction and pervasive computing
2 papers
Accessibility and assistive technology · 54% User interface design and tools · 46%

Topics — the 11 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › multimodal understanding
multimodal document understanding
1.722025
Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding · EMNLP 2025
A Simple yet Effective Layout Token in Large Language Models for Document Understanding · CVPR 2025
Natural language and speech › Information extraction and text analysis
document understanding
1.522025
A Simple yet Effective Layout Token in Large Language Models for Document Understanding · CVPR 2025
LORE: Logical Location Regression Network for Table Structure Recognition · AAAI 2023
Natural language and speech › Language models and text generation
instruction tuning
0.912025
BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks · EMNLP 2025
Machine learning › Trustworthy machine learning › multimodal trustworthiness
multimodal knowledge conflict
0.912025
Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding · EMNLP 2025
Natural language and speech › Information extraction and text analysis › document analysis
document parsing
0.812024
DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing · EMNLP 2024
Computer vision › Vision and language › multimodal understanding
multimodal web understanding
0.712023
GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render Tree · EMNLP 2023
Natural language and speech › Information extraction and text analysis › document understanding › table recognition
table structure recognition
0.712023
LORE: Logical Location Regression Network for Table Structure Recognition · AAAI 2023
Natural language and speech › Information extraction and text analysis
web information extraction
0.712023
GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render Tree · EMNLP 2023
Computer vision › Vision and language › visual question answering
document visual question answering
0.312025
Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding · EMNLP 2025
Machine learning › Deep learning architectures and training
positional encoding
0.312025
A Simple yet Effective Layout Token in Large Language Models for Document Understanding · CVPR 2025
Natural language and speech › Question answering and dialogue systems › open-domain question answering
web question answering
0.212023
GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render Tree · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

large language model · 2.6pretraining objective · 0.9positional encoding · 0.9multimodal large language model · 0.9fine-tuning · 0.9OCR · 0.9multi-page document parsing · 0.8spatial location regression · 0.7logical location regression · 0.7gestalt psychology-inspired modeling · 0.7
YearPublicationVenuePosition
2025 A Simple yet Effective Layout Token in Large Language Models for Document Understanding
abstract
Recent methods that integrate spatial layouts with text for document understanding in large language models (LLMs) have shown promising results. A commonly used method is to represent layout information as text tokens and interleave them with text content as inputs to the LLMs. However, such a method still demonstrates limitations, as it requires additional position IDs for tokens that are used to represent layout information. Due to the constraint on max position IDs, assigning them to layout information reduces those available for text content, reducing the capacity for the model to learn from the text during training, while also introducing a large number of potentially untrained position IDs during long-context inference, which can hinder performance on document understanding tasks. To address these issues, we propose LayTokenLLM, a simple yet effective method for document understanding. LayTokenLLM represents layout information as a single token per text segment and uses a specialized positional encoding scheme. It shares position IDs between text and layout tokens, eliminating the need for additional position IDs. This design maintains the model’s capacity to learn from text while mitigating long-context issues during inference. Furthermore, a novel pre-training objective called Next Interleaved Text and Layout Token Prediction (NTLP) is devised to enhance cross-modality learning between text and layout tokens. Extensive experiments show that LayTokenLLM outperforms existing layout-integrated LLMs and MLLMs of similar scales on multi-page document understanding tasks, as well as most single-page tasks.
Zhaoqing Zhu, Chuwei Luo, Zirui Shao, Feiyu Gao, Hangdi Xing, Qi Zheng 0002, Ji Zhang 0011
CVPR5
2025 BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks
abstract
Tianyuan Huang, Zepeng Zhu, Hangdi Xing, Zirui Shao, Zhi Yu, Chaoxiong Yang, Jiaxian He, Xiaozhong Liu, Jiajun Bu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zepeng Zhu, Hangdi Xing, Zirui Shao, Chaoxiong Yang, Jiaxian He, Xiaozhong Liu 0001, Jiajun Bu
EMNLP3
2025 Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding
abstract
Multimodal large language models (MLLMs) have shown impressive capabilities in document understanding, a rapidly growing research area with significant industrial demand.As a multimodal task, document understanding requires models to possess both perceptual and cognitive abilities.However, due to different types of annotation noise in training, current MLLMs often face conflicts between perception and cognition.Taking a document VQA task (cognition) as an example, an MLLM might generate answers that do not match the corresponding visual content identified by its OCR (perception).This conflict suggests that the MLLM might struggle to establish an intrinsic connection between the information it "sees" and what it "understands".Such conflicts challenge the intuitive notion that cognition is consistent with perception, hindering the performance and explainability of MLLMs.In this paper, we define the conflicts between cognition and perception as Cognition and Perception (C&P) knowledge conflicts, a form of multimodal knowledge conflicts, and systematically assess them with a focus on document understanding.Our analysis reveals that even GPT-4o, a leading MLLM, achieves only 75.26% C&P consistency.To mitigate the C&P knowledge conflicts, we propose a novel method called Multimodal Knowledge Consistency Fine-tuning.Our method reduces C&P knowledge conflicts across all tested MLLMs and enhances their performance in both cognitive and perceptual tasks.
Zirui Shao, Feiyu Gao, Zhaoqing Zhu, Chuwei Luo, Hangdi Xing, Qi Zheng 0002, Ming Yan 0008, Jiajun Bu
EMNLP5
2025 LORE++: Logical location regression network for table structure recognition with pre-training
Rujiao Long, Hangdi Xing, Zhibo Yang 0003, Qi Zheng 0002, Fei Huang 0002, Cong Yao
Pattern Recognit.2
2024 WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation
Zirui Shao, Feiyu Gao, Hangdi Xing, Zepeng Zhu, Jiajun Bu, Qi Zheng 0002, Cong Yao
ECCV (8)3
2024 DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing
abstract
Parsing documents from pixels, such as pictures and scanned PDFs, into hierarchical structures is extensively demanded in the daily routines of data storage, retrieval and understanding.However, previously the research on this topic has been largely hindered since most existing datasets are small-scale, or contain documents of only a single type, which are characterized by a lack of document diversity.Moreover, there is a significant discrepancy in the annotation standards across datasets.In this paper, we introduce a large and diverse document hierarchy parsing (DHP) dataset to compensate for the data scarcity and inconsistency problem.We aim to set a new standard as a more practical, long-standing benchmark.Meanwhile, we present a new DHP framework designed to grasp both fine-grained text content and coarsegrained pattern at layout element level, enhancing the capacity of pre-trained text-layout models in handling the multi-page and multi-level challenges in DHP.Through exhaustive experiments, we validate the effectiveness of our proposed dataset and method 1 .
Hangdi Xing, Changxu Cheng, Feiyu Gao, Zirui Shao, Jiajun Bu, Qi Zheng 0002, Cong Yao
EMNLP1
2023 LORE: Logical Location Regression Network for Table Structure Recognition
abstract
Table structure recognition (TSR) aims at extracting tables in images into machine-understandable formats. Recent methods solve this problem by predicting the adjacency relations of detected cell boxes, or learning to generate the corresponding markup sequences from the table images. However, they either count on additional heuristic rules to recover the table structures, or require a huge amount of training data and time-consuming sequential decoders. In this paper, we propose an alternative paradigm. We model TSR as a logical location regression problem and propose a new TSR framework called LORE, standing for LOgical location REgression network, which for the first time combines logical location regression together with spatial location regression of table cells. Our proposed LORE is conceptually simpler, easier to train and more accurate than previous TSR models of other paradigms. Experiments on standard benchmarks demonstrate that LORE consistently outperforms prior arts. Code is available at https:// github.com/AlibabaResearch/AdvancedLiterateMachinery/tree/main/DocumentUnderstanding/LORE-TSR.
Hangdi Xing, Feiyu Gao, Rujiao Long, Jiajun Bu, Qi Zheng 0002, Liangcheng Li, Cong Yao
AAAI1
2023 GEM: Gestalt Enhanced Markup Language Model for Web Understanding via Render Tree
abstract
Inexhaustible web content carries abundant perceptible information beyond text.Unfortunately, most prior efforts in pre-trained Language Models (LMs) ignore such cyberrichness, while few of them only employ plain HTMLs, and crucial information in the rendered web, such as visual, layout, and style, are excluded.Intuitively, those perceptible web information can provide essential intelligence to facilitate content understanding tasks.This study presents an innovative Gestalt Enhanced Markup (GEM) Language Model inspired by Gestalt psychological theory for hosting heterogeneous visual information from the render tree into the language model without requiring additional visual input.Comprehensive experiments on multiple downstream tasks, i.e., web question answering and web information extraction, validate GEM superiority.
Zirui Shao, Feiyu Gao, Zhongda Qi, Hangdi Xing, Jiajun Bu, Qi Zheng 0002, Xiaozhong Liu 0001
EMNLP4