EDBT 2026 Demo / reviewers in the wild / expert
Zening Lin
dblp:359/4747
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Information extraction and text analysis · 57% Vision and language · 33% Efficient and distributed learning · 10% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.0 | 1 | 2026 | URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding · AAAI 2026 |
Natural language and speech › Information extraction and text analysis › document analysis
document information extraction |
0.8 | 1 | 2024 | PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction · ACM Multimedia 2024 |
Natural language and speech › Information extraction and text analysis
entity linking |
0.8 | 1 | 2024 | PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction · ACM Multimedia 2024 |
Machine learning › Efficient and distributed learning › model compression
token compression |
0.3 | 1 | 2026 | URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding · AAAI 2026 |
Natural language and speech › Information extraction and text analysis › document analysis
document structure extraction |
0.2 | 1 | 2024 | PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction · ACM Multimedia 2024 |
Methods — techniques the papers use, named apart from their topics
evidence localization · 2.0cross-modal retrieval · 2.0transformer · 0.8LiLT · 0.8LayoutLMv3 · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document UnderstandingabstractRecent multimodal large language models (MLLMs) still struggle with long document understanding due to two fundamental challenges: information interference from abundant irrelevant content, and the quadratic computational cost of Transformer-based architectures. Existing approaches primarily fall into two categories: token compression, which sacrifices fine-grained details; and introducing external retrievers, which increase system complexity and prevent end-to-end optimization. To address these issues, we conduct an in-depth analysis and observe that MLLMs exhibit a human-like coarse-to-fine reasoning pattern: early Transformer layers attend broadly across the document, while deeper layers focus on relevant evidence pages. Motivated by this insight, we posit that the inherent evidence localization capabilities of MLLMs can be explicitly leveraged to perform retrieval during the reasoning process, facilitating efficient long document understanding. To this end, we propose URaG, a simple-yet-effective framework that Unifies Retrieval and Generation within a single MLLM. URaG introduces a lightweight cross-modal retrieval module that converts the early Transformer layers into an efficient evidence selector, identifying and preserving the most relevant pages while discarding irrelevant content. This design enables the deeper layers to concentrate computational resources on pertinent information, improving both accuracy and efficiency. Extensive experiments demonstrate that URaG achieves state-of-the-art performance while reducing computational overhead by 44-56%. Yongxin Shi, Zeyu Shan, Dezhi Peng, Zening Lin |
AAAI | 5 |
| 2025 | LiLTv2: Language-substitutable Layout-image Transformer for Visual Information ExtractionabstractVisual Information Extraction (VIE) has experienced substantial growth and heightened interest due to its pivotal role in intelligent document processing. However, most existing related pre-trained models typically can only process the data from a certain (set of) language(s)—often just English, representing a distinct limitation. To solve it, we present a L anguage-subst i tutable L ayout-image T ransformer (LiLTv2). It can be pre-trained just once on mono-lingual documents and then collaborate with off-the-shelf textual models in other languages during fine-tuning. Firstly, LiLTv2 utilizes a new dual-stream model architecture, one stream for substitutable text information and the other for layout and image information. Then, LiLTv2 has improved upon the optimization strategy and the diverse tasks adopted in the pre-training stage. Finally, we innovatively propose a teacher-student knowledge distillation learning with segment-level multi-modal features named SegKD. Extensive experimental results on widely used benchmarks can demonstrate the superior effectiveness of our method. Jiapeng Wang 0003, Zening Lin, Dayi Huang, Longfei Xiong |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | ROISER: Towards Real World Semantic Entity Recognition from Visually-Rich Documents
Zening Lin, Wenhui Liao, Weicong Dai, Longfei Xiong |
ICPR (31) | 1 |
| 2024 | PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair ExtractionabstractDocument pair extraction aims to identify key and value entities as well as their relationships from visually-rich documents. Most existing methods divide it into two separate tasks: semantic entity recognition (SER) and relation extraction (RE). However, simply concatenating SER and RE serially can lead to severe error propagation, and it fails to handle cases like multi-line entities in real scenarios. To address these issues, this paper introduces a novel framework, PEneo (Pair Extraction new decoder option), which performs document pair extraction in a unified pipeline, incorporating three concurrent sub-tasks: line extraction, line grouping, and entity linking. This approach alleviates the error accumulation problem and can handle the case of multi-line entities. Furthermore, to better evaluate the model's performance and to facilitate future research on pair extraction, we introduce RFUND, a re-annotated version of the commonly used FUNSD and XFUND datasets, to make them more accurate and cover realistic situations. Experiments on various benchmarks demonstrate PEneo's superiority over previous pipelines, boosting the performance by a large margin (e.g., 19.89%-22.91% F1 score on RFUND-EN) when combined with various backbones like LiLT and LayoutLMv3, showing its effectiveness and generality. Codes and the new annotations are available at https://github.com/ZeningLin/PEneo. Zening Lin, Jiapeng Wang 0003, Wenhui Liao, Dayi Huang, Longfei Xiong |
ACM Multimedia | 1 |