VLDB 2026 Research / reviewers in the wild / expert
Jiapeng Wang 0003
dblp:240/3481-3
· DBLP profile ↗
12ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-2060-3488ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document UnderstandingabstractText-rich document understanding (TDU) requires comprehensive analysis of documents containing substantial textual content and complex layouts. While Multimodal Large Language Models (MLLMs) have achieved fast progress in this domain, existing approaches either demand significant computational resources or struggle with effective multi-modal integration. In this paper, we introduce DocLayLLM, an efficient multi-modal extension of LLMs specifically designed for TDU. By lightly integrating visual patch tokens and 2D positional tokens into LLMs’ input and encoding the document content using the LLMs themselves, we fully take advantage of the document comprehension capability of LLMs and enhance their perception of OCR information. We have also deeply considered the role of chain-of-thought (CoT) and innovatively proposed the techniques of CoT Pre-training and CoT Annealing. Our DocLayLLM can achieve remarkable performances with lightweight training settings, showcasing its efficiency and effectiveness. Experimental results demonstrate that our DocLayLLM outperforms existing OCR-dependent methods and OCR-free competitors. Code and model are available at https://github.com/whlscut/DocLayLLM. Wenhui Liao, Jiapeng Wang 0003, Chengyu Wang 0001, Jun Huang 0007 |
CVPR | 2 |
| 2025 | Hallucination-Aware Prompt Optimization for Text-to-Video SynthesisabstractThe rapid advancements in AI-generated content (AIGC) have led to extensive research and application of deep text-to-video (T2V) synthesis models, such as OpenAI's Sora. These models typically rely on high-quality prompt-video pairs and detailed text prompts for model training in order to produce high-quality videos. To boost the effectiveness of Sora-like T2V models, we introduce VidPrompter, an innovative large multi-modal model supporting T2V applications with three key functionalities: (1) generating detailed prompts from raw videos, (2) enhancing prompts from videos grounded with short descriptions, and (3) refining simple user-provided prompts to elevate T2V video quality. We train VidPrompter using a hybrid multi-task paradigm and propose the hallucination-aware direct preference optimization (HDPO) technique to improve the multi-modal, multi-task prompt optimization process. Experiments on various tasks show our method surpasses strong baselines and other competitors. Jiapeng Wang 0003, Chengyu Wang 0001, Jun Huang 0007 |
IJCAI | 1 |
| 2025 | LiLTv2: Language-substitutable Layout-image Transformer for Visual Information ExtractionabstractVisual Information Extraction (VIE) has experienced substantial growth and heightened interest due to its pivotal role in intelligent document processing. However, most existing related pre-trained models typically can only process the data from a certain (set of) language(s)—often just English, representing a distinct limitation. To solve it, we present a L anguage-subst i tutable L ayout-image T ransformer (LiLTv2). It can be pre-trained just once on mono-lingual documents and then collaborate with off-the-shelf textual models in other languages during fine-tuning. Firstly, LiLTv2 utilizes a new dual-stream model architecture, one stream for substitutable text information and the other for layout and image information. Then, LiLTv2 has improved upon the optimization strategy and the diverse tasks adopted in the pre-training stage. Finally, we innovatively propose a teacher-student knowledge distillation learning with segment-level multi-modal features named SegKD. Extensive experimental results on widely used benchmarks can demonstrate the superior effectiveness of our method. Jiapeng Wang 0003, Zening Lin, Dayi Huang, Longfei Xiong |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP ModelsabstractContrastive Language-Image Pre-training (CLIP) has been widely studied and applied in numerous applications.However, the emphasis on brief summary texts during pre-training prevents CLIP from understanding long descriptions.This issue is particularly acute regarding videos given that videos often contain abundant detailed contents.In this paper, we propose the VideoCLIP-XL (eXtra Length) model, which aims to unleash the long-description understanding capability of video CLIP models.Firstly, we establish an automatic data collection system and gather a large-scale VILD pre-training dataset 1 with VIdeo and Long-Description pairs.Then, we propose Text-similarity-guided Primary Component Matching (TPCM) to better learn the distribution of feature space while expanding the long description capability.We also introduce two new tasks namely Detail-aware Description Ranking (DDR) and Hallucination-aware Description Ranking (HDR) for further understanding improvement.Finally, we construct a Long Video Description Ranking (LVDR) benchmark 2 for evaluating the long-description capability more comprehensively.Extensive experimental results on widely-used text-video retrieval benchmarks with both short and long descriptions and our LVDR benchmark can fully demonstrate the effectiveness of our method.3 Jiapeng Wang 0003, Chengyu Wang 0001, Kunzhe Huang, Jun Huang 0007 |
EMNLP | 1 |
| 2024 | PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair ExtractionabstractDocument pair extraction aims to identify key and value entities as well as their relationships from visually-rich documents. Most existing methods divide it into two separate tasks: semantic entity recognition (SER) and relation extraction (RE). However, simply concatenating SER and RE serially can lead to severe error propagation, and it fails to handle cases like multi-line entities in real scenarios. To address these issues, this paper introduces a novel framework, PEneo (Pair Extraction new decoder option), which performs document pair extraction in a unified pipeline, incorporating three concurrent sub-tasks: line extraction, line grouping, and entity linking. This approach alleviates the error accumulation problem and can handle the case of multi-line entities. Furthermore, to better evaluate the model's performance and to facilitate future research on pair extraction, we introduce RFUND, a re-annotated version of the commonly used FUNSD and XFUND datasets, to make them more accurate and cover realistic situations. Experiments on various benchmarks demonstrate PEneo's superiority over previous pipelines, boosting the performance by a large margin (e.g., 19.89%-22.91% F1 score on RFUND-EN) when combined with various backbones like LiLT and LayoutLMv3, showing its effectiveness and generality. Codes and the new annotations are available at https://github.com/ZeningLin/PEneo. Zening Lin, Jiapeng Wang 0003, Wenhui Liao, Dayi Huang, Longfei Xiong |
ACM Multimedia | 2 |
| 2023 | Towards Better Translations from Classical to Modern Chinese: A New Dataset and a New Method
Zongyuan Jiang, Jiapeng Wang 0003, Jiahuan Cao, Xue Gao |
NLPCC (1) | 2 |
| 2022 | CMT-Co: Contrastive Learning with Character Movement Task for Handwritten Text Recognition
Jiapeng Wang 0003, Yujin Ren, Yang Xue 0001 |
ACCV (7) | 2 |
| 2022 | LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document UnderstandingabstractStructured document understanding has attracted considerable attention and made significant progress recently, owing to its crucial role in intelligent document processing.However, most existing related models can only deal with the document data of specific language(s) (typically English) included in the pre-training collection, which is extremely limited.To address this issue, we propose a simple yet effective Language-independent Layout Transformer (LiLT) for structured document understanding.LiLT can be pre-trained on the structured documents of a single language and then directly fine-tuned on other languages with the corresponding off-the-shelf monolingual/multilingual pre-trained textual models.Experimental results on eight languages have shown that LiLT can achieve competitive or even superior performance on diverse widely-used downstream benchmarks, which enables language-independent benefit from the pre-training of document layout structure.Code and model are publicly available at https://github.com/jpWang/LiLT. Jiapeng Wang 0003, Kai Ding 0009 |
ACL (1) | 1 |
| 2022 | ChaCo: Character Contrastive Learning for Handwritten Text Recognition
Jiapeng Wang 0003, Canjie Luo, Yang Xue 0001 |
ICFHR | 3 |
| 2021 | Improving Machine Understanding of Human Intent in Charts
Sihang Wu, Canyu Xie, Guozhi Tang, Qianying Liao, Jiapeng Wang 0003, Bangdong Chen, Xinfeng Chang, Kai Ding 0009, Yichao Huang |
ICDAR (3) | 6 |
| 2021 | Tag, Copy or Predict: A Unified Weakly-Supervised Learning Framework for Visual Information Extraction using SequencesabstractVisual information extraction (VIE) has attracted increasing attention in recent years. The existing methods usually first organized optical character recognition (OCR) results in plain texts and then utilized token-level category annotations as supervision to train a sequence tagging model. However, it expends great annotation costs and may be exposed to label confusion, the OCR errors will also significantly affect the final performance. In this paper, we propose a unified weakly-supervised learning framework called TCPNet (Tag, Copy or Predict Network), which introduces 1) an efficient encoder to simultaneously model the semantic and layout information in 2D OCR results, 2) a weakly-supervised training method that utilizes only sequence-level supervision; and 3) a flexible and switchable decoder which contains two inference modes: one (Copy or Predict Mode) is to output key information sequences of different categories by copying a token from the input or predicting one in each time step, and the other (Tag Mode) is to directly tag the input sequence in a single forward pass. Our method shows new state-of-the-art performance on several public benchmarks, which fully proves its effectiveness. Jiapeng Wang 0003, Guozhi Tang, Weihong Ma, Kai Ding 0009, Yichao Huang |
IJCAI | 1 |
| 2020 | Precise detection of Chinese characters in historical documents with deep reinforcement learning
Sihang Wu, Jiapeng Wang 0003, Weihong Ma |
Pattern Recognit. | 2 |