VLDB 2026 Research / reviewers in the wild / expert
Pengfei Hu 0006
dblp:71/9969-6
· DBLP profile ↗
23ranked-venue papers
4as first author
23since 2021 · last 2026
0009-0005-3345-605XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | See then tell: Enhancing key information extraction with vision grounding
Shuhang Liu, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Qing Wang 0008, Jianshu Zhang 0001 |
Neurocomputing | 3 |
| 2026 | Two-stage decomposition network for handwritten Chinese character error correction
Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jianshu Zhang 0001, Jianqing Gao, Qingfeng Liu |
Pattern Recognit. | 1 |
| 2025 | RFL: Simplifying Chemical Structure Recognition with Ring-Free LanguageabstractThe primary objective of Optical Chemical Structure Recognition is to identify chemical structure images into corresponding markup sequences. However, the complex two-dimensional structures of molecules, particularly those with rings and multiple branches, present significant challenges for current end-to-end methods to learn one-dimensional markup directly. To overcome this limitation, we propose a novel Ring-Free Language (RFL), which utilizes a divide-and-conquer strategy to describe chemical structures in a hierarchical form. RFL allows complex molecular structures to be decomposed into multiple parts, ensuring both uniqueness and conciseness while enhancing readability. This approach significantly reduces the learning difficulty for recognition models. Leveraging RFL, we propose a universal Molecular Skeleton Decoder (MSD), which comprises a skeleton generation module that progressively predicts the molecular skeleton and individual rings, along with a branch classification module for predicting branch information. Experimental results demonstrate that the proposed RFL and MSD can be applied to various mainstream methods, achieving superior performance compared to state-of-the-art approaches in both printed and handwritten scenarios. Qikai Chang, Mingjun Chen, Changpeng Pi, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jinshui Hu |
AAAI | 4 |
| 2025 | DocMamba: Efficient Document Pre-training with State Space ModelabstractIn recent years, visually-rich document understanding has attracted increasing attention. Transformer-based pre-trained models have become the mainstream approach, yielding significant performance gains in this field. However, the self-attention mechanism's quadratic computational complexity hinders their efficiency and ability to process long documents. In this paper, we present DocMamba, a novel framework based on the state space model. It is designed to reduce computational complexity to linear while preserving global modeling capabilities. To further enhance its effectiveness in document processing, we introduce the Segment-First Bidirectional Scan (SFBS) to capture contiguous semantic information. Experimental results demonstrate that DocMamba achieves new state-of-the-art results on downstream datasets such as FUNSD, CORD, and SORIE, while significantly improving speed and reducing memory usage. Notably, experiments on the HRDoc confirm DocMamba's potential for length extrapolation. Pengfei Hu 0006, Jiefeng Ma, Shuhang Liu, Jun Du 0002, Jianshu Zhang 0001 |
AAAI | 1 |
| 2025 | Adaptive Radical Similarity Learning for Chinese Character Recognition
Zhongyuan Han, Jun Du 0002, Pengfei Hu 0006, Mobai Xue |
ICDAR (5) | 3 |
| 2025 | SPS-CG: Shape, Pronunciation, and Semantic Joint Modeling for Chinese Character Generation
Mobai Xue, Jun Du 0002, Pengfei Hu 0006 |
ICDAR (2) | 3 |
| 2025 | DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking head Video GenerationabstractTalking head generation intends to produce vivid and realistic talking head videos from a single portrait and speech audio clip. Although significant progress has been made in diffusion-based talking head generation, almost all methods rely on autoregressive strategies, which suffer from limited context utilization beyond the current generation step, error accumulation, and slower generation speed. To address these challenges, we present DAWN (\textbf{D}ynamic frame \textbf{A}vatar \textbf{W}ith \textbf{N}on-autoregressive diffusion), a framework that enables all-at-once generation of dynamic-length video sequences. Specifically, it consists of two main components: (1) audio-driven holistic facial dynamics generation in the latent motion space, and (2) audio-driven head pose and blink generation. Extensive experiments demonstrate that our method generates authentic and vivid videos with precise lip motions, and natural pose/blink movements. Additionally, with a high generation speed, DAWN possesses strong extrapolation capabilities, ensuring the stable production of high-quality long videos. These results highlight the considerable promise and potential impact of DAWN in the field of talking head video generation. Furthermore, we hope that DAWN sparks further exploration of non-autoregressive approaches in diffusion models. Our code will be publicly available at \url{https://github.com/Hanbo-Cheng/DAWN-pytorch}. Hanbo Cheng, Limin Lin, Pengcheng Xia 0002, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002 |
ICLR | 5 |
| 2025 | Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural IntegrationabstractRecent advances in Multimodal Large Language Models (MLLMs) have achieved remarkable progress in general domains and demonstrated promise in multimodal mathematical reasoning. However, applying MLLMs to geometry problem solving (GPS) remains challenging due to lack of accurate step-by-step solution data and severe hallucinations during reasoning. In this paper, we propose GeoGen, a pipeline that can automatically generates step-wise reasoning paths for geometry diagrams. By leveraging the precise symbolic reasoning, GeoGen produces large-scale, high-quality question-answer pairs. To further enhance the logical reasoning ability of MLLMs, we train GeoLogic, a Large Language Model (LLM) using synthetic data generated by GeoGen. Serving as a bridge between natural language and symbolic systems, GeoLogic enables symbolic tools to help verifying MLLM outputs, making the reasoning process more rigorous and alleviating hallucinations. Experimental results show that our approach consistently improves the performance of MLLMs, achieving remarkable results on benchmarks for geometric reasoning tasks. This improvement stems from our integration of the strengths of LLMs and symbolic systems, which enables a more reliable and interpretable approach for the GPS task. Codes are available at https://github.com/ycpNotFound/GeoGen. Yicheng Pan 0004, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jianshu Zhang 0001, Jianqing Gao |
ACM Multimedia | 3 |
| 2025 | Bidirectional trained tree-structured decoder for Handwritten Mathematical Expression Recognition
Hanbo Cheng, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002 |
Pattern Recognit. | 3 |
| 2025 | Count, decompose and correct: A new approach to handwritten Chinese character error correction
Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jianshu Zhang 0001 |
Pattern Recognit. | 1 |
| 2024 | Viewing Writing as Video: Optical Flow based Multi-Modal Handwritten Mathematical Expression RecognitionabstractHandwritten Mathematical Expression Recognition (HMER) forms a crucial task in the domain of document intelligence. It encompasses online and offline modalities, which utilize the trajectory sequence and static image as input, respectively. It is intuitive to utilize both online and offline modalities to build a more powerful recognition system. However, a formidable challenge arises as a result of the substantial heterogeneity between the online and offline modalities, which consequently leads to considerable obstacles in their alignment and fusion. In this work, we perceive the writing process as a video and introduce the Aggregated Optical Flow Map (AOFM) to represent the online modality, which is more compatible with the offline modality. Additionally, we propose the Optical Flow Aware Network (OFAN) in order to automatically extract, align, and fuse the features across online and offline modalities. Through experiment analysis, our method can be seamlessly applied to multiple existing offline HMER models, thereby yielding stable and substantial enhancements across CROHME 2014, 2016, and 2019 datasets. The code in this work is available at https: //github.com/Hanbo-Cheng/OFAN.git. Hanbo Cheng, Jun Du 0002, Pengfei Hu 0006, Jiefeng Ma, Mobai Xue |
ICASSP | 3 |
| 2024 | ICDAR 2024 Competition on Recognition of Chemical Structures
Mingjun Chen, Hao Wu 0090, Qikai Chang, Hanbo Cheng, Jiefeng Ma, Pengfei Hu 0006, Changpeng Pi, Jinshui Hu, Cong Liu 0006, Jun Du 0002 |
ICDAR (6) | 6 |
| 2024 | Radical Similarity Based Model Optimization and Post-correction for Chinese Character Recognition
Zhongyuan Han, Jun Du 0002, Mobai Xue, Jiefeng Ma, Pengfei Hu 0006 |
ICDAR (1) | 5 |
| 2024 | Maths: Multimodal Transformer-Based Human-Readable SolverabstractMultimodal mathematical reasoning has gained increasing attention in recent times. However, previous effective methods have not tried to reason in the form of natural language. In this paper, we introduce a model named MATHS (MultimodAl Transformer-based Human-readable Solver) for visual arithmetic and geometry problems in multimodal mathematical reasoning tasks. Drawing inspiration from Multimodal Large Language Models (MLLMs), our approach involves generating problem-solving processes expressed in natural language, in order to leverage the inherent reasoning capabilities embedded within language models. To address the challenge of precise calculations for language models, our work proposes a Math-Constrained Generation (MCG) method to impose hard constraints on generated outputs. Extensive experiments demonstrate our model excels in visual arithmetic task, and achieves results that are either better or comparable to existing methods in geometry problems. Code is available at https://github.com/ycpNotFound/MATHS. Yicheng Pan 0004, Jiefeng Ma, Pengfei Hu 0006, Jun Du 0002, Qing Wang 0008, Jianshu Zhang 0001, Dan Liu 0008, Si Wei |
ICME | 4 |
| 2024 | Exploring Audio-Visual Information Fusion for Sound Event Localization and Detection In Low-Resource Realistic ScenariosabstractThis study presents an audio-visual information fusion approach to sound event localization and detection (SELD) in low-resource scenarios. We aim at utilizing audio and video modality information through cross-modal learning and multi-modal fusion. First, we propose a cross-modal teacher-student learning (TSL) framework to transfer information from an audio-only teacher model, trained on a rich collection of audio data with multiple data augmentation techniques, to an audiovisual student model trained with only a limited set of multimodal data. Next, we propose a two-stage audio-visual fusion strategy, consisting of an early feature fusion and a late video-guided decision fusion to exploit synergies between audio and video modalities. Finally, we introduce an innovative video pixel swapping (VPS) technique to extend an audio channel swapping (ACS) method to an audio-visual joint augmentation. Evaluation results on the Detection and Classification of Acoustic Scenes and Events (DCASE) 2023 Challenge data set demonstrate significant improvements in SELD performances. Furthermore, our submission to the SELD task of the DCASE 2023 Challenge ranks first place by effectively integrating the proposed techniques into a model ensemble. Ya Jiang, Qing Wang 0008, Jun Du 0002, Maocheng Hu, Pengfei Hu 0006, Zeyan Liu, Shi Cheng 0001, Zhaoxu Nian, Mingqi Cai, Chin-Hui Lee 0001 |
ICME | 5 |
| 2024 | SEMv3: A Fast and Robust Approach to Table Separation Line Detection
Chunxia Qin, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002 |
IJCAI | 3 |
| 2024 | SRFUND: A Multi-Granularity Hierarchical Structure Reconstruction Benchmark in Form UnderstandingabstractAccurately identifying and organizing textual content is crucial for the automation of document processing in the field of form understanding. Existing datasets, such as FUNSD and XFUND, support entity classification and relationship prediction tasks but are typically limited to local and entity-level annotations. This limitation overlooks the hierarchically structured representation of documents, constraining comprehensive understanding of complex forms. To address this issue, we present the SRFUND, a hierarchically structured multi-task form understanding benchmark. SRFUND provides refined annotations on top of the original FUNSD and XFUND datasets, encompassing five tasks: (1) word to text-line merging, (2) text-line to entity merging, (3) entity category classification, (4) item table localization, and (5) entity-based full-document hierarchical structure recovery. We meticulously supplemented the original dataset with missing annotations at various levels of granularity and added detailed annotations for multi-item table regions within the forms. Additionally, we introduce global hierarchical structure dependencies for entity relation prediction tasks, surpassing traditional local key-value associations. The SRFUND dataset includes eight languages including English, Chinese, Japanese, German, French, Spanish, Italian, and Portuguese, making it a powerful tool for cross-lingual form understanding. Extensive experimental results demonstrate that the SRFUND dataset presents new challenges and significant opportunities in handling diverse layouts and global hierarchical structures of forms, thus providing deep insights into the field of form understanding. The original dataset and implementations of baseline methods are available at https://sprateam-ustc.github.io/SRFUND. Jiefeng Ma, Jun Du 0002, Yu Hu 0003, Pengfei Hu 0006, Qing Wang 0008, Jianshu Zhang 0001 |
NeurIPS | 7 |
| 2024 | Generate, transform, and clean: the role of GANs and transformers in palm leaf manuscript generation and enhancement
Nimol Thuon, Jun Du 0002, Jiefeng Ma, Pengfei Hu 0006 |
Int. J. Document Anal. Recognit. | 5 |
| 2024 | SEMv2: Table separation line detection based on instance segmentation
Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jianshu Zhang 0001, Cong Liu 0006 |
Pattern Recognit. | 2 |
| 2023 | HRDoc: Dataset and Baseline Method toward Hierarchical Reconstruction of Document StructuresabstractThe problem of document structure reconstruction refers to converting digital or scanned documents into corresponding semantic structures. Most existing works mainly focus on splitting the boundary of each element in a single document page, neglecting the reconstruction of semantic structure in multi-page documents. This paper introduces hierarchical reconstruction of document structures as a novel task suitable for NLP and CV fields. To better evaluate the system performance on the new task, we built a large-scale dataset named HRDoc, which consists of 2,500 multi-page documents with nearly 2 million semantic units. Every document in HRDoc has line-level annotations including categories and relations obtained from rule-based extractors and human annotators. Moreover, we proposed an encoder-decoder-based hierarchical document structure parsing system (DSPS) to tackle this problem. By adopting a multi-modal bidirectional encoder and a structure-aware GRU decoder with soft-mask operation, the DSPS model surpass the baseline method by a large margin. All scripts and datasets will be made publicly available at https://github.com/jfma-USTC/HRDoc. Jiefeng Ma, Jun Du 0002, Pengfei Hu 0006, Jianshu Zhang 0001, Cong Liu 0006 |
AAAI | 3 |
| 2023 | Group, Contrast and Recognize: A Self-supervised Method for Chinese Character Recognition
Xinzhe Jiang, Jun Du 0002, Pengfei Hu 0006, Mobai Xue, Jiefeng Ma, Jiajia Wu 0003, Jianshu Zhang 0001 |
ICDAR (4) | 3 |
| 2023 | Hierarchical Audio-Visual Information Fusion with Multi-label Joint Decoding for MER 2023abstractIn this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as robust acoustic and visual representations of raw video. Three different structures based on attention-guided feature gathering (AFG) are designed for deep feature fusion. Then, we introduce a joint decoding structure for emotion classification and valence regression in the decoding stage. A multi-task loss based on uncertainty is also designed to optimize the whole process. Finally, by combining three different structures on the posterior probability level, we obtain the final predictions of discrete and dimensional emotions. When tested on the dataset of multimodal emotion recognition challenge (MER 2023), the proposed framework yields consistent improvements in both emotion classification and valence regression. Our final system achieves state-of-the-art performance and ranks third on the leaderboard on MER-MULTI sub-challenge. Yuxuan Xi, Hang Chen 0001, Jun Du 0002, Yan Song 0001, Qing Wang 0008, Hengshun Zhou, Jiefeng Ma, Pengfei Hu 0006, Ya Jiang, Shi Cheng 0001, Jie Zhang 0042, Yuzhe Weng |
ACM Multimedia | 10 |
| 2022 | Multimodal Tree Decoder for Table of Contents Extraction in Document ImagesabstractTable of contents (ToC) extraction aims to extract headings of different levels in documents to better understand the outline of the contents, which can be widely used for document understanding and information retrieval. Existing works often use hand-crafted features and predefined rule-based functions to detect headings and resolve the hierarchical relationship between headings. Both the benchmark and research based on deep learning are still limited. Accordingly, in this paper, we first introduce a standard dataset, HierDoc, including image samples from 650 documents of scientific papers with their content labels. Then we propose a novel end-to-end model by using the multimodal tree decoder (MTD) for ToC as a benchmark for HierDoc. The MTD model is mainly composed of three parts, namely encoder, classifier, and decoder. The encoder fuses the multimodality features of vision, text, and layout information for each entity of the document. Then the classifier recognizes and selects the heading entities. Next, to parse the hierarchical relationship between the heading entities, a tree-structured decoder is designed. To evaluate the performance, both the metric of tree-edit-distance similarity (TEDS) and F1-Measure are adopted. Finally, our MTD approach achieves an average TEDS of 87.2% and an average F1-Measure of 88.1% on the test set of HierDoc. The code and dataset will be released at: https://github.com/Pengfei-Hu/MTD. Pengfei Hu 0006, Jianshu Zhang 0001, Jun Du 0002, Jiajia Wu 0003 |
ICPR | 1 |