Jiefeng Ma

dblp:302/9166 · DBLP profile ↗
← Back
25ranked-venue papers
3as first author
25since 2021 · last 2026
0000-0003-2416-3720ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 2 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 13 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 See then tell: Enhancing key information extraction with vision grounding
Shuhang Liu, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Qing Wang 0008, Jianshu Zhang 0001
Neurocomputing4
2026 Two-stage decomposition network for handwritten Chinese character error correction
Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jianshu Zhang 0001, Jianqing Gao, Qingfeng Liu
Pattern Recognit.3
2025 RFL: Simplifying Chemical Structure Recognition with Ring-Free Language
abstract
The primary objective of Optical Chemical Structure Recognition is to identify chemical structure images into corresponding markup sequences. However, the complex two-dimensional structures of molecules, particularly those with rings and multiple branches, present significant challenges for current end-to-end methods to learn one-dimensional markup directly. To overcome this limitation, we propose a novel Ring-Free Language (RFL), which utilizes a divide-and-conquer strategy to describe chemical structures in a hierarchical form. RFL allows complex molecular structures to be decomposed into multiple parts, ensuring both uniqueness and conciseness while enhancing readability. This approach significantly reduces the learning difficulty for recognition models. Leveraging RFL, we propose a universal Molecular Skeleton Decoder (MSD), which comprises a skeleton generation module that progressively predicts the molecular skeleton and individual rings, along with a branch classification module for predicting branch information. Experimental results demonstrate that the proposed RFL and MSD can be applied to various mainstream methods, achieving superior performance compared to state-of-the-art approaches in both printed and handwritten scenarios.
Qikai Chang, Mingjun Chen, Changpeng Pi, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jinshui Hu
AAAI6
2025 DocMamba: Efficient Document Pre-training with State Space Model
abstract
In recent years, visually-rich document understanding has attracted increasing attention. Transformer-based pre-trained models have become the mainstream approach, yielding significant performance gains in this field. However, the self-attention mechanism's quadratic computational complexity hinders their efficiency and ability to process long documents. In this paper, we present DocMamba, a novel framework based on the state space model. It is designed to reduce computational complexity to linear while preserving global modeling capabilities. To further enhance its effectiveness in document processing, we introduce the Segment-First Bidirectional Scan (SFBS) to capture contiguous semantic information. Experimental results demonstrate that DocMamba achieves new state-of-the-art results on downstream datasets such as FUNSD, CORD, and SORIE, while significantly improving speed and reducing memory usage. Notably, experiments on the HRDoc confirm DocMamba's potential for length extrapolation.
Pengfei Hu 0006, Jiefeng Ma, Shuhang Liu, Jun Du 0002, Jianshu Zhang 0001
AAAI3
2025 EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion
abstract
Diffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In this research, we propose an EmotiveTalk framework to address these issues. Firstly, to realize better control over the generation of lip movement and facial expression, a Vision-guided Audio Information Decoupling (V-AID) approach is designed to generate audio-based decoupled representations aligned with lip movements and expression. Specifically, to achieve alignment between audio and facial expression representation spaces, we present a Diffusion-based Co-speech Temporal Expansion (Di-CTE) module within V-AID to generate expression-related representations under multi-source emotion condition constraints. Then we propose a well-designed Emotional Talking Head Diffusion (ETHD) backbone to efficiently generate highly expressive talking head videos, which contains an Expression Decoupling Injection (EDI) module to automatically decouple the expressions from reference portraits while integrating the target expression information, achieving more expressive generation performance. Experimental results show that EmotiveTalk can generate expressive talking head videos, ensuring the promised controllability of emotions and metric stability during long-time generation, yielding state-of-the-art performance compared to existing methods. The main page of our paper can be found in https://emotivetalk.github.io/.
Yuzhe Weng, Zilu Guo, Jun Du 0002, Shutong Niu, Jiefeng Ma, Cong Liu 0006, Qingfeng Liu
CVPR7
2025 Latent Swap Joint Diffusion for 2D Long-Form Latent Generation
Yusheng Dai, Jun Du 0002, Lei Sun 0010, Jianqing Gao, Ruoyu Wang 0029, Jiefeng Ma
ICCV10
2025 DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking head Video Generation
abstract
Talking head generation intends to produce vivid and realistic talking head videos from a single portrait and speech audio clip. Although significant progress has been made in diffusion-based talking head generation, almost all methods rely on autoregressive strategies, which suffer from limited context utilization beyond the current generation step, error accumulation, and slower generation speed. To address these challenges, we present DAWN (\textbf{D}ynamic frame \textbf{A}vatar \textbf{W}ith \textbf{N}on-autoregressive diffusion), a framework that enables all-at-once generation of dynamic-length video sequences. Specifically, it consists of two main components: (1) audio-driven holistic facial dynamics generation in the latent motion space, and (2) audio-driven head pose and blink generation. Extensive experiments demonstrate that our method generates authentic and vivid videos with precise lip motions, and natural pose/blink movements. Additionally, with a high generation speed, DAWN possesses strong extrapolation capabilities, ensuring the stable production of high-quality long videos. These results highlight the considerable promise and potential impact of DAWN in the field of talking head video generation. Furthermore, we hope that DAWN sparks further exploration of non-autoregressive approaches in diffusion models. Our code will be publicly available at \url{https://github.com/Hanbo-Cheng/DAWN-pytorch}.
Hanbo Cheng, Limin Lin, Pengcheng Xia 0002, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002
ICLR6
2025 Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural Integration
abstract
Recent advances in Multimodal Large Language Models (MLLMs) have achieved remarkable progress in general domains and demonstrated promise in multimodal mathematical reasoning. However, applying MLLMs to geometry problem solving (GPS) remains challenging due to lack of accurate step-by-step solution data and severe hallucinations during reasoning. In this paper, we propose GeoGen, a pipeline that can automatically generates step-wise reasoning paths for geometry diagrams. By leveraging the precise symbolic reasoning, GeoGen produces large-scale, high-quality question-answer pairs. To further enhance the logical reasoning ability of MLLMs, we train GeoLogic, a Large Language Model (LLM) using synthetic data generated by GeoGen. Serving as a bridge between natural language and symbolic systems, GeoLogic enables symbolic tools to help verifying MLLM outputs, making the reasoning process more rigorous and alleviating hallucinations. Experimental results show that our approach consistently improves the performance of MLLMs, achieving remarkable results on benchmarks for geometric reasoning tasks. This improvement stems from our integration of the strengths of LLMs and symbolic systems, which enables a more reliable and interpretable approach for the GPS task. Codes are available at https://github.com/ycpNotFound/GeoGen.
Yicheng Pan 0004, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jianshu Zhang 0001, Jianqing Gao
ACM Multimedia4
2025 Bidirectional trained tree-structured decoder for Handwritten Mathematical Expression Recognition
Hanbo Cheng, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002
Pattern Recognit.5
2025 Count, decompose and correct: A new approach to handwritten Chinese character error correction
Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jianshu Zhang 0001
Pattern Recognit.2
2024 Viewing Writing as Video: Optical Flow based Multi-Modal Handwritten Mathematical Expression Recognition
abstract
Handwritten Mathematical Expression Recognition (HMER) forms a crucial task in the domain of document intelligence. It encompasses online and offline modalities, which utilize the trajectory sequence and static image as input, respectively. It is intuitive to utilize both online and offline modalities to build a more powerful recognition system. However, a formidable challenge arises as a result of the substantial heterogeneity between the online and offline modalities, which consequently leads to considerable obstacles in their alignment and fusion. In this work, we perceive the writing process as a video and introduce the Aggregated Optical Flow Map (AOFM) to represent the online modality, which is more compatible with the offline modality. Additionally, we propose the Optical Flow Aware Network (OFAN) in order to automatically extract, align, and fuse the features across online and offline modalities. Through experiment analysis, our method can be seamlessly applied to multiple existing offline HMER models, thereby yielding stable and substantial enhancements across CROHME 2014, 2016, and 2019 datasets. The code in this work is available at https: //github.com/Hanbo-Cheng/OFAN.git.
Hanbo Cheng, Jun Du 0002, Pengfei Hu 0006, Jiefeng Ma, Mobai Xue
ICASSP4
2024 ICDAR 2024 Competition on Recognition of Chemical Structures
Mingjun Chen, Hao Wu 0090, Qikai Chang, Hanbo Cheng, Jiefeng Ma, Pengfei Hu 0006, Changpeng Pi, Jinshui Hu, Cong Liu 0006, Jun Du 0002
ICDAR (6)5
2024 Radical Similarity Based Model Optimization and Post-correction for Chinese Character Recognition
Zhongyuan Han, Jun Du 0002, Mobai Xue, Jiefeng Ma, Pengfei Hu 0006
ICDAR (1)4
2024 Maths: Multimodal Transformer-Based Human-Readable Solver
abstract
Multimodal mathematical reasoning has gained increasing attention in recent times. However, previous effective methods have not tried to reason in the form of natural language. In this paper, we introduce a model named MATHS (MultimodAl Transformer-based Human-readable Solver) for visual arithmetic and geometry problems in multimodal mathematical reasoning tasks. Drawing inspiration from Multimodal Large Language Models (MLLMs), our approach involves generating problem-solving processes expressed in natural language, in order to leverage the inherent reasoning capabilities embedded within language models. To address the challenge of precise calculations for language models, our work proposes a Math-Constrained Generation (MCG) method to impose hard constraints on generated outputs. Extensive experiments demonstrate our model excels in visual arithmetic task, and achieves results that are either better or comparable to existing methods in geometry problems. Code is available at https://github.com/ycpNotFound/MATHS.
Yicheng Pan 0004, Jiefeng Ma, Pengfei Hu 0006, Jun Du 0002, Qing Wang 0008, Jianshu Zhang 0001, Dan Liu 0008, Si Wei
ICME3
2024 SEMv3: A Fast and Robust Approach to Table Separation Line Detection
Chunxia Qin, Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002
IJCAI5
2024 SRFUND: A Multi-Granularity Hierarchical Structure Reconstruction Benchmark in Form Understanding
abstract
Accurately identifying and organizing textual content is crucial for the automation of document processing in the field of form understanding. Existing datasets, such as FUNSD and XFUND, support entity classification and relationship prediction tasks but are typically limited to local and entity-level annotations. This limitation overlooks the hierarchically structured representation of documents, constraining comprehensive understanding of complex forms. To address this issue, we present the SRFUND, a hierarchically structured multi-task form understanding benchmark. SRFUND provides refined annotations on top of the original FUNSD and XFUND datasets, encompassing five tasks: (1) word to text-line merging, (2) text-line to entity merging, (3) entity category classification, (4) item table localization, and (5) entity-based full-document hierarchical structure recovery. We meticulously supplemented the original dataset with missing annotations at various levels of granularity and added detailed annotations for multi-item table regions within the forms. Additionally, we introduce global hierarchical structure dependencies for entity relation prediction tasks, surpassing traditional local key-value associations. The SRFUND dataset includes eight languages including English, Chinese, Japanese, German, French, Spanish, Italian, and Portuguese, making it a powerful tool for cross-lingual form understanding. Extensive experimental results demonstrate that the SRFUND dataset presents new challenges and significant opportunities in handling diverse layouts and global hierarchical structures of forms, thus providing deep insights into the field of form understanding. The original dataset and implementations of baseline methods are available at https://sprateam-ustc.github.io/SRFUND.
Jiefeng Ma, Jun Du 0002, Yu Hu 0003, Pengfei Hu 0006, Qing Wang 0008, Jianshu Zhang 0001
NeurIPS1
2024 Generate, transform, and clean: the role of GANs and transformers in palm leaf manuscript generation and enhancement
Nimol Thuon, Jun Du 0002, Jiefeng Ma, Pengfei Hu 0006
Int. J. Document Anal. Recognit.4
2024 SEMv2: Table separation line detection based on instance segmentation
Pengfei Hu 0006, Jiefeng Ma, Jun Du 0002, Jianshu Zhang 0001, Cong Liu 0006
Pattern Recognit.3
2023 HRDoc: Dataset and Baseline Method toward Hierarchical Reconstruction of Document Structures
abstract
The problem of document structure reconstruction refers to converting digital or scanned documents into corresponding semantic structures. Most existing works mainly focus on splitting the boundary of each element in a single document page, neglecting the reconstruction of semantic structure in multi-page documents. This paper introduces hierarchical reconstruction of document structures as a novel task suitable for NLP and CV fields. To better evaluate the system performance on the new task, we built a large-scale dataset named HRDoc, which consists of 2,500 multi-page documents with nearly 2 million semantic units. Every document in HRDoc has line-level annotations including categories and relations obtained from rule-based extractors and human annotators. Moreover, we proposed an encoder-decoder-based hierarchical document structure parsing system (DSPS) to tackle this problem. By adopting a multi-modal bidirectional encoder and a structure-aware GRU decoder with soft-mask operation, the DSPS model surpass the baseline method by a large margin. All scripts and datasets will be made publicly available at https://github.com/jfma-USTC/HRDoc.
Jiefeng Ma, Jun Du 0002, Pengfei Hu 0006, Jianshu Zhang 0001, Cong Liu 0006
AAAI1
2023 Group, Contrast and Recognize: A Self-supervised Method for Chinese Character Recognition
Xinzhe Jiang, Jun Du 0002, Pengfei Hu 0006, Mobai Xue, Jiefeng Ma, Jiajia Wu 0003, Jianshu Zhang 0001
ICDAR (4)5
2023 Hierarchical Audio-Visual Information Fusion with Multi-label Joint Decoding for MER 2023
abstract
In this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as robust acoustic and visual representations of raw video. Three different structures based on attention-guided feature gathering (AFG) are designed for deep feature fusion. Then, we introduce a joint decoding structure for emotion classification and valence regression in the decoding stage. A multi-task loss based on uncertainty is also designed to optimize the whole process. Finally, by combining three different structures on the posterior probability level, we obtain the final predictions of discrete and dimensional emotions. When tested on the dataset of multimodal emotion recognition challenge (MER 2023), the proposed framework yields consistent improvements in both emotion classification and valence regression. Our final system achieves state-of-the-art performance and ranks third on the leaderboard on MER-MULTI sub-challenge.
Yuxuan Xi, Hang Chen 0001, Jun Du 0002, Yan Song 0001, Qing Wang 0008, Hengshun Zhou, Jiefeng Ma, Pengfei Hu 0006, Ya Jiang, Shi Cheng 0001, Jie Zhang 0042, Yuzhe Weng
ACM Multimedia9
2023 Multimodal Pre-Training Based on Graph Attention Network for Document Understanding
abstract
Document intelligence as a relatively new research topic supports many business applications. Its main task is to automatically read, understand, and analyze documents. However, due to the diversity of formats (invoices, reports, forms, etc.) and layouts in documents, it is difficult to make machines understand documents. In this paper, we present the GraphDoc, a multimodal graph attention-based model for various document understanding tasks. GraphDoc is pre-trained in a multimodal framework by utilizing text, layout, and image information simultaneously. In a document, a text block relies heavily on its surrounding contexts, accordingly we inject the graph structure into the attention mechanism to form a graph attention layer so that each input node can only attend to its neighborhoods. The input nodes of each graph attention layer are composed of textual, visual, and positional features from semantically meaningful regions in a document image. We do the multimodal feature fusion of each node by the gate fusion layer. The contextualization between each node is modeled by the graph attention layer. GraphDoc learns a generic representation from only 320k unlabeled documents via the Masked Sentence Modeling task. Extensive experimental results on the publicly available datasets show that GraphDoc achieves state-of-the-art performance, which demonstrates the effectiveness of our proposed method.
Jiefeng Ma, Jun Du 0002, Jianshu Zhang 0001
IEEE Trans. Multim.2
2022 Query-driven Generative Network for Document Information Extraction in the Wild
abstract
This paper focuses on solving Document Information Extraction (DIE) in the wild problem, which is rarely explored before. In contrast to existing studies mainly tailored for document cases in known templates with predefined layouts and keys under the ideal input without OCR errors involved, we aim to build up a more practical DIE paradigm for real-world scenarios where input document images may contain unknown layouts and keys in the scenes of the problematic OCR results. To achieve this goal, we propose a novel architecture, termed Query-driven Generative Network (QGN), which is equipped with two consecutive modules, i.e., Layout Context-aware Module (LCM) and Structured Generation Module (SGM). Given a document image with unseen layouts and fields, the former LCM yields the value prefix candidates serving as the query prompts for the SGM to generate the final key-value pairs even with OCR noise. To further investigate the potential of our method, we create a new large-scale dataset, named LArge-scale STructured Documents (LastDoc4000), containing 4,000 documents with 1,511 layouts and 3,500 different keys. In experiments, we demonstrate that our QGN consistently achieves the best F1-score on the new LastDoc4000 dataset by at most 30.32% absolute improvement. A more comprehensive experimental analysis and experiments on other public benchmarks also verify the effectiveness and robustness of our proposed method for the wild DIE task.
Haoyu Cao 0001, Xin Li 0118, Jiefeng Ma, Deqiang Jiang, Antai Guo, Yiqing Hu, Hao Liu 0003, Yinsong Liu, Bo Ren 0002
ACM Multimedia3
2022 GMN: Generative Multi-modal Network for Practical Document Information Extraction
abstract
Haoyu Cao, Jiefeng Ma, Antai Guo, Yiqing Hu, Hao Liu, Deqiang Jiang, Yinsong Liu, Bo Ren. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Haoyu Cao 0001, Jiefeng Ma, Antai Guo, Yiqing Hu, Hao Liu 0003, Deqiang Jiang, Yinsong Liu, Bo Ren 0002
NAACL-HLT2
2021 An Open-Source Library of 2D-GMM-HMM Based on Kaldi Toolkit and Its Application to Handwritten Chinese Character Recognition
Jiefeng Ma, Jun Du 0002
ICIG (1)1