VLDB 2026 Research / reviewers in the wild / expert
Jiaxin Zhang 0003
dblp:32/7698-3
· DBLP profile ↗
22ranked-venue papers
6as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoScaler: Self scale alignment for handwritten mathematical expression recognition
Wentao Yang 0003, Jiaxin Zhang 0003 |
Pattern Recognit. | 2 |
| 2026 | DocAligner: Automating the annotation of photographed documents through real-virtual alignment
Jiaxin Zhang 0003, Peirong Zhang 0001, Huiyi Cheng, Kai Ding 0009 |
Pattern Recognit. | 1 |
| 2025 | DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual SlimmingabstractCurrent multimodal large language models (MLLMs) face significant challenges in visual document understanding (VDU) tasks due to the high resolution, dense text, and complex layouts typical of document images. These characteristics demand a high level of detail perception ability from MLLMs. While increasing input resolution improves detail perception capability, it also leads to longer sequences of visual tokens, increasing computational costs and straining the models' ability to handle long contexts. To address these challenges, we introduce DocKylin, a document-centric MLLM that performs visual content slimming at both the pixel and token levels, thereby reducing token sequence length in VDU scenarios. We introduce an Adaptive Pixel Slimming (APS) preprocessing module to perform pixel-level slimming, increasing the proportion of informative pixels. Moreover, we propose a novel Dynamic Token Slimming (DTS) module to conduct token-level slimming, filtering essential tokens and removing others to adaptively create a more compact visual sequence. Experiments demonstrate DocKylin's promising performance across various VDU benchmarks and the effectiveness of each component. Jiaxin Zhang 0003, Wentao Yang 0003, Songxuan Lai, Zecheng Xie |
AAAI | 1 |
| 2025 | Towards Real-World Document Specular Highlight Removal: The DocHighlight Dataset and DocSHRNet Method
Jiaxin Zhang 0003, Hiuyi Cheng, Peirong Zhang 0001, Xuhan Zheng |
PRCV (7) | 2 |
| 2025 | Smaller But Better: Unifying Layout Generation with Smaller Large Language Models
Peirong Zhang 0001, Jiaxin Zhang 0003, Jiahuan Cao |
Int. J. Comput. Vis. | 2 |
| 2025 | Enhancing document dewarping evaluation: A new metric with improved accuracy and efficiency
Jiaxin Zhang 0003, Peirong Zhang 0001, Dezhi Peng |
Pattern Recognit. Lett. | 1 |
| 2024 | DocRes: A Generalist Model Toward Unifying Document Image Restoration TasksabstractDocument image restoration is a crucial aspect of Document AI systems, as the quality of document images significantly influences the overall performance. Prevailing methods address distinct restoration tasks independently, leading to intricate systems and the incapability to harness the potential synergies of multi-task learning. To overcome this challenge, we propose DocRes, a generalist model that unifies five document image restoration tasks including dewarping, deshadowing, appearance enhancement, deblurring, and binarization. To instruct DocRes to perform various restoration tasks, we propose a novel visual prompt approach called Dynamic Task-Specific Prompt (DTSPrompt). The DTSPrompt for different tasks comprises distinct prior features, which are additional characteristics extracted from the input image. Beyond its role as a cue for task-specific execution, DTSPrompt can also serve as supplementary information to enhance the model's performance. Moreover, DTSPrompt is more flexible than prior visual prompt approaches as it can be seamlessly applied and adapted to inputs with high and variable resolutions. Experimental results demonstrate that DocRes achieves competitive or superior performance compared to existing state-of-the-art task-specific models. This under-scores the potential of DocRes across a broader spectrum of document image restoration tasks. The source code is publicly available at https://github.com/ZZZHANGjx/DocRes. Jiaxin Zhang 0003, Dezhi Peng, Chongyu Liu, Peirong Zhang 0001 |
CVPR | 1 |
| 2024 | UPOCR: Towards Unified Pixel-Level OCR InterfaceabstractExisting optical character recognition (OCR) methods rely on task-specific designs with divergent paradigms, architectures, and training strategies, which significantly increases the complexity of research and maintenance and hinders the fast deployment in applications. To this end, we propose UPOCR, a simple-yet-effective generalist model for Unified Pixel-level OCR interface. Specifically, the UPOCR unifies the paradigm of diverse OCR tasks as image-to-image transformation and the architecture as a vision Transformer (ViT)-based encoder-decoder with learnable task prompts. The prompts push the general feature representations extracted by the encoder towards task-specific spaces, endowing the decoder with task awareness. Moreover, the model training is uniformly aimed at minimizing the discrepancy between the predicted and ground-truth images regardless of the inhomogeneity among tasks. Experiments are conducted on three pixel-level OCR tasks including text removal, text segmentation, and tampered text detection. Without bells and whistles, the experimental results showcase that the proposed method can simultaneously achieve state-of-the-art performance on three tasks with a unified single model, which provides valuable strategies and insights for future research on generalist OCR models. Code is available at https://github.com/shannanyinxiang/UPOCR. Dezhi Peng, Zhenhua Yang, Jiaxin Zhang 0003, Chongyu Liu, Yongxin Shi, Kai Ding 0009, Fengjun Guo |
ICML | 3 |
| 2024 | Irregular text block recognition via decoupling visual, linguistic, and positional information
Chengquan Zhang, Jiaxin Zhang 0003, Zecheng Xie, Pengyuan Lv |
Pattern Recognit. | 4 |
| 2023 | M6Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout AnalysisabstractDocument layout analysis is a crucial prerequisite for document understanding, including document retrieval and conversion. Most public datasets currently contain only PDF documents and lack realistic documents. Models trained on these datasets may not generalize well to real-world scenarios. Therefore, this paper introduces a large and diverse document layout analysis dataset called M6Doc. The M6 designation represents six properties: (1) Multi-Format (including scanned, photographed, and PDF documents); (2) Multi-Type (such as scientific articles, textbooks, books, test papers, magazines, newspapers, and notes); (3) Multi-Layout (rectangular, Manhattan, non-Manhattan, and multi-column Manhattan); (4) Multi-Language (Chinese and English); (5) Multi-Annotation Category (74 types of annotation labels with 237,116 annotation instances in 9,080 manually annotated pages); and (6) Modern documents. Additionally, we propose a transformer-based document layout analysis method called TransDLANet, which leverages an adaptive element matching mechanism that enables query embedding to better match ground truth to improve recall, and constructs a segmentation branch for more precise document image instance segmentation. We conduct a comprehensive evaluation of M6Doc with various layout analysis methods and demonstrate its effectiveness. TransDLANet achieves state-of-the-art performance on M6 Doc with 64.5% mAP. The M6Doc dataset will be available at https://github.com/HcIILAB/M6Doc. Hiuyi Cheng, Peirong Zhang 0001, Sihang Wu, Jiaxin Zhang 0003, Qiyuan Zhu, Zecheng Xie, Jing Li 0036, Kai Ding 0009 |
CVPR | 4 |
| 2023 | ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in TransformerabstractIn recent years, end-to-end scene text spotting approaches are evolving to the Transformer-based framework. While previous studies have shown the crucial importance of the intrinsic synergy between text detection and recognition, recent advances in Transformer-based methods usually adopt an implicit synergy strategy with shared query, which can not fully realize the potential of these two interactive tasks. In this paper, we argue that the explicit synergy considering distinct characteristics of text detection and recognition can significantly improve the performance text spotting. To this end, we introduce a new model named Explicit Synergy-based Text Spotting Transformer framework (ESTextSpotter), which achieves explicit synergy by modeling discriminative and interactive features for text detection and recognition within a single decoder. Specifically, we decompose the conventional shared query into task-aware queries for text polygon and content, respectively. Through the decoder with the proposed vision-language communication module, the queries interact with each other in an explicit manner while preserving discriminative patterns of text detection and recognition, thus improving performance significantly. Additionally, we propose a task-aware query initialization scheme to ensure stable training. Experimental results demonstrate that our model significantly outperforms previous state-of-the-art methods. Code is available at https://github.com/mxin262/ESTextSpotter. Mingxin Huang, Jiaxin Zhang 0003, Dezhi Peng, Hao Lu 0003, Can Huang 0002, Xiang Bai |
ICCV | 2 |
| 2023 | SPTS v2: Single-Point Scene Text SpottingabstractEnd-to-end scene text spotting has made significant progress due to its intrinsic synergy between text detection and recognition. Previous methods commonly regard manual annotations such as horizontal rectangles, rotated rectangles, quadrangles, and polygons as a prerequisite, which are much more expensive than using single-point. Our new framework, SPTS v2, allows us to train high-performing text-spotting models using a single-point annotation. SPTS v2 reserves the advantage of the auto-regressive Transformer with an Instance Assignment Decoder (IAD) through sequentially predicting the center points of all text instances inside the same predicting sequence, while with a Parallel Recognition Decoder (PRD) for text recognition in parallel, which significantly reduces the requirement of the length of the sequence. These two decoders share the same parameters and are interactively connected with a simple but effective information transmission process to pass the gradient and information. Comprehensive experiments on various existing benchmark datasets demonstrate the SPTS v2 can outperform previous state-of-the-art single-point text spotters with fewer parameters while achieving 19× faster inference speed. Within the context of our SPTS v2 framework, our experiments suggest a potential preference for single-point representation in scene text spotting when compared to other representations. Such an attempt provides a significant opportunity for scene text spotting applications beyond the realms of existing paradigms. Jiaxin Zhang 0003, Dezhi Peng, Mingxin Huang, Xinyu Wang 0010, Jingqun Tang, Can Huang 0002, Dahua Lin, Chunhua Shen, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Scene table structure recognition with segmentation collaboration and alignment
Hongyi Wang 0008, Yang Xue 0001, Jiaxin Zhang 0003 |
Pattern Recognit. Lett. | 3 |
| 2022 | Looking from a Higher-Level Perspective: Attention and Recognition Enhanced Multi-scale Scene Text Segmentation
Yujin Ren, Jiaxin Zhang 0003, Bangdong Chen |
ACCV (7) | 2 |
| 2022 | Complex Table Structure Recognition in the Wild Using Transformer and Identity Matrix-Based Augmentation
Bangdong Chen, Dezhi Peng, Jiaxin Zhang 0003, Yujin Ren |
ICFHR | 3 |
| 2022 | TextSRNet: Scene Text Super-Resolution Based on Contour Prior and Atrous ConvolutionabstractLow-resolution (LR) scene text images are often encountered in real application scenarios. It is difficult to recognize these scene texts due to the loss of their character details. Super-resolution (SR) is an intuitive alternative to solve this problem. Unlike generic SR, which mainly focuses on image quality, text SR focuses on the accuracy of the downstream recognition tasks. Although several scene text SR methods have been recently proposed, they ignore retaining character contours that are meaningful for character readability and poorly perform on long texts. Therefore, we proposes a new scene text SR network called TextSRNet. In order to get fine character details, we adopt the segmentation maps of scene text images as the prior knowledge of character contours and embed it into the proposed TextSRNet. In addition, we incorporate Sobel loss to enhance the shape boundaries of characters. To improve the robustness of the model for long texts, we adopt a text atrous spatial pyramid pooling module. The proposed TextSRNet is trained and evaluated on TextZoom, a real-world scene text SR dataset. Extensive experiments demonstrate that our TextSRNet outperforms existing methods in terms of both recognition accuracy and image quality. Jizhao Ma, Jiaxin Zhang 0003, Yang Xue 0001, Mengchao He |
ICPR | 3 |
| 2022 | Marior: Margin Removal and Iterative Content Rectification for Document Dewarping in the WildabstractCamera-captured document images usually suffer from perspective and geometric deformations. It is of great value to rectify them when considering poor visual aesthetics and the deteriorated performance of OCR systems. Recent learning-based methods intensively focus on the accurately cropped document image. However, this might not be sufficient for overcoming practical challenges, including document images either with large marginal regions or without margins. Due to this impracticality, users struggle to crop documents precisely when they encounter large marginal regions. Simultaneously, dewarping images without margins is still an insurmountable problem. To the best of our knowledge, there is still no complete and effective pipeline for rectifying document images in the wild. To address this issue, we propose a novel approach called Marior (Margin Removal and Iterative Content Rectification). Marior follows a progressive strategy to iteratively improve the dewarping quality and readability in a coarse-to-fine manner. Specifically, we divide the pipeline into two modules: margin removal module (MRM) and iterative content rectification module (ICRM). First, we predict the segmentation mask of the input image to remove the margin, thereby obtaining a preliminary result. Then we refine the image further by producing dense displacement flows to achieve content-aware rectification. We determine the number of refinement iterations adaptively. Experiments demonstrate the state-of-the-art performance of our method on public benchmarks. The resources are available at https://github.com/ZZZHANG-jx/Marior for further comparison. Jiaxin Zhang 0003, Canjie Luo, Fengjun Guo, Kai Ding 0009 |
ACM Multimedia | 1 |
| 2022 | SPTS: Single-Point Text SpottingabstractExisting scene text spotting (i.e., end-to-end text detection and recognition) methods rely on costly bounding box annotations (e.g., text-line, word-level, or character-level bounding boxes). For the first time, we demonstrate that training scene text spotting models can be achieved with an extremely low-cost annotation of a single-point for each instance. We propose an end-to-end scene text spotting method that tackles scene text spotting as a sequence prediction task. Given an image as input, we formulate the desired detection and recognition results as a sequence of discrete tokens and use an auto-regressive Transformer to predict the sequence. The proposed method is simple yet effective, which can achieve state-of-the-art results on widely used benchmarks. Most significantly, we show that the performance is not very sensitive to the positions of the point annotation, meaning that it can be much easier to be annotated or even be automatically generated than the bounding box that requires precise positions. We believe that such a pioneer attempt indicates a significant opportunity for scene text spotting applications of a much larger scale than previously possible. The code is available at https://github.com/shannanyinxiang/SPTS. Dezhi Peng, Xinyu Wang 0010, Jiaxin Zhang 0003, Mingxin Huang, Songxuan Lai, Jing Li 0036, Shenggao Zhu, Dahua Lin, Chunhua Shen, Xiang Bai |
ACM Multimedia | 4 |
| 2022 | Forgery-free signature verification with stroke-aware cycle-consistent generative adversarial network
Songxuan Lai, Yecheng Zhu, Jiaxin Zhang 0003, Bangdong Chen |
Neurocomputing | 5 |
| 2021 | Towards Robust Visual Information Extraction in Real World: New Dataset and Novel SolutionabstractVisual Information Extraction (VIE) has attracted considerable attention recently owing to its various advanced applications such as document understanding, automatic marking and intelligent education. Most existing works decoupled this problem into several independent sub-tasks of text spotting (text detection and recognition) and information extraction, which completely ignored the high correlation among them during optimization. In this paper, we propose a robust Visual Information Extraction System (VIES) towards real-world scenarios, which is an unified end-to-end trainable framework for simultaneous text detection, recognition and information extraction by taking a single document image as input and outputting the structured information. Specifically, the information extraction branch collects abundant visual and semantic representations from text spotting for multimodal feature fusion and conversely, provides higher-level semantic clues to contribute to the optimization of text spotting. Moreover, regarding the shortage of public benchmarks, we construct a fully-annotated dataset called EPHOIE (https://github.com/HCIILAB/EPHOIE), which is the first Chinese benchmark for both text spotting and visual information extraction. EPHOIE consists of 1,494 images of examination paper head with complex layouts and background, including a total of 15,771 Chinese handwritten or printed text instances. Compared with the state-of-the-art methods, our VIES shows significant superior performance on the EPHOIE dataset and achieves a 9.01% F-score gain on the widely used SROIE dataset under the end-to-end scenario. Chongyu Liu, Guozhi Tang, Jiaxin Zhang 0003, Shuaitao Zhang, Qianying Wang 0002, Yaqiang Wu, Mingxiang Cai |
AAAI | 5 |
| 2021 | A Multi-level Progressive Rectification Mechanism for Irregular Scene Text Recognition
Qianying Liao, Qingxiang Lin, Canjie Luo, Jiaxin Zhang 0003, Dezhi Peng |
ICDAR (4) | 5 |
| 2020 | SaHAN: Scale-aware hierarchical attention network for scene text recognition
Jiaxin Zhang 0003, Canjie Luo, Weiying Zhou |
Pattern Recognit. Lett. | 1 |