VLDB 2026 Research / reviewers in the wild / expert
Dezhi Peng
dblp:217/2342
· DBLP profile ↗
36ranked-venue papers
8as first author
33since 2021 · last 2026
0000-0002-3263-3449ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 6 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 16 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document UnderstandingabstractRecent multimodal large language models (MLLMs) still struggle with long document understanding due to two fundamental challenges: information interference from abundant irrelevant content, and the quadratic computational cost of Transformer-based architectures. Existing approaches primarily fall into two categories: token compression, which sacrifices fine-grained details; and introducing external retrievers, which increase system complexity and prevent end-to-end optimization. To address these issues, we conduct an in-depth analysis and observe that MLLMs exhibit a human-like coarse-to-fine reasoning pattern: early Transformer layers attend broadly across the document, while deeper layers focus on relevant evidence pages. Motivated by this insight, we posit that the inherent evidence localization capabilities of MLLMs can be explicitly leveraged to perform retrieval during the reasoning process, facilitating efficient long document understanding. To this end, we propose URaG, a simple-yet-effective framework that Unifies Retrieval and Generation within a single MLLM. URaG introduces a lightweight cross-modal retrieval module that converts the early Transformer layers into an efficient evidence selector, identifying and preserving the most relevant pages while discarding irrelevant content. This design enables the deeper layers to concentrate computational resources on pertinent information, improving both accuracy and efficiency. Extensive experiments demonstrate that URaG achieves state-of-the-art performance while reducing computational overhead by 44-56%. Yongxin Shi, Zeyu Shan, Dezhi Peng, Zening Lin |
AAAI | 4 |
| 2025 | Predicting the Original Appearance of Damaged Historical DocumentsabstractHistorical documents encompass a wealth of cultural treasures but suffer from severe damages including character missing, paper damage, and ink erosion over time. However, existing document processing methods primarily focus on binarization, enhancement, etc., neglecting the repair of these damages. To this end, we present a new task, termed Historical Document Repair (HDR), which aims to predict the original appearance of damaged historical documents. To fill the gap in this field, we propose a large-scale dataset HDR28K and a diffusion-based network DiffHDR for historical document repair. Specifically, HDR28K contains 28,552 damaged-repaired image pairs with character-level annotations and multi-style degradations. Moreover, DiffHDR augments the vanilla diffusion framework with semantic and spatial information and a meticulously designed character perceptual loss for contextual and visual coherence. Experimental results demonstrate that the proposed DiffHDR trained on HDR28K significantly surpasses existing approaches and exhibits remarkable performance in handling real scenarios. Notably, DiffHDR can also be extended to document editing and text block generation, showcasing its high flexibility and generalization capacity. We believe this study could pioneer a new direction of document processing and contribute to the inheritance of invaluable cultures and civilizations. Zhenhua Yang, Dezhi Peng, Yongxin Shi, Yuyi Zhang 0002, Chongyu Liu |
AAAI | 2 |
| 2025 | SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
Mingxin Huang, Dezhi Peng, Zhenghao Peng, Chongyu Liu, Dahua Lin, Xiang Bai |
Int. J. Comput. Vis. | 2 |
| 2025 | QT-TextSR: Enhancing scene text image super-resolution via efficient interaction with text recognition using a Query-aware TransformerabstractScene text image super-resolution (STISR) has obtained widespread attention in recent years due to its ability to enhance text recognition performance. Many previous methods proposed to incorporate text prior knowledge into the super-resolution architecture for reconstructing high-quality text images. However, these text priors are typically derived from pretrained text recognition models, and the inaccurate recognition feedback will hinder overall performance. In this paper, we propose a novel model, QT-TextSR, which promotes scene text image super-resolution by introducing efficient interaction with text recognition to release the inaccurate text feedback through a Query-aware Transformer. Specifically, QT-TextSR decomposes scene text image super-resolution and scene text recognition into different sets of queries within a Vision-Language Cooperation Module, explicitly modeling discriminative and interactive features between text recognition and text image super-resolution tasks. By employing two separate yet simultaneous projection heads on the corresponding features, QT-TextSR can recover the low-quality text image meanwhile obtain the recognition results. Additionally, to mitigate the limitations caused by recognition errors and enhance text structure preservation, we introduce a strong texture prior through self-supervised pre-training, leveraging visual cues more effectively. Experiments on public dataset, TextZoom demonstrate that our QT-TextSR significantly outperforms previous state-of-the-art methods in the metrics of Recognition Accuracy (68% v s . 65.5%), PSNR (22.51 v s . 22.10), and SSIM (0.7960 v s . 7930). The code for QT-TextSR is available at https://github.com/lcy0604/QT-TextSR . Chongyu Liu, Dezhi Peng, Yuxin Kong, Jiaixin Zhang, Longfei Xiong, Jiwei Duan |
Neurocomputing | 3 |
| 2025 | EGO-LM: An efficient, generic, and out-of-the-box language model for handwritten text recognition
Dezhi Peng |
Pattern Recognit. | 2 |
| 2025 | HierCode: A lightweight hierarchical codebook for zero-shot Chinese text recognition
Yuyi Zhang 0002, Dezhi Peng, Peirong Zhang 0001, Zhenhua Yang, Zhibo Yang 0003, Cong Yao |
Pattern Recognit. | 3 |
| 2025 | Enhancing document dewarping evaluation: A new metric with improved accuracy and efficiency
Jiaxin Zhang 0003, Peirong Zhang 0001, Dezhi Peng |
Pattern Recognit. Lett. | 3 |
| 2025 | CTRNet++: Dual-Path Learning with Local-Global Context Modeling for Scene Text RemovalabstractRecent advances in scene text removal have attracted growing research interest due to its applications on privacy protection, document restoration, and text editing. While deep learning and generative adversarial network have shown significant progress, existing methods often struggle to generate consistent and plausible textures when erasing texts on complex backgrounds. To address this challenge, we propose a Contextual-guided Text Removal Network (CTRNet). CTRNet utilizes Low-level/High-level Contextual Guidance blocks (LCG, HCG) to explore both low-level structure and high-level discriminative context features from existing data to guide the text erasure and background restoration process. We further extend CTRNet to CTRNet++ by incorporate an auto-encoder architecture as a novel and effective HCG block, which serves as an additional image-inpainting branch, providing more accurate texture and context clues with the assistance of a large volume of natural images. Then we introduce a Context Embedding and Content Feature Modeling (CECFM) block that combines depth-wise CNN and Transformer layers to capture local features and establish long-term relationships among pixels globally. In addition, an efficient Progressive Feature Fusion Module (PFFM) is proposed to fully utilize multi-scale features from different branches. Experiments on benchmark datasets, SCUT-EnsText and SCUT-Syn, demonstrate that CTRNet++ significantly outperforms existing state-of-the-art methods and exhibits a stronger ability for complex background reconstruction. The code is available at https://github.com/lcy0604/CTRNet-plus . Chongyu Liu, Dezhi Peng |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM PretrainingabstractScene text removal (STR) aims at replacing text strokes in natural scenes with visually coherent backgrounds. Recent STR approaches rely on iterative refinements or explicit text masks, resulting in high complexity and sensitivity to the accuracy of text localization. Moreover, most existing STR methods adopt convolutional architectures while the potential of vision Transformers (ViTs) remains largely unexplored. In this paper, we propose a simple-yet-effective ViT-based text eraser, dubbed ViTEraser. Following a concise encoder-decoder framework, ViTEraser can easily incorporate various ViTs to enhance long-range modeling. Specifically, the encoder hierarchically maps the input image into the hidden space through ViT blocks and patch embedding layers, while the decoder gradually upsamples the hidden features to the text-erased image with ViT blocks and patch splitting layers. As ViTEraser implicitly integrates text localization and inpainting, we propose a novel end-to-end pretraining method, termed SegMIM, which focuses the encoder and decoder on the text box segmentation and masked image modeling tasks, respectively. Experimental results demonstrate that ViTEraser with SegMIM achieves state-of-the-art performance on STR by a substantial margin and exhibits strong generalization ability when extended to other tasks, e.g., tampered scene text detection. Furthermore, we comprehensively explore the architecture, pretraining, and scalability of the ViT-based encoder-decoder for STR, which provides deep insights into the application of ViT to the STR field. Code is available at https://github.com/shannanyinxiang/ViTEraser. Dezhi Peng, Chongyu Liu |
AAAI | 1 |
| 2024 | FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive LearningabstractAutomatic font generation is an imitation task, which aims to create a font library that mimics the style of reference images while preserving the content from source images. Although existing font generation methods have achieved satisfactory performance, they still struggle with complex characters and large style variations. To address these issues, we propose FontDiffuser, a diffusion-based image-to-image one-shot font generation method, which innovatively models the font imitation task as a noise-to-denoise paradigm. In our method, we introduce a Multi-scale Content Aggregation (MCA) block, which effectively combines global and local content cues across different scales, leading to enhanced preservation of intricate strokes of complex characters. Moreover, to better manage the large variations in style transfer, we propose a Style Contrastive Refinement (SCR) module, which is a novel structure for style representation learning. It utilizes a style extractor to disentangle styles from images, subsequently supervising the diffusion model via a meticulously designed style contrastive loss. Extensive experiments demonstrate FontDiffuser's state-of-the-art performance in generating diverse characters and styles. It consistently excels on complex characters and large style changes compared to previous methods. The code is available at https://github.com/yeungchenwa/FontDiffuser. Zhenhua Yang, Dezhi Peng, Yuxin Kong, Yuyi Zhang 0002, Cong Yao |
AAAI | 2 |
| 2024 | DocRes: A Generalist Model Toward Unifying Document Image Restoration TasksabstractDocument image restoration is a crucial aspect of Document AI systems, as the quality of document images significantly influences the overall performance. Prevailing methods address distinct restoration tasks independently, leading to intricate systems and the incapability to harness the potential synergies of multi-task learning. To overcome this challenge, we propose DocRes, a generalist model that unifies five document image restoration tasks including dewarping, deshadowing, appearance enhancement, deblurring, and binarization. To instruct DocRes to perform various restoration tasks, we propose a novel visual prompt approach called Dynamic Task-Specific Prompt (DTSPrompt). The DTSPrompt for different tasks comprises distinct prior features, which are additional characteristics extracted from the input image. Beyond its role as a cue for task-specific execution, DTSPrompt can also serve as supplementary information to enhance the model's performance. Moreover, DTSPrompt is more flexible than prior visual prompt approaches as it can be seamlessly applied and adapted to inputs with high and variable resolutions. Experimental results demonstrate that DocRes achieves competitive or superior performance compared to existing state-of-the-art task-specific models. This under-scores the potential of DocRes across a broader spectrum of document image restoration tasks. The source code is publicly available at https://github.com/ZZZHANGjx/DocRes. Jiaxin Zhang 0003, Dezhi Peng, Chongyu Liu, Peirong Zhang 0001 |
CVPR | 2 |
| 2024 | Towards Modern Image Manipulation Localization: A Large-Scale Dataset and Novel MethodsabstractIn recent years, image manipulation localization has attracted increasing attention due to its pivotal role in guaranteeing social media security. However, how to accurately identify the forged regions remains an open challenge. One of the main bottlenecks lies in the severe scarcity of high-quality data, due to its costly creation process. To address this limitation, we propose a novel paradigm, termed as CAAA, to automatically and precisely annotate the numerous manually forged images from the web at the pixel level. We further propose a novel metric QES to facilitate the automatic filtering of unreliable annotations. With CAAA and QES, we construct a large-scale, diverse, and high-quality dataset comprising 123,150 manually forged images with mask annotations. Besides, we develop a new model APSC-Net for accurate image manipulation localization. According to extensive experiments, our dataset significantly improves the performance of various models on the widely-used benchmarks and such improvements are attributed to our proposed effective methods. The dataset and code are publicly available at https://github.com/qcf-568/MIML. Chenfan Qu, Yiwu Zhong, Chongyu Liu, Guitao Xu, Dezhi Peng, Fengjun Guo |
CVPR | 5 |
| 2024 | DTSM: Toward Dense Table Structure Recognition with Text Query Encoder and Adjacent Feature Aggregator
Xinhong Chen 0005, Bangdong Chen, Chenfan Qu, Dezhi Peng, Chongyu Liu |
ICDAR (1) | 4 |
| 2024 | UPOCR: Towards Unified Pixel-Level OCR InterfaceabstractExisting optical character recognition (OCR) methods rely on task-specific designs with divergent paradigms, architectures, and training strategies, which significantly increases the complexity of research and maintenance and hinders the fast deployment in applications. To this end, we propose UPOCR, a simple-yet-effective generalist model for Unified Pixel-level OCR interface. Specifically, the UPOCR unifies the paradigm of diverse OCR tasks as image-to-image transformation and the architecture as a vision Transformer (ViT)-based encoder-decoder with learnable task prompts. The prompts push the general feature representations extracted by the encoder towards task-specific spaces, endowing the decoder with task awareness. Moreover, the model training is uniformly aimed at minimizing the discrepancy between the predicted and ground-truth images regardless of the inhomogeneity among tasks. Experiments are conducted on three pixel-level OCR tasks including text removal, text segmentation, and tampered text detection. Without bells and whistles, the experimental results showcase that the proposed method can simultaneously achieve state-of-the-art performance on three tasks with a unified single model, which provides valuable strategies and insights for future research on generalist OCR models. Code is available at https://github.com/shannanyinxiang/UPOCR. Dezhi Peng, Zhenhua Yang, Jiaxin Zhang 0003, Chongyu Liu, Yongxin Shi, Kai Ding 0009, Fengjun Guo |
ICML | 1 |
| 2024 | SideNet: Learning representations from interactive side information for zero-shot Chinese character recognition
Dezhi Peng, Mengchao He |
Pattern Recognit. | 3 |
| 2023 | Building A Mobile Text Recognizer via Truncated SVD-based Knowledge Distillation-Guided NAS
Weifeng Lin, Canyu Xie, Dezhi Peng, Cong Yao, Mengchao He |
BMVC | 3 |
| 2023 | Towards Robust Tampered Text Detection in Document Image: New Dataset and New SolutionabstractRecently, tampered text detection in document image has attracted increasingly attention due to its essential role on information security. However, detecting visually consistent tampered text in photographed document images is still a main challenge. In this paper, we propose a novel framework to capture more fine-grained clues in complex scenarios for tampered text detection, termed as Document Tampering Detector (DTD), which consists of a Frequency Perception Head (FPH) to compensate the deficiencies caused by the inconspicuous visual features, and a Multi-view Iterative Decoder (MID) for fully utilizing the information of features in different scales. In addition, we design a new training paradigm, termed as Curriculum Learning for Tampering Detection (CLTD), which can address the confusion during the training procedure and thus to improve the robustness for image compression and the ability to generalize. To further facilitate the tampered text detection in document images, we construct a large-scale document image dataset, termed as DocTamper, which contains 170,000 document images of various types. Experiments demonstrate that our proposed DTD outperforms previous state-of-the-art by 9.2%, 26.3% and 12.3% in terms of F-measure on the DocTamper testing set, and the crossdomain testing sets of DocTamper-FCD and DocTamper-SCD, respectively. Codes and dataset will be available at https://github.com/qcf-568/DocTamper. Chenfan Qu, Chongyu Liu, Xinhong Chen 0005, Dezhi Peng, Fengjun Guo |
CVPR | 5 |
| 2023 | ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in TransformerabstractIn recent years, end-to-end scene text spotting approaches are evolving to the Transformer-based framework. While previous studies have shown the crucial importance of the intrinsic synergy between text detection and recognition, recent advances in Transformer-based methods usually adopt an implicit synergy strategy with shared query, which can not fully realize the potential of these two interactive tasks. In this paper, we argue that the explicit synergy considering distinct characteristics of text detection and recognition can significantly improve the performance text spotting. To this end, we introduce a new model named Explicit Synergy-based Text Spotting Transformer framework (ESTextSpotter), which achieves explicit synergy by modeling discriminative and interactive features for text detection and recognition within a single decoder. Specifically, we decompose the conventional shared query into task-aware queries for text polygon and content, respectively. Through the decoder with the proposed vision-language communication module, the queries interact with each other in an explicit manner while preserving discriminative patterns of text detection and recognition, thus improving performance significantly. Additionally, we propose a task-aware query initialization scheme to ensure stable training. Experimental results demonstrate that our model significantly outperforms previous state-of-the-art methods. Code is available at https://github.com/mxin262/ESTextSpotter. Mingxin Huang, Jiaxin Zhang 0003, Dezhi Peng, Hao Lu 0003, Can Huang 0002, Xiang Bai |
ICCV | 3 |
| 2023 | Revisiting Scene Text Recognition: A Data PerspectiveabstractThis paper aims to re-assess scene text recognition (STR) from a data-oriented perspective. We begin by revisiting the six commonly used benchmarks in STR and observe a trend of performance saturation, whereby only 2.91% of the benchmark images cannot be accurately recognized by an ensemble of 13 representative models. While these results are impressive and suggest that STR could be considered solved, however, we argue that this is primarily due to the less challenging nature of the common benchmarks, thus concealing the underlying issues that STR faces. To this end, we consolidate a large-scale real STR dataset, namely Union14M, which comprises 4 million labeled images and 10 million unlabeled images, to assess the performance of STR models in more complex real-world scenarios. Our experiments demonstrate that the 13 models can only achieve an average accuracy of 66.53% on the 4 million labeled images, indicating that STR still faces numerous challenges in the real world. By analyzing the error patterns of the 13 models, we identify seven open challenges in STR and develop a challenge-driven benchmark consisting of eight distinct subsets to facilitate further progress in the field. Our exploration demonstrates that STR is far from being solved and leveraging data may be a promising solution. In this regard, we find that utilizing the 10 million unlabeled images through self-supervised pre-training can significantly improve the robustness of STR model in real-world scenarios and leads to state-of-the-art performance. Code and dataset is available at https://github.com/Mountchicken/Union14M. Dezhi Peng, Chongyu Liu |
ICCV | 3 |
| 2023 | EnsExam: A Dataset for Handwritten Text Erasure on Examination Papers
Liufeng Huang, Bangdong Chen, Chongyu Liu, Dezhi Peng, Weiying Zhou, Yaqiang Wu, Hao Ni 0001 |
ICDAR (3) | 4 |
| 2023 | SegCTC: Offline Handwritten Chinese Text Recognition via Better Fusion Between Explicit and Implicit Segmentation
Jiarong Huang, Dezhi Peng, Hao Ni 0001 |
ICDAR (4) | 2 |
| 2023 | Read Ten Lines at One Glance: Line-Aware Semi-Autoregressive Transformer for Multi-Line Handwritten Mathematical Expression RecognitionabstractHandwritten Mathematical Expression Recognition (HMER) plays a critical role in various applications, such as digitized education and scientific research. Although existing methods have achieved promising performance on publicly available datasets, they still struggle to recognize multi-line mathematical expressions (MEs), suffering from complex structures and slow inference speed. To address these issues, we propose a Line-Aware Semi-autoregressive Transformer (LAST) that treats multi-line mathematical expression sequences as two-dimensional dual-end structures. The proposed LAST utilizes a line-wise dual-end decoding strategy to decode multi-line mathematical expressions in parallel and perform dual-end decoding within each line. Specifically, we introduce a line-aware positional encoding module and a line-partitioned dual-end mask to endow LAST with line order awareness and directionality. Additionally, we adopt a shared-task optimization strategy to train LAST in both autoregressive and semi-autoregressive tasks. To evaluate the effectiveness of our approach in real-world scenarios, we have built a new Multi-line Mathematical Expression dataset (M2E), which, to the best of our knowledge, is the first of its kind and boasts with the largest character category, the largest samples of characters, and the longest average sequence length, compared to existing ME datasets. Experimental results on both the M2E dataset and publicly available datasets demonstrate the effectiveness of our proposed method. Notably, our semi-autoregressive decoding approach achieves significantly faster decoding speeds while still achieving state-of-the-art performance compared to the existing methods. Wentao Yang 0003, Zhe Li 0046, Dezhi Peng, Mengchao He, Cong Yao |
ACM Multimedia | 3 |
| 2023 | M5HisDoc: A Large-scale Multi-style Chinese Historical Document Analysis BenchmarkabstractRecognizing and organizing text in correct reading order plays a crucial role in historical document analysis and preservation. While existing methods have shown promising performance, they often struggle with challenges such as diverse layouts, low image quality, style variations, and distortions. This is primarily due to the lack of consideration for these issues in the current benchmarks, which hinders the development and evaluation of historical document analysis and recognition (HDAR) methods in complex real-world scenarios. To address this gap, this paper introduces a complex multi-style Chinese historical document analysis benchmark, named M5HisDoc. The M5 indicates five properties of style, ie., Multiple layouts, Multiple document types, Multiple calligraphy styles, Multiple backgrounds, and Multiple challenges. The M5HisDoc dataset consists of two subsets, M5HisDoc-R (Regular) and M5HisDoc-H (Hard). The M5HisDoc-R subset comprises 4,000 historical document images. To ensure high-quality annotations, we meticulously perform manual annotation and triple-checking. To replicate real-world conditions for historical document analysis applications, we incorporate image rotation, distortion, and resolution reduction into M5HisDoc-R subset to form a new challenging subset named M5HisDoc-H, which contains the same number of images as M5HisDoc-R. The dataset exhibits diverse styles, significant scale variations, dense texts, and an extensive character set. We conduct benchmarking experiments on five tasks: text line detection, text line recognition, character detection, character recognition, and reading order prediction. We also conduct cross-validation with other benchmarks. Experimental results demonstrate that the M5HisDoc dataset can offer new challenges and great opportunities for future research in this field, thereby providing deep insights into the solution for HDAR. The dataset is available at https://github.com/HCIILAB/M5HisDoc. Yongxin Shi, Chongyu Liu, Dezhi Peng, Cheng Jian, Jiarong Huang |
NeurIPS | 3 |
| 2023 | SPTS v2: Single-Point Scene Text SpottingabstractEnd-to-end scene text spotting has made significant progress due to its intrinsic synergy between text detection and recognition. Previous methods commonly regard manual annotations such as horizontal rectangles, rotated rectangles, quadrangles, and polygons as a prerequisite, which are much more expensive than using single-point. Our new framework, SPTS v2, allows us to train high-performing text-spotting models using a single-point annotation. SPTS v2 reserves the advantage of the auto-regressive Transformer with an Instance Assignment Decoder (IAD) through sequentially predicting the center points of all text instances inside the same predicting sequence, while with a Parallel Recognition Decoder (PRD) for text recognition in parallel, which significantly reduces the requirement of the length of the sequence. These two decoders share the same parameters and are interactively connected with a simple but effective information transmission process to pass the gradient and information. Comprehensive experiments on various existing benchmark datasets demonstrate the SPTS v2 can outperform previous state-of-the-art single-point text spotters with fewer parameters while achieving 19× faster inference speed. Within the context of our SPTS v2 framework, our experiments suggest a potential preference for single-point representation in scene text spotting when compared to other representations. Such an attempt provides a significant opportunity for scene text spotting applications beyond the realms of existing paradigms. Jiaxin Zhang 0003, Dezhi Peng, Mingxin Huang, Xinyu Wang 0010, Jingqun Tang, Can Huang 0002, Dahua Lin, Chunhua Shen, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Recognition of Handwritten Chinese Text by Segmentation: A Segment-Annotation-Free ApproachabstractOnline and offline handwritten Chinese text recognition (HTCR) has been studied for decades. Early methods adopted oversegmentation-based strategies but suffered from low speed, insufficient accuracy, and high cost of character segmentation annotations. Recently, segmentation-free methods based on connectionist temporal classification (CTC) and attention mechanism, have dominated the field of HCTR. However, people actually read text character by character, especially for ideograms such as Chinese. This raises the question: are segmentation-free strategies really the best solution to HCTR? To explore this issue, we propose a new segmentation-based method for recognizing handwritten Chinese text that is implemented using a simple yet efficient fully convolutional network. A novel weakly supervised learning method is proposed to enable the network to be trained using only transcript annotations; thus, the expensive character segmentation annotations required by previous segmentation-based methods can be avoided. Owing to the lack of context modeling in fully convolutional networks, we propose a contextual regularization method to integrate contextual information into the network during the training stage, which can further improve the recognition performance. Extensive experiments conducted on four widely used benchmarks, namely CASIA-HWDB, CASIA-OLHWDB, ICDAR2013, and SCUT-HCCDoc, show that our method significantly surpasses existing methods on both online and offline HCTR, and exhibits a considerably higher inference speed than CTC/attention-based approaches. Dezhi Peng, Weihong Ma, Canyu Xie, Hesuo Zhang, Shenggao Zhu |
IEEE Trans. Multim. | 1 |
| 2023 | SLOGAN: Handwriting Style Synthesis for Arbitrary-Length and Out-of-Vocabulary TextabstractLarge amounts of labeled data are urgently required for the training of robust text recognizers. However, collecting handwriting data of diverse styles, along with an immense lexicon, is considerably expensive. Although data synthesis is a promising way to relieve data hunger, two key issues of handwriting synthesis, namely, style representation and content embedding, remain unsolved. To this end, we propose a novel method that can synthesize parameterized and controllable handwriting S tyles for arbitrary-Length and O ut-of-vocabulary text based on a G enerative A dversarial N etwork (GAN), termed SLOGAN. Specifically, we propose a style bank to parameterize specific handwriting styles as latent vectors, which are input to a generator as style priors to achieve the corresponding handwritten styles. The training of the style bank requires only writer identification of the source images, rather than attribute annotations. Moreover, we embed the text content by providing an easily obtainable printed style image, so that the diversity of the content can be flexibly achieved by changing the input printed image. Finally, the generator is guided by dual discriminators to handle both the handwriting characteristics that appear as separated characters and in a series of cursive joins. Our method can synthesize words that are not included in the training vocabulary and with various new styles. Extensive experiments have shown that high-quality text images with great style diversity and rich vocabulary can be synthesized using our method, thereby enhancing the robustness of the recognizer. Canjie Luo, Zhe Li 0046, Dezhi Peng |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Complex Table Structure Recognition in the Wild Using Transformer and Identity Matrix-Based Augmentation
Bangdong Chen, Dezhi Peng, Jiaxin Zhang 0003, Yujin Ren |
ICFHR | 2 |
| 2022 | SPTS: Single-Point Text SpottingabstractExisting scene text spotting (i.e., end-to-end text detection and recognition) methods rely on costly bounding box annotations (e.g., text-line, word-level, or character-level bounding boxes). For the first time, we demonstrate that training scene text spotting models can be achieved with an extremely low-cost annotation of a single-point for each instance. We propose an end-to-end scene text spotting method that tackles scene text spotting as a sequence prediction task. Given an image as input, we formulate the desired detection and recognition results as a sequence of discrete tokens and use an auto-regressive Transformer to predict the sequence. The proposed method is simple yet effective, which can achieve state-of-the-art results on widely used benchmarks. Most significantly, we show that the performance is not very sensitive to the positions of the point annotation, meaning that it can be much easier to be annotated or even be automatically generated than the bounding box that requires precise positions. We believe that such a pioneer attempt indicates a significant opportunity for scene text spotting applications of a much larger scale than previously possible. The code is available at https://github.com/shannanyinxiang/SPTS. Dezhi Peng, Xinyu Wang 0010, Jiaxin Zhang 0003, Mingxin Huang, Songxuan Lai, Jing Li 0036, Shenggao Zhu, Dahua Lin, Chunhua Shen, Xiang Bai |
ACM Multimedia | 1 |
| 2022 | PageNet: Towards End-to-End Weakly Supervised Page-Level Handwritten Chinese Text Recognition
Dezhi Peng, Canjie Luo, Songxuan Lai |
Int. J. Comput. Vis. | 1 |
| 2021 | Implicit Feature Alignment: Learn To Convert Text Recognizer to Text SpotterabstractText recognition is a popular research subject with many associated challenges. Despite the considerable progress made in recent years, the text recognition task itself is still constrained to solve the problem of reading cropped line text images and serves as a subtask of optical character recognition (OCR) systems. As a result, the final text recognition result is limited by the performance of the text detector. In this paper, we propose a simple, elegant and effective paradigm called Implicit Feature Alignment (IFA), which can be easily integrated into current text recognizers, resulting in a novel inference mechanism called IFA- inference. This enables an ordinary text recognizer to process multi-line text such that text detection can be completely freed. Specifically, we integrate IFA into the two most prevailing text recognition streams (attention-based and CTC-based) and propose attention-guided dense prediction (ADP) and Extended CTC (ExCTC). Furthermore, the Wasserstein-based Hollow Aggregation Cross-Entropy (WH-ACE) is proposed to suppress negative predictions to assist in training ADP and ExCTC. We experimentally demonstrate that IFA achieves state-of-the-art performance on end-to-end document recognition tasks while maintaining the fastest speed, and ADP and ExCTC complement each other on the perspective of different application scenarios. Code will be available at https://github.com/Wang-Tianwei/Implicit-feature-alignment. Dezhi Peng, Zhe Li 0046, Mengchao He, Yongpan Wang, Canjie Luo |
CVPR | 4 |
| 2021 | Zero-Shot Chinese Text Recognition via Matching Class Embedding
Dezhi Peng |
ICDAR (3) | 3 |
| 2021 | A Multi-level Progressive Rectification Mechanism for Irregular Scene Text Recognition
Qianying Liao, Qingxiang Lin, Canjie Luo, Jiaxin Zhang 0003, Dezhi Peng |
ICDAR (4) | 6 |
| 2021 | Towards Fast, Accurate and Compact Online Handwritten Chinese Text Recognition
Dezhi Peng, Canyu Xie, Zecheng Xie, Kai Ding 0009, Yichao Huang, Yaqiang Wu |
ICDAR (3) | 1 |
| 2019 | A Fast and Accurate Fully Convolutional Network for End-to-End Handwritten Chinese Text Segmentation and RecognitionabstractHandwritten Chinese Text Recognition (HCTR) is a challenging problem due to its high complexity. Previous methods based on over-segmentation, hidden Markov model (HMM) or long short-term memory recurrent neural network (LSTM-RNN) have achieved great success in recognition results. However, all of them, including over-segmentation based methods, are incompetent in accurate segmentation of single character. To solve this problem, we propose a fast and accurate fully convolutional network for end-to-end segmentation and recognition of handwritten Chinese text. Experiments on CASIA-HWDB datasets and ICDAR 2013 competition dataset show that our method achieves a competitive performance on recognition and produces great character segmentation results. Moreover, our model reaches a real-time speed of 70 fps, which is fast enough for various applications. Dezhi Peng, Yaqiang Wu, Zhepeng Wang 0002, Mingxiang Cai |
ICDAR | 1 |
| 2018 | Scale Mapping and Dynamic Re-Detecting in Dense Head DetectionabstractConvolutional neural networks (CNNs) have demonstrated a strong ability to extract semantics from images during object detection; however, the extracted semantics are typically have strong scale priors for a specific circumstance. In this paper, we investigate the influence of head scale and contextual information, and then propose a scale-invariant method for head detection. Our method can dynamically detect heads depending on the complexity of the image. It uses an extra feature map to represent the scale information of the spatial relationship, and then uses this feature map for auxiliary detection. Particularly, we exploit several new techniques, including contextual information, scale invariance, and hard example mining. We evaluated our method on three head datasets and achieved state-of-the-art results for the Brainwash dataset, HollywoodHeads dataset, and SCUT-HEAD dataset. Zikai Sun, Dezhi Peng, Zirui Cai, Zirong Chen |
ICIP | 2 |
| 2018 | Detecting Heads using Feature Refine Net and Cascaded Multi-scale ArchitectureabstractThis paper presents a method that can accurately detect heads especially small heads under the indoor scene. To achieve this, we propose a novel method, Feature Refine Net (FRN), and a cascaded multi-scale architecture. FRN exploits the multi-scale hierarchical features created by deep convolutional neural networks. The proposed channel weighting method enables FRN to make use of features alternatively and effectively. To improve the performance of small head detection, we propose a cascaded multi-scale architecture which has two detectors. One called global detector is responsible for detecting large objects and acquiring the global distribution information. The other called local detector is designed for small objects detection and makes use of the information provided by global detector. Due to the lack of head detection datasets, we have collected and labeled a new large dataset named SCUT-HEAD which includes 4405 images with 111251 heads annotated. Experiments show that our method has achieved state-of-the-art performance on SCUT-HEAD. Dezhi Peng, Zikai Sun, Zirong Chen, Zirui Cai, Lele Xie |
ICPR | 1 |