Peirong Zhang 0001

dblp:306/3180-1 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-1857-5473ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 10 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Based Scalable Typography
abstract
Commercial-grade poster design demands the seamless integration of aesthetic appeal with precise, informative content delivery. Current automated poster generation systems face significant limitations, including incomplete design workflows, poor text rendering accuracy, and insufficient flexibility for commercial applications. To address these challenges, we propose PosterVerse, a full-workflow, commercial-grade poster generation method that seamlessly automates the entire design process while delivering high-density and scalable text rendering. PosterVerse replicates professional design through three key stages: (1) blueprint creation using fine-tuned LLMs to extract key design elements from user requirements, (2) graphical background generation via customized diffusion models to create visually appealing imagery, and (3) unified layout-text rendering with an MLLM-powered HTML engine to guarantee high text accuracy and flexible customization. In addition, we introduce PosterDNA, a commercial-grade, HTML-based dataset tailored for training and validating poster design models. To the best of our knowledge, PosterDNA is the first Chinese poster generation dataset to introduce HTML typography files, enabling scalable text rendering and fundamentally solving the challenges of rendering small and high-density text. Experimental results demonstrate that PosterVerse consistently produces commercial-grade posters with appealing visuals, accurate text alignment, and customizable layouts, making it a promising solution for automating commercial poster design.
Junle Liu, Peirong Zhang 0001, Yuyi Zhang 0002, Pengyu Yan, Xinyue Zhou, Fengjun Guo
AAAI2
2026 Frequency Mining Empowered by Text Aggregation: A New Perspective on Document Image Tampering Detection
abstract
Document image tampering detection faces significant challenges due to the subtle and spatially dispersed nature of tampering traces, which are often confined to localized regions within tampered text. While existing methods leverage frequency domain information to reveal hidden artifacts, they fail to fully exploit the rich frequency spectrum and lack effective mechanisms for aggregating scattered tampering evidence across extended text regions. To overcome these limitations, we propose the Text Aggregation and multi-Frequency Enhancement Network (TAFE-Net). Specifically, to capture more subtle tampering traces, we design a Multi-Frequency Feature Extractor that comprehensively utilizes various proven effective frequency information. In addition, the Visual-Frequency Integration Module and Direction-aware Frequency Decoupling Enhancement module are introduced to aggregate text features in both horizontal and vertical directions within the frequency domain, from coarse to fine granularity, addressing the incomplete detection of tampered text caused by dispersed tampering traces. Experiments on the DocTamper and RTM datasets demonstrate that our approach establishes new state-of-the-art results and maintains superior robustness against various degradations.
Ziqi Yi, Guitao Xu, Shihang Wu, Peirong Zhang 0001
AAAI4
2026 Draft, Verify, Restore: Self-Refining Historical Inscription Restoration with a Unified MLLM
abstract
Inscriptions are invaluable cultural heritage, yet centuries of degradation (e.g., fractures, erosion, oxidation) have rendered many partially illegible. Existing Historical Inscription Restoration (HIR) methods rely on task-separated pipelines with irreversible error accumulation and patch-based generation that sacrifices page-level consistency. Therefore, we present UniHIR, the first unified MLLM for end-to-end historical inscription restoration. It integrates two novel designs, Draft-Guided Localization and Hierarchical Self-Refinement, to enable accurate damage localization and illegible-content prediction via iterative reasoning and self-correction. This unified approach enables true page-level restoration with consistent typography and style. To support training under high-resolution inputs and long sequences, we design UHIRFactory and construct HIRBench, enabling step-wise, memory-efficient instruction tuning with step-aware annotations for intermediate drafts and refinements. Experiments demonstrate that UniHIR achieves superior performance in both text restoration accuracy and appearance restoration quality, validating that HIR can be effectively tackled by a standalone model in a unified manner. The model and code are available at https://github.com/ZZXF11/UniHIR.
Yuyi Zhang 0002, Junle Liu, Peirong Zhang 0001, Jianliang Liu, Zhenhua Yang
ACL (1)3
2026 DocAligner: Automating the annotation of photographed documents through real-virtual alignment
Jiaxin Zhang 0003, Peirong Zhang 0001, Huiyi Cheng, Kai Ding 0009
Pattern Recognit.2
2026 Bridging the Reality Gap in Tampered Text Detection: A Human-Crafted Real-World Dataset and a Text-Centric Approach
Guitao Xu, Peirong Zhang 0001
IEEE Trans. Inf. Forensics Secur.2
2025 Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document Restoration
abstract
Yuyi Zhang, Peirong Zhang, Zhenhua Yang, Pengyu Yan, Yongxin Shi, Pengwei Liu, Fengjun Guo, Lianwen Jin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yuyi Zhang 0002, Peirong Zhang 0001, Zhenhua Yang, Pengyu Yan, Yongxin Shi, Pengwei Liu, Fengjun Guo
ACL (1)2
2025 Generalizable Audio Deepfake Detection via Hierarchical Structure Learning and Feature Whitening in Poincaré sphere
Mingru Yang, Yanmei Gu, Qianhua He, Yanxiong Li, Peirong Zhang 0001, Huijia Zhu, Weiqiang Wang 0002
INTERSPEECH5
2025 TongGu-VL: Advancing Visual-Language Understanding in Chinese Classical Studies through Parameter Sensitivity-Guided Instruction Tuning
abstract
Chinese Classical Studies (CCS) is a pivotal gateway to ancient Chinese culture. Spanning ancient texts, illustrations, paintings, and calligraphy, CCS presents significant challenges for non-specialists due to its language and visual complexity. While Large Language Models (LLMs) have been explored to facilitate CCS, current methods primarily focus on textual analysis, overlooking the rich visual information intrinsic to classical materials. To bridge this gap, we propose TongGu-VL, a pioneering specialized MLLM designed for CCS applications. Our contributions are threefold. First, we construct CCS358K, a comprehensive multimodal instruction dataset to enhance MLLMs' CCS capabilities. Second, we propose Parameter Sensitivity-Guided Instruction Tuning (PSG-IT), a novel method that mitigates catastrophic forgetting without data replay. It effectively preserves TongGu-VL's general skills, while optimizing its CCS performance. Third, we design a Visual-Text Early Fusion (VTEF) module, which harnesses MLLMs' modality alignment to generate instruction-aware visual representations, thus improving language modeling. Extensive experimental results demonstrate that our model outperforms existing MLLMs on a broad range of CCS tasks, while maintaining general capabilities that benefit other domains beyond CCS. Our model and dataset will be publicly available.
Jiahuan Cao, Yang Liu 0353, Peirong Zhang 0001, Yongxin Shi, Kai Ding 0009
ACM Multimedia3
2025 From Pixels to Semantics: A Novel MLLM-Driven Approach for Explainable Tampered Text Detection
abstract
The spread of tampered text poses a critical challenge to information security. Previous methods for tampered text detection (TTD) primarily relied on visual artifacts as clues, while overlooking potential semantic inconsistencies introduced during text manipulation. To address this limitation, we propose TVSIP (Tampered text Visual-Semantic InterPreter), a novel framework leveraging Multimodal Large Language Models (MLLMs) to integrate both visual and semantic clues for comprehensive tampered text analysis and verification. TVSIP consists of a Locator and an Interpreter. The Locator combines the visual detection ability of existing expert models with the semantic comprehension capabilities of MLLMs to create precise tampering masks. Subsequently, the Interpreter provides comprehensive descriptions and explanations based on identified tampered regions. To train and evaluate TVSIP, we construct the TextDDLE benchmark using GPT-4o. Extensive experiments demonstrate that TVSIP outperforms expert models in pixel-level localization and advanced MLLMs in interpretability. Furthermore, it maintains robustness against image degradation and exhibits strong generalization ability on out-of-domain datasets. Our work highlights the crucial role of semantic inconsistencies in TTD and establishes a more reliable verification system for ensuring document authenticity in the digital age.
Guitao Xu, Ziqi Yi, Peirong Zhang 0001, Jiahuan Cao, Shihang Wu
ACM Multimedia3
2025 Generalizable Audio Deepfake Detection via Risk-Aware Style Alignment and Structural Empirical Risk Minimization
abstract
With the rapid advancement of AIGC technologies, audio deepfakes have become increasingly realistic, posing serious threats to information security and biometric authentication. Therefore, audio deepfake detection (ADD) has emerged as a critical and fast-evolving research area, particularly requiring superior generalization in out-of-domain scenarios. However, existing ADD methods suffer from constrained generalization and limited access to target data. To address these challenges, we propose Risk-Aware Style Alignment (RASA), a novel generalizable ADD framework that projects the style of any input feature into a shared style space through similarity-based projection. This alignment reduces both inter-domain and intra-source discrepancies without requiring target data during training. In addition, we adopt Structural Empirical Risk Minimization (SERM) in the Poincaré ball model to capture the hierarchical structure of the data and further minimize source risk. By jointly optimizing RASA and SERM, the proposed method effectively tightens the theoretical upper bound of target risk across three key dimensions: source risk, inter-domain divergence, and intra-source discrepancy. Extensive experiments demonstrate that our approach achieves superior generalization and outperforms existing state-of-the-art methods.
Mingru Yang, Yanmei Gu, Qianhua He, Peirong Zhang 0001, Haolin He, Huijia Zhu, Weiqiang Wang 0002
ACM Multimedia4
2025 Capturing More: Learning Multi-Domain Representations for Robust Online Handwriting Verification
Peirong Zhang 0001, Kai Ding 0009
ACM Multimedia1
2025 Towards Real-World Document Specular Highlight Removal: The DocHighlight Dataset and DocSHRNet Method
Jiaxin Zhang 0003, Hiuyi Cheng, Peirong Zhang 0001, Xuhan Zheng
PRCV (7)4
2025 Smaller But Better: Unifying Layout Generation with Smaller Large Language Models
Peirong Zhang 0001, Jiaxin Zhang 0003, Jiahuan Cao
Int. J. Comput. Vis.1
2025 Privacy-Preserving Biometric Verification With Handwritten Random Digit String
abstract
Handwriting verification has stood as a steadfast identity authentication method for decades. However, this technique risks potential privacy breaches due to the inclusion of personal information in handwritten biometrics such as signatures. To address this concern, we propose using the Random Digit String (RDS) for privacy-preserving handwriting verification. This approach allows users to authenticate themselves by writing an arbitrary digit sequence, effectively ensuring privacy protection. To evaluate the effectiveness of RDS, we construct a new HRDS4BV dataset composed of online naturally handwritten RDS. Unlike conventional handwriting, RDS encompasses unconstrained and variable content, posing significant challenges for modeling consistent personal writing style. To surmount this, we propose the Pattern Attentive VErification Network (PAVENet), along with a Discriminative Pattern Mining (DPM) module. DPM adaptively enhances the recognition of consistent and discriminative writing patterns, thus refining handwriting style representation. Through comprehensive evaluations, we scrutinize the applicability of online RDS verification and showcase a pronounced outperformance of our model over existing methods. Furthermore, we discover a noteworthy forgery phenomenon that deviates from prior findings and discuss its positive impact in countering malicious impostor attacks. Substantially, our work underscores the feasibility of privacy-preserving biometric verification and propels the prospects of its broader acceptance and application.
Peirong Zhang 0001, Songxuan Lai
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 MegaHan97K: A large-scale dataset for mega-category Chinese character recognition with over 97K categories
Yuyi Zhang 0002, Yongxin Shi, Peirong Zhang 0001, Yixin Zhao, Zhenhua Yang
Pattern Recognit.3
2025 HierCode: A lightweight hierarchical codebook for zero-shot Chinese text recognition
Yuyi Zhang 0002, Dezhi Peng, Peirong Zhang 0001, Zhenhua Yang, Zhibo Yang 0003, Cong Yao
Pattern Recognit.4
2025 Enhancing document dewarping evaluation: A new metric with improved accuracy and efficiency
Jiaxin Zhang 0003, Peirong Zhang 0001, Dezhi Peng
Pattern Recognit. Lett.2
2024 DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
abstract
Document image restoration is a crucial aspect of Document AI systems, as the quality of document images significantly influences the overall performance. Prevailing methods address distinct restoration tasks independently, leading to intricate systems and the incapability to harness the potential synergies of multi-task learning. To overcome this challenge, we propose DocRes, a generalist model that unifies five document image restoration tasks including dewarping, deshadowing, appearance enhancement, deblurring, and binarization. To instruct DocRes to perform various restoration tasks, we propose a novel visual prompt approach called Dynamic Task-Specific Prompt (DTSPrompt). The DTSPrompt for different tasks comprises distinct prior features, which are additional characteristics extracted from the input image. Beyond its role as a cue for task-specific execution, DTSPrompt can also serve as supplementary information to enhance the model's performance. Moreover, DTSPrompt is more flexible than prior visual prompt approaches as it can be seamlessly applied and adapted to inputs with high and variable resolutions. Experimental results demonstrate that DocRes achieves competitive or superior performance compared to existing state-of-the-art task-specific models. This under-scores the potential of DocRes across a broader spectrum of document image restoration tasks. The source code is publicly available at https://github.com/ZZZHANGjx/DocRes.
Jiaxin Zhang 0003, Dezhi Peng, Chongyu Liu, Peirong Zhang 0001
CVPR4
2024 Online Writer Retrieval With Chinese Handwritten Phrases: A Synergistic Temporal-Frequency Representation Learning Approach
abstract
Currently, the prevalence of online handwriting has spurred a critical need for effective retrieval systems to accurately search relevant handwriting instances from specific writers, known as online writer retrieval. Despite the growing demand, this field suffers from a scarcity of well-established methodologies and public large-scale datasets. This paper tackles these challenges with a focus on Chinese handwritten phrases. First, we propose DOLPHIN, a novel retrieval model designed to enhance handwriting representations through synergistic temporal-frequency analysis. For frequency feature learning, we propose the HFGA block, which performs gated cross-attention between the vanilla temporal handwriting sequence and its high-frequency sub-bands to amplify salient writing details. For temporal feature learning, we propose the CAIR block, tailored to promote channel interaction and reduce channel redundancy. Second, to address data deficit, we introduce OLIWER, a large-scale online writer retrieval dataset encompassing over 670,000 Chinese handwritten phrases from 1,731 individuals. Through extensive evaluations, we demonstrate the superior performance of DOLPHIN over existing methods. In addition, we explore cross-domain writer retrieval and reveal the pivotal role of increasing feature alignment in bridging the distributional gap between different handwriting data. Our findings emphasize the significance of point sampling frequency and pressure features in improving handwriting representation quality and retrieval performance. Code and dataset are available athttps:// github.com/SCUT-DLVCLab/DOLPHIN.
Peirong Zhang 0001
IEEE Trans. Inf. Forensics Secur.1
2023 M6Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout Analysis
abstract
Document layout analysis is a crucial prerequisite for document understanding, including document retrieval and conversion. Most public datasets currently contain only PDF documents and lack realistic documents. Models trained on these datasets may not generalize well to real-world scenarios. Therefore, this paper introduces a large and diverse document layout analysis dataset called M6Doc. The M6 designation represents six properties: (1) Multi-Format (including scanned, photographed, and PDF documents); (2) Multi-Type (such as scientific articles, textbooks, books, test papers, magazines, newspapers, and notes); (3) Multi-Layout (rectangular, Manhattan, non-Manhattan, and multi-column Manhattan); (4) Multi-Language (Chinese and English); (5) Multi-Annotation Category (74 types of annotation labels with 237,116 annotation instances in 9,080 manually annotated pages); and (6) Modern documents. Additionally, we propose a transformer-based document layout analysis method called TransDLANet, which leverages an adaptive element matching mechanism that enables query embedding to better match ground truth to improve recall, and constructs a segmentation branch for more precise document image instance segmentation. We conduct a comprehensive evaluation of M6Doc with various layout analysis methods and demonstrate its effectiveness. TransDLANet achieves state-of-the-art performance on M6 Doc with 64.5% mAP. The M6Doc dataset will be available at https://github.com/HcIILAB/M6Doc.
Hiuyi Cheng, Peirong Zhang 0001, Sihang Wu, Jiaxin Zhang 0003, Qiyuan Zhu, Zecheng Xie, Jing Li 0036, Kai Ding 0009
CVPR2