EDBT 2026 Demo / reviewers in the wild / expert
Fengjun Guo
dblp:74/9103
· DBLP profile ↗
16ranked-venue papers
1as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 11 · 11 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Based Scalable TypographyabstractCommercial-grade poster design demands the seamless integration of aesthetic appeal with precise, informative content delivery. Current automated poster generation systems face significant limitations, including incomplete design workflows, poor text rendering accuracy, and insufficient flexibility for commercial applications. To address these challenges, we propose PosterVerse, a full-workflow, commercial-grade poster generation method that seamlessly automates the entire design process while delivering high-density and scalable text rendering. PosterVerse replicates professional design through three key stages: (1) blueprint creation using fine-tuned LLMs to extract key design elements from user requirements, (2) graphical background generation via customized diffusion models to create visually appealing imagery, and (3) unified layout-text rendering with an MLLM-powered HTML engine to guarantee high text accuracy and flexible customization. In addition, we introduce PosterDNA, a commercial-grade, HTML-based dataset tailored for training and validating poster design models. To the best of our knowledge, PosterDNA is the first Chinese poster generation dataset to introduce HTML typography files, enabling scalable text rendering and fundamentally solving the challenges of rendering small and high-density text. Experimental results demonstrate that PosterVerse consistently produces commercial-grade posters with appealing visuals, accurate text alignment, and customizable layouts, making it a promising solution for automating commercial poster design. Junle Liu, Peirong Zhang 0001, Yuyi Zhang 0002, Pengyu Yan, Xinyue Zhou, Fengjun Guo |
AAAI | 7 |
| 2026 | DocIQ: A Benchmark Dataset and Feature Fusion Network for Document Image Quality AssessmentabstractDocument image quality assessment (DIQA) is an important component for various applications, including optical character recognition (OCR), document restoration, and the evaluation of document image processing systems. In this paper, we introduce a subjective DIQA dataset DIQA-5000. The DIQA-5000 dataset comprises 5,000 document images, generated by applying multiple document enhancement techniques to 500 real-world images with diverse distortions. Each enhanced image was rated by 15 subjects across three rating dimensions: overall quality, sharpness, and color fidelity. Furthermore, we propose a specialized no-reference DIQA model that exploits document layout features to maintain quality perception at reduced resolutions to lower computational cost. Recognizing that image quality is influenced by both low-level and high-level visual features, we designed a feature fusion module to extract and integrate multi-level features from document images. To generate multi-dimensional scores, our model employs independent quality heads for each dimension to predict score distributions, allowing it to learn distinct aspects of document image quality. Experimental results demonstrate that our method outperforms current state-of-the-art general-purpose IQA models on both DIQA-5000 and an additional document image dataset focused on OCR accuracy. Fengjun Guo, Guangtao Zhai, Xiongkuo Min |
ISCAS | 4 |
| 2025 | Revisiting Tampered Scene Text Detection in the Era of Generative AIabstractThe rapid advancements of generative AI have fueled the potential of generative text image editing, meanwhile escalating the threat of misinformation spreading. However, existing forensics methods struggle to detect unseen forgery types that they have not been trained on, underscoring the need for a model capable of generalized detection of tampered scene text. To tackle this, we propose a novel task: open-set tampered scene text detection, which evaluates forensics models on their ability to identify both seen and previously unseen forgery types. We have curated a comprehensive, high-quality dataset, featuring the texts tampered by eight text editing models, to thoroughly assess the open-set generalization capabilities. Further, we introduce a novel and effective pre-training paradigm that subtly alters the texture of selected texts within an image and trains the model to identify these regions. This approach not only mitigates the scarcity of high-quality training data but also enhances models' fine-grained perception and open-set generalization abilities. Additionally, we present DAF, a novel framework that improves open-set generalization by distinguishing between the features of authentic and tampered text, rather than focusing solely on the tampered text's features. Our extensive experiments validate the remarkable efficacy of our methods. For example, our zero-shot performance can even beat the previous state-of-the-art full-shot model by a large margin. Chenfan Qu, Yiwu Zhong, Fengjun Guo |
AAAI | 3 |
| 2025 | Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document RestorationabstractYuyi Zhang, Peirong Zhang, Zhenhua Yang, Pengyu Yan, Yongxin Shi, Pengwei Liu, Fengjun Guo, Lianwen Jin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yuyi Zhang 0002, Peirong Zhang 0001, Zhenhua Yang, Pengyu Yan, Yongxin Shi, Pengwei Liu, Fengjun Guo |
ACL (1) | 7 |
| 2024 | Towards Modern Image Manipulation Localization: A Large-Scale Dataset and Novel MethodsabstractIn recent years, image manipulation localization has attracted increasing attention due to its pivotal role in guaranteeing social media security. However, how to accurately identify the forged regions remains an open challenge. One of the main bottlenecks lies in the severe scarcity of high-quality data, due to its costly creation process. To address this limitation, we propose a novel paradigm, termed as CAAA, to automatically and precisely annotate the numerous manually forged images from the web at the pixel level. We further propose a novel metric QES to facilitate the automatic filtering of unreliable annotations. With CAAA and QES, we construct a large-scale, diverse, and high-quality dataset comprising 123,150 manually forged images with mask annotations. Besides, we develop a new model APSC-Net for accurate image manipulation localization. According to extensive experiments, our dataset significantly improves the performance of various models on the widely-used benchmarks and such improvements are attributed to our proposed effective methods. The dataset and code are publicly available at https://github.com/qcf-568/MIML. Chenfan Qu, Yiwu Zhong, Chongyu Liu, Guitao Xu, Dezhi Peng, Fengjun Guo |
CVPR | 6 |
| 2024 | Leveraging Text Localization for Scene Text Removal via Text-Aware Masked Image Modeling
Zixiao Wang 0002, Hongtao Xie 0001, Yuxin Wang 0002, Yadong Qu, Fengjun Guo, Pengwei Liu |
ECCV (66) | 5 |
| 2024 | Coarse-to-Fine Document Image Registration for Dewarping
Qiufeng Wang 0001, Kaizhu Huang, Xiaomeng Gu, Fengjun Guo |
ICDAR (4) | 5 |
| 2024 | UPOCR: Towards Unified Pixel-Level OCR InterfaceabstractExisting optical character recognition (OCR) methods rely on task-specific designs with divergent paradigms, architectures, and training strategies, which significantly increases the complexity of research and maintenance and hinders the fast deployment in applications. To this end, we propose UPOCR, a simple-yet-effective generalist model for Unified Pixel-level OCR interface. Specifically, the UPOCR unifies the paradigm of diverse OCR tasks as image-to-image transformation and the architecture as a vision Transformer (ViT)-based encoder-decoder with learnable task prompts. The prompts push the general feature representations extracted by the encoder towards task-specific spaces, endowing the decoder with task awareness. Moreover, the model training is uniformly aimed at minimizing the discrepancy between the predicted and ground-truth images regardless of the inhomogeneity among tasks. Experiments are conducted on three pixel-level OCR tasks including text removal, text segmentation, and tampered text detection. Without bells and whistles, the experimental results showcase that the proposed method can simultaneously achieve state-of-the-art performance on three tasks with a unified single model, which provides valuable strategies and insights for future research on generalist OCR models. Code is available at https://github.com/shannanyinxiang/UPOCR. Dezhi Peng, Zhenhua Yang, Jiaxin Zhang 0003, Chongyu Liu, Yongxin Shi, Kai Ding 0009, Fengjun Guo |
ICML | 7 |
| 2024 | Document Registration: Towards Automated Labeling of Pixel-Level Alignment Between Warped-Flat Documents
Qiufeng Wang 0001, Kaizhu Huang, Xiaowei Huang 0001, Fengjun Guo, Xiaomeng Gu |
ACM Multimedia | 5 |
| 2023 | Towards Robust Tampered Text Detection in Document Image: New Dataset and New SolutionabstractRecently, tampered text detection in document image has attracted increasingly attention due to its essential role on information security. However, detecting visually consistent tampered text in photographed document images is still a main challenge. In this paper, we propose a novel framework to capture more fine-grained clues in complex scenarios for tampered text detection, termed as Document Tampering Detector (DTD), which consists of a Frequency Perception Head (FPH) to compensate the deficiencies caused by the inconspicuous visual features, and a Multi-view Iterative Decoder (MID) for fully utilizing the information of features in different scales. In addition, we design a new training paradigm, termed as Curriculum Learning for Tampering Detection (CLTD), which can address the confusion during the training procedure and thus to improve the robustness for image compression and the ability to generalize. To further facilitate the tampered text detection in document images, we construct a large-scale document image dataset, termed as DocTamper, which contains 170,000 document images of various types. Experiments demonstrate that our proposed DTD outperforms previous state-of-the-art by 9.2%, 26.3% and 12.3% in terms of F-measure on the DocTamper testing set, and the crossdomain testing sets of DocTamper-FCD and DocTamper-SCD, respectively. Codes and dataset will be available at https://github.com/qcf-568/DocTamper. Chenfan Qu, Chongyu Liu, Xinhong Chen 0005, Dezhi Peng, Fengjun Guo |
CVPR | 6 |
| 2023 | Deep Image Harmonization with Globally Guided Feature Transformation and Relation DistillationabstractGiven a composite image, image harmonization aims to adjust the foreground illumination to be consistent with background. Previous methods have explored transforming foreground features to achieve competitive performance. In this work, we show that using global information to guide foreground feature transformation could achieve significant improvement. Besides, we propose to transfer the foreground-background relation from real images to composite images, which can provide intermediate supervision for the transformed encoder features. Additionally, considering the drawbacks of existing harmonization datasets, we also contribute a ccHarmony dataset which simulates the natural illumination variation. Extensive experiments on iHarmony4 and our contributed dataset demonstrate the superiority of our method. Our ccHarmony dataset is released at https://github.com/bcmi/Image-HarmonizationDataset-ccHarmony. Li Niu 0002, Linfeng Tan, Xinhao Tao, Junyan Cao, Fengjun Guo, Liqing Zhang 0001 |
ICCV | 5 |
| 2022 | Inharmonious Region Localization by Magnifying Domain DiscrepancyabstractInharmonious region localization aims to localize the region in a synthetic image which is incompatible with surrounding background. The inharmony issue is mainly attributed to the color and illumination inconsistency produced by image editing techniques. In this work, we tend to transform the input image to another color space to magnify the domain discrepancy between inharmonious region and background, so that the model can identify the inharmonious region more easily. To this end, we present a novel framework consisting of a color mapping module and an inharmonious region localization network, in which the former is equipped with a novel domain discrepancy magnification loss and the latter could be an arbitrary localization network. Extensive experiments on image harmonization dataset show the superiority of our designed framework. Jing Liang 0007, Li Niu 0002, Penghao Wu, Fengjun Guo |
AAAI | 4 |
| 2022 | Don't Forget Me: Accurate Background Recovery for Text Removal via Modeling Local-Global Context
Chongyu Liu, Canjie Luo, Bangdong Chen, Fengjun Guo, Kai Ding 0009 |
ECCV (28) | 6 |
| 2022 | Marior: Margin Removal and Iterative Content Rectification for Document Dewarping in the WildabstractCamera-captured document images usually suffer from perspective and geometric deformations. It is of great value to rectify them when considering poor visual aesthetics and the deteriorated performance of OCR systems. Recent learning-based methods intensively focus on the accurately cropped document image. However, this might not be sufficient for overcoming practical challenges, including document images either with large marginal regions or without margins. Due to this impracticality, users struggle to crop documents precisely when they encounter large marginal regions. Simultaneously, dewarping images without margins is still an insurmountable problem. To the best of our knowledge, there is still no complete and effective pipeline for rectifying document images in the wild. To address this issue, we propose a novel approach called Marior (Margin Removal and Iterative Content Rectification). Marior follows a progressive strategy to iteratively improve the dewarping quality and readability in a coarse-to-fine manner. Specifically, we divide the pipeline into two modules: margin removal module (MRM) and iterative content rectification module (ICRM). First, we predict the segmentation mask of the input image to remove the margin, thereby obtaining a preliminary result. Then we refine the image further by producing dense displacement flows to achieve content-aware rectification. We determine the number of refinement iterations adaptively. Experiments demonstrate the state-of-the-art performance of our method on public benchmarks. The resources are available at https://github.com/ZZZHANG-jx/Marior for further comparison. Jiaxin Zhang 0003, Canjie Luo, Fengjun Guo, Kai Ding 0009 |
ACM Multimedia | 4 |
| 2021 | Visible Watermark Removal via Self-calibrated Localization and Background RefinementabstractSuperimposing visible watermarks on images provides a powerful weapon to cope with the copyright issue. Watermark removal techniques, which can strengthen the robustness of visible watermarks in an adversarial way, have attracted increasing research interest. Modern watermark removal methods perform watermark localization and background restoration simultaneously, which could be viewed as a multi-task learning problem. However, existing approaches suffer from incomplete detected watermark and degraded texture quality of restored background. Therefore, we design a two-stage multi-task network to address the above issues. The coarse stage consists of a watermark branch and a background branch, in which the watermark branch self-calibrates the roughly estimated mask and passes the calibrated mask to background branch to reconstruct the watermarked area. In the refinement stage, we integrate multi-level features to improve the texture quality of watermarked area. Extensive experiments on two datasets demonstrate the effectiveness of our proposed method. Jing Liang 0007, Li Niu 0002, Fengjun Guo, Liqing Zhang 0001 |
ACM Multimedia | 3 |
| 2010 | Gesture Recognition Techniques in Handwriting Recognition ApplicationabstractHandwriting-gesture recognition has been widely implemented in handwriting input application. Usually, gestures are used to conduct edit operations or be set as short-cut of an application. In this paper, we compare several handwriting-gesture recognition methods, and address their different user cases. These methods include pixel-matching method, rule based method and discriminant-function based method. For discriminant-function based method, we describe 2 sub-methods. They are prototypes based method and training based method. We not only analyze recognition accuracy of gestures for these methods, but also analyze their distinguished capability when recognizing gestures and alphanumeric in same recognizing mode. Experiments results show that, if the gesture-samples are enough, training based method achieves the highest accuracy. Furthermore, when recognizing mixed input of gestures and other handwriting symbols, training based method almost doesn't degrade accuracy of these symbols. Fengjun Guo |
ICFHR | 1 |