Zhaoran Zhao

dblp:276/3114 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-3406-2598ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Stroke-Based Cyclic Amplifier: Image Super-Resolution at Arbitrary Ultra-Large Scales
abstract
Prior Arbitrary-Scale Image Super-Resolution (ASISR) methods often experience a significant performance decline when the upsampling factor exceeds the range covered by the training data, introducing substantial blurring. To address this issue, we propose a unified model, Stroke-based Cyclic Amplifier (SbCA), for ultra-large upsampling tasks. The key of SbCA is the stroke vector amplifier, which decomposes the image into a series of strokes represented as vector graphics for magnification. Then, the detail completion module also restores missing details, ensuring high-fidelity image reconstruction. Our cyclic strategy achieves ultra-large upsampling by iteratively refining details with this unified SbCA model, trained only once for all, while keeping sub-scales within the training range. Our approach effectively addresses the distribution drift issue and eliminates artifacts, noise and blurring, producing high-quality, high-resolution super-resolved images. Experimental validations on both synthetic and real-world datasets demonstrate that our approach significantly outperforms existing methods in ultra-large upsampling tasks (e.g. $\times 100$ ), delivering visual quality far superior to state-of-the-art techniques.
Wenhao Guo 0003, Peng Lu 0007, Xujun Peng, Zhaoran Zhao, Sheng Li 0008
IEEE Trans. Image Process.4
2025 Can Machines Understand Composition? Dataset and Benchmark for Photographic Image Composition Embedding and Understanding
abstract
With the rapid growth of social media and digital photography, visually appealing images have become essential for effective communication and emotional engagement. Among the factors influencing aesthetic appeal, composition—the arrangement of visual elements within a frame—plays a crucial role. In recent years, specialized models for photographic composition have achieved impressive results across various aesthetic tasks. Meanwhile, rapidly advancing multimodal large language models (MLLMs) have excelled in several visual perception tasks. However, their ability to embed and understand compositional information remains underexplored, primarily due to the lack of suitable evaluation datasets. To address this gap, we introduce the Photographic Image Composition Dataset (PICD), a large-scale dataset consisting of 36,857 images categorized into 24 composition categories across 355 diverse scenes. We demonstrate the advantages of PICD over existing datasets in terms of data scale, composition category, label quality, and scene diversity. Building on PICD, we establish benchmarks to evaluate the composition embedding capabilities of specialized models and the compositional understanding ability of MLLMs. To enable efficient and effective evaluation, we propose a novel Composition Discrimination Accuracy (CDA) metric. Our evaluation highlights the limitations of current models and provides insights into directions for improving their ability to embed and understand composition.
Zhaoran Zhao, Peng Lu 0007, Peipei Li 0002, Xuannan Liu, Shiyi Chen, Wenhao Guo 0003
CVPR1
2025 Learnable adaptive bilateral filter for improved generalization in Single Image Super-Resolution
Wenhao Guo 0003, Peng Lu 0007, Xujun Peng, Zhaoran Zhao
Pattern Recognit.4
2025 Self-Supervised Photographic Image Layout Representation Learning
abstract
Image layout representation learning, which converts layouts into compact vectors, is essential for tasks such as image retrieval, editing, and generation. However, existing methods—especially those applied to photographic images—face several challenges: supervised methods rely on expensive labeled datasets, weakly-supervised methods struggle with generalization, and self-supervised methods are limited in handling the diversity of photographic layouts. To address these issues, we propose a novel heterogeneous layout graph that efficiently captures the layout information in images. The vertices of this graph represent the compositional primitives of the image, capturing their attributes, while the edges encode the relationships between these primitives. We also design effective pretext tasks to guide a layout encoder-decoder in self-supervised training, ultimately generating the layout graph embedding vector. Additionally, we introduce a new layout evaluation dataset—LODB—which features a richer variety of layout categories, significantly better label quality than existing datasets, and a more balanced distribution of semantic scenes across layout categories, providing a comprehensive benchmark for evaluation. Experiments on the LODB dataset demonstrate that our method outperforms existing approaches in representing photographic image layouts.
Zhaoran Zhao, Peng Lu 0007, Xujun Peng, Wenhao Guo 0003
IEEE Trans. Multim.1
2024 BCSCN: Reducing Domain Gap through Bézier Curve basis-based Sparse Coding Network for Single-Image Super-Resolution
abstract
Single Image Super-Resolution (SISR) is a pivotal challenge in computer vision, aiming to restore high-resolution (HR) images from their low-resolution (LR) counterparts. The presence of diverse degradation kernels creates a significant domain gap, limiting the effective generalization of models in real-world scenarios. This study introduces the Bézier Curve basis-based Sparse Coding Network (BCSCN), a preprocessing network designed to mitigate input distribution discrepancies between the training and testing phases of super-resolution networks. BCSCN achieves this by removing visual defects associated with the degradation kernel in LR images, such as artifacts, residual structures, and noise. Additionally, we propose a set of rewards to guide the search for basis coefficients in BCSCN, enhancing the preservation of main content while eliminating information related to degradation. The experimental results highlight the importance of BCSCN, showcasing its capacity to effectively reduce domain gaps and enhance the generalization of super-resolution networks.
Wenhao Guo 0003, Peng Lu 0007, Xujun Peng, Zhaoran Zhao, Xiangtao Dong
ACM Multimedia4
2024 Learning Realistic Sketching: A Dual-agent Reinforcement Learning Approach
abstract
This paper presents a pioneering method for teaching computer sketching that transforms input images into sequential, parameterized strokes. However, two challenges are raised for this sketching task: weak stimuli during stroke decomposition and maintaining semantic correctness, stylistic consistency, and detail integrity in the final drawings. To tackle the challenge of weak stimuli, our method incorporates an attention agent, which enhances the algorithm's sensitivity to subtle canvas changes by focusing on smaller, magnified areas. Moreover, in enhancing the perceived quality of drawing outcomes, we integrate a sketching style feature extractor to seamlessly capture semantic information and execute style adaptation at feature level, alongside a drawing agent that decomposes strokes under the guidance of a fine-grained reward, thereby ensuring the integrity of sketch details. Based on dual intelligent agents, we have constructed an efficient sketching model. Experimental results attest to the superiority of our approach in both visual effects and perceptual metrics when compared to state-of-the-art techniques, confirming its efficacy in achieving realistic sketching.
Peng Lu 0007, Xujun Peng, Wenhao Guo 0003, Zhaoran Zhao, Xiangtao Dong
ACM Multimedia5
2023 RLSCNet: A Residual Line-Shaped Convolutional Network for Vanishing Point Detection
Peng Lu 0007, Xujun Peng, Wang Yin, Zhaoran Zhao
MMM (2)5
2021 Yes, "Attention Is All You Need", for Exemplar based Colorization
abstract
Conventional exemplar based image colorization tends to transfer colors from reference image only to grayscale image based on the semantic correspondence between them. But their practical capabilities are limited when semantic correspondence can hardly be found. To overcome this issue, additional information, such as colors from the database is normally introduced. However, it's a great challenge to consider color information from reference image and database simultaneously because there lacks a unified framework to model different color information and the multi-modal ambiguity in database cannot be removed easily. Also, it is difficult to fuse different color information effectively. Thus, a general attention based colorization framework is proposed in this work, where the color histogram of reference image is adopted as a prior to eliminate the ambiguity in database. Moreover, a sparse loss is designed to guarantee the success of information fusion. Both qualitative and quantitative experimental results show that the proposed approach achieves better colorization performance compared with the state-of-the-art methods on public databases with different quality metrics.
Wang Yin, Peng Lu 0007, Zhaoran Zhao, Xujun Peng
ACM Multimedia3
2020 Gray2ColorNet: Transfer More Colors from Reference Image
abstract
Image colorization is an effective approach to provide plausible colors for grayscale images, which can achieve better and pleasing visual qualities. Although exemplar based colorization approaches provide promising results, they are relied on semantic colors or global colors only from the reference images. For the former situation, when the correspondence between the input grayscale image and reference image is not established, the colors of the reference image cannot be transferred to the input grayscale image successfully. With the later circumstance, because only global colors are considered, it is hard to produce a color image whose objects have the same color as the reference image when they are semantically related. Thus, an end-to-end colorization network Gray2ColorNet is proposed in this work, where an attention gating mechanism based color fusion network is designed to accomplish the colorization tasks. Relied on the proposed method, the semantic colors and global color distribution from the reference image are fused effectively, which are transferred to the final color images along with the prior knowledge of colors contained in the training data. The experimental results demonstrate the superior colorization performances of the proposed method compared to other state-of-the-art approaches.
Peng Lu 0007, Jinbei Yu, Xujun Peng, Zhaoran Zhao, Xiaojie Wang 0006
ACM Multimedia4