Peng Lu 0007

dblp:86/241-7 · DBLP profile ↗
← Back
25ranked-venue papers
11as first author
11since 2021 · last 2026
0000-0002-0162-6449ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 10 first-author · 10 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Stroke-Based Cyclic Amplifier: Image Super-Resolution at Arbitrary Ultra-Large Scales
abstract
Prior Arbitrary-Scale Image Super-Resolution (ASISR) methods often experience a significant performance decline when the upsampling factor exceeds the range covered by the training data, introducing substantial blurring. To address this issue, we propose a unified model, Stroke-based Cyclic Amplifier (SbCA), for ultra-large upsampling tasks. The key of SbCA is the stroke vector amplifier, which decomposes the image into a series of strokes represented as vector graphics for magnification. Then, the detail completion module also restores missing details, ensuring high-fidelity image reconstruction. Our cyclic strategy achieves ultra-large upsampling by iteratively refining details with this unified SbCA model, trained only once for all, while keeping sub-scales within the training range. Our approach effectively addresses the distribution drift issue and eliminates artifacts, noise and blurring, producing high-quality, high-resolution super-resolved images. Experimental validations on both synthetic and real-world datasets demonstrate that our approach significantly outperforms existing methods in ultra-large upsampling tasks (e.g. $\times 100$ ), delivering visual quality far superior to state-of-the-art techniques.
Wenhao Guo 0003, Peng Lu 0007, Xujun Peng, Zhaoran Zhao, Sheng Li 0008
IEEE Trans. Image Process.2
2025 Can Machines Understand Composition? Dataset and Benchmark for Photographic Image Composition Embedding and Understanding
abstract
With the rapid growth of social media and digital photography, visually appealing images have become essential for effective communication and emotional engagement. Among the factors influencing aesthetic appeal, composition—the arrangement of visual elements within a frame—plays a crucial role. In recent years, specialized models for photographic composition have achieved impressive results across various aesthetic tasks. Meanwhile, rapidly advancing multimodal large language models (MLLMs) have excelled in several visual perception tasks. However, their ability to embed and understand compositional information remains underexplored, primarily due to the lack of suitable evaluation datasets. To address this gap, we introduce the Photographic Image Composition Dataset (PICD), a large-scale dataset consisting of 36,857 images categorized into 24 composition categories across 355 diverse scenes. We demonstrate the advantages of PICD over existing datasets in terms of data scale, composition category, label quality, and scene diversity. Building on PICD, we establish benchmarks to evaluate the composition embedding capabilities of specialized models and the compositional understanding ability of MLLMs. To enable efficient and effective evaluation, we propose a novel Composition Discrimination Accuracy (CDA) metric. Our evaluation highlights the limitations of current models and provides insights into directions for improving their ability to embed and understand composition.
Zhaoran Zhao, Peng Lu 0007, Peipei Li 0002, Xuannan Liu, Shiyi Chen, Wenhao Guo 0003
CVPR2
2025 Learnable adaptive bilateral filter for improved generalization in Single Image Super-Resolution
Wenhao Guo 0003, Peng Lu 0007, Xujun Peng, Zhaoran Zhao
Pattern Recognit.2
2025 Self-Supervised Photographic Image Layout Representation Learning
abstract
Image layout representation learning, which converts layouts into compact vectors, is essential for tasks such as image retrieval, editing, and generation. However, existing methods—especially those applied to photographic images—face several challenges: supervised methods rely on expensive labeled datasets, weakly-supervised methods struggle with generalization, and self-supervised methods are limited in handling the diversity of photographic layouts. To address these issues, we propose a novel heterogeneous layout graph that efficiently captures the layout information in images. The vertices of this graph represent the compositional primitives of the image, capturing their attributes, while the edges encode the relationships between these primitives. We also design effective pretext tasks to guide a layout encoder-decoder in self-supervised training, ultimately generating the layout graph embedding vector. Additionally, we introduce a new layout evaluation dataset—LODB—which features a richer variety of layout categories, significantly better label quality than existing datasets, and a more balanced distribution of semantic scenes across layout categories, providing a comprehensive benchmark for evaluation. Experiments on the LODB dataset demonstrate that our method outperforms existing approaches in representing photographic image layouts.
Zhaoran Zhao, Peng Lu 0007, Xujun Peng, Wenhao Guo 0003
IEEE Trans. Multim.2
2024 COLORFLOW: A Conditional Normalizing Flow for Image Colorization
abstract
Image colorization is an ill-posed task, as objects within grayscale images can correspond to multiple colors, motivating researchers to establish a one-to-many relationship between objects and colors. Previous work mostly could only create an insufficient deterministic relationship. Normalizing flow can fully capture the color diversity from natural image manifold. However, classical flow often overlooks the color correlations between different objects, resulting in generating unrealistic color. To solve this issue, we propose a conditional normalizing flow, named ColorFlow, to jointly learn the one-to-many relationships between objects and colors, and the color correlations between different objects within image. To represent these color correlations in flow, we design a color distribution predictor to estimate the global color histogram of grayscale image as global tones, which is utilized as the mean value of flow’s latent variables. Experiments results show that ColorFlow outperforms state-of-the-art methods.
Wang Yin, Peng Lu 0007, Xujun Peng
ICASSP2
2024 BCSCN: Reducing Domain Gap through Bézier Curve basis-based Sparse Coding Network for Single-Image Super-Resolution
abstract
Single Image Super-Resolution (SISR) is a pivotal challenge in computer vision, aiming to restore high-resolution (HR) images from their low-resolution (LR) counterparts. The presence of diverse degradation kernels creates a significant domain gap, limiting the effective generalization of models in real-world scenarios. This study introduces the Bézier Curve basis-based Sparse Coding Network (BCSCN), a preprocessing network designed to mitigate input distribution discrepancies between the training and testing phases of super-resolution networks. BCSCN achieves this by removing visual defects associated with the degradation kernel in LR images, such as artifacts, residual structures, and noise. Additionally, we propose a set of rewards to guide the search for basis coefficients in BCSCN, enhancing the preservation of main content while eliminating information related to degradation. The experimental results highlight the importance of BCSCN, showcasing its capacity to effectively reduce domain gaps and enhance the generalization of super-resolution networks.
Wenhao Guo 0003, Peng Lu 0007, Xujun Peng, Zhaoran Zhao, Xiangtao Dong
ACM Multimedia2
2024 Learning Realistic Sketching: A Dual-agent Reinforcement Learning Approach
abstract
This paper presents a pioneering method for teaching computer sketching that transforms input images into sequential, parameterized strokes. However, two challenges are raised for this sketching task: weak stimuli during stroke decomposition and maintaining semantic correctness, stylistic consistency, and detail integrity in the final drawings. To tackle the challenge of weak stimuli, our method incorporates an attention agent, which enhances the algorithm's sensitivity to subtle canvas changes by focusing on smaller, magnified areas. Moreover, in enhancing the perceived quality of drawing outcomes, we integrate a sketching style feature extractor to seamlessly capture semantic information and execute style adaptation at feature level, alongside a drawing agent that decomposes strokes under the guidance of a fine-grained reward, thereby ensuring the integrity of sketch details. Based on dual intelligent agents, we have constructed an efficient sketching model. Experimental results attest to the superiority of our approach in both visual effects and perceptual metrics when compared to state-of-the-art techniques, confirming its efficacy in achieving realistic sketching.
Peng Lu 0007, Xujun Peng, Wenhao Guo 0003, Zhaoran Zhao, Xiangtao Dong
ACM Multimedia2
2023 Learning to Draw Through A Multi-Stage Environment Model Based Reinforcement Learning
abstract
Machine drawing has gradually become a hot research topic in computer vision and robotics domains recently. However, decomposing a given target image from raster space into an ordered sequence and reconstructing those strokes is a challenging task. In this work, we focus on the drawing task for the images in various styles where the distribution of stroke parameters differs. We propose a multi-stage environment model based reinforcement learning (RL) drawing framework with fine-grained perceptual reward to guide the agent under this framework to draw details and an overall outline of the target image accurately. The experiments show that the visual quality of our method slightly outperforms SOTA method in nature and doodle style, while it outperforms the SOTA approaches by a large margin with high efficiency in sketch style.
Peng Lu 0007, Xujun Peng
ICIP2
2023 RLSCNet: A Residual Line-Shaped Convolutional Network for Vanishing Point Detection
Peng Lu 0007, Xujun Peng, Wang Yin, Zhaoran Zhao
MMM (2)2
2021 Yes, "Attention Is All You Need", for Exemplar based Colorization
abstract
Conventional exemplar based image colorization tends to transfer colors from reference image only to grayscale image based on the semantic correspondence between them. But their practical capabilities are limited when semantic correspondence can hardly be found. To overcome this issue, additional information, such as colors from the database is normally introduced. However, it's a great challenge to consider color information from reference image and database simultaneously because there lacks a unified framework to model different color information and the multi-modal ambiguity in database cannot be removed easily. Also, it is difficult to fuse different color information effectively. Thus, a general attention based colorization framework is proposed in this work, where the color histogram of reference image is adopted as a prior to eliminate the ambiguity in database. Moreover, a sparse loss is designed to guarantee the success of information fusion. Both qualitative and quantitative experimental results show that the proposed approach achieves better colorization performance compared with the state-of-the-art methods on public databases with different quality metrics.
Wang Yin, Peng Lu 0007, Zhaoran Zhao, Xujun Peng
ACM Multimedia2
2021 Learning the Relation Between Interested Objects and Aesthetic Region for Image Cropping
abstract
As one of the fundamental techniques for image editing, image cropping discards irrelevant contents and remains the pleasing portions of the image to enhance the overall composition and achieve better visual/aesthetic perception. In this paper, we primarily focus on improving the efficiency of automatic image cropping, and on further exploring its potential in public datasets with high accuracy. From this perspective, we propose a deep learning based framework to learn the objects composition from photos with high aesthetic qualities, where an interested object region is detected through a convolutional neural network (CNN) based on the saliency map. The features of the detected interested objects are then fed into a regression network to obtain the final cropping result. Unlike the conventional methods that multiple candidates are proposed and evaluated iteratively, only a single interested object region is produced in our model, which is mapped to the final output directly. Thus, low computational resources are required for the proposed approach. The experimental results on the public datasets show that as a weakly supervised method, the proposed network outperforms the other weakly supervised methods on FLMS and FCD datasets and achieves comparable results to the existing methods on CUHK dataset. Furthermore, the proposed method is more efficient than these methods, where the processing speed is as fast as 20 ms per image.
Peng Lu 0007, Xujun Peng, Xiaofu Jin
IEEE Trans. Multim.1
2020 Weakly Supervised Real-time Image Cropping based on Aesthetic Distributions
abstract
Image cropping is an effective tool to edit and manipulate images to achieve better aesthetic quality. Most existing cropping approaches rely on the two-step paradigm where multiple candidate cropping areas are proposed initially and the optimal cropping window is determined based on some quality criteria for these candidates afterwards. The obvious disadvantage of this mechanism is its low efficiency due to the huge searching space of candidate crops. In order to tackle this problem, a weakly supervised cropping framework is proposed, where the distribution dissimilarity between high quality images and cropped images is used to guide the coordinate predictor's training and the ground truths of cropping windows are not required by the proposed method. Meanwhile, to improve the cropping performance, a saliency loss is also designed in the proposed framework to force the neural network to focus more on the interested objects in the image. Under this framework, the images can be cropped effectively by the trained coordinate predictor in a one-pass favor without multiple candidates proposals, which ensures the high efficiency of the proposed system . Also, based on the proposed framework, many existing distribution dissimilarity measurements can be applied to train the image cropping system with high flexibility, such as likelihood based and divergence based distribution dissimilarity measure proposed in this work. The experiments on the public databases show that the proposed cropping method achieves the state-of-the-art accuracy, and the high computation efficiency as fast as 285 FPS is also obtained.
Peng Lu 0007, Xujun Peng, Xiaojie Wang 0006
ACM Multimedia1
2020 Gray2ColorNet: Transfer More Colors from Reference Image
abstract
Image colorization is an effective approach to provide plausible colors for grayscale images, which can achieve better and pleasing visual qualities. Although exemplar based colorization approaches provide promising results, they are relied on semantic colors or global colors only from the reference images. For the former situation, when the correspondence between the input grayscale image and reference image is not established, the colors of the reference image cannot be transferred to the input grayscale image successfully. With the later circumstance, because only global colors are considered, it is hard to produce a color image whose objects have the same color as the reference image when they are semantically related. Thus, an end-to-end colorization network Gray2ColorNet is proposed in this work, where an attention gating mechanism based color fusion network is designed to accomplish the colorization tasks. Relied on the proposed method, the semantic colors and global color distribution from the reference image are fused effectively, which are transferred to the final color images along with the prior knowledge of colors contained in the training data. The experimental results demonstrate the superior colorization performances of the proposed method compared to other state-of-the-art approaches.
Peng Lu 0007, Jinbei Yu, Xujun Peng, Zhaoran Zhao, Xiaojie Wang 0006
ACM Multimedia1
2019 Gated CNN for visual quality assessment based on color perception
Peng Lu 0007, Xujun Peng, Jinbei Yu
Signal Process. Image Commun.1
2019 Aesthetic guided deep regression network for image cropping
Peng Lu 0007, Xujun Peng
Signal Process. Image Commun.1
2018 Sobel Heuristic Kernel for Aerial Semantic Segmentation
abstract
Misclassification in semantic segmentation mostly occurs in the pixels around the semantic contour. In this work, we address the task of aerial image segmentation by borrowing the kernel prior from classical edge detecting operator. We propose a module called Sobel Heuristic Kernel(SHK). Our work makes several main contributions and experimentally shows good performance. To the best of our knowledge, we are the first to combine traditional edge detection method and deep learning method in semantic segmentation. Our SHK module reaches state of the art in the Inria Aerial Image Labeling dataset.
Yao Wang 0018, Yisong Chen, Peng Lu 0007
ICIP4
2018 Deep Conditional Color Harmony Model for Image Aesthetic Assessment
abstract
As one of the important features, color provides plenty useful information to represent images. Thus, color harmony, which is defined as “two or more colors are sensed together as a single, pleasing, collective impression” [1], can also be served as a fundamental feature and plays a key role to determine the aesthetics quality of images. To reveal the inherent color harmony attribute within patch and the harmonious relations between image patches which construct the pleasing colorful images, we designed a conditional random field (CRF) based color harmony model in this paper to accomplish the image aesthetic assessment tasks. Unlike the previous learning based color harmony models, we used deep neural networks to obtain the coherence properties between original image patch pairs, and embedded these relations along with each patch's own color harmony characteristic into a CRF to measure the harmony scores of the entire image. The experimental results on a public dataset show that the proposed deep conditional color harmony model is superior to the existing color harmony models in respect of the image aesthetic assessment.
Peng Lu 0007, Jinbei Yu, Xujun Peng
ICPR1
2018 Large-Scale Structure from Motion with Semantic Constraints of Aerial Images
Yao Wang 0018, Peng Lu 0007, Yisong Chen
PRCV (1)3
2016 Image color harmony modeling through neighbored co-occurrence colors
Peng Lu 0007, Xujun Peng, Caixia Yuan, Ruifan Li, Xiaojie Wang 0006
Neurocomputing1
2016 An EL-LDA based general color harmony model for photo aesthetics assessment
Peng Lu 0007, Xujun Peng, Xinshan Zhu, Ruifan Li
Signal Process.1
2015 Towards aesthetics of image: A Bayesian framework for color harmony modeling
Peng Lu 0007, Xujun Peng, Ruifan Li, Xiaojie Wang 0006
Signal Process. Image Commun.1
2015 Finding more relevance: Propagating similarity on Markov random field for object retrieval
Peng Lu 0007, Xujun Peng, Xinshan Zhu, Ruifan Li
Signal Process. Image Commun.1
2014 Discovering Harmony: A Hierarchical Colour Harmony Model for Aesthetics Assessment
Peng Lu 0007, Zhijie Kuang, Xujun Peng, Ruifan Li
ACCV (3)1
2014 Object Ranking on Deformable Part Models with Bagged LambdaMART
Chaobo Sun, Xiaojie Wang 0006, Peng Lu 0007
ACCV (2)3
2014 Sentiment analysis of microblog combining dictionary and rules
abstract
Microblog has become a daily communication tool in recent years. Researches on microblog have drawn more and more attention. Microblogging emotional classification is a major research of user intent analysis based on User-Generated Content (UGC). This paper focuses on the discrimination on two emotional tendencies: positive and negative. Firstly, the system cleared the noisy elements in the microblog, then extracted the features of the microblog and finally classified the microblog using Support Vector Machine (SVM). Furthermore, we improve the algorithms of feature extraction and weight computing combining dictionary approach and rule based approach. The result of experiment shows that the method is effective.
Yanquan Zhou, Ruifan Li, Peng Lu 0007
ASONAM4