VLDB 2026 Research / reviewers in the wild / expert
Wenhao Guo 0003
dblp:09/4580-3
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0006-0143-6238ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
5 papers |
Image and video processing · 52% Visual content generation and editing · 24% Computational photography and imaging · 13% | |
| Artificial intelligence
3 papers |
Vision and language · 39% Representation and self-supervised learning · 39% Graph learning · 12% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing › super-resolution
image super-resolution |
1.8 | 2 | 2026 | Stroke-Based Cyclic Amplifier: Image Super-Resolution at Arbitrary Ultra-Large Scales · IEEE Trans. Image Process. 2026 BCSCN: Reducing Domain Gap through Bézier Curve basis-based Sparse Coding Network for Single-Image Super-Resolution · ACM Multimedia 2024 |
Image and video processing › super-resolution › image super-resolution
arbitrary-scale super-resolution |
1.0 | 1 | 2026 | Stroke-Based Cyclic Amplifier: Image Super-Resolution at Arbitrary Ultra-Large Scales · IEEE Trans. Image Process. 2026 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | Can Machines Understand Composition? Dataset and Benchmark for Photographic Image Composition Embedding and Understanding · CVPR 2025 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
pretext task |
0.9 | 1 | 2025 | Self-Supervised Photographic Image Layout Representation Learning · IEEE Trans. Multim. 2025 |
Computational photography and imaging
image aesthetics |
0.9 | 1 | 2025 | Can Machines Understand Composition? Dataset and Benchmark for Photographic Image Composition Embedding and Understanding · CVPR 2025 |
Image and video processing › super-resolution › image super-resolution
single image super-resolution |
0.8 | 1 | 2024 | BCSCN: Reducing Domain Gap through Bézier Curve basis-based Sparse Coding Network for Single-Image Super-Resolution · ACM Multimedia 2024 |
Visual content generation and editing
sketch generation |
0.8 | 1 | 2024 | Learning Realistic Sketching: A Dual-agent Reinforcement Learning Approach · ACM Multimedia 2024 |
Rendering
stroke-based rendering |
0.8 | 1 | 2024 | Learning Realistic Sketching: A Dual-agent Reinforcement Learning Approach · ACM Multimedia 2024 |
Machine learning › Graph learning
network embedding |
0.3 | 1 | 2025 | Self-Supervised Photographic Image Layout Representation Learning · IEEE Trans. Multim. 2025 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.2 | 1 | 2024 | Learning Realistic Sketching: A Dual-agent Reinforcement Learning Approach · ACM Multimedia 2024 |
Methods — techniques the papers use, named apart from their topics
multimodal large language model · 1.7layout encoder-decoder · 1.7heterogeneous layout graph · 1.7reinforcement learning · 1.5attention mechanism · 1.5stroke vector amplification · 1.0cyclic refinement · 1.0style feature extraction · 0.8sparse coding · 0.8reinforcement-guided coefficient search · 0.8bézier curve basis · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Stroke-Based Cyclic Amplifier: Image Super-Resolution at Arbitrary Ultra-Large ScalesabstractPrior Arbitrary-Scale Image Super-Resolution (ASISR) methods often experience a significant performance decline when the upsampling factor exceeds the range covered by the training data, introducing substantial blurring. To address this issue, we propose a unified model, Stroke-based Cyclic Amplifier (SbCA), for ultra-large upsampling tasks. The key of SbCA is the stroke vector amplifier, which decomposes the image into a series of strokes represented as vector graphics for magnification. Then, the detail completion module also restores missing details, ensuring high-fidelity image reconstruction. Our cyclic strategy achieves ultra-large upsampling by iteratively refining details with this unified SbCA model, trained only once for all, while keeping sub-scales within the training range. Our approach effectively addresses the distribution drift issue and eliminates artifacts, noise and blurring, producing high-quality, high-resolution super-resolved images. Experimental validations on both synthetic and real-world datasets demonstrate that our approach significantly outperforms existing methods in ultra-large upsampling tasks (e.g. $\times 100$ ), delivering visual quality far superior to state-of-the-art techniques. Wenhao Guo 0003, Peng Lu 0007, Xujun Peng, Zhaoran Zhao, Sheng Li 0008 |
IEEE Trans. Image Process. | 1 |
| 2025 | Can Machines Understand Composition? Dataset and Benchmark for Photographic Image Composition Embedding and UnderstandingabstractWith the rapid growth of social media and digital photography, visually appealing images have become essential for effective communication and emotional engagement. Among the factors influencing aesthetic appeal, composition—the arrangement of visual elements within a frame—plays a crucial role. In recent years, specialized models for photographic composition have achieved impressive results across various aesthetic tasks. Meanwhile, rapidly advancing multimodal large language models (MLLMs) have excelled in several visual perception tasks. However, their ability to embed and understand compositional information remains underexplored, primarily due to the lack of suitable evaluation datasets. To address this gap, we introduce the Photographic Image Composition Dataset (PICD), a large-scale dataset consisting of 36,857 images categorized into 24 composition categories across 355 diverse scenes. We demonstrate the advantages of PICD over existing datasets in terms of data scale, composition category, label quality, and scene diversity. Building on PICD, we establish benchmarks to evaluate the composition embedding capabilities of specialized models and the compositional understanding ability of MLLMs. To enable efficient and effective evaluation, we propose a novel Composition Discrimination Accuracy (CDA) metric. Our evaluation highlights the limitations of current models and provides insights into directions for improving their ability to embed and understand composition. Zhaoran Zhao, Peng Lu 0007, Peipei Li 0002, Xuannan Liu, Shiyi Chen, Wenhao Guo 0003 |
CVPR | 10 |
| 2025 | Learnable adaptive bilateral filter for improved generalization in Single Image Super-Resolution
Wenhao Guo 0003, Peng Lu 0007, Xujun Peng, Zhaoran Zhao |
Pattern Recognit. | 1 |
| 2025 | Self-Supervised Photographic Image Layout Representation LearningabstractImage layout representation learning, which converts layouts into compact vectors, is essential for tasks such as image retrieval, editing, and generation. However, existing methods—especially those applied to photographic images—face several challenges: supervised methods rely on expensive labeled datasets, weakly-supervised methods struggle with generalization, and self-supervised methods are limited in handling the diversity of photographic layouts. To address these issues, we propose a novel heterogeneous layout graph that efficiently captures the layout information in images. The vertices of this graph represent the compositional primitives of the image, capturing their attributes, while the edges encode the relationships between these primitives. We also design effective pretext tasks to guide a layout encoder-decoder in self-supervised training, ultimately generating the layout graph embedding vector. Additionally, we introduce a new layout evaluation dataset—LODB—which features a richer variety of layout categories, significantly better label quality than existing datasets, and a more balanced distribution of semantic scenes across layout categories, providing a comprehensive benchmark for evaluation. Experiments on the LODB dataset demonstrate that our method outperforms existing approaches in representing photographic image layouts. Zhaoran Zhao, Peng Lu 0007, Xujun Peng, Wenhao Guo 0003 |
IEEE Trans. Multim. | 4 |
| 2024 | BCSCN: Reducing Domain Gap through Bézier Curve basis-based Sparse Coding Network for Single-Image Super-ResolutionabstractSingle Image Super-Resolution (SISR) is a pivotal challenge in computer vision, aiming to restore high-resolution (HR) images from their low-resolution (LR) counterparts. The presence of diverse degradation kernels creates a significant domain gap, limiting the effective generalization of models in real-world scenarios. This study introduces the Bézier Curve basis-based Sparse Coding Network (BCSCN), a preprocessing network designed to mitigate input distribution discrepancies between the training and testing phases of super-resolution networks. BCSCN achieves this by removing visual defects associated with the degradation kernel in LR images, such as artifacts, residual structures, and noise. Additionally, we propose a set of rewards to guide the search for basis coefficients in BCSCN, enhancing the preservation of main content while eliminating information related to degradation. The experimental results highlight the importance of BCSCN, showcasing its capacity to effectively reduce domain gaps and enhance the generalization of super-resolution networks. Wenhao Guo 0003, Peng Lu 0007, Xujun Peng, Zhaoran Zhao, Xiangtao Dong |
ACM Multimedia | 1 |
| 2024 | Learning Realistic Sketching: A Dual-agent Reinforcement Learning ApproachabstractThis paper presents a pioneering method for teaching computer sketching that transforms input images into sequential, parameterized strokes. However, two challenges are raised for this sketching task: weak stimuli during stroke decomposition and maintaining semantic correctness, stylistic consistency, and detail integrity in the final drawings. To tackle the challenge of weak stimuli, our method incorporates an attention agent, which enhances the algorithm's sensitivity to subtle canvas changes by focusing on smaller, magnified areas. Moreover, in enhancing the perceived quality of drawing outcomes, we integrate a sketching style feature extractor to seamlessly capture semantic information and execute style adaptation at feature level, alongside a drawing agent that decomposes strokes under the guidance of a fine-grained reward, thereby ensuring the integrity of sketch details. Based on dual intelligent agents, we have constructed an efficient sketching model. Experimental results attest to the superiority of our approach in both visual effects and perceptual metrics when compared to state-of-the-art techniques, confirming its efficacy in achieving realistic sketching. Peng Lu 0007, Xujun Peng, Wenhao Guo 0003, Zhaoran Zhao, Xiangtao Dong |
ACM Multimedia | 4 |