VLDB 2026 Research / reviewers in the wild / expert
Pu Cao
dblp:169/2437
· DBLP profile ↗
12ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring Position Encoding Mechanism in Diffusion U-Net for Training-free High-resolution Image GenerationabstractDenoising higher-resolution latents using a pre-trained U-Net often results in repetitive and disordered image patterns. In this work, we are motivated to reveal the intrinsic cause of such pattern disruption in high-resolution image generation. Through theoretical analysis and empirical studies, we reveal that the pre-trained U-Net fails to provide sufficient positional information for tokens at high-resolution. Specifically, 1) zero-padding serves as a critical mechanism for position encoding but lacks robustness across varying resolutions; and 2) tokens located farther from the feature map boundaries have increasing difficulty acquiring positional awareness, leading to pattern disruptions. Inspired by these findings, we propose a novel training-free approach for high-resolution generation, introducing a Progressive Boundary Complement (PBC) method. It creates dynamic virtual image boundaries inside the feature map to supplement position information at high resolution, enabling high-quality and rich-content high-resolution image synthesis. Extensive experiments show that our method significantly improves high-resolution image synthesis in terms of visual quality and content richness, achieving state-of-the-art performance. Pu Cao, Yiyang Ma, Lu Yang 0006, Yonghao Dang, Jianqin Yin |
AAAI | 2 |
| 2026 | Controllable Generation With Text-to-Image Diffusion Models: A SurveyabstractIn the rapidly advancing realm of visual generation, diffusion models have revolutionized the landscape, marking a significant shift in capabilities with their impressive text-guided generative functions. However, relying solely on text for conditioning these models does not fully cater to the varied and complex requirements of different applications and scenarios. Acknowledging this shortfall, a variety of studies aim to control pre-trained text-to-image (T2I) models to support novel conditions. In this survey, we undertake a thorough review of the literature on controllable generation with T2I diffusion models, covering both the theoretical foundations and practical advancements in this domain. Our review begins with a brief introduction to the basics of denoising diffusion probabilistic models (DDPMs) and widely used T2I diffusion models. Additionally, we provide a detailed overview of research in this area, categorizing it from the condition perspective into three directions: generation with specific conditions, generation with multiple conditions, and universal controllable generation. For each category, we analyze the underlying control mechanisms and review representative methods based on their core techniques. Pu Cao, Qing Song 0006, Lu Yang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Large-Scale Omnidirectional Person Positioning
Lu Yang 0006, Liulei Li, Jianan Wei, Pu Cao, Wenguan Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | ParsingFormer for pixel-wise hierarchical human representation learning
Pu Cao, Zhixiang Lv, Junyi Ji, Shan Li 0001, Qing Song 0006, Lu Yang 0006 |
Pattern Recognit. | 1 |
| 2026 | Quality transformer for human parsing
Lu Yang 0006, Pu Cao, Shan Li 0001, Qing Song 0006 |
Pattern Recognit. | 3 |
| 2026 | Joint Trajectory and Power Optimization for Dynamic Spectrum Control-Assisted Secure UAV CommunicationsabstractUnmanned aerial vehicles (UAVs) play a crucial role in modern communication systems owing to their high mobility and broad coverage. However, due to the inherent open nature of the wireless channels, UAV-to-ground links are facing significant security threats from eavesdroppers and malicious jammers. To address these challenges, we propose a dynamic spectrum control (DSC) scheme integrating joint UAV trajectory and transmit power optimization to enhance UAV communication security in this paper. This scheme divides transmission channels from time and frequency dimensions and intelligently generates secure decision sequences using cryptographic principles based on real-time channel states, enabling transmissions for legitimate users without intra-cell interference. Based on a rapid-flooding time synchronization protocol, we analyze inter-cell collision probability (CP) and formulate an optimization problem for the secrecy rate. To further enhance security, we conduct a joint UAV trajectory and transmit power optimization. Through the successive convex approximation (SCA) method, we transform the non-convex optimization problem into a tractable convex form, obtaining a suboptimal solution. Simulations demonstrate that our proposed scheme significantly enhances security compared to conventional UAV communication methods. Pu Cao, Zan Li 0001, Haixia Peng, Chuan Zhang 0003 |
IEEE Trans. Commun. | 1 |
| 2026 | OMEGAS: Object Mesh Extraction From Large Scenes Guided by Gaussian SegmentationabstractRecent advancements in 3D reconstruction technologies have paved the way for high-quality and real-time rendering of complex 3D scenes. Despite these achievements, a notable challenge persists: it is difficult to precisely reconstruct specific objects from large scenes. Current scene reconstruction techniques frequently result in the loss of object detail textures and are unable to reconstruct object portions that are occluded or unseen in views. To address this challenge, we delve into the meticulous 3D reconstruction of specific objects within large scenes and propose a framework termed OMEGAS: Object Mesh Extraction from Large Scenes Guided by GAussian Segmentation. Specifically, we propose a novel 3D target segmentation technique based on 2D Gaussian Splatting, which segments 3D consistent target masks in multi-view scene images and generates a preliminary target model. Moreover, to reconstruct the unseen portions of the target, we propose a novel target replenishment technique driven by large-scale generative diffusion priors. We demonstrate that our method can accurately reconstruct specific targets from large scenes, both quantitatively and qualitatively. Our experiments show that OMEGAS significantly outperforms existing reconstruction methods across various scenarios. Pu Cao, Jianqin Yin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Image is All You Need to Empower Large-scale Diffusion Models for In-Domain GenerationabstractIn-domain generation aims to perform a variety of tasks within a specific domain, such as unconditional generation, text-to-image, image editing, 3D generation, and more. Early research typically required training specialized generators for each unique task and domain, often relying on fully-labeled data. Motivated by the powerful generative capabilities and broad applications of diffusion models, we are driven to explore leveraging label-free data to empower these models for in-domain generation. Fine-tuning a pre-trained generative model on domain data is an intuitive but challenging way and often requires complex manual hyper-parameter adjustments since the limited diversity of the training data can easily disrupt the model’s original generative capabilities. To address this challenge, we propose a guidance-decoupled prior preservation mechanism to achieve high generative quality and controllability by image-only data, inspired by preserving the pre-trained model from a denoising guidance perspective. We decouple domain-related guidance from the conditional guidance used in classifier-free guidance mechanisms to preserve open-world control guidance and unconditional guidance from the pre-trained model. We further propose an efficient domain knowledge learning technique to train an additional text-free UNet copy to predict domain guidance. Besides, we theoretically illustrate a multi-guidance in-domain generation pipeline for a variety of generative tasks, leveraging multiple guidances from distinct diffusion models and conditions. Extensive experiments demonstrate the superiority of our method in domain-specific synthesis and its compatibility with various diffusion-based control methods and applications. Pu Cao, Lu Yang 0006, Tianrui Huang, Qing Song 0006 |
CVPR | 1 |
| 2025 | E4C: Enhance Editability for Text-Based Image Editing by Harnessing Efficient CLIP GuidanceabstractDiffusion-based image editing involves both preserving the source image content and generating new content or applying modifications. Although current editing approaches have made improvements under text guidance, they have two key drawbacks: overemphasis on retaining original image info, neglecting editability and text alignment, and inability to handle both structure-consistent and non-rigid editing tasks. In this paper, we propose a zero-shot image editing method, named Enhance Editability for text-based image Editing via Efficient CLIP guidance (E4C), which presents an innovative adaptive feature sharing mechanism to enable multi-task editing. Additionally, a novel random gateway mechanism is designed to efficiently introduce CLIP guidance into the multi-step sampling of diffusion, achieving high congruence between editing results and target text. Comprehensive quantitative and qualitative experiments demonstrate that our method effectively resolves the text alignment issues prevalent in existing methods while maintaining the fidelity to the source image, and performs well across a wide range of editing tasks. Tianrui Huang, Pu Cao, Lu Yang 0006, Chun Liu 0004, Mengjie Hu 0002, Zhiwei Liu 0004, Qing Song 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | What Decreases Editing Capability? Domain-Specific Hybrid Refinement for Improved GAN InversionabstractRecently, inversion methods have been exploring the incorporation of additional high-rate information from pretrained generators (such as weights or intermediate features) to improve the refinement of inversion and editing results from embedded latent codes. While such techniques have shown reasonable improvements in reconstruction, they often lead to a decrease in editing capability, especially when dealing with complex images that contain occlusions, detailed backgrounds, and artifacts. To address this problem, we propose a novel refinement mechanism called Domain-Specific Hybrid Refinement (DHR), which draws on the advantages and disadvantages of two mainstream refinement techniques. We find that the weight modulation can gain favorable editing results but is vulnerable to these complex image areas and feature modulation is efficient at reconstructing. Hence, we divide the image into two domains and process them with these two methods separately. We first propose a Domain-Specific Segmentation module to automatically segment images into in-domain and out-of-domain parts according to their invertibility and editability without additional data annotation, where our hybrid refinement process aims to maintain the editing capability for in-domain areas and improve fidelity for both of them. We achieve this through Hybrid Modulation Refinement, which respectively refines these two domains by weight modulation and feature modulation. Our proposed method is compatible with all latent code embedding methods. Extension experiments demonstrate that our approach achieves state-of-the-art in real image inversion and editing. Code is available at https: //github.com/caopulan/Domain-Specific_ Hybrid_Refinement_Inversion. Pu Cao, Lu Yang 0006, Dongxv Liu, Xiaoya Yang, Tianrui Huang, Qing Song 0006 |
WACV | 1 |
| 2024 | Frequency-Based Matcher for Long-Tailed Semantic SegmentationabstractThe successful application of semantic segmentation technology in the real world has been among the most exciting achievements in the computer vision community over the past decade. Although the long-tailed phenomenon has been investigated in many fields,e.g., classification and object detection, it has not received enough attention in semantic segmentation and has become a nonnegligible obstacle to applying semantic segmentation technology in autonomous driving and virtual reality. Therefore, in this work, we focus on a relatively underexplored task setting,long-tailed semantic segmentation(LTSS). We first establish three representative datasets from different aspects, i.e., scene, object, and human. We further propose a dual-metric evaluation system and construct the LTSS benchmark to demonstrate the performance of semantic segmentation methods and long-tailed solutions. We also propose a transformer-based algorithm to improve LTSS,frequency-based matcher, which solves the oversuppression problem by one-to-many matching and automatically determines the number of matching queries for each class. Given the comprehensiveness of this work and the importance of the issues revealed, this work aims to promote the empirical study of semantic segmentation tasks. Our datasets, codes, and models will be publicly available. Shan Li 0001, Lu Yang 0006, Pu Cao, Liulei Li, Huadong Ma |
IEEE Trans. Multim. | 3 |
| 2020 | Deep Relevance Feature Clustering for Discovering Visual Representation of Tourism Destination
Qiannan Wang, Zhaoyan Zhu, Xuefeng Liang, Huiwen Shi, Pu Cao |
PRCV (3) | 5 |