Peng Zhou 0010

dblp:23/5823-10 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-0674-9296ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Text-Driven Relation Manipulation of Diffusion Imagery
abstract
Text-guided image manipulation has recently attracted significant attention. Prevailing algorithms predominantly focus on modifying the appearances of existing instances, such as texture and attribute editing, while they often fail to address the interactions between different instances or achieve fundamental structural changes, such as multi-object editing. This paper introduces a novel text-guided manipulation task named "relation manipulation", aimed at fundamentally altering the structure of images. This task is capable of modifying the quantity of instances and, more importantly, enhancing the understanding and editing of interactions among diverse instances. Our approach comprises two main components: relation customization and multi-region guided diffusion. Relation customization fine-tunes specific relationships using a compact dataset of exemplary relations, facilitating nuanced understanding and implementation of instance interactions. Multi-region guided diffusion employs gradient optimization to update the generation process across multiple regions, integrating a fine-grained attention control strategy to minimize regional interference and conflict. Additionally, the demonstrated applications of our method in multi-region inversion underline its potential in practical scenarios, such as relation manipulation of real images and consecutive image manipulation. Compatible with different variants of Stable Diffusion models, our approach seamlessly integrates into the Stable Diffusion WebUI, enabling high-quality image generation and exceptional control over extensive manipulation. This makes it a robust tool for both academic research and creative industries. Code is available at https://github.com/liyiming09/RMD.
Peng Zhou 0010, Hongwei Hu, Xiaokang Qin, Jun Sun 0005, Yi Xu 0001
IEEE Trans. Image Process.2
2025 Position-LoRA: Enhanced Relation Customization through Structural Prior in Initial Latent Noise
abstract
Recent advancements in concept customization via diffusion models have significantly enhanced controllability and quality. However, precise relation customization, which controls the position of interactions among multiple instances, remains challenging due to unpredictable initial latent noise. Existing methods primarily rely on conditional prompts and attention control, overlooking the structured potential of initial noise. This paper introduces Position-LoRA, a novel framework leveraging structural prior in initial noise to improve relation customization and layout control. Position-LoRA employs a differential fine-tuning scheme and a latent noise encoder. The guided fine-tuning enhances generation tendencies from structured initial noise, embedding explicit relationship-specific spatial information. The latent noise encoder dynamically manipulates latent noises, enabling precise spatial control and flexibility in relational image generation. Furthermore, a fine-grained guidance and control strategy is employed during generation to enhance the image-text alignment and layout alignment. Experiments demonstrate that Position-LoRA improves stability, controllability, and fidelity in relational image generation with layout control, surpassing existing concept customization and layout-to-image methods in qualitative and quantitative evaluations. Code is available at https://github.com/liyiming09/Position-LoRA.
Peng Zhou 0010, Xiaokang Qin, Hongwei Hu, Jun Sun 0005, Yi Xu 0001
ACM Multimedia2
2024 FocalDreamer: Text-Driven 3D Editing via Focal-Fusion Assembly
abstract
While text-3D editing has made significant strides in leveraging score distillation sampling, emerging approaches still fall short in delivering separable, precise and consistent outcomes that are vital to content creation. In response, we introduce FocalDreamer, a framework that merges base shape with editable parts according to text prompts for fine-grained editing within desired regions. Specifically, equipped with geometry union and dual-path rendering, FocalDreamer assembles independent 3D parts into a complete object, tailored for convenient instance reuse and part-wise control. We propose geometric focal loss and style consistency regularization, which encourage focal fusion and congruent overall appearance. Furthermore, FocalDreamer generates high-fidelity geometry and PBR textures which are compatible with widely-used graphics engines. Extensive experiments have highlighted the superior editing capabilities of FocalDreamer in both quantitative and qualitative evaluations.
Yuhan Li 0003, Yishun Dou, Xuanhong Chen, Peng Zhou 0010, Bingbing Ni
AAAI7
2024 Puff-Net: Efficient Style Transfer with Pure Content and Style Feature Fusion Network
abstract
Style transfer aims to render an image with the artistic features of a style image, while maintaining the origi-nal structure. Various methods have been put forward for this task, but some challenges still exist. For instance, it is difficult for CNN-based methods to handle global information and long-range dependencies between input images, for which transformer-based methods have been proposed. Although transformers can better model the relationship between content and style images, they require high-cost hard-ware and time-consuming inference. To address these is-sues, we design a novel transformer model that includes only the encoder, thus significantly reducing the computational cost. In addition, we also find that existing style transfer methods may lead to images under-stylied or missing content. In order to achieve better stylization, we de-sign a content feature extractor and a style feature extrac-tor, based on which pure content and style images can be fed to the transformer. Finally, we propose a novel network termed Puff-Net, i.e., pure content and style feature fusion network. Through qualitative and quantitative experiments, we demonstrate the advantages of our model compared to state-of-the-art ones in the literature. The code is available at https://github.com/ZszYmy9/Puff-Net.
Sizhe Zheng, Pan Gao 0001, Peng Zhou 0010, Jie Qin 0004
CVPR3
2024 Edit3D: Elevating 3D Scene Editing with Attention-Driven Multi-Turn Interactivity
abstract
With the rise of new 3D representations like NeRF and 3D Gaussian splatting, creating realistic 3D scenes is easier than ever before. However, the incompatibility of these 3D representations with existing editing software has also introduced unprecedented challenges to 3D editing tasks. Although recent advances in text-to-image generative models have made some progress in 3D editing, these methods either lack precision or require users to manually specify the editing areas in 3D space, complicating the editing process. To overcome these issues, we propose Edit3D, an innovative 3D editing method designed to enhance editing quality. Specifically, we propose a multi-turn editing framework and introduce an attention-driven open-set segmentation (ADSS) technique within this framework. ADSS allows for more precise segmentation of parts, which enhances the editing precision and minimizes interference with pixels in areas that are not being edited. Additionally, we propose a fine-tuning phase, intended to further improve the overall editing quality without compromising the training efficiency. Experiments demonstrate that Edit3D effectively adjusts 3D scenes based on textual instructions. Through continuous and multiple turns of editing, it achieves more intricate combinations, enhancing the diversity of 3D editing effects. Code is available at https://github.com/PeterouZh/Edit3D.
Peng Zhou 0010, Dunbo Cai, Yujian Du, Runqing Zhang, Bingbing Ni, Jie Qin 0004, Ling Qian
ACM Multimedia1
2023 CIPS-3D++: End-to-End Real-Time High-Resolution 3D-Aware GANs for GAN Inversion and Stylization
abstract
Style-based GANs achieve state-of-the-art results for generating high-quality images, but lack explicit and precise control over camera poses. Recently proposed NeRF-based GANs have made great progress towards 3D-aware image generation. However, the methods either rely on convolution operators which are not rotationally invariant, or utilize complex yet suboptimal training procedures to integrate both NeRF and CNN sub-structures, yielding un-robust, low-quality images with a large computational burden. This article presents an upgraded version called CIPS-3D++, aiming at high-robust, high-resolution and high-efficiency 3D-aware GANs. On the one hand, our basic model CIPS-3D, encapsulated in a style-based architecture, features a shallow NeRF-based 3D shape encoder as well as a deep MLP-based 2D image decoder, achieving robust image generation/editing with rotation-invariance. On the other hand, our proposed CIPS-3D++, inheriting the rotational invariance of CIPS-3D, together with geometric regularization and upsampling operations, encourages high-resolution high-quality image generation/editing with great computational efficiency. Trained on raw single-view images, without any bells and whistles, CIPS-3D++ sets new records for 3D-aware image synthesis, with an impressive FID of 3.2 on FFHQ at the 1024×1024 resolution. In the meantime, CIPS-3D++ runs efficiently and enjoys a low GPU memory footprint so that it can be trained end-to-end on high-resolution images directly, in contrast to previous alternate/progressive methods. Based on the infrastructure of CIPS-3D++, we propose a 3D-aware GAN inversion algorithm named FlipInversion, which can reconstruct the 3D object from a single-view image. We also provide a 3D-aware stylization method for real images based on CIPS-3D++ and FlipInversion. In addition, we analyze the problem of mirror symmetry suffered in training, and solve it by introducing an auxiliary discriminator for the NeRF network. Overall, CIPS-3D++ provides a strong base model that can serve as a testbed for transferring GAN-based image editing methods from 2D to 3D.
Peng Zhou 0010, Lingxi Xie, Bingbing Ni, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 HRInversion: High-Resolution GAN Inversion for Cross-Domain Image Synthesis
abstract
We investigate GAN inversion problems of using pre-trained GANs to reconstruct real images. Recent methods for such problems typically employ a VGG perceptual loss to measure the difference between images. While the perceptual loss has achieved remarkable success in various computer vision tasks, it may cause unpleasant artifacts and is sensitive to changes in input scale. This paper delivers an important message that algorithm details are crucial for achieving satisfying performance. In particular, we propose two important but undervalued design principles: (i) not down-sampling the input of the perceptual loss to avoid high-frequency artifacts; and (ii) calculating the perceptual loss using convolutional features which are robust to scale. Integrating these designs derives the proposed framework, HRInversion, that achieves superior performance in reconstructing image details. We validate the effectiveness of HRInversion on a cross-domain image synthesis task and propose a post-processing approach named local style optimization (LSO) to synthesize clean and controllable stylized images. For the evaluation of the cross-domain images, we introduce a metric named ID retrieval which captures the similarity of face identities of stylized images to content images. We also test HRInversion on non-square images. Equipped with implicit neural representation, HRInversion applies to ultra-high resolution images with more than 10 million pixels. Furthermore, we show applications of style transfer and 3D-aware GAN inversion, paving the way for extending the application range of HRInversion.
Peng Zhou 0010, Lingxi Xie, Bingbing Ni, Lin Liu 0016, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Searching Towards Class-Aware Generators for Conditional Generative Adversarial Networks
abstract
Conditional generative adversarial networks (cGANs) are designed to generate images based on the provided conditions,e.g., class-level distributions, semantic label maps,etc. Existing methods have used the same generator architecture for all classes. This paper presents an idea that adopts neural architecture search (NAS) to find a class-aware architecture for each class. The search space contains regular and class-modulated convolutions, where the latter is designed to introduce class-specific information while avoiding the reduction of training data for each class generator. The search algorithm follows a weight-sharing pipeline with mixed-architecture optimization so that the search cost does not grow with the number of classes. To learn the sampling policy, a Markov decision process is embedded into the search algorithm, and a moving average is applied for better stability. Class-aware generators show advantages over class-agnostic architectures experimentally. Moreover, we discover two intriguing phenomena that are inspirational to craft cGANs by hand.
Peng Zhou 0010, Lingxi Xie, Bingbing Ni, Qi Tian 0001
IEEE Signal Process. Lett.1
2021 Omni-GAN: On the Secrets of cGANs and Beyond
abstract
The conditional generative adversarial network (cGAN) is a powerful tool of generating high-quality images, but existing approaches mostly suffer unsatisfying performance or the risk of mode collapse. This paper presents Omni-GAN, a variant of cGAN that reveals the devil in designing a proper discriminator for training the model. The key is to ensure that the discriminator receives strong supervision to perceive the concepts and moderate regularization to avoid collapse. Omni-GAN is easily implemented and freely integrated with off-the-shelf encoding methods (e.g., implicit neural representation, INR). Experiments validate the superior performance of Omni-GAN and Omni-INR-GAN in a wide range of image generation and restoration tasks. In particular, Omni-INR-GAN sets new records on the ImageNet dataset with impressive Inception scores of 262.85 and 343.22 for the image sizes of 128 and 256, respectively, surpassing the previous records by 100+ points. Moreover, leveraging the generator prior, Omni-INR-GAN can extrapolate low-resolution images to arbitrary resolution, even up to ×60+ higher resolution. Code is available1.
Peng Zhou 0010, Lingxi Xie, Bingbing Ni, Cong Geng, Qi Tian 0001
ICCV1
2018 Pose Transferrable Person Re-Identification
abstract
Person re-identification (ReID) is an important task in the field of intelligent security. A key challenge is how to capture human pose variations, while existing benchmarks (i.e., Market1501, DukeMTMC-reID, CUHK03, etc.) do NOT provide sufficient pose coverage to train a robust ReID system. To address this issue, we propose a pose-transferrable person ReID framework which utilizes pose-transferred sample augmentations (i.e., with ID supervision) to enhance ReID model training. On one hand, novel training samples with rich pose variations are generated via transferring pose instances from MARS dataset, and they are added into the target dataset to facilitate robust training. On the other hand, in addition to the conventional discriminator of GAN (i.e., to distinguish between REAL/FAKE samples), we propose a novel guider sub-network which encourages the generated sample (i.e., with novel pose) towards better satisfying the ReID loss (i.e., cross-entropy ReID loss, triplet ReID loss). In the meantime, an alternative optimization procedure is proposed to train the proposed Generator-Guider-Discriminator network. Experimental results on Market-1501, DukeMTMC-reID and CUHK03 show that our method achieves great performance improvement, and outperforms most state-of-the-art methods without elaborate designing the ReID model.
Jinxian Liu, Bingbing Ni, Yichao Yan, Peng Zhou 0010, Jianguo Hu
CVPR4
2018 Scale-Transferrable Object Detection
abstract
Scale problem lies in the heart of object detection. In this work, we develop a novel Scale-Transferrable Detection Network (STDN) for detecting multi-scale objects in images. In contrast to previous methods that simply combine object predictions from multiple feature maps from different network depths, the proposed network is equipped with embedded super-resolution layers (named as scale-transfer layer/module in this work) to explicitly explore the interscale consistency nature across multiple detection scales. Scale-transfer module naturally fits the base network with little computational cost. This module is further integrated with a dense convolutional network (DenseNet) to yield a one-stage object detector. We evaluate our proposed architecture on PASCAL VOC 2007 and MS COCO benchmark tasks and STDN obtains significant improvements over the comparable state-of-the-art detection models.
Peng Zhou 0010, Bingbing Ni, Cong Geng, Jianguo Hu, Yi Xu 0001
CVPR1
2018 Uniface: A Unified Network for Face Detection and Recognition
abstract
Typically, cropped and aligned face images are required as the input of a face recognition model. In contrast, popular object detectors based on deep convolutional network usually locate and classify objects simultaneously, which eliminates redundant computation. This work presents a single-network model called Uniface network for simultaneous face detection, landmark localization and recognition. We develop a feature sharing infrastructure for seamlessly integrate both the detection/localization module and the recognition module. To facilitate large-scale end-to-end training, we propose a method by encouraging top-level features of our model to mimic those of a well-trained single-task face recognition model. Comprehensive experiments on face detection, landmark localization and verification tasks demonstrate that the proposed network achieves competing performance in both face recognition benchmark (99.0% on LFW for a single model) and face detection benchmark (86.4% against 2000 false positives on FDDB for a single model).
Zhouyingcheng Liao, Peng Zhou 0010, Qinlong Wu, Bingbing Ni
ICPR2
2018 Live Face Verification with Multiple Instantialized Local Homographic Parameterization
abstract
State-of-the-art live face verification methods would easily be attacked by recorded facial expression sequence. This work directly addresses this issue via proposing a patch-wise motion parameterization based verification network infrastructure. This method directly explores the underlying subtle motion difference between the facial movements re-captured from a planer screen (e.g., a pad) and those from a real face; therefore interactive facial expression is no longer required. Furthermore, inspired by the fact that ?a fake facial movement sequence MUST contains many patch-wise fake sequences?, we embed our network into a multiple instance learning framework, which further enhance the recall rate of the proposed technique. Extensive experimental results on several face benchmarks well demonstrate the superior performance of our method.
Chen Lin 0001, Zhouyingcheng Liao, Peng Zhou 0010, Jianguo Hu, Bingbing Ni
IJCAI3