Zongze Wu 0002

dblp:125/6476-2 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0001-9190-1717ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2025 TurboFill: Adapting Few-step Text-to-image Model for Fast Image Inpainting
abstract
This paper introduces TurboFill, a fast image inpainting model that enhances a few-step text-to-image diffusion model with an inpainting adapter for high-quality and efficient inpainting. While standard diffusion models generate high-quality results, they incur high computational costs. We overcome this by training an inpainting adapter on a few-step distilled text-to-image model, DMD2, using a novel 3-step adversarial training scheme to ensure realistic, structurally consistent, and visually harmonious inpainted regions. To evaluate TurboFill, we propose two benchmarks: DilationBench, which tests performance across mask sizes, and HumanBench, based on human feedback for complex prompts. Experiments show that TurboFill outperforms both multi-step BrushNet and few-step inpainting methods, setting a new benchmark for high-performance inpainting tasks. The project page is available here.
Liangbin Xie, Daniil Pakhomov, Zongze Wu 0002, Yuqian Zhou, Haitian Zheng, Zhe Lin 0001, Jiantao Zhou 0001, Chao Dong 0005
CVPR4
2025 SliderSpace: Decomposing the Visual Capabilities of Diffusion Models
abstract
We present SliderSpace, a framework for automatically decomposing the visual capabilities of diffusion models into controllable and human-understandable directions. Unlike existing control methods that require a user to specify attributes for each edit direction individually, SliderSpace discovers multiple interpretable and diverse directions simultaneously from a single text prompt. Each direction is trained as a low-rank adaptor, enabling compositional control and the discovery of surprising possibilities in the model's latent space. Through extensive experiments on state-of-the-art diffusion models, we demonstrate SliderSpace's effectiveness across three applications: concept decomposition, artistic style exploration, and diversity enhancement. Our quantitative evaluation shows that SliderSpace-discovered directions decompose the visual structure of model's knowledge effectively, offering insights into the latent capabilities encoded within diffusion models. User studies further validate that our method produces more diverse and useful variations compared to baselines. Our code, data and trained weights are available at https://sliderspace.baulab.info
Rohit Gandikota, Zongze Wu 0002, Richard Zhang 0001, David Bau, Eli Shechtman, Nicholas I. Kolkin
ICCV2
2025 Family-based continual learning for multi-domain pattern analysis in federated frameworks with GCN and ViT
Saeed Iqbal, Xiaopin Zhong, Muhammad Attique Khan, Zongze Wu 0002, Dina Abdulaziz Alhammadi, Weixiang Liu
Neural Networks4
2024 Lazy Diffusion Transformer for Interactive Image Editing
Yotam Nitzan, Zongze Wu 0002, Richard Zhang 0001, Eli Shechtman, Daniel Cohen-Or, Taesung Park, Michaël Gharbi
ECCV (24)2
2024 TurboEdit: Instant Text-Based Image Editing
Zongze Wu 0002, Nicholas I. Kolkin, Jonathan Brandt, Richard Zhang 0001, Eli Shechtman
ECCV (80)1
2022 StyleAlign: Analysis and Applications of Aligned StyleGAN Models
Zongze Wu 0002, Yotam Nitzan, Eli Shechtman, Dani Lischinski
ICLR1
2021 StyleSpace Analysis: Disentangled Controls for StyleGAN Image Generation
abstract
We explore and analyze the latent style space of Style-GAN2, a state-of-the-art architecture for image generation, using models pretrained on several different datasets. We first show that StyleSpace, the space of channel-wise style parameters, is significantly more disentangled than the other intermediate latent spaces explored by previous works. Next, we describe a method for discovering a large collection of style channels, each of which is shown to control a distinct visual attribute in a highly localized and dis-entangled manner. Third, we propose a simple method for identifying style channels that control a specific attribute, using a pretrained classifier or a small number of example images. Manipulation of visual attributes via these StyleSpace controls is shown to be better disentangled than via those proposed in previous works. To show this, we make use of a newly proposed Attribute Dependency metric. Finally, we demonstrate the applicability of StyleSpace controls to the manipulation of real images. Our findings pave the way to semantically meaningful and well-disentangled image manipulations via simple and intuitive interfaces.
Zongze Wu 0002, Dani Lischinski, Eli Shechtman
CVPR1
2021 StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery
abstract
Inspired by the ability of StyleGAN to generate highly realistic images in a variety of domains, much recent work has focused on understanding how to use the latent spaces of StyleGAN to manipulate generated and real images. However, discovering semantically meaningful latent manipulations typically involves painstaking human examination of the many degrees of freedom, or an annotated collection of images for each desired manipulation. In this work, we explore leveraging the power of recently introduced Contrastive Language-Image Pre-training (CLIP) models in order to develop a text-based interface for StyleGAN image manipulation that does not require such manual effort. We first introduce an optimization scheme that utilizes a CLIP-based loss to modify an input latent vector in response to a user-provided text prompt. Next, we describe a latent mapper that infers a text-guided latent manipulation step for a given input image, allowing faster and more stable text-based manipulation. Finally, we present a method for mapping text prompts to input-agnostic directions in StyleGAN’s style space, enabling interactive text-driven image manipulation. Extensive results and comparisons demonstrate the effectiveness of our approaches.
Or Patashnik, Zongze Wu 0002, Eli Shechtman, Daniel Cohen-Or, Dani Lischinski
ICCV2
2021 Fine-grained Foreground Retrieval via Teacher-Student Learning
abstract
Foreground image retrieval is a challenging computer vision task. Given a background scene image with a bounding box indicating a target location, the goal is to retrieve a set of images of foreground objects from a given category, which are semantically compatible with the background. We formulate foreground retrieval as a self-supervised domain adaptation task, where the source domain consists of foreground images and the target domain of background images. Specifically, given pretrained object feature extraction networks that serve as teachers, we train a student network to infer compatible foreground features from background images. Thus, foregrounds and backgrounds are effectively mapped into a common feature space, enabling retrieval of the foregrounds that are closest to the target background in that space. A notable feature of our approach is that our training strategy does not require instance segmentation, unlike current state-of-the-art methods. Thus, our method may be applied to diverse foreground categories and background scene types and enables us to retrieve the foreground in a fine-grained manner, which is closer to the requirements of real world applications.
Zongze Wu 0002, Dani Lischinski, Eli Shechtman
WACV1