VLDB 2026 Research / reviewers in the wild / expert
Zhiwen Zuo
dblp:234/7977
· DBLP profile ↗
22ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0002-0062-6308ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 3 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance CustomizationabstractGarment-centric fashion image generation aims to synthesize realistic and controllable human models dressing a given garment, which has attracted growing interest due to its practical applications in e-commerce. The key challenges of the task lie in two aspects: (1) faithfully preserving the garment details, and (2) gaining fine-grained controllability over the model's appearance. Existing methods typically require performing garment deformation in the generation process, which often leads to garment texture distortions. Also, they fail to control the fine-grained attributes of the generated models, due to the lack of specifically designed mechanisms. To address these issues, we propose FashionMAC, a novel diffusion-based deformation-free framework that achieves high-quality and controllable fashion showcase image generation. The core idea of our framework is to eliminate the need for performing garment deformation and directly outpaint the garment segmented from a dressed person, which enables faithful preservation of the intricate garment details. Moreover, we propose a novel region-adaptive decoupled attention (RADA) mechanism along with a chained mask injection strategy to achieve fine-grained appearance controllability over the synthesized human models. Specifically, RADA adaptively predicts the generated regions for each fine-grained text attribute and enforces the text attribute to focus on the predicted regions by a chained mask injection strategy, significantly enhancing the visual fidelity and the controllability. Extensive experiments validate the superior performance of our framework compared to existing state-of-the-art methods. Jinxiao Li, Jingnan Wang, Zhiwen Zuo, Jianfeng Dong, Wei Li 0111, Chi Wang 0004, Weiwei Xu 0003, Xun Wang 0007 |
AAAI | 4 |
| 2026 | ConceptCraft: One-Shot Personalized Text-to-Image Generation via Object-Background DisentanglementabstractPersonalized text-to-image generation aims to learn new concepts from user-provided images and subsequently generate diverse scenes or styles of the concepts from input prompts. Most existing methods usually require a set of images (typically 3-5) for each concept, which can be cumbersome. Although several methods allow personalized generation with a single reference image, they often require heavy model training and suffer from many issues such as domain-specific applicability, insufficient fidelity, and limited editability. To address these problems, we propose a novel one-shot personalized text-to-image generation method called ConceptCraft, which explicitly separates the reference image into object and background regions and treats them as two distinct concepts to learn, significantly improving the personalization performance. Specifically, we incorporate two unique identifiers into the text prompts: one followed by the object’s class name and the other by the word “background”. To bind these two identifiers to the reference image’s object and background respectively, we introduce a mask-aware object preservation loss and a mask-aware background preservation loss to optimize their corresponding token embeddings under well-designed text conditions, enabling both object and background personalization. In addition, we also develop an identifier regularization scheme to enhance our editability, allowing the synthesis of personalized images across a broader range of scenes and styles without changing the identity. Extensive qualitative and quantitative experiments are conducted to verify the effectiveness and superiority of our method. Haibo Chen 0006, Zhiwen Zuo, Lei Zhao 0011, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Rethink arbitrary style transfer with transformer and contrastive learning
Zhanjie Zhang, Jiakai Sun, Lei Zhao 0011, Quanwei Zhang, Zehua Lan, Haolin Yin, Huaizhong Lin, Zhiwen Zuo |
Comput. Vis. Image Underst. | 10 |
| 2024 | Statistics Enhancement Generative Adversarial Networks for Diverse Conditional Image SynthesisabstractConditional generative adversarial networks (cGANs) aim to synthesize diverse images given the input conditions and the latent codes, but they are prone to map an input to a single output regardless of the variations in latent code, which is also well known as the mode collapse problem of cGANs. To alleviate the problem, in this paper, we investigate explicitly enhancing the statistical dependency between the latent code and the synthesized image in cGANs by utilizing mutual information neural estimators to estimate and maximize the conditional mutual information (CMI) between them given the input condition. The method provides a new perspective from information theory to improve diversity for cGANs and can facilitate many existing conditional image synthesis frameworks with a simple neural estimator extension. Moreover, our studies show that several key designs, including the neural estimator choice, the neural estimator’s network design, and the sampling strategy, are crucial to the success of the method. Extensive experiments on four popular conditional image synthesis tasks, including class-conditioned image generation, paired and unpaired image-to-image translation, and text-to-image generation, demonstrate the effectiveness and superiority of the proposed method. Zhiwen Zuo, Ailin Li, Zhizhong Wang, Lei Zhao 0011, Jianfeng Dong, Xun Wang 0007, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | MicroAST: Towards Super-fast Ultra-Resolution Arbitrary Style TransferabstractArbitrary style transfer (AST) transfers arbitrary artistic styles onto content images. Despite the recent rapid progress, existing AST methods are either incapable or too slow to run at ultra-resolutions (e.g., 4K) with limited resources, which heavily hinders their further applications. In this paper, we tackle this dilemma by learning a straightforward and lightweight model, dubbed MicroAST. The key insight is to completely abandon the use of cumbersome pre-trained Deep Convolutional Neural Networks (e.g., VGG) at inference. Instead, we design two micro encoders (content and style encoders) and one micro decoder for style transfer. The content encoder aims at extracting the main structure of the content image. The style encoder, coupled with a modulator, encodes the style image into learnable dual-modulation signals that modulate both intermediate features and convolutional filters of the decoder, thus injecting more sophisticated and flexible style signals to guide the stylizations. In addition, to boost the ability of the style encoder to extract more distinct and representative style signals, we also introduce a new style signal contrastive loss in our model. Compared to the state of the art, our MicroAST not only produces visually superior results but also is 5-73 times smaller and 6-18 times faster, for the first time enabling super-fast (about 0.5 seconds) AST at 4K ultra-resolutions. Zhizhong Wang, Lei Zhao 0011, Zhiwen Zuo, Ailin Li, Haibo Chen 0006, Wei Xing 0001, Dongming Lu |
AAAI | 3 |
| 2023 | Generative Image Inpainting with Segmentation Confusion Adversarial Training and Contrastive LearningabstractThis paper presents a new adversarial training framework for image inpainting with segmentation confusion adversarial training (SCAT) and contrastive learning. SCAT plays an adversarial game between an inpainting generator and a segmentation network, which provides pixel-level local training signals and can adapt to images with free-form holes. By combining SCAT with standard global adversarial training, the new adversarial training framework exhibits the following three advantages simultaneously: (1) the global consistency of the repaired image, (2) the local fine texture details of the repaired image, and (3) the flexibility of handling images with free-form holes. Moreover, we propose the textural and semantic contrastive learning losses to stabilize and improve our inpainting model's training by exploiting the feature representation space of the discriminator, in which the inpainting images are pulled closer to the ground truth images but pushed farther from the corrupted images. The proposed contrastive losses better guide the repaired images to move from the corrupted image data points to the real image data points in the feature representation space, resulting in more realistic completed images. We conduct extensive experiments on two benchmark datasets, demonstrating our model's effectiveness and superiority both qualitatively and quantitatively. Zhiwen Zuo, Lei Zhao 0011, Ailin Li, Zhizhong Wang, Zhanjie Zhang, Jiafu Chen, Wei Xing 0001, Dongming Lu |
AAAI | 1 |
| 2023 | CRFAST: Clip-Based Reference-Guided Facial Image Semantic TransferabstractThis paper presents a new task for CLIP-based reference-guided facial image semantic transfer: the source facial image is translated to the output image with the high-level semantic attributes from the reference image while maintaining identity preservation. To this end, we employ the powerful generative capability of StyleGAN generator and the rich semantic knowledge of CLIP encoder to accomplish such a task. Additionally, a novel contrastive loss is designed to comprehensively explore the rich semantic information of CLIP for facial semantic concepts. This loss guides the semantic transfer toward desired directions from different perspectives in the pre-defined CLIP space. Besides, a simple yet effective semantic-preserved modulation module is proposed to explicitly map CLIP embeddings of reference image to the latent space. Experiments demonstrate that our approach achieves realistic facial image semantic transfer driven by reference images with various facial semantics. Ailin Li, Lei Zhao 0011, Zhiwen Zuo, Zhizhong Wang, Wei Xing 0001, Dongming Lu |
ICASSP | 3 |
| 2023 | Rethinking Fast Fourier Convolution in Image InpaintingabstractRecently proposed LaMa [25] introduce Fast Fourier Convolution (FFC) [4] into image inpainting. FFC empowers the fully convolutional network to have a global receptive field in its early layers, and have the ability to produce robust repeating texture. However, LaMa has difficulty in generating clear and sharp complex content. In this paper, we analyze the fundamental flaws of using FFC in image inpainting, which are 1) spectrum shifting, 2) unexpected spatial activation, and 3) limited frequency receptive field. Such flaws make FFC-based inpainting framework difficult in generating complicated texture and performing faithful reconstruction. Based on the above analysis, we propose a novel Unbiased Fast Fourier Convolution (UFFC) module. UFFC is constructed by modifying the vanilla FFC module with 1) range transform and inverse transform, 2) absolute position embedding, 3) dynamic skip connection, and 4) adaptive clip, to overcome the above flaws. UFFC captures frequency information efficiently and realize reconstruction without introducing additional artifacts, achieving better inpainting results and more efficient training. In addition, we propose two novel perceptual losses for better generation quality and more robust training. Extensive experiments on several benchmark datasets demonstrate the effectiveness of our method, outperforming the state-of-the-art methods in both texture-capturing ability and expressiveness. Tianyi Chu, Jiafu Chen, Jiakai Sun, Shuobin Lian, Zhizhong Wang, Zhiwen Zuo, Lei Zhao 0011, Wei Xing 0001, Dongming Lu |
ICCV | 6 |
| 2023 | TeSTNeRF: Text-Driven 3D Style Transfer via Cross-Modal LearningabstractText-driven 3D style transfer aims at stylizing a scene according to the text and generating arbitrary novel views with consistency. Simply combining image/video style transfer methods and novel view synthesis methods results in flickering when changing viewpoints, while existing 3D style transfer methods learn styles from images instead of texts. To address this problem, we for the first time design an efficient text-driven model for 3D style transfer, named TeSTNeRF, to stylize the scene using texts via cross-modal learning: we leverage an advanced text encoder to embed the texts in order to control 3D style transfer and align the input text and output stylized images in latent space. Furthermore, to obtain better visual results, we introduce style supervision, learning feature statistics from style images and utilizing 2D stylization results to rectify abrupt color spill. Extensive experiments demonstrate that TeSTNeRF significantly outperforms existing methods and provides a new way to guide 3D style transfer. Jiafu Chen, Boyan Ji, Zhanjie Zhang, Tianyi Chu, Zhiwen Zuo, Lei Zhao 0011, Wei Xing 0001, Dongming Lu |
IJCAI | 5 |
| 2023 | Towards Interactive Facial Image Inpainting by Text or Exemplar Image
Ailin Li, Lei Zhao 0011, Zhiwen Zuo, Zhizhong Wang, Wei Xing 0001, Dongming Lu |
MMM (1) | 3 |
| 2023 | MIGT: Multi-modal image inpainting guided with text
Ailin Li, Lei Zhao 0011, Zhiwen Zuo, Zhizhong Wang, Wei Xing 0001, Dongming Lu |
Neurocomputing | 3 |
| 2022 | Texture Reformer: Towards Fast and Universal Interactive Texture TransferabstractIn this paper, we present the texture reformer, a fast and universal neural-based framework for interactive texture transfer with user-specified guidance. The challenges lie in three aspects: 1) the diversity of tasks, 2) the simplicity of guidance maps, and 3) the execution efficiency. To address these challenges, our key idea is to use a novel feed-forward multi-view and multi-stage synthesis procedure consisting of I) a global view structure alignment stage, II) a local view texture refinement stage, and III) a holistic effect enhancement stage to synthesize high-quality results with coherent structures and fine texture details in a coarse-to-fine fashion. In addition, we also introduce a novel learning-free view-specific texture reformation (VSTR) operation with a new semantic map guidance strategy to achieve more accurate semantic-guided and structure-preserved texture transfer. The experimental results on a variety of application scenarios demonstrate the effectiveness and superiority of our framework. And compared with the state-of-the-art interactive texture transfer algorithms, it not only achieves higher quality results but, more remarkably, also is 2-5 orders of magnitude faster. Zhizhong Wang, Lei Zhao 0011, Haibo Chen 0006, Ailin Li, Zhiwen Zuo, Wei Xing 0001, Dongming Lu |
AAAI | 5 |
| 2022 | DivSwapper: Towards Diversified Patch-based Arbitrary Style TransferabstractGram-based and patch-based approaches are two important research lines of style transfer. Recent diversified Gram-based methods have been able to produce multiple and diverse stylized outputs for the same content and style images. However, as another widespread research interest, the diversity of patch-based methods remains challenging due to the stereotyped style swapping process based on nearest patch matching. To resolve this dilemma, in this paper, we dive into the crux of existing patch-based methods and propose a universal and efficient module, termed DivSwapper, for diversified patch-based arbitrary style transfer. The key insight is to use an essential intuition that neural patches with higher activation values could contribute more to diversity. Our DivSwapper is plug-and-play and can be easily integrated into existing patch-based and Gram-based methods to generate diverse results for arbitrary styles. We conduct theoretical analyses and extensive experiments to demonstrate the effectiveness of our method, and compared with state-of-the-art algorithms, it shows superiority in diversity, quality, and efficiency. Zhizhong Wang, Lei Zhao 0011, Haibo Chen 0006, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
IJCAI | 4 |
| 2022 | Style Fader Generative Adversarial Networks for Style Degree Controllable Artistic Style TransferabstractArtistic style transfer is the task of synthesizing content images with learned artistic styles. Recent studies have shown the potential of Generative Adversarial Networks (GANs) for producing artistically rich stylizations. Despite the promising results, they usually fail to control the generated images' style degree, which is inflexible and limits their applicability for practical use. To address the issue, in this paper, we propose a novel method that for the first time allows adjusting the style degree for existing GAN-based artistic style transfer frameworks in real time after training. Our method introduces two novel modules into existing GAN-based artistic style transfer frameworks: a Style Scaling Injection (SSI) module and a Style Degree Interpretation (SDI) module. The SSI module accepts the value of Style Degree Factor (SDF) as the input and outputs parameters that scale the feature activations in existing models, offering control signals to alter the style degrees of the stylizations. And the SDI module interprets the output probabilities of a multi-scale content-style binary classifier as the style degrees, providing a mechanism to parameterize the style degree of the stylizations. Moreover, we show that after training our method can enable existing GAN-based frameworks to produce over-stylizations. The proposed method can facilitate many existing GAN-based artistic style transfer frameworks with marginal extra training overheads and modifications. Extensive qualitative evaluations on two typical GAN-based style transfer models demonstrate the effectiveness of the proposed method for gaining style degree control for them. Zhiwen Zuo, Lei Zhao 0011, Shuobin Lian, Haibo Chen 0006, Zhizhong Wang, Ailin Li, Wei Xing 0001, Dongming Lu |
IJCAI | 1 |
| 2022 | AesUST: Towards Aesthetic-Enhanced Universal Style TransferabstractRecent studies have shown remarkable success in universal style transfer which transfers arbitrary visual styles to content images. However, existing approaches suffer from the aesthetic-unrealistic problem that introduces disharmonious patterns and evident artifacts, making the results easy to spot from real paintings. To address this limitation, we propose AesUST, a novel Aesthetic-enhanced Universal Style Transfer approach that can generate aesthetically more realistic and pleasing results for arbitrary styles. Specifically, our approach introduces an aesthetic discriminator to learn the universal human-delightful aesthetic features from a large corpus of artist-created paintings. Then, the aesthetic features are incorporated to enhance the style transfer process via a novel Aesthetic-aware Style-Attention (AesSA) module. Such an AesSA module enables our AesUST to efficiently and flexibly integrate the style patterns according to the global aesthetic channel distribution of the style image and the local semantic spatial distribution of the content image. Moreover, we also develop a new two-stage transfer training strategy with two aesthetic regularizations to train our model more effectively, further improving stylization performance. Extensive experiments and user studies demonstrate that our approach synthesizes aesthetically more harmonious and realistic results than state of the art, greatly narrowing the disparity with real artist-created paintings. Our code is available at https://github.com/EndyWon/AesUST. Zhizhong Wang, Zhanjie Zhang, Lei Zhao 0011, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
ACM Multimedia | 4 |
| 2022 | Dual distribution matching GAN
Zhiwen Zuo, Lei Zhao 0011, Ailin Li, Zhizhong Wang, Haibo Chen 0006, Wi Xing, Dongming Lu |
Neurocomputing | 1 |
| 2021 | DualAST: Dual Style-Learning Networks for Artistic Style TransferabstractArtistic style transfer is an image editing task that aims at repainting everyday photographs with learned artistic styles. Existing methods learn styles from either a single style example or a collection of artworks. Accordingly, the stylization results are either inferior in visual quality or limited in style controllability. To tackle this problem, we propose a novel Dual Style-Learning Artistic Style Transfer (DualAST) framework to learn simultaneously both the holistic artist-style (from a collection of artworks) and the specific artwork-style (from a single style image): the artist-style sets the tone (i.e., the overall feeling) for the stylized image, while the artwork-style determines the details of the stylized image, such as color and texture. Moreover, we introduce a Style-Control Block (SCB) to adjust the styles of generated images with a set of learnable style-control factors. We conduct extensive experiments to evaluate the performance of the proposed framework, the results of which confirm the superiority of our method. Haibo Chen 0006, Lei Zhao 0011, Zhizhong Wang, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
CVPR | 5 |
| 2021 | Diverse Image Style Transfer via Invertible Cross-Space MappingabstractImage style transfer aims to transfer the styles of artworks onto arbitrary photographs to create novel artistic images. Although style transfer is inherently an underdetermined problem, existing approaches usually assume a deterministic solution, thus failing to capture the full distribution of possible outputs. To address this limitation, we propose a Diverse Image Style Transfer (DIST) framework which achieves significant diversity by enforcing an invertible cross-space mapping. Specifically, the framework consists of three branches: disentanglement branch, inverse branch, and stylization branch. Among them, the disentanglement branch factorizes artworks into content space and style space; the inverse branch encourages the invertible mapping between the latent space of input noise vectors and the style space of generated artistic images; the stylization branch renders the input content image with the style of an artist. Armed with these three branches, our approach is able to synthesize significantly diverse stylized images without loss of quality. We conduct extensive experiments and comparisons to evaluate our approach qualitatively and quantitatively. The experimental results demonstrate the effectiveness of our method. Haibo Chen 0006, Lei Zhao 0011, Zhizhong Wang, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
ICCV | 5 |
| 2021 | Artistic Style Transfer with Internal-external Learning and Contrastive LearningabstractAlthough existing artistic style transfer methods have achieved significant improvement with deep neural networks, they still suffer from artifacts such as disharmonious colors and repetitive patterns. Motivated by this, we propose an internal-external style transfer method with two contrastive losses. Specifically, we utilize internal statistics of a single style image to determine the colors and texture patterns of the stylized image, and in the meantime, we leverage the external information of the large-scale style dataset to learn the human-aware style information, which makes the color distributions and texture patterns in the stylized image more reasonable and harmonious. In addition, we argue that existing style transfer methods only consider the content-to-stylization and style-to-stylization relations, neglecting the stylization-to-stylization relations. To address this issue, we introduce two contrastive losses, which pull the multiple stylization embeddings closer to each other when they share the same content or style, but push far away otherwise. We conduct extensive experiments, showing that our proposed method can not only produce visually more harmonious and satisfying artistic images, but also promote the stability and consistency of rendered video clips. Haibo Chen 0006, Lei Zhao 0011, Zhizhong Wang, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
NeurIPS | 5 |
| 2021 | Diversified text-to-image generation via deep mutual information estimation
Ailin Li, Lei Zhao 0011, Zhiwen Zuo, Zhizhong Wang, Haibo Chen 0006, Dongming Lu, Wei Xing 0001 |
Comput. Vis. Image Underst. | 3 |
| 2021 | Evaluate and improve the quality of neural style transfer
Zhizhong Wang, Lei Zhao 0011, Haibo Chen 0006, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
Comput. Vis. Image Underst. | 4 |
| 2020 | UCTGAN: Diverse Image Inpainting Based on Unsupervised Cross-Space TranslationabstractAlthough existing image inpainting approaches have been able to produce visually realistic and semantically correct results, they produce only one result for each masked input. In order to produce multiple and diverse reasonable solutions, we present Unsupervised Cross-space Translation Generative Adversarial Network (called UCTGAN) which mainly consists of three network modules: conditional encoder module, manifold projection module and generation module. The manifold projection module and the generation module are combined to learn one-to-one image mapping between two spaces in an unsupervised way by projecting instance image space and conditional completion image space into common low-dimensional manifold space, which can greatly improve the diversity of the repaired samples. For understanding of global information, we also introduce a new cross semantic attention layer that exploits the long-range dependencies between the known parts and the completed parts, which can improve realism and appearance consistency of repaired samples. Extensive experiments on various datasets such as CelebA-HQ, Places2, Paris Street View and ImageNet clearly demonstrate that our method not only generates diverse inpainting solutions from the same image to be repaired, but also has high image quality. Lei Zhao 0011, Qihang Mo, Sihuan Lin, Zhizhong Wang, Zhiwen Zuo, Haibo Chen 0006, Wei Xing 0001, Dongming Lu |
CVPR | 5 |