VLDB 2026 Research / reviewers in the wild / expert
Lei Zhao 0011
dblp:87/734-11
· DBLP profile ↗
56ranked-venue papers
2as first author
51since 2021 · last 2026
0000-0003-4791-454XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 1 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 2 first-author · 33 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Inpaint-Anywhere: Zero-Shot Multi-Identity Inpainting with Efficient Diffusion TransformerabstractSubject-driven generation, which aims to synthesize visual content for a given identity V* with specific attributes, has garnered increasing attention in recent years. While existing methods demonstrate impressive identity consistency for both single and multiple identities, they often lack user-specified spatial control. Recent approaches, such as OminiControl-2 and EasyControl, enable inpainting conditioned on a single identity but fall short in multi-identity scenarios. In this paper, we introduce BoundID, a dataset synthesis pipeline for generating multi-identity images with bounding box annotations, and introduce Inpaint-Anywhere, a diffusion transformer framework for multi-identity inpainting. Given multiple identity references and corresponding masks, our method simultaneously generates all desired identities at precise locations while achieving both high identity and prompt fidelity. Extensive experiments show that Inpaint-Anywhere achieves state-of-the-art performance in multi-identity inpainting. Junsheng Luan, Lei Zhao 0011, Wei Xing 0001 |
AAAI | 2 |
| 2026 | Towards arbitrary-scale image super-resolution with prompting and diffusion prior
Xinyue Tu, Zhanjie Zhang, Wei Xing 0001, Lei Zhao 0011, Yuanxing Liu 0005 |
Knowl. Based Syst. | 5 |
| 2026 | ConceptCraft: One-Shot Personalized Text-to-Image Generation via Object-Background DisentanglementabstractPersonalized text-to-image generation aims to learn new concepts from user-provided images and subsequently generate diverse scenes or styles of the concepts from input prompts. Most existing methods usually require a set of images (typically 3-5) for each concept, which can be cumbersome. Although several methods allow personalized generation with a single reference image, they often require heavy model training and suffer from many issues such as domain-specific applicability, insufficient fidelity, and limited editability. To address these problems, we propose a novel one-shot personalized text-to-image generation method called ConceptCraft, which explicitly separates the reference image into object and background regions and treats them as two distinct concepts to learn, significantly improving the personalization performance. Specifically, we incorporate two unique identifiers into the text prompts: one followed by the object’s class name and the other by the word “background”. To bind these two identifiers to the reference image’s object and background respectively, we introduce a mask-aware object preservation loss and a mask-aware background preservation loss to optimize their corresponding token embeddings under well-designed text conditions, enabling both object and background personalization. In addition, we also develop an identifier regularization scheme to enhance our editability, allowing the synthesis of personalized images across a broader range of scenes and styles without changing the identity. Extensive qualitative and quantitative experiments are conducted to verify the effectiveness and superiority of our method. Haibo Chen 0006, Zhiwen Zuo, Lei Zhao 0011, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | IP-Controller: Decomposition and Optimization of Cross-Attention Maps for Accurate Subject-Driven Text-to-Image Diffusion GenerationabstractAlthough large pretrained stable diffusion (SD) models can generate high-quality images from prompts, they cannot generate images that are consistent with the fine-grained characteristics of a specific identityV∗(e.g., an anime character). Subject-driven generation focuses on exploring and leveraging the prior knowledge within a model to achieve the goals of ID and context preservation. There have been efforts, such as DreamBooth, to conduct subject-driven generation; however, they suffer from ID and context mistakes. An ID mistake means a feature loss ofV∗, and a context mistake means that the generated image does not align with the given prompt. To rectify these problems, in this paper, we propose masked fine-tuning for efficient feature learning ofV∗, then propose IP-Controller for decomposing and optimizing cross-attention maps ofV∗and prompt words other thanV∗. Specifically, we generate the cross-attention map using a vanilla input prompt and decompose it into an ID cross-attention map (matchingV∗) and a context cross-attention map (matching prompt words other thanV∗). Next, we generate fitter ID and context cross-attention maps on the basis of the input ID and context prompts, respectively. We optimize the ID and context cross-attention maps with the fitter ID and context cross-attention maps, respectively, so that the diffusion process pays fitter attention for specific contents. Experiments show that IP-Controller correctly integrates the core features ofV∗and the semantic context of the prompt words other thanV∗and generates high-quality images for the given prompt. Junsheng Luan, Zhanjie Zhang, Lei Zhao 0011, Wei Xing 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Fast and Robust Deformable 3D Gaussian Splattingabstract3D Gaussian Splatting has demonstrated remarkable real-time rendering capabilities and superior visual quality in novel view synthesis for static scenes. Building upon these advantages, researchers have progressively extended 3D Gaussians to dynamic scene reconstruction. Deformation field-based methods have emerged as a promising approach among various techniques. These methods maintain 3D Gaussian attributes in a canonical field and employ the deformation field to transform this field across temporal sequences. Nevertheless, these approaches frequently encounter challenges such as suboptimal rendering speeds, significant dependence on initial point clouds, and vulnerability to local optima in dim scenes. To overcome these limitations, we present FRoG, an efficient and robust framework for high-quality dynamic scene reconstruction. FRoG integrates per-Gaussian embedding with a coarse-to-fine temporal embedding strategy, accelerating rendering through the early fusion of temporal embeddings. Moreover, to enhance robustness against sparse initializations, we introduce a novel depth- and error-guided sampling strategy. This strategy populates the canonical field with new 3D Gaussians at low-deviation initial positions, significantly reducing the optimization burden on the deformation field and improving detail reconstruction in both static and dynamic regions. Furthermore, by modulating opacity variations, we mitigate the local optima problem in dim scenes, improving color fidelity. Comprehensive experimental results validate that our method achieves accelerated rendering speeds while maintaining state-of-the-art visual quality. Han Jiao 0005, Jiakai Sun, Lei Zhao 0011, Zhanjie Zhang, Wei Xing 0001, Huaizhong Lin |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Cascaded Diffusion Models for Virtual Try-On: Improving Control and ResolutionabstractPrevious virtual try-on methods have employed ControlNet architecture in exemplar-based inpainting diffusion models to guide the generation of try-on images, preserving the garment's features and enhancing the realism of the generated images. While these methods have maintained the identity of the garment and improved the naturalness of the generated images, they still face the following limitations: (1) For garments with complex features, such as intricate text, patterns, and uncommon styles, they struggle to retain these detailed features in the generated try-on images. (2) They are limited to generating try-on images at a maximum resolution of 1K, which may not meet the demands of real-world scenarios, where higher resolutions might be required. To address the aforementioned issues, in this paper, we propose a Cascaded Diffusion Model for virtual try-on to enhance both image controllability and resolution. We call it CDM-VTON. Specifically, we design two diffusion models: the Multi-Conditioned Diffusion Model (MC-DM) and the Super-Resolution Diffusion Model (SR-DM). The former generates low-resolution try-on images while preserving the garment's complex features, and the latter enhances the resolution of these images. Additionally, we incorporate a multi-control integration module in the MC-DM, which injects multiple control conditions into a frozen denoising U-Net to ensure that the generated try-on images retain complex garment features. Our experimental results demonstrate that our method outperforms previous approaches in preserving garment details and generating authentic virtual try-on images, both qualitatively and quantitatively. Junsheng Luan, Lei Zhao 0011, Wei Xing 0001, Huaizhong Lin, Binkai Ou |
AAAI | 4 |
| 2025 | Liquid-State Drive: A Case for DNA Block Device for Enormous Data
Mingkai Dong 0002, Fei Wang 0130, Jingyao Zeng, Lei Zhao 0011, Chunhai Fan, Haibo Chen 0001 |
FAST | 5 |
| 2025 | UATST: Towards unpaired arbitrary text-guided style transfer with cross-space modulation
Haibo Chen 0006, Lei Zhao 0011 |
Comput. Vis. Image Underst. | 2 |
| 2025 | Personalized text-to-image generation with Large Language and Vision Assistant enhanced training
Junsheng Luan, Zhanjie Zhang, Wei Xing 0001, Lei Zhao 0011 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | VectorSketcher: Learning to create a vector-based free-hand sketch
Zhanjie Zhang, Quanwei Zhang, Junsheng Luan, Mengyuan Yang 0002, Yun Wang 0053, Lei Zhao 0011 |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | PaintDiffusion: Towards text-driven painting variation via collaborative diffusion guidance
Haibo Chen 0006, Lei Zhao 0011, Jun Li 0027, Jian Yang 0003 |
Neurocomputing | 3 |
| 2025 | LGAST: Towards high-quality arbitrary style transfer with local-global style learning
Zhanjie Zhang, Ruichen Xia 0002, Mengyuan Yang 0002, Yun Wang 0053, Lei Zhao 0011, Wei Xing 0001 |
Neurocomputing | 6 |
| 2025 | DyArtbank: Diverse artistic style transfer via pre-trained stable diffusion and dynamic style prompt Artbank
Zhanjie Zhang, Quanwei Zhang, Junsheng Luan, Mengyuan Yang 0002, Yun Wang 0053, Lei Zhao 0011 |
Knowl. Based Syst. | 7 |
| 2025 | SPAST: Arbitrary style transfer with style priors via pre-trained large-scale model
Zhanjie Zhang, Quanwei Zhang, Junsheng Luan, Mengyuan Yang 0002, Yun Wang 0053, Lei Zhao 0011 |
Neural Networks | 6 |
| 2025 | TRTST: Arbitrary High-Quality Text-Guided Style Transfer With TransformersabstractText-guided style transfer aims to repaint a content image with the target style described by a text prompt, offering greater flexibility and creativity compared to traditional image-guided style transfer. Despite the potential, existing text-guided style transfer methods often suffer from many issues, including insufficient visual quality, poor generalization ability, or a reliance on large amounts of paired training data. To address these limitations, we leverage the inherent strengths of transformers in handling multimodal data and propose a novel transformer-based framework called TRTST that not only achieves unpaired arbitrary text-guided style transfer but also significantly improves the visual quality. Specifically, TRTST explores combining a text transformer encoder with an image transformer encoder to project the input text prompt and content image into a joint embedding space and extract the desired style and content features. These features are then input into a multimodal co-attention module to stylize the image sequence based on the text sequence. We also propose a new adaptive parametric positional encoding (APPE) scheme which can adaptively produce different positional encodings to optimally match different inputs with a position encoder. In addition, to further improve content preservation, we introduce a text-guided identity loss to our model. Extensive results and comparisons are conducted to demonstrate the effectiveness and superiority of our method. Haibo Chen 0006, Zhoujie Wang, Lei Zhao 0011, Jun Li 0027, Jian Yang 0003 |
IEEE Trans. Image Process. | 3 |
| 2024 | PNeSM: Arbitrary 3D Scene Stylization via Prompt-Based Neural Style Mappingabstract3D scene stylization refers to transform the appearance of a 3D scene to match a given style image, ensuring that images rendered from different viewpoints exhibit the same style as the given style image, while maintaining the 3D consistency of the stylized scene. Several existing methods have obtained impressive results in stylizing 3D scenes. However, the mod- els proposed by these methods need to be re-trained when applied to a new scene. In other words, their models are cou- pled with a specific scene and cannot adapt to arbitrary other scenes. To address this issue, we propose a novel 3D scene stylization framework to transfer an arbitrary style to an ar- bitrary scene, without any style-related or scene-related re- training. Concretely, we first map the appearance of the 3D scene into a 2D style pattern space, which realizes complete disentanglement of the geometry and appearance of the 3D scene and makes our model be generalized to arbitrary 3D scenes. Then we stylize the appearance of the 3D scene in the 2D style pattern space via a prompt-based 2D stylization al- gorithm. Experimental results demonstrate that our proposed framework is superior to SOTA methods in both visual qual- ity and generalization. Jiafu Chen, Wei Xing 0001, Jiakai Sun, Tianyi Chu, Boyan Ji, Lei Zhao 0011, Huaizhong Lin, Haibo Chen 0006, Zhizhong Wang |
AAAI | 7 |
| 2024 | Attack Deterministic Conditional Image Generative Models for Diverse and Controllable GenerationabstractExisting generative adversarial network (GAN) based conditional image generative models typically produce fixed output for the same conditional input, which is unreasonable for highly subjective tasks, such as large-mask image inpainting or style transfer. On the other hand, GAN-based diverse image generative methods require retraining/fine-tuning the network or designing complex noise injection functions, which is computationally expensive, task-specific, or struggle to generate high-quality results. Given that many deterministic conditional image generative models have been able to produce high-quality yet fixed results, we raise an intriguing question: is it possible for pre-trained deterministic conditional image generative models to generate diverse results without changing network structures or parameters? To answer this question, we re-examine the conditional image generation tasks from the perspective of adversarial attack and propose a simple and efficient plug-in projected gradient descent (PGD) like method for diverse and controllable image generation. The key idea is attacking the pre-trained deterministic generative models by adding a micro perturbation to the input condition. In this way, diverse results can be generated without any adjustment of network structures or fine-tuning of the pre-trained models. In addition, we can also control the diverse results to be generated by specifying the attack direction according to a reference text or image. Our work opens the door to applying adversarial attack to low-level vision tasks, and experiments on various conditional image generation tasks demonstrate the effectiveness and superiority of the proposed method. Tianyi Chu, Wei Xing 0001, Jiafu Chen, Zhizhong Wang, Jiakai Sun, Lei Zhao 0011, Haibo Chen 0006, Huaizhong Lin |
AAAI | 6 |
| 2024 | ArtBank: Artistic Style Transfer with Pre-trained Diffusion Model and Implicit Style Prompt BankabstractArtistic style transfer aims to repaint the content image with the learned artistic style. Existing artistic style transfer methods can be divided into two categories: small model-based approaches and pre-trained large-scale model-based approaches. Small model-based approaches can preserve the content strucuture, but fail to produce highly realistic stylized images and introduce artifacts and disharmonious patterns; Pre-trained large-scale model-based approaches can generate highly realistic stylized images but struggle with preserving the content structure. To address the above issues, we propose ArtBank, a novel artistic style transfer framework, to generate highly realistic stylized images while preserving the content structure of the content images. Specifically, to sufficiently dig out the knowledge embedded in pre-trained large-scale models, an Implicit Style Prompt Bank (ISPB), a set of trainable parameter matrices, is designed to learn and store knowledge from the collection of artworks and behave as a visual prompt to guide pre-trained large-scale models to generate highly realistic stylized images while preserving content structure. Besides, to accelerate training the above ISPB, we propose a novel Spatial-Statistical-based self-Attention Module (SSAM). The qualitative and quantitative experiments demonstrate the superiority of our proposed method over state-of-the-art artistic style transfer methods. Code is available at https://github.com/Jamie-Cheung/ArtBank. Zhanjie Zhang, Quanwei Zhang, Wei Xing 0001, Lei Zhao 0011, Jiakai Sun, Zehua Lan, Junsheng Luan, Huaizhong Lin |
AAAI | 5 |
| 2024 | Rethinking Diffusion Model for Multi-Contrast MRI Super-ResolutionabstractRecently, diffusion models (DM) have been applied in magnetic resonance imaging (MRI) super-resolution (SR) reconstruction, exhibiting impressive performance, especially with regard to detailed reconstruction. However, the current DM-based SR reconstruction methods still face the following issues: (1) They require a large number of iterations to reconstruct the final image, which is inefficient and consumes a significant amount of computational re-sources. (2) The results reconstructed by these methods are often misaligned with the real high-resolution images, leading to remarkable distortion in the reconstructed MR images. To address the aforementioned issues, we propose an efficient diffusion model for multi-contrast MRI SR, named as DiffMSR. Specifically, we apply DM in a highly compact low-dimensional latent space to generate prior knowledge with high-frequency detail information. The highly compact latent space ensures that DM requires only a few simple iterations to produce accurate prior knowledge. In addition, we design the Prior-Guide Large Window Trans-former (PLWformer) as the decoder for DM, which can ex-tend the receptive field while fully utilizing the prior knowledge generated by DM to ensure that the reconstructed MR image remains undistorted. Extensive experiments on public and clinical datasets demonstrate that our DiffMSR11Code: https://github.com/GuangYuanKK/DiffMSR outperforms state-of-the-art methods. Chen Rao, Juncheng Mo, Zhanjie Zhang, Wei Xing 0001, Lei Zhao 0011 |
CVPR | 6 |
| 2024 | 3DGStream: On-the-Fly Training of 3D Gaussians for Efficient Streaming of Photo-Realistic Free-Viewpoint VideosabstractConstructing photo-realistic Free-Viewpoint Videos (FVVs) of dynamic scenes from multi-view videos remains a challenging endeavor. Despite the remarkable advance- ments achieved by current neural rendering techniques, these methods generally require complete video sequences for offline training and are not capable of real-time rendering. To address these constraints, we introduce 3DGStream, a method designed for efficient FVV streaming of real-world dynamic scenes. Our method achieves fast on-the-fly per- frame reconstruction within 12 seconds and real-time ren- dering at 200 FPS. Specifically, we utilize 3D Gaussians (3DGs) to represent the scene. Instead of the naï ve ap- proach of directly optimizing 3DGs per-frame, we employ a compact Neural Transformation Cache (NTC) to model the translations and rotations of 3DGs, markedly reducing the training time and storage required for each FVV frame. Furthermore, we propose an adaptive 3DG addition strat- egy to handle emerging objects in dynamic scenes. Exper- iments demonstrate that 3DGStream achieves competitive performance in terms of rendering speed, image quality, training time, and model storage when compared with state- of-the-art methods. Jiakai Sun, Han Jiao 0005, Zhanjie Zhang, Lei Zhao 0011, Wei Xing 0001 |
CVPR | 5 |
| 2024 | Single-Mask Inpainting for Voxel-Based Neural Radiance Fields
Jiafu Chen, Tianyi Chu, Jiakai Sun, Wei Xing 0001, Lei Zhao 0011 |
ECCV (57) | 5 |
| 2024 | Rethinking Video Deblurring with Wavelet-Aware Dynamic Transformer and Diffusion Model
Chen Rao, Zehua Lan, Jiakai Sun, Junsheng Luan, Wei Xing 0001, Lei Zhao 0011, Huaizhong Lin, Jianfeng Dong, Dalong Zhang |
ECCV (45) | 7 |
| 2024 | Towards Highly Realistic Artistic Style Transfer via Stable Diffusion with Step-aware and Layer-aware Prompt
Zhanjie Zhang, Quanwei Zhang, Huaizhong Lin, Wei Xing 0001, Juncheng Mo, Shuaicheng Huang, Jinheng Xie, Junsheng Luan, Lei Zhao 0011, Dalong Zhang, Lixia Chen |
IJCAI | 10 |
| 2024 | PKD-Net: Distillation of prior knowledge for image completion by multi-level semantic attentionabstractSummary Prior knowledge plays a crucial role in image completion. Although almost all of the existing image completion methods use prior knowledge to complete the image to be repaired from different perspectives, the learning and modeling of the prior knowledge is still a challenging problem. In order to address this issue, we propose a novel prior knowledge distillation framework (PKD‐Net) which could distill prior knowledge of structure and style from multiple semantic space and generates not only plausible content but also consistent style with surrounding image area. Our PKD‐Net replaces the skip connection in the vanilla U‐Net with a semantic shift attention module. The semantic shift attention module takes features from encoder layer and those from decoder layer as input pairs to output shifted features which take into account the long‐range dependency of encoder layer features and corresponding decoder layer features from the perspective of local structure and style. Semantic shift attention module models the global interdependencies in local spatial structures (patches centered at each position) and style (appearance texture) dimensions respectively, which could implement distillation of prior knowledge from two aspects: structure and style. Experiments on multiple datasets including faces (CelebA, CelebA‐HQ) and natural images (ImageNet, Places2, Paris Street View) demonstrate that our proposed approach generates higher quality completion results than existing ones. Qiong Lu, Huaizhong Lin, Wei Xing 0001, Lei Zhao 0011, Jingjing Chen 0002 |
Concurr. Comput. Pract. Exp. | 4 |
| 2024 | Rethink arbitrary style transfer with transformer and contrastive learning
Zhanjie Zhang, Jiakai Sun, Lei Zhao 0011, Quanwei Zhang, Zehua Lan, Haolin Yin, Huaizhong Lin, Zhiwen Zuo |
Comput. Vis. Image Underst. | 4 |
| 2024 | Statistics Enhancement Generative Adversarial Networks for Diverse Conditional Image SynthesisabstractConditional generative adversarial networks (cGANs) aim to synthesize diverse images given the input conditions and the latent codes, but they are prone to map an input to a single output regardless of the variations in latent code, which is also well known as the mode collapse problem of cGANs. To alleviate the problem, in this paper, we investigate explicitly enhancing the statistical dependency between the latent code and the synthesized image in cGANs by utilizing mutual information neural estimators to estimate and maximize the conditional mutual information (CMI) between them given the input condition. The method provides a new perspective from information theory to improve diversity for cGANs and can facilitate many existing conditional image synthesis frameworks with a simple neural estimator extension. Moreover, our studies show that several key designs, including the neural estimator choice, the neural estimator’s network design, and the sampling strategy, are crucial to the success of the method. Extensive experiments on four popular conditional image synthesis tasks, including class-conditioned image generation, paired and unpaired image-to-image translation, and text-to-image generation, demonstrate the effectiveness and superiority of the proposed method. Zhiwen Zuo, Ailin Li, Zhizhong Wang, Lei Zhao 0011, Jianfeng Dong, Xun Wang 0007, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | MicroAST: Towards Super-fast Ultra-Resolution Arbitrary Style TransferabstractArbitrary style transfer (AST) transfers arbitrary artistic styles onto content images. Despite the recent rapid progress, existing AST methods are either incapable or too slow to run at ultra-resolutions (e.g., 4K) with limited resources, which heavily hinders their further applications. In this paper, we tackle this dilemma by learning a straightforward and lightweight model, dubbed MicroAST. The key insight is to completely abandon the use of cumbersome pre-trained Deep Convolutional Neural Networks (e.g., VGG) at inference. Instead, we design two micro encoders (content and style encoders) and one micro decoder for style transfer. The content encoder aims at extracting the main structure of the content image. The style encoder, coupled with a modulator, encodes the style image into learnable dual-modulation signals that modulate both intermediate features and convolutional filters of the decoder, thus injecting more sophisticated and flexible style signals to guide the stylizations. In addition, to boost the ability of the style encoder to extract more distinct and representative style signals, we also introduce a new style signal contrastive loss in our model. Compared to the state of the art, our MicroAST not only produces visually superior results but also is 5-73 times smaller and 6-18 times faster, for the first time enabling super-fast (about 0.5 seconds) AST at 4K ultra-resolutions. Zhizhong Wang, Lei Zhao 0011, Zhiwen Zuo, Ailin Li, Haibo Chen 0006, Wei Xing 0001, Dongming Lu |
AAAI | 2 |
| 2023 | Generative Image Inpainting with Segmentation Confusion Adversarial Training and Contrastive LearningabstractThis paper presents a new adversarial training framework for image inpainting with segmentation confusion adversarial training (SCAT) and contrastive learning. SCAT plays an adversarial game between an inpainting generator and a segmentation network, which provides pixel-level local training signals and can adapt to images with free-form holes. By combining SCAT with standard global adversarial training, the new adversarial training framework exhibits the following three advantages simultaneously: (1) the global consistency of the repaired image, (2) the local fine texture details of the repaired image, and (3) the flexibility of handling images with free-form holes. Moreover, we propose the textural and semantic contrastive learning losses to stabilize and improve our inpainting model's training by exploiting the feature representation space of the discriminator, in which the inpainting images are pulled closer to the ground truth images but pushed farther from the corrupted images. The proposed contrastive losses better guide the repaired images to move from the corrupted image data points to the real image data points in the feature representation space, resulting in more realistic completed images. We conduct extensive experiments on two benchmark datasets, demonstrating our model's effectiveness and superiority both qualitatively and quantitatively. Zhiwen Zuo, Lei Zhao 0011, Ailin Li, Zhizhong Wang, Zhanjie Zhang, Jiafu Chen, Wei Xing 0001, Dongming Lu |
AAAI | 2 |
| 2023 | CRFAST: Clip-Based Reference-Guided Facial Image Semantic TransferabstractThis paper presents a new task for CLIP-based reference-guided facial image semantic transfer: the source facial image is translated to the output image with the high-level semantic attributes from the reference image while maintaining identity preservation. To this end, we employ the powerful generative capability of StyleGAN generator and the rich semantic knowledge of CLIP encoder to accomplish such a task. Additionally, a novel contrastive loss is designed to comprehensively explore the rich semantic information of CLIP for facial semantic concepts. This loss guides the semantic transfer toward desired directions from different perspectives in the pre-defined CLIP space. Besides, a simple yet effective semantic-preserved modulation module is proposed to explicitly map CLIP embeddings of reference image to the latent space. Experiments demonstrate that our approach achieves realistic facial image semantic transfer driven by reference images with various facial semantics. Ailin Li, Lei Zhao 0011, Zhiwen Zuo, Zhizhong Wang, Wei Xing 0001, Dongming Lu |
ICASSP | 2 |
| 2023 | Rethinking Fast Fourier Convolution in Image InpaintingabstractRecently proposed LaMa [25] introduce Fast Fourier Convolution (FFC) [4] into image inpainting. FFC empowers the fully convolutional network to have a global receptive field in its early layers, and have the ability to produce robust repeating texture. However, LaMa has difficulty in generating clear and sharp complex content. In this paper, we analyze the fundamental flaws of using FFC in image inpainting, which are 1) spectrum shifting, 2) unexpected spatial activation, and 3) limited frequency receptive field. Such flaws make FFC-based inpainting framework difficult in generating complicated texture and performing faithful reconstruction. Based on the above analysis, we propose a novel Unbiased Fast Fourier Convolution (UFFC) module. UFFC is constructed by modifying the vanilla FFC module with 1) range transform and inverse transform, 2) absolute position embedding, 3) dynamic skip connection, and 4) adaptive clip, to overcome the above flaws. UFFC captures frequency information efficiently and realize reconstruction without introducing additional artifacts, achieving better inpainting results and more efficient training. In addition, we propose two novel perceptual losses for better generation quality and more robust training. Extensive experiments on several benchmark datasets demonstrate the effectiveness of our method, outperforming the state-of-the-art methods in both texture-capturing ability and expressiveness. Tianyi Chu, Jiafu Chen, Jiakai Sun, Shuobin Lian, Zhizhong Wang, Zhiwen Zuo, Lei Zhao 0011, Wei Xing 0001, Dongming Lu |
ICCV | 7 |
| 2023 | Rethinking Multi-Contrast MRI Super-Resolution: Rectangle-Window Cross-Attention Transformer and Arbitrary-Scale UpsamplingabstractRecently, several methods have explored the potential of multi-contrast magnetic resonance imaging (MRI) super-resolution (SR) and obtain results superior to single-contrast SR methods. However, existing approaches still have two shortcomings: (1) They can only address fixed integer upsampling scales, such as 2×, 3×, and 4×, which require training and storing the corresponding model separately for each upsampling scale in clinic. (2) They lack direct interaction among different windows as they adopt the square window (e.g., 8×8) transformer network architecture, which results in inadequate modelling of longer-range dependencies. Moreover, the relationship between reference images and target images is not fully mined. To address these issues, we develop a novel network for multi-contrast MRI arbitrary-scale SR, dubbed as McASSR. Specifically, we design a rectangle-window cross-attention transformer to establish longer-range dependencies in MR images without increasing computational complexity and fully use reference information. Besides, we propose the reference-aware implicit attention as an upsampling module, achieving arbitrary-scale super-resolution via implicit neural representation, further fusing supplementary information of the reference image. Extensive and comprehensive experiments on both public and clinical datasets show that our McASSR yields superior performance over SOTA methods, demonstrating its great potential to be applied in clinical practice. Code will be available at https://github.com/GuangYuanKK/McASSR. Lei Zhao 0011, Jiakai Sun, Zehua Lan, Zhanjie Zhang, Jiafu Chen, Huaizhong Lin, Wei Xing 0001 |
ICCV | 2 |
| 2023 | StyleDiffusion: Controllable Disentangled Style Transfer via Diffusion ModelsabstractContent and style (C-S) disentanglement is a fundamental problem and critical challenge of style transfer. Existing approaches based on explicit definitions (e.g., Gram matrix) or implicit learning (e.g., GANs) are neither interpretable nor easy to control, resulting in entangled representations and less satisfying results. In this paper, we propose a new C-S disentangled framework for style transfer without using previous assumptions. The key insight is to explicitly extract the content information and implicitly learn the complementary style information, yielding interpretable and controllable C-S disentanglement and style transfer. A simple yet effective CLIP-based style disentanglement loss coordinated with a style reconstruction prior is introduced to disentangle C-S in the CLIP image space. By further leveraging the powerful style removal and generative ability of diffusion models, our framework achieves superior results than state of the art and flexible C-S disentanglement and trade-off control. Our work provides new insights into the C-S disentanglement in style transfer and demonstrates the potential of diffusion models for learning well-disentangled C-S characteristics. Zhizhong Wang, Lei Zhao 0011, Wei Xing 0001 |
ICCV | 2 |
| 2023 | TeSTNeRF: Text-Driven 3D Style Transfer via Cross-Modal LearningabstractText-driven 3D style transfer aims at stylizing a scene according to the text and generating arbitrary novel views with consistency. Simply combining image/video style transfer methods and novel view synthesis methods results in flickering when changing viewpoints, while existing 3D style transfer methods learn styles from images instead of texts. To address this problem, we for the first time design an efficient text-driven model for 3D style transfer, named TeSTNeRF, to stylize the scene using texts via cross-modal learning: we leverage an advanced text encoder to embed the texts in order to control 3D style transfer and align the input text and output stylized images in latent space. Furthermore, to obtain better visual results, we introduce style supervision, learning feature statistics from style images and utilizing 2D stylization results to rectify abrupt color spill. Extensive experiments demonstrate that TeSTNeRF significantly outperforms existing methods and provides a new way to guide 3D style transfer. Jiafu Chen, Boyan Ji, Zhanjie Zhang, Tianyi Chu, Zhiwen Zuo, Lei Zhao 0011, Wei Xing 0001, Dongming Lu |
IJCAI | 6 |
| 2023 | VGOS: Voxel Grid Optimization for View Synthesis from Sparse InputsabstractNeural Radiance Fields (NeRF) has shown great success in novel view synthesis due to its state-of-the-art quality and flexibility. However, NeRF requires dense input views (tens to hundreds) and a long training time (hours to days) for a single scene to generate high-fidelity images. Although using the voxel grids to represent the radiance field can significantly accelerate the optimization process, we observe that for sparse inputs, the voxel grids are more prone to overfitting to the training views and will have holes and floaters, which leads to artifacts. In this paper, we propose VGOS, an approach for fast (3-5 minutes) radiance field reconstruction from sparse inputs (3-10 views) to address these issues. To improve the performance of voxel-based radiance field in sparse input scenarios, we propose two methods: (a) We introduce an incremental voxel training strategy, which prevents overfitting by suppressing the optimization of peripheral voxels in the early stage of reconstruction. (b) We use several regularization techniques to smooth the voxels, which avoids degenerate solutions. Experiments demonstrate that VGOS achieves state-of-the-art performance for sparse inputs with super-fast convergence. Code will be available at https://github.com/SJoJoK/VGOS. Jiakai Sun, Zhanjie Zhang, Jiafu Chen, Boyan Ji, Lei Zhao 0011, Wei Xing 0001 |
IJCAI | 6 |
| 2023 | TSSAT: Two-Stage Statistics-Aware Transformation for Artistic Style TransferabstractArtistic style transfer aims to create new artistic images by rendering a given photograph with the target artistic style. Existing methods learn styles simply based on global statistics or local patches, lacking careful consideration of the drawing process in practice. Consequently, the stylization results either fail to capture abundant and diversified local style patterns, or contain undesired semantic information of the style image and deviate from the global style distribution. To address this issue, we imitate the drawing process of humans and propose a Two-Stage Statistics-Aware Transformation (TSSAT) module, which first builds the global style foundation by aligning the global statistics of content and style features and then further enriches local style details by swapping the local statistics (instead of local features) in a patch-wise manner, significantly improving the stylization effects. Moreover, to further enhance both content and style representations, we introduce two novel losses: an attention-based content loss and a patch-based style loss, where the former enables better content preservation by enforcing the semantic relation in the content image to be retained during stylization, and the latter focuses on increasing the local style similarity between the style and stylized images. Extensive qualitative and quantitative experiments verify the effectiveness of our method. Haibo Chen 0006, Lei Zhao 0011, Jun Li 0027, Jian Yang 0003 |
ACM Multimedia | 2 |
| 2023 | Self-Reference Image Super-Resolution via Pre-trained Diffusion Large Model and Window Adjustable TransformerabstractCurrently, reference-based super-resolution (RefSR) techniques leverage high-resolution (HR) reference images to provide useful content and texture information for low-resolution (LR) images during the super-resolution (SR) process. Nevertheless, it is time-consuming, laborious, and even impossible in some cases to find high-quality reference images. To tackle this problem, we propose a brand-new self-reference image super-resolution approach using a pre-trained diffusion large model and a window adjustable transformer, termed DWTrans. Our proposed method does not require explicitly inputting manually acquired reference images during training and inference. Specifically, we feed the degraded LR images into a pre-trained stable diffusion large model to automatically generate corresponding high-quality self-reference (SRef) images that provide valuable high-frequency details for the LR images in the process of SR. To extract valuable high-frequency information in SRef images, we design a window adjustable transformer with both non-adjustable window layer (NWL) and adjustable window layer (AWL). The NWL learns local features from LR images using a dense window, while the AWL acquires global features from the SRef images using a random sparse window. Furthermore, to fully utilize the high-frequency features in the SRef image, we introduce the adaptive deformable fusion module to adaptively fuse the features of the LR and SRef images. Experimental results validate that our proposed DWTrans outperforms state-of-the-art methods on various benchmark datasets both quantitatively and visually. Wei Xing 0001, Lei Zhao 0011, Zehua Lan, Jiakai Sun, Zhanjie Zhang, Quanwei Zhang, Huaizhong Lin |
ACM Multimedia | 3 |
| 2023 | DuDoINet: Dual-Domain Implicit Network for Multi-Modality MR Image Arbitrary-scale Super-ResolutionabstractCompared to single-modality magnetic resonance (MR) image super-resolution (SR) methods, multi-modality MR image methods can utilize high-resolution reference modality (e.g., T1 modality) to provide valuable complementary information for low-resolution target modality (e.g., T2 modality) in SR reconstruction, which can further improve the quality of the SR images. Although they have achieved impressive results, these methods still suffer from the following drawbacks: (1) They can only handle fixed integer upsampling factors, such as 2X, 3X, and 4X, and require training and storing corresponding models for each upsampling factor, which is infeasible in clinical practice; (2) They only perform feature extraction and reconstruction in the image domain. However, the aliasing artifacts produced in the image domain are structural and non-local. Therefore, using only the image domain cannot effectively reconstruct high-quality aliasing-free SR images. To address these issues, we develop a brand-new Dual-Domain Implicit Network (DuDoINet) for multi-modality MR image arbitrary-scale SR. Specifically, we propose a dual-domain learning scheme for multi-modality MR image SR, which allows the network to sufficiently exploit the frequency and image domain information in MR images. In addition, we design implicit attention to achieve arbitrary-scale upsampling of MR images, which utilizes a continuously differentiable function that generates pixel values from pixel coordinates. Furthermore, we designed a deformable cross-modality attention mechanism that can adaptively transfer high-frequency details from the T1 to the T2 modality, better integrating valuable complementary information from the T1 modality. Extensive and comprehensive experiments on healthy subjects and patient datasets demonstrate that our DuDoINet outperforms SOTA methods, demonstrating its great potential for clinical practice. Wei Xing 0001, Lei Zhao 0011, Zehua Lan, Zhanjie Zhang, Jiakai Sun, Haolin Yin, Huaizhong Lin |
ACM Multimedia | 3 |
| 2023 | Towards Interactive Facial Image Inpainting by Text or Exemplar Image
Ailin Li, Lei Zhao 0011, Zhiwen Zuo, Zhizhong Wang, Wei Xing 0001, Dongming Lu |
MMM (1) | 2 |
| 2023 | MIGT: Multi-modal image inpainting guided with text
Ailin Li, Lei Zhao 0011, Zhiwen Zuo, Zhizhong Wang, Wei Xing 0001, Dongming Lu |
Neurocomputing | 2 |
| 2023 | Caster: Cartoon style transfer via dynamic cartoon style casting
Zhanjie Zhang, Jiakai Sun, Jiafu Chen, Lei Zhao 0011, Boyan Ji, Zehua Lan, Wei Xing 0001, Duanqing Xu |
Neurocomputing | 4 |
| 2022 | Texture Reformer: Towards Fast and Universal Interactive Texture TransferabstractIn this paper, we present the texture reformer, a fast and universal neural-based framework for interactive texture transfer with user-specified guidance. The challenges lie in three aspects: 1) the diversity of tasks, 2) the simplicity of guidance maps, and 3) the execution efficiency. To address these challenges, our key idea is to use a novel feed-forward multi-view and multi-stage synthesis procedure consisting of I) a global view structure alignment stage, II) a local view texture refinement stage, and III) a holistic effect enhancement stage to synthesize high-quality results with coherent structures and fine texture details in a coarse-to-fine fashion. In addition, we also introduce a novel learning-free view-specific texture reformation (VSTR) operation with a new semantic map guidance strategy to achieve more accurate semantic-guided and structure-preserved texture transfer. The experimental results on a variety of application scenarios demonstrate the effectiveness and superiority of our framework. And compared with the state-of-the-art interactive texture transfer algorithms, it not only achieves higher quality results but, more remarkably, also is 2-5 orders of magnitude faster. Zhizhong Wang, Lei Zhao 0011, Haibo Chen 0006, Ailin Li, Zhiwen Zuo, Wei Xing 0001, Dongming Lu |
AAAI | 2 |
| 2022 | DivSwapper: Towards Diversified Patch-based Arbitrary Style TransferabstractGram-based and patch-based approaches are two important research lines of style transfer. Recent diversified Gram-based methods have been able to produce multiple and diverse stylized outputs for the same content and style images. However, as another widespread research interest, the diversity of patch-based methods remains challenging due to the stereotyped style swapping process based on nearest patch matching. To resolve this dilemma, in this paper, we dive into the crux of existing patch-based methods and propose a universal and efficient module, termed DivSwapper, for diversified patch-based arbitrary style transfer. The key insight is to use an essential intuition that neural patches with higher activation values could contribute more to diversity. Our DivSwapper is plug-and-play and can be easily integrated into existing patch-based and Gram-based methods to generate diverse results for arbitrary styles. We conduct theoretical analyses and extensive experiments to demonstrate the effectiveness of our method, and compared with state-of-the-art algorithms, it shows superiority in diversity, quality, and efficiency. Zhizhong Wang, Lei Zhao 0011, Haibo Chen 0006, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
IJCAI | 2 |
| 2022 | Style Fader Generative Adversarial Networks for Style Degree Controllable Artistic Style TransferabstractArtistic style transfer is the task of synthesizing content images with learned artistic styles. Recent studies have shown the potential of Generative Adversarial Networks (GANs) for producing artistically rich stylizations. Despite the promising results, they usually fail to control the generated images' style degree, which is inflexible and limits their applicability for practical use. To address the issue, in this paper, we propose a novel method that for the first time allows adjusting the style degree for existing GAN-based artistic style transfer frameworks in real time after training. Our method introduces two novel modules into existing GAN-based artistic style transfer frameworks: a Style Scaling Injection (SSI) module and a Style Degree Interpretation (SDI) module. The SSI module accepts the value of Style Degree Factor (SDF) as the input and outputs parameters that scale the feature activations in existing models, offering control signals to alter the style degrees of the stylizations. And the SDI module interprets the output probabilities of a multi-scale content-style binary classifier as the style degrees, providing a mechanism to parameterize the style degree of the stylizations. Moreover, we show that after training our method can enable existing GAN-based frameworks to produce over-stylizations. The proposed method can facilitate many existing GAN-based artistic style transfer frameworks with marginal extra training overheads and modifications. Extensive qualitative evaluations on two typical GAN-based style transfer models demonstrate the effectiveness of the proposed method for gaining style degree control for them. Zhiwen Zuo, Lei Zhao 0011, Shuobin Lian, Haibo Chen 0006, Zhizhong Wang, Ailin Li, Wei Xing 0001, Dongming Lu |
IJCAI | 2 |
| 2022 | AesUST: Towards Aesthetic-Enhanced Universal Style TransferabstractRecent studies have shown remarkable success in universal style transfer which transfers arbitrary visual styles to content images. However, existing approaches suffer from the aesthetic-unrealistic problem that introduces disharmonious patterns and evident artifacts, making the results easy to spot from real paintings. To address this limitation, we propose AesUST, a novel Aesthetic-enhanced Universal Style Transfer approach that can generate aesthetically more realistic and pleasing results for arbitrary styles. Specifically, our approach introduces an aesthetic discriminator to learn the universal human-delightful aesthetic features from a large corpus of artist-created paintings. Then, the aesthetic features are incorporated to enhance the style transfer process via a novel Aesthetic-aware Style-Attention (AesSA) module. Such an AesSA module enables our AesUST to efficiently and flexibly integrate the style patterns according to the global aesthetic channel distribution of the style image and the local semantic spatial distribution of the content image. Moreover, we also develop a new two-stage transfer training strategy with two aesthetic regularizations to train our model more effectively, further improving stylization performance. Extensive experiments and user studies demonstrate that our approach synthesizes aesthetically more harmonious and realistic results than state of the art, greatly narrowing the disparity with real artist-created paintings. Our code is available at https://github.com/EndyWon/AesUST. Zhizhong Wang, Zhanjie Zhang, Lei Zhao 0011, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
ACM Multimedia | 3 |
| 2022 | Dual distribution matching GAN
Zhiwen Zuo, Lei Zhao 0011, Ailin Li, Zhizhong Wang, Haibo Chen 0006, Wi Xing, Dongming Lu |
Neurocomputing | 2 |
| 2022 | Dual-constraint burst image denoising methodabstractDeep learning has proven to be an effective mechanism for computer vision tasks, especially for image denoising and burst image denoising. In this paper, we focus on solving the burst image denoising problem and aim to generate a single clean image from a burst of noisy images. We propose to combine the power of block matching and 3D filtering (BM3D) and a convolutional neural network (CNN) for burst image denoising. In particular, we design a CNN with a divide-and-conquer strategy. First, we employ BM3D to preprocess the noisy burst images. Then, the preprocessed images and noisy images are fed separately into two parallel CNN branches. The two branches produce somewhat different results. Finally, we use a light CNN block to combine the two outputs. In addition, we improve the performance by optimizing the two branches using two different constraints: a signal constraint and a noise constraint. One maps a clean signal, and the other maps the noise distribution. In addition, we adopt block matching in the network to avoid frame misalignment. Experimental results on synthetic and real noisy images show that our algorithm is competitive with other algorithms. Lei Zhao 0011, Duanqing Xu, Dongming Lu |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2021 | DualAST: Dual Style-Learning Networks for Artistic Style TransferabstractArtistic style transfer is an image editing task that aims at repainting everyday photographs with learned artistic styles. Existing methods learn styles from either a single style example or a collection of artworks. Accordingly, the stylization results are either inferior in visual quality or limited in style controllability. To tackle this problem, we propose a novel Dual Style-Learning Artistic Style Transfer (DualAST) framework to learn simultaneously both the holistic artist-style (from a collection of artworks) and the specific artwork-style (from a single style image): the artist-style sets the tone (i.e., the overall feeling) for the stylized image, while the artwork-style determines the details of the stylized image, such as color and texture. Moreover, we introduce a Style-Control Block (SCB) to adjust the styles of generated images with a set of learnable style-control factors. We conduct extensive experiments to evaluate the performance of the proposed framework, the results of which confirm the superiority of our method. Haibo Chen 0006, Lei Zhao 0011, Zhizhong Wang, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
CVPR | 2 |
| 2021 | Diverse Image Style Transfer via Invertible Cross-Space MappingabstractImage style transfer aims to transfer the styles of artworks onto arbitrary photographs to create novel artistic images. Although style transfer is inherently an underdetermined problem, existing approaches usually assume a deterministic solution, thus failing to capture the full distribution of possible outputs. To address this limitation, we propose a Diverse Image Style Transfer (DIST) framework which achieves significant diversity by enforcing an invertible cross-space mapping. Specifically, the framework consists of three branches: disentanglement branch, inverse branch, and stylization branch. Among them, the disentanglement branch factorizes artworks into content space and style space; the inverse branch encourages the invertible mapping between the latent space of input noise vectors and the style space of generated artistic images; the stylization branch renders the input content image with the style of an artist. Armed with these three branches, our approach is able to synthesize significantly diverse stylized images without loss of quality. We conduct extensive experiments and comparisons to evaluate our approach qualitatively and quantitatively. The experimental results demonstrate the effectiveness of our method. Haibo Chen 0006, Lei Zhao 0011, Zhizhong Wang, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
ICCV | 2 |
| 2021 | Artistic Style Transfer with Internal-external Learning and Contrastive LearningabstractAlthough existing artistic style transfer methods have achieved significant improvement with deep neural networks, they still suffer from artifacts such as disharmonious colors and repetitive patterns. Motivated by this, we propose an internal-external style transfer method with two contrastive losses. Specifically, we utilize internal statistics of a single style image to determine the colors and texture patterns of the stylized image, and in the meantime, we leverage the external information of the large-scale style dataset to learn the human-aware style information, which makes the color distributions and texture patterns in the stylized image more reasonable and harmonious. In addition, we argue that existing style transfer methods only consider the content-to-stylization and style-to-stylization relations, neglecting the stylization-to-stylization relations. To address this issue, we introduce two contrastive losses, which pull the multiple stylization embeddings closer to each other when they share the same content or style, but push far away otherwise. We conduct extensive experiments, showing that our proposed method can not only produce visually more harmonious and satisfying artistic images, but also promote the stability and consistency of rendered video clips. Haibo Chen 0006, Lei Zhao 0011, Zhizhong Wang, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
NeurIPS | 2 |
| 2021 | Diversified text-to-image generation via deep mutual information estimation
Ailin Li, Lei Zhao 0011, Zhiwen Zuo, Zhizhong Wang, Haibo Chen 0006, Dongming Lu, Wei Xing 0001 |
Comput. Vis. Image Underst. | 2 |
| 2021 | Evaluate and improve the quality of neural style transfer
Zhizhong Wang, Lei Zhao 0011, Haibo Chen 0006, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
Comput. Vis. Image Underst. | 2 |
| 2020 | Diversified Arbitrary Style Transfer via Deep Feature PerturbationabstractImage style transfer is an underdetermined problem, where a large number of solutions can satisfy the same constraint (the content and style). Although there have been some efforts to improve the diversity of style transfer by introducing an alternative diversity loss, they have restricted generalization, limited diversity and poor scalability. In this paper, we tackle these limitations and propose a simple yet effective method for diversified arbitrary style transfer. The key idea of our method is an operation called deep feature perturbation (DFP), which uses an orthogonal random noise matrix to perturb the deep image feature maps while keeping the original style information unchanged. Our DFP operation can be easily integrated into many existing WCT (whitening and coloring transform)-based methods, and empower them to generate diverse results for arbitrary styles. Experimental results demonstrate that this learning-free and universal method can greatly increase the diversity while maintaining the quality of stylization. Zhizhong Wang, Lei Zhao 0011, Haibo Chen 0006, Lihong Qiu, Qihang Mo, Sihuan Lin, Wei Xing 0001, Dongming Lu |
CVPR | 2 |
| 2020 | UCTGAN: Diverse Image Inpainting Based on Unsupervised Cross-Space TranslationabstractAlthough existing image inpainting approaches have been able to produce visually realistic and semantically correct results, they produce only one result for each masked input. In order to produce multiple and diverse reasonable solutions, we present Unsupervised Cross-space Translation Generative Adversarial Network (called UCTGAN) which mainly consists of three network modules: conditional encoder module, manifold projection module and generation module. The manifold projection module and the generation module are combined to learn one-to-one image mapping between two spaces in an unsupervised way by projecting instance image space and conditional completion image space into common low-dimensional manifold space, which can greatly improve the diversity of the repaired samples. For understanding of global information, we also introduce a new cross semantic attention layer that exploits the long-range dependencies between the known parts and the completed parts, which can improve realism and appearance consistency of repaired samples. Extensive experiments on various datasets such as CelebA-HQ, Places2, Paris Street View and ImageNet clearly demonstrate that our method not only generates diverse inpainting solutions from the same image to be repaired, but also has high image quality. Lei Zhao 0011, Qihang Mo, Sihuan Lin, Zhizhong Wang, Zhiwen Zuo, Haibo Chen 0006, Wei Xing 0001, Dongming Lu |
CVPR | 1 |
| 2020 | SpatialGAN: Progressive Image Generation Based on Spatial Recursive Adversarial ExpansionabstractThe image generation model based on generative adversarial networks has recently received significant attention and can produce diverse, sharp, and realistic images. However, generating high-resolution images has long been a challenge. In this paper, we propose a progressive spatial recursive adversarial expansion model(called SpatialGAN) capable of producing high-quality samples of the natural image. Our approach uses a cascade of convolutional networks to progressively generate images in a part-to-whole fashion. At each level of spatial expansion, a separate image-to-image spatial adversarial expansion network (conditional GAN) is recursively trained based on context image generated by previous GAN or CGAN. Unlike other coarse-to-fine generative methods that constraint on generative process either by multi-scale resolution or by hierarchical feature, the SpatialGAN decomposes image space into multiple subspaces and gradually resolves uncertainties in the local-to-whole generative process. The SpatialGAN greatly stabilizes and speeds up the training, which allows us to produce images of high quality. Based on visual Inception Score and Fréchet Inception Distance, we demonstrate that the quality of images generated by SpatialGAN on several typical datasets is better than that of images generated by GANs without cascading and comparative with the state of art methods with cascading. Lei Zhao 0011, Sihuan Lin, Ailin Li, Huaizhong Lin, Wei Xing 0001, Dongming Lu |
ACM Multimedia | 1 |
| 2020 | Creative and diverse artwork generation using adversarial networksabstractExisting style transfer methods have achieved great success in artwork generation by transferring artistic styles onto everyday photographs while keeping their contents unchanged. Despite this success, these methods have one inherent limitation: they cannot produce newly created image contents, lacking creativity and flexibility. On the other hand, generative adversarial networks (GANs) can synthesise images with new content, whereas cannot specify the artistic style of these images. The authors consider combining style transfer with convolutional GANs to generate more creative and diverse artworks. Instead of simply concatenating these two networks: the first for synthesising new content and the second for transferring artistic styles, which is inefficient and inconvenient, they design an end‐to‐end network called ArtistGAN to perform these two operations at the same time and achieve visually better results. Moreover, to generate images of higher quality, they propose the bi‐discriminator GAN containing a pixel discriminator and a feature discriminator that constrain the generated image from pixel level and feature level, respectively. They conduct extensive experiments and comparisons to evaluate their methods quantitatively and qualitatively. The experimental results verify the effectiveness of their methods. Haibo Chen 0006, Lei Zhao 0011, Lihong Qiu, Zhizhong Wang, Wei Xing 0001, Dongming Lu |
IET Comput. Vis. | 2 |
| 2020 | GLStyleNet: exquisite style transfer combining global and local pyramid featuresabstractRecent studies using deep neural networks have shown remarkable success in style transfer, especially for artistic and photo‐realistic images. However, these methods cannot solve more sophisticated problems. The approaches using global statistics fail to capture small, intricate textures and maintain correct texture scales of the artworks, and the others based on local patches are defective on global effect. To address these issues, this study presents a unified model [global and local style network (GLStyleNet)] to achieve exquisite style transfer with higher quality. Specifically, a simple yet effective perceptual loss is proposed to consider the information of global semantic‐level structure, local patch‐level style, and global channel‐level effect at the same time. This could help transfer not just large‐scale, obvious style cues but also subtle, exquisite ones, and dramatically improve the quality of style transfer. Besides, the authors introduce a novel deep pyramid feature fusion module to provide a more flexible style expression and a more efficient transfer process. This could help retain both high‐frequency pixel information and low‐frequency construct information. They demonstrate the effectiveness and superiority of their approach on numerous style transfer tasks, especially the Chinese ancient painting style transfer. Experimental results indicate that their unified approach improves image style transfer quality over previous state‐of‐the‐art methods. Zhizhong Wang, Lei Zhao 0011, Sihuan Lin, Qihang Mo, Wei Xing 0001, Dongming Lu |
IET Comput. Vis. | 2 |