EDBT 2026 Demo / reviewers in the wild / expert
Zhizhong Wang
dblp:26/4017
· DBLP profile ↗
39ranked-venue papers
9as first author
24since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 7 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 17 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SCSA: A Plug-and-Play Semantic Continuous-Sparse Attention for Arbitrary Semantic Style TransferabstractAttention-based arbitrary style transfer methods, including CNN-based, Transformer-based, and Diffusion-based, have flourished and produced high-quality stylized images. However, they perform poorly on the content and style images with the same semantics, i.e., the style of the corresponding semantic region of the generated stylized image is inconsistent with that of the style image. We argue that the root cause lies in their failure to consider the relationship between local regions and semantic regions. To address this issue, we propose a plug-and-play semantic continuous-sparse attention, dubbed SCSA, for arbitrary semantic style transfer—each query point considers certain key points in the corresponding semantic region. Specifically, semantic continuous attention ensures each query point fully attends to all the continuous key points in the same semantic region that reflect the overall style characteristics of that region; Semantic sparse attention allows each query point to focus on the most similar sparse key point in the same semantic region that exhibits the specific stylistic texture of that region. By combining the two modules, the resulting SCSA aligns the overall style of the corresponding semantic regions while transferring the vivid textures of these regions. Qualitative and quantitative results demonstrate that SCSA enables attention-based arbitrary style transfer methods to produce high-quality semantic stylized images. The code can be found in https://github.com/scn-00/SCSA. Chunnan Shang, Zhizhong Wang, Xiangming Meng |
CVPR | 2 |
| 2024 | PNeSM: Arbitrary 3D Scene Stylization via Prompt-Based Neural Style Mappingabstract3D scene stylization refers to transform the appearance of a 3D scene to match a given style image, ensuring that images rendered from different viewpoints exhibit the same style as the given style image, while maintaining the 3D consistency of the stylized scene. Several existing methods have obtained impressive results in stylizing 3D scenes. However, the mod- els proposed by these methods need to be re-trained when applied to a new scene. In other words, their models are cou- pled with a specific scene and cannot adapt to arbitrary other scenes. To address this issue, we propose a novel 3D scene stylization framework to transfer an arbitrary style to an ar- bitrary scene, without any style-related or scene-related re- training. Concretely, we first map the appearance of the 3D scene into a 2D style pattern space, which realizes complete disentanglement of the geometry and appearance of the 3D scene and makes our model be generalized to arbitrary 3D scenes. Then we stylize the appearance of the 3D scene in the 2D style pattern space via a prompt-based 2D stylization al- gorithm. Experimental results demonstrate that our proposed framework is superior to SOTA methods in both visual qual- ity and generalization. Jiafu Chen, Wei Xing 0001, Jiakai Sun, Tianyi Chu, Boyan Ji, Lei Zhao 0011, Huaizhong Lin, Haibo Chen 0006, Zhizhong Wang |
AAAI | 10 |
| 2024 | Attack Deterministic Conditional Image Generative Models for Diverse and Controllable GenerationabstractExisting generative adversarial network (GAN) based conditional image generative models typically produce fixed output for the same conditional input, which is unreasonable for highly subjective tasks, such as large-mask image inpainting or style transfer. On the other hand, GAN-based diverse image generative methods require retraining/fine-tuning the network or designing complex noise injection functions, which is computationally expensive, task-specific, or struggle to generate high-quality results. Given that many deterministic conditional image generative models have been able to produce high-quality yet fixed results, we raise an intriguing question: is it possible for pre-trained deterministic conditional image generative models to generate diverse results without changing network structures or parameters? To answer this question, we re-examine the conditional image generation tasks from the perspective of adversarial attack and propose a simple and efficient plug-in projected gradient descent (PGD) like method for diverse and controllable image generation. The key idea is attacking the pre-trained deterministic generative models by adding a micro perturbation to the input condition. In this way, diverse results can be generated without any adjustment of network structures or fine-tuning of the pre-trained models. In addition, we can also control the diverse results to be generated by specifying the attack direction according to a reference text or image. Our work opens the door to applying adversarial attack to low-level vision tasks, and experiments on various conditional image generation tasks demonstrate the effectiveness and superiority of the proposed method. Tianyi Chu, Wei Xing 0001, Jiafu Chen, Zhizhong Wang, Jiakai Sun, Lei Zhao 0011, Haibo Chen 0006, Huaizhong Lin |
AAAI | 4 |
| 2024 | DKNN: deep kriging neural network for interpretable geospatial interpolationabstractGeospatial interpolation plays a pivotal role in spatial analysis because it provides high-quality data support for various spatiotemporal data mining (STDM) tasks. However, statistical methods, such as kriging, face challenges in dealing with complex geo-big data. Additionally, deep-learning-based methods, despite their exceptional performance, suffer from limitations, such as poor interpretability. To harness the complementary advantages of these statistical methods and deep learning approaches, this study proposes a novel geospatial artificial intelligence (GeoAI) framework called deep kriging neural network (DKNN). The primary contribution lies in the development of an asymmetric encoder-decoder structure, which includes a deep-learning-based spatial encoder and a geostatistics-based kriging decoder. The spatial encoder consists of three specialized neural networks, whereas the kriging decoder relies on the proposed unified kriging system. During forward propagation, the kriging decoder leverage messages from the spatial encoder to generate interpolation weights for prediction. Conversely, during backward propagation, the kriging decoder guides the spatial encoder in learning interpretable knowledge. Experiments were conducted using both synthetic and practical datasets. The results demonstrate an average improvement of 20.18% in MAE, 25.04% in RMSE and 24.06% in MAPE when compared to the best-performing baseline method. Furthermore, these results confirm the superior interpretability of our DKNN framework. Enbo Liu, Xiaoyong Tan, Jiaoju Wang, Yan Shi 0007, Zhizhong Wang |
Int. J. Geogr. Inf. Sci. | 7 |
| 2024 | Statistics Enhancement Generative Adversarial Networks for Diverse Conditional Image SynthesisabstractConditional generative adversarial networks (cGANs) aim to synthesize diverse images given the input conditions and the latent codes, but they are prone to map an input to a single output regardless of the variations in latent code, which is also well known as the mode collapse problem of cGANs. To alleviate the problem, in this paper, we investigate explicitly enhancing the statistical dependency between the latent code and the synthesized image in cGANs by utilizing mutual information neural estimators to estimate and maximize the conditional mutual information (CMI) between them given the input condition. The method provides a new perspective from information theory to improve diversity for cGANs and can facilitate many existing conditional image synthesis frameworks with a simple neural estimator extension. Moreover, our studies show that several key designs, including the neural estimator choice, the neural estimator’s network design, and the sampling strategy, are crucial to the success of the method. Extensive experiments on four popular conditional image synthesis tasks, including class-conditioned image generation, paired and unpaired image-to-image translation, and text-to-image generation, demonstrate the effectiveness and superiority of the proposed method. Zhiwen Zuo, Ailin Li, Zhizhong Wang, Lei Zhao 0011, Jianfeng Dong, Xun Wang 0007, Meng Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | MicroAST: Towards Super-fast Ultra-Resolution Arbitrary Style TransferabstractArbitrary style transfer (AST) transfers arbitrary artistic styles onto content images. Despite the recent rapid progress, existing AST methods are either incapable or too slow to run at ultra-resolutions (e.g., 4K) with limited resources, which heavily hinders their further applications. In this paper, we tackle this dilemma by learning a straightforward and lightweight model, dubbed MicroAST. The key insight is to completely abandon the use of cumbersome pre-trained Deep Convolutional Neural Networks (e.g., VGG) at inference. Instead, we design two micro encoders (content and style encoders) and one micro decoder for style transfer. The content encoder aims at extracting the main structure of the content image. The style encoder, coupled with a modulator, encodes the style image into learnable dual-modulation signals that modulate both intermediate features and convolutional filters of the decoder, thus injecting more sophisticated and flexible style signals to guide the stylizations. In addition, to boost the ability of the style encoder to extract more distinct and representative style signals, we also introduce a new style signal contrastive loss in our model. Compared to the state of the art, our MicroAST not only produces visually superior results but also is 5-73 times smaller and 6-18 times faster, for the first time enabling super-fast (about 0.5 seconds) AST at 4K ultra-resolutions. Zhizhong Wang, Lei Zhao 0011, Zhiwen Zuo, Ailin Li, Haibo Chen 0006, Wei Xing 0001, Dongming Lu |
AAAI | 1 |
| 2023 | Generative Image Inpainting with Segmentation Confusion Adversarial Training and Contrastive LearningabstractThis paper presents a new adversarial training framework for image inpainting with segmentation confusion adversarial training (SCAT) and contrastive learning. SCAT plays an adversarial game between an inpainting generator and a segmentation network, which provides pixel-level local training signals and can adapt to images with free-form holes. By combining SCAT with standard global adversarial training, the new adversarial training framework exhibits the following three advantages simultaneously: (1) the global consistency of the repaired image, (2) the local fine texture details of the repaired image, and (3) the flexibility of handling images with free-form holes. Moreover, we propose the textural and semantic contrastive learning losses to stabilize and improve our inpainting model's training by exploiting the feature representation space of the discriminator, in which the inpainting images are pulled closer to the ground truth images but pushed farther from the corrupted images. The proposed contrastive losses better guide the repaired images to move from the corrupted image data points to the real image data points in the feature representation space, resulting in more realistic completed images. We conduct extensive experiments on two benchmark datasets, demonstrating our model's effectiveness and superiority both qualitatively and quantitatively. Zhiwen Zuo, Lei Zhao 0011, Ailin Li, Zhizhong Wang, Zhanjie Zhang, Jiafu Chen, Wei Xing 0001, Dongming Lu |
AAAI | 4 |
| 2023 | CRFAST: Clip-Based Reference-Guided Facial Image Semantic TransferabstractThis paper presents a new task for CLIP-based reference-guided facial image semantic transfer: the source facial image is translated to the output image with the high-level semantic attributes from the reference image while maintaining identity preservation. To this end, we employ the powerful generative capability of StyleGAN generator and the rich semantic knowledge of CLIP encoder to accomplish such a task. Additionally, a novel contrastive loss is designed to comprehensively explore the rich semantic information of CLIP for facial semantic concepts. This loss guides the semantic transfer toward desired directions from different perspectives in the pre-defined CLIP space. Besides, a simple yet effective semantic-preserved modulation module is proposed to explicitly map CLIP embeddings of reference image to the latent space. Experiments demonstrate that our approach achieves realistic facial image semantic transfer driven by reference images with various facial semantics. Ailin Li, Lei Zhao 0011, Zhiwen Zuo, Zhizhong Wang, Wei Xing 0001, Dongming Lu |
ICASSP | 4 |
| 2023 | Rethinking Fast Fourier Convolution in Image InpaintingabstractRecently proposed LaMa [25] introduce Fast Fourier Convolution (FFC) [4] into image inpainting. FFC empowers the fully convolutional network to have a global receptive field in its early layers, and have the ability to produce robust repeating texture. However, LaMa has difficulty in generating clear and sharp complex content. In this paper, we analyze the fundamental flaws of using FFC in image inpainting, which are 1) spectrum shifting, 2) unexpected spatial activation, and 3) limited frequency receptive field. Such flaws make FFC-based inpainting framework difficult in generating complicated texture and performing faithful reconstruction. Based on the above analysis, we propose a novel Unbiased Fast Fourier Convolution (UFFC) module. UFFC is constructed by modifying the vanilla FFC module with 1) range transform and inverse transform, 2) absolute position embedding, 3) dynamic skip connection, and 4) adaptive clip, to overcome the above flaws. UFFC captures frequency information efficiently and realize reconstruction without introducing additional artifacts, achieving better inpainting results and more efficient training. In addition, we propose two novel perceptual losses for better generation quality and more robust training. Extensive experiments on several benchmark datasets demonstrate the effectiveness of our method, outperforming the state-of-the-art methods in both texture-capturing ability and expressiveness. Tianyi Chu, Jiafu Chen, Jiakai Sun, Shuobin Lian, Zhizhong Wang, Zhiwen Zuo, Lei Zhao 0011, Wei Xing 0001, Dongming Lu |
ICCV | 5 |
| 2023 | StyleDiffusion: Controllable Disentangled Style Transfer via Diffusion ModelsabstractContent and style (C-S) disentanglement is a fundamental problem and critical challenge of style transfer. Existing approaches based on explicit definitions (e.g., Gram matrix) or implicit learning (e.g., GANs) are neither interpretable nor easy to control, resulting in entangled representations and less satisfying results. In this paper, we propose a new C-S disentangled framework for style transfer without using previous assumptions. The key insight is to explicitly extract the content information and implicitly learn the complementary style information, yielding interpretable and controllable C-S disentanglement and style transfer. A simple yet effective CLIP-based style disentanglement loss coordinated with a style reconstruction prior is introduced to disentangle C-S in the CLIP image space. By further leveraging the powerful style removal and generative ability of diffusion models, our framework achieves superior results than state of the art and flexible C-S disentanglement and trade-off control. Our work provides new insights into the C-S disentanglement in style transfer and demonstrates the potential of diffusion models for learning well-disentangled C-S characteristics. Zhizhong Wang, Lei Zhao 0011, Wei Xing 0001 |
ICCV | 1 |
| 2023 | Towards Interactive Facial Image Inpainting by Text or Exemplar Image
Ailin Li, Lei Zhao 0011, Zhiwen Zuo, Zhizhong Wang, Wei Xing 0001, Dongming Lu |
MMM (1) | 4 |
| 2023 | MIGT: Multi-modal image inpainting guided with text
Ailin Li, Lei Zhao 0011, Zhiwen Zuo, Zhizhong Wang, Wei Xing 0001, Dongming Lu |
Neurocomputing | 4 |
| 2022 | Texture Reformer: Towards Fast and Universal Interactive Texture TransferabstractIn this paper, we present the texture reformer, a fast and universal neural-based framework for interactive texture transfer with user-specified guidance. The challenges lie in three aspects: 1) the diversity of tasks, 2) the simplicity of guidance maps, and 3) the execution efficiency. To address these challenges, our key idea is to use a novel feed-forward multi-view and multi-stage synthesis procedure consisting of I) a global view structure alignment stage, II) a local view texture refinement stage, and III) a holistic effect enhancement stage to synthesize high-quality results with coherent structures and fine texture details in a coarse-to-fine fashion. In addition, we also introduce a novel learning-free view-specific texture reformation (VSTR) operation with a new semantic map guidance strategy to achieve more accurate semantic-guided and structure-preserved texture transfer. The experimental results on a variety of application scenarios demonstrate the effectiveness and superiority of our framework. And compared with the state-of-the-art interactive texture transfer algorithms, it not only achieves higher quality results but, more remarkably, also is 2-5 orders of magnitude faster. Zhizhong Wang, Lei Zhao 0011, Haibo Chen 0006, Ailin Li, Zhiwen Zuo, Wei Xing 0001, Dongming Lu |
AAAI | 1 |
| 2022 | Universal Video Style Transfer via Crystallization, Separation, and BlendingabstractUniversal video style transfer aims to migrate arbitrary styles to input videos. However, how to maintain the temporal consistency of videos while achieving high-quality arbitrary style transfer is still a hard nut to crack. To resolve this dilemma, in this paper, we propose the CSBNet which involves three key modules: 1) the Crystallization (Cr) Module that generates several orthogonal crystal nuclei, representing hierarchical stability-aware content and style components, from raw VGG features; 2) the Separation (Sp) Module that separates these crystal nuclei to generate the stability-enhanced content and style features; 3) the Blending (Bd) Module to cross-blend these stability-enhanced content and style features, producing more stable and higher-quality stylized videos. Moreover, we also introduce a new pair of component enhancement losses to improve network performance. Extensive qualitative and quantitative experiments are conducted to demonstrate the effectiveness and superiority of our CSBNet. Compared with the state-of-the-art models, it not only produces temporally more consistent and stable results for arbitrary videos but also achieves higher-quality stylizations for arbitrary images. Haofei Lu, Zhizhong Wang |
IJCAI | 2 |
| 2022 | DivSwapper: Towards Diversified Patch-based Arbitrary Style TransferabstractGram-based and patch-based approaches are two important research lines of style transfer. Recent diversified Gram-based methods have been able to produce multiple and diverse stylized outputs for the same content and style images. However, as another widespread research interest, the diversity of patch-based methods remains challenging due to the stereotyped style swapping process based on nearest patch matching. To resolve this dilemma, in this paper, we dive into the crux of existing patch-based methods and propose a universal and efficient module, termed DivSwapper, for diversified patch-based arbitrary style transfer. The key insight is to use an essential intuition that neural patches with higher activation values could contribute more to diversity. Our DivSwapper is plug-and-play and can be easily integrated into existing patch-based and Gram-based methods to generate diverse results for arbitrary styles. We conduct theoretical analyses and extensive experiments to demonstrate the effectiveness of our method, and compared with state-of-the-art algorithms, it shows superiority in diversity, quality, and efficiency. Zhizhong Wang, Lei Zhao 0011, Haibo Chen 0006, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
IJCAI | 1 |
| 2022 | Style Fader Generative Adversarial Networks for Style Degree Controllable Artistic Style TransferabstractArtistic style transfer is the task of synthesizing content images with learned artistic styles. Recent studies have shown the potential of Generative Adversarial Networks (GANs) for producing artistically rich stylizations. Despite the promising results, they usually fail to control the generated images' style degree, which is inflexible and limits their applicability for practical use. To address the issue, in this paper, we propose a novel method that for the first time allows adjusting the style degree for existing GAN-based artistic style transfer frameworks in real time after training. Our method introduces two novel modules into existing GAN-based artistic style transfer frameworks: a Style Scaling Injection (SSI) module and a Style Degree Interpretation (SDI) module. The SSI module accepts the value of Style Degree Factor (SDF) as the input and outputs parameters that scale the feature activations in existing models, offering control signals to alter the style degrees of the stylizations. And the SDI module interprets the output probabilities of a multi-scale content-style binary classifier as the style degrees, providing a mechanism to parameterize the style degree of the stylizations. Moreover, we show that after training our method can enable existing GAN-based frameworks to produce over-stylizations. The proposed method can facilitate many existing GAN-based artistic style transfer frameworks with marginal extra training overheads and modifications. Extensive qualitative evaluations on two typical GAN-based style transfer models demonstrate the effectiveness of the proposed method for gaining style degree control for them. Zhiwen Zuo, Lei Zhao 0011, Shuobin Lian, Haibo Chen 0006, Zhizhong Wang, Ailin Li, Wei Xing 0001, Dongming Lu |
IJCAI | 5 |
| 2022 | AesUST: Towards Aesthetic-Enhanced Universal Style TransferabstractRecent studies have shown remarkable success in universal style transfer which transfers arbitrary visual styles to content images. However, existing approaches suffer from the aesthetic-unrealistic problem that introduces disharmonious patterns and evident artifacts, making the results easy to spot from real paintings. To address this limitation, we propose AesUST, a novel Aesthetic-enhanced Universal Style Transfer approach that can generate aesthetically more realistic and pleasing results for arbitrary styles. Specifically, our approach introduces an aesthetic discriminator to learn the universal human-delightful aesthetic features from a large corpus of artist-created paintings. Then, the aesthetic features are incorporated to enhance the style transfer process via a novel Aesthetic-aware Style-Attention (AesSA) module. Such an AesSA module enables our AesUST to efficiently and flexibly integrate the style patterns according to the global aesthetic channel distribution of the style image and the local semantic spatial distribution of the content image. Moreover, we also develop a new two-stage transfer training strategy with two aesthetic regularizations to train our model more effectively, further improving stylization performance. Extensive experiments and user studies demonstrate that our approach synthesizes aesthetically more harmonious and realistic results than state of the art, greatly narrowing the disparity with real artist-created paintings. Our code is available at https://github.com/EndyWon/AesUST. Zhizhong Wang, Zhanjie Zhang, Lei Zhao 0011, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
ACM Multimedia | 1 |
| 2022 | Automated localization and severity period prediction of myocardial infarction with clinical interpretability based on deep learning and knowledge graph
Chuang Han, Shihao Pan, Wenge Que, Zhizhong Wang, Yunkai Zhai |
Expert Syst. Appl. | 4 |
| 2022 | Dual distribution matching GAN
Zhiwen Zuo, Lei Zhao 0011, Ailin Li, Zhizhong Wang, Haibo Chen 0006, Wi Xing, Dongming Lu |
Neurocomputing | 4 |
| 2021 | DualAST: Dual Style-Learning Networks for Artistic Style TransferabstractArtistic style transfer is an image editing task that aims at repainting everyday photographs with learned artistic styles. Existing methods learn styles from either a single style example or a collection of artworks. Accordingly, the stylization results are either inferior in visual quality or limited in style controllability. To tackle this problem, we propose a novel Dual Style-Learning Artistic Style Transfer (DualAST) framework to learn simultaneously both the holistic artist-style (from a collection of artworks) and the specific artwork-style (from a single style image): the artist-style sets the tone (i.e., the overall feeling) for the stylized image, while the artwork-style determines the details of the stylized image, such as color and texture. Moreover, we introduce a Style-Control Block (SCB) to adjust the styles of generated images with a set of learnable style-control factors. We conduct extensive experiments to evaluate the performance of the proposed framework, the results of which confirm the superiority of our method. Haibo Chen 0006, Lei Zhao 0011, Zhizhong Wang, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
CVPR | 3 |
| 2021 | Diverse Image Style Transfer via Invertible Cross-Space MappingabstractImage style transfer aims to transfer the styles of artworks onto arbitrary photographs to create novel artistic images. Although style transfer is inherently an underdetermined problem, existing approaches usually assume a deterministic solution, thus failing to capture the full distribution of possible outputs. To address this limitation, we propose a Diverse Image Style Transfer (DIST) framework which achieves significant diversity by enforcing an invertible cross-space mapping. Specifically, the framework consists of three branches: disentanglement branch, inverse branch, and stylization branch. Among them, the disentanglement branch factorizes artworks into content space and style space; the inverse branch encourages the invertible mapping between the latent space of input noise vectors and the style space of generated artistic images; the stylization branch renders the input content image with the style of an artist. Armed with these three branches, our approach is able to synthesize significantly diverse stylized images without loss of quality. We conduct extensive experiments and comparisons to evaluate our approach qualitatively and quantitatively. The experimental results demonstrate the effectiveness of our method. Haibo Chen 0006, Lei Zhao 0011, Zhizhong Wang, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
ICCV | 4 |
| 2021 | Artistic Style Transfer with Internal-external Learning and Contrastive LearningabstractAlthough existing artistic style transfer methods have achieved significant improvement with deep neural networks, they still suffer from artifacts such as disharmonious colors and repetitive patterns. Motivated by this, we propose an internal-external style transfer method with two contrastive losses. Specifically, we utilize internal statistics of a single style image to determine the colors and texture patterns of the stylized image, and in the meantime, we leverage the external information of the large-scale style dataset to learn the human-aware style information, which makes the color distributions and texture patterns in the stylized image more reasonable and harmonious. In addition, we argue that existing style transfer methods only consider the content-to-stylization and style-to-stylization relations, neglecting the stylization-to-stylization relations. To address this issue, we introduce two contrastive losses, which pull the multiple stylization embeddings closer to each other when they share the same content or style, but push far away otherwise. We conduct extensive experiments, showing that our proposed method can not only produce visually more harmonious and satisfying artistic images, but also promote the stability and consistency of rendered video clips. Haibo Chen 0006, Lei Zhao 0011, Zhizhong Wang, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
NeurIPS | 3 |
| 2021 | Diversified text-to-image generation via deep mutual information estimation
Ailin Li, Lei Zhao 0011, Zhiwen Zuo, Zhizhong Wang, Haibo Chen 0006, Dongming Lu, Wei Xing 0001 |
Comput. Vis. Image Underst. | 4 |
| 2021 | Evaluate and improve the quality of neural style transfer
Zhizhong Wang, Lei Zhao 0011, Haibo Chen 0006, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
Comput. Vis. Image Underst. | 1 |
| 2020 | Diversified Arbitrary Style Transfer via Deep Feature PerturbationabstractImage style transfer is an underdetermined problem, where a large number of solutions can satisfy the same constraint (the content and style). Although there have been some efforts to improve the diversity of style transfer by introducing an alternative diversity loss, they have restricted generalization, limited diversity and poor scalability. In this paper, we tackle these limitations and propose a simple yet effective method for diversified arbitrary style transfer. The key idea of our method is an operation called deep feature perturbation (DFP), which uses an orthogonal random noise matrix to perturb the deep image feature maps while keeping the original style information unchanged. Our DFP operation can be easily integrated into many existing WCT (whitening and coloring transform)-based methods, and empower them to generate diverse results for arbitrary styles. Experimental results demonstrate that this learning-free and universal method can greatly increase the diversity while maintaining the quality of stylization. Zhizhong Wang, Lei Zhao 0011, Haibo Chen 0006, Lihong Qiu, Qihang Mo, Sihuan Lin, Wei Xing 0001, Dongming Lu |
CVPR | 1 |
| 2020 | UCTGAN: Diverse Image Inpainting Based on Unsupervised Cross-Space TranslationabstractAlthough existing image inpainting approaches have been able to produce visually realistic and semantically correct results, they produce only one result for each masked input. In order to produce multiple and diverse reasonable solutions, we present Unsupervised Cross-space Translation Generative Adversarial Network (called UCTGAN) which mainly consists of three network modules: conditional encoder module, manifold projection module and generation module. The manifold projection module and the generation module are combined to learn one-to-one image mapping between two spaces in an unsupervised way by projecting instance image space and conditional completion image space into common low-dimensional manifold space, which can greatly improve the diversity of the repaired samples. For understanding of global information, we also introduce a new cross semantic attention layer that exploits the long-range dependencies between the known parts and the completed parts, which can improve realism and appearance consistency of repaired samples. Extensive experiments on various datasets such as CelebA-HQ, Places2, Paris Street View and ImageNet clearly demonstrate that our method not only generates diverse inpainting solutions from the same image to be repaired, but also has high image quality. Lei Zhao 0011, Qihang Mo, Sihuan Lin, Zhizhong Wang, Zhiwen Zuo, Haibo Chen 0006, Wei Xing 0001, Dongming Lu |
CVPR | 4 |
| 2020 | Creative and diverse artwork generation using adversarial networksabstractExisting style transfer methods have achieved great success in artwork generation by transferring artistic styles onto everyday photographs while keeping their contents unchanged. Despite this success, these methods have one inherent limitation: they cannot produce newly created image contents, lacking creativity and flexibility. On the other hand, generative adversarial networks (GANs) can synthesise images with new content, whereas cannot specify the artistic style of these images. The authors consider combining style transfer with convolutional GANs to generate more creative and diverse artworks. Instead of simply concatenating these two networks: the first for synthesising new content and the second for transferring artistic styles, which is inefficient and inconvenient, they design an end‐to‐end network called ArtistGAN to perform these two operations at the same time and achieve visually better results. Moreover, to generate images of higher quality, they propose the bi‐discriminator GAN containing a pixel discriminator and a feature discriminator that constrain the generated image from pixel level and feature level, respectively. They conduct extensive experiments and comparisons to evaluate their methods quantitatively and qualitatively. The experimental results verify the effectiveness of their methods. Haibo Chen 0006, Lei Zhao 0011, Lihong Qiu, Zhizhong Wang, Wei Xing 0001, Dongming Lu |
IET Comput. Vis. | 4 |
| 2020 | GLStyleNet: exquisite style transfer combining global and local pyramid featuresabstractRecent studies using deep neural networks have shown remarkable success in style transfer, especially for artistic and photo‐realistic images. However, these methods cannot solve more sophisticated problems. The approaches using global statistics fail to capture small, intricate textures and maintain correct texture scales of the artworks, and the others based on local patches are defective on global effect. To address these issues, this study presents a unified model [global and local style network (GLStyleNet)] to achieve exquisite style transfer with higher quality. Specifically, a simple yet effective perceptual loss is proposed to consider the information of global semantic‐level structure, local patch‐level style, and global channel‐level effect at the same time. This could help transfer not just large‐scale, obvious style cues but also subtle, exquisite ones, and dramatically improve the quality of style transfer. Besides, the authors introduce a novel deep pyramid feature fusion module to provide a more flexible style expression and a more efficient transfer process. This could help retain both high‐frequency pixel information and low‐frequency construct information. They demonstrate the effectiveness and superiority of their approach on numerous style transfer tasks, especially the Chinese ancient painting style transfer. Experimental results indicate that their unified approach improves image style transfer quality over previous state‐of‐the‐art methods. Zhizhong Wang, Lei Zhao 0011, Sihuan Lin, Qihang Mo, Wei Xing 0001, Dongming Lu |
IET Comput. Vis. | 1 |
| 2020 | Small target detection based on bird's visual information processing mechanism
Zhizhong Wang, Donghaisheng Liu, Yuehui Lei, Xiaoke Niu, Songwei Wang |
Multim. Tools Appl. | 1 |
| 2016 | Prediction of categorical spatial data via Bayesian updatingabstractThis study introduces a transition probability-based Bayesian updating (BU) approach for spatial classification through expert system. Transition probabilities are interpreted as expert opinions for updating the prior marginal probabilities of categorical response variables. The main objective of this paper is to provide a spatial categorical variable prediction method which has a solid theoretical foundation and yields relatively higher classification accuracy compared with conventional ones. The basic idea is to first build a linear Bayesian updating (LBU) model that corresponds to an application of Bayes’ theorem. Since the linear opinion pool is intrinsically suboptimal and underconfident, the beta-transformed Bayesian updating (BBU) model is proposed to overcome this limitation. Another type of BU approach, conditional independent Bayesian updating (CIBU), is derived based on conditional independent experts. It is shown that traditional Markovian-type categorical prediction (MCP) is equivalent to a particular CIBU model with specific parameters. As three variants of the BU method, these techniques are illustrated in synthetic and real-world case studies, comparison results with both the LBU and MCP favor the BBU model. Xiang Huang 0002, Zhizhong Wang |
Int. J. Geogr. Inf. Sci. | 2 |
| 2010 | GMRVVm-SVR model for financial time series forecasting
Hui Jiang 0003, Zhizhong Wang |
Expert Syst. Appl. | 2 |
| 2010 | Selecting discriminant eigenfaces by using binary feature selection
Wangxin Yu, Zhizhong Wang, Weiting Chen |
Neural Comput. Appl. | 2 |
| 2009 | A new framework to combine vertical and horizontal information for face recognition
Wangxin Yu, Zhizhong Wang, Weiting Chen |
Neurocomputing | 2 |
| 2008 | Direct simplification for kernel regression machines
Wenwu He, Zhizhong Wang |
Neurocomputing | 2 |
| 2008 | Model optimizing and feature selecting for support vector regression in time series forecasting
Wenwu He, Zhizhong Wang, Hui Jiang 0003 |
Neurocomputing | 2 |
| 2005 | Multiple Feature Domains Information Fusion for Computer-Aided Clinical Electromyography
Hongbo Xie, Zhizhong Wang |
CAIP | 3 |
| 1999 | Fixed-point error analysis and an efficient array processor design of two-dimensional sliding DFT
Yisheng Zhu, Zhizhong Wang |
Signal Process. | 4 |
| 1996 | An unbiased parametric imaging algorithm for nonuniformly sampled biomedical system parameter estimationabstractAn unbiased algorithm of generalized linear least squares (GLLS) for parameter estimation of nonuniformly sampled biomedical systems is proposed. The basic theory and detailed derivation of the algorithm are given. This algorithm removes the initial values required and computational burden of nonlinear least regression and achieves a comparable estimation quality in terms of the estimates' bias and standard deviation. Therefore, this algorithm is particular useful in image-wide (pixel-by-pixel based) parameter estimation, e.g., to generate parametric images from tracer dynamic studies with positron emission tomography. An example is presented to demonstrate the performance of this new technique. This algorithm is also generally applicable to other continuous system parameter estimation. David Dagan Feng, Sung-Cheng Huang, Zhizhong Wang, Dino Ho |
IEEE Trans. Medical Imaging | 3 |
| 1993 | A study on statistically reliable and computationally efficient algorithms for generating local cerebral blood flow parametric images with positron emission tomographyabstractWith the advent of positron emission tomography (PET), a variety of techniques have been developed to measure local cerebral blood flow (LCBF) noninvasively in humans. A potential class of techniques, which includes linear least squares (LS), linear weighted least squares (WLS), linear generalized least squares (GLS), and linear generalized weighted least squares (GWLS), is proposed. The statistical characteristics of these methods are examined by computer simulation. The authors present a comparison of these four methods with two other rapid estimation techniques developed by Huang et al. (1982) and Alpert (1984), and two classical methods, the unweighted and weighted nonlinear least squares regression. The results show that these methods can take full advantage of the contribution from the fine temporal sampling data of modern tomographs, and thus provide statistically reliable estimates that are comparable to those obtained from nonlinear LS regression. These methods also have high computational efficiency, and the parameters can be estimated directly from operational equations in one single step. Therefore, they can potentially be used in image-wide estimation of local cerebral blood flow and distribution volume with PET. David Dagan Feng, Zhizhong Wang, Sung-Cheng Huang |
IEEE Trans. Medical Imaging | 2 |