Jiafu Chen

dblp:344/1013 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Pansharpening Based on Multiresolution Panchromatic Feature Guidance
abstract
Pansharpening, as a critical image fusion technique, aims to synthesize high-resolution multispectral (HRMS) images by integrating panchromatic (PAN) images with low-resolution multispectral (LRMS) data, thereby enhancing spatial details while maintaining spectral fidelity. Current methodologies exhibit limitations in comprehensively modeling inter-feature dependencies, often leading to suboptimal fusion performance characterized by spectral distortion or insufficient spatial enhancement. To address these challenges, we propose a novel multi-resolution PAN-guided pansharpening framework that systematically coordinates feature extraction, texture refinement, and spectral-spatial reconstruction. The proposed methodology is initiated with the development of a multi-resolution feature extraction network, founded on the spatial frequency Transformer module, which is designed to comprehensively extract spatial feature information at varying scales of both PAN and multispectral (MS) images. Subsequently, to effectively constrain and optimise the extracted spatial feature information, an interactive self-attentive texture injection module is employed to further extract detailed texture features that are more compatible with the spectral information of the MS image. These texture features are then injected into the MS reconstruction network in a step-by-step manner to guide the sharpening process of the MS image, while progressively up-sampling the size of the MS image to the PAN image to achieve more refined spatial features. Finally, the integration of multi-resolution feature information connected by residuals in the MS reconstruction network is achieved by means of an adaptive channel self-attention module. This module models the dependencies between feature maps in the channel dimension, and ultimately generates HRMS images. Experimental validation results demonstrate that the proposed method achieves excellent qualitative and quantitative evaluation results. Notably, it effectively reduces spatial blurring and edge distortion, significantly enhancing the visual effect and detail representation of the image.
Jian Wang 0090, Jiafu Chen, Lihui Zhou
IEEE Trans. Geosci. Remote. Sens.3
2024 PNeSM: Arbitrary 3D Scene Stylization via Prompt-Based Neural Style Mapping
abstract
3D scene stylization refers to transform the appearance of a 3D scene to match a given style image, ensuring that images rendered from different viewpoints exhibit the same style as the given style image, while maintaining the 3D consistency of the stylized scene. Several existing methods have obtained impressive results in stylizing 3D scenes. However, the mod- els proposed by these methods need to be re-trained when applied to a new scene. In other words, their models are cou- pled with a specific scene and cannot adapt to arbitrary other scenes. To address this issue, we propose a novel 3D scene stylization framework to transfer an arbitrary style to an ar- bitrary scene, without any style-related or scene-related re- training. Concretely, we first map the appearance of the 3D scene into a 2D style pattern space, which realizes complete disentanglement of the geometry and appearance of the 3D scene and makes our model be generalized to arbitrary 3D scenes. Then we stylize the appearance of the 3D scene in the 2D style pattern space via a prompt-based 2D stylization al- gorithm. Experimental results demonstrate that our proposed framework is superior to SOTA methods in both visual qual- ity and generalization.
Jiafu Chen, Wei Xing 0001, Jiakai Sun, Tianyi Chu, Boyan Ji, Lei Zhao 0011, Huaizhong Lin, Haibo Chen 0006, Zhizhong Wang
AAAI1
2024 Attack Deterministic Conditional Image Generative Models for Diverse and Controllable Generation
abstract
Existing generative adversarial network (GAN) based conditional image generative models typically produce fixed output for the same conditional input, which is unreasonable for highly subjective tasks, such as large-mask image inpainting or style transfer. On the other hand, GAN-based diverse image generative methods require retraining/fine-tuning the network or designing complex noise injection functions, which is computationally expensive, task-specific, or struggle to generate high-quality results. Given that many deterministic conditional image generative models have been able to produce high-quality yet fixed results, we raise an intriguing question: is it possible for pre-trained deterministic conditional image generative models to generate diverse results without changing network structures or parameters? To answer this question, we re-examine the conditional image generation tasks from the perspective of adversarial attack and propose a simple and efficient plug-in projected gradient descent (PGD) like method for diverse and controllable image generation. The key idea is attacking the pre-trained deterministic generative models by adding a micro perturbation to the input condition. In this way, diverse results can be generated without any adjustment of network structures or fine-tuning of the pre-trained models. In addition, we can also control the diverse results to be generated by specifying the attack direction according to a reference text or image. Our work opens the door to applying adversarial attack to low-level vision tasks, and experiments on various conditional image generation tasks demonstrate the effectiveness and superiority of the proposed method.
Tianyi Chu, Wei Xing 0001, Jiafu Chen, Zhizhong Wang, Jiakai Sun, Lei Zhao 0011, Haibo Chen 0006, Huaizhong Lin
AAAI3
2024 Single-Mask Inpainting for Voxel-Based Neural Radiance Fields
Jiafu Chen, Tianyi Chu, Jiakai Sun, Wei Xing 0001, Lei Zhao 0011
ECCV (57)1
2024 DPH-YOLOv8: Improved YOLOv8 Based on Double Prediction Heads for the UAV Image Object Detection
abstract
Object detection on unmanned aerial vehicle (UAV) images has been a hot research topic recently. However, object detection models for general scenarios struggle with UAV images due to the challenges of detecting small targets and handling complex image backgrounds. To solve the two issues, we proposed DPH-YOLOv8, an enhanced version of YOLOv8 tailored for UAV scenarios. First, we improved the prediction head from three prediction heads to double prediction heads (DPHs), reducing the model’s parameters by 42.9% while improving the mAP for small targets. Second, we designed a tiny-path feature fusion (TP-Fusion) module to fuse richer detailed information, enabling the detector to accurately match targets of different sizes and shapes. Third, we introduced a coordinate attention (CA) module to help the model reduce the interference of complex background information and focus on detecting foreground targets. Finally, bottleneck modules were added before the prediction heads to enhance the extraction of small target features. Extensive experimental results on both the VisDrone2021 and UAVDT benchmarks demonstrated that DPH-YOLOv8 not only improved the mAP by 4.5% on VisDrone2021 and 2.2% on UAVDT but also reduced localization error by 0.53%, confusion with objects by 0.69%, and confusion with background by 0.45%. These enhancements, along with a reduction in the model’s parameters, make DPH-YOLOv8 more suitable for UAV scenarios.
Jian Wang 0090, Jiafu Chen, Lihui Zhou, Linyang Guo
IEEE Trans. Geosci. Remote. Sens.3
2024 Adaptive Receptive Field Enhancement Network Based on Attention Mechanism for Detecting the Small Target in the Aerial Image
abstract
To address the problem of insufficient semantic feature information caused by small objects in aerial images, an adaptive receptive field enhancement network based on attention mechanism suitable for aerial scene is proposed. First, to make the model more suitable for deployment on unmanned aerial vehicle (UAV) platforms with limited resources, the receptive field block (RFB) is reconstructed by a series branch with feature multiplexing, which can make the output feature map of the front branch pass through the convolutional layer of the back branch for further feature extraction, improving the utilization of convolutional resources. Second, to expand the receptive field with as little loss of local contextual information as possible, Kronecker convolution (K-conv) is employed in the RFB branch for feature extraction, which can augment the original image covered by a single pixel on the feature map, expand its global semantic information and local contextual information, and improve the integrity of object extraction for small targets. Finally, to solve the problem of a fixed receptive field size of neurons at each layer in the network caused by direct aggregation feature maps, the selective convolutional module based on attention mechanism is added to the network in this article, so that neurons at each layer in the network can adaptively adjust the receptive field size, so as to make the different sizes of the detector head to play the best detection ability. In this article, experiments are carried out on a homemade aerial photography small target dataset to verify the effectiveness of the adaptive field enhancement network. The experimental results show that the adaptive field enhancement network proposed in this article can effectively improve the detection accuracy of the algorithm on targets with low resolution and insufficient semantic feature information in aerial images and is more suitable for use in aerial scene.
Jian Wang 0090, Lihui Zhou, Jiafu Chen, Linyang Guo
IEEE Trans. Geosci. Remote. Sens.4
2023 Generative Image Inpainting with Segmentation Confusion Adversarial Training and Contrastive Learning
abstract
This paper presents a new adversarial training framework for image inpainting with segmentation confusion adversarial training (SCAT) and contrastive learning. SCAT plays an adversarial game between an inpainting generator and a segmentation network, which provides pixel-level local training signals and can adapt to images with free-form holes. By combining SCAT with standard global adversarial training, the new adversarial training framework exhibits the following three advantages simultaneously: (1) the global consistency of the repaired image, (2) the local fine texture details of the repaired image, and (3) the flexibility of handling images with free-form holes. Moreover, we propose the textural and semantic contrastive learning losses to stabilize and improve our inpainting model's training by exploiting the feature representation space of the discriminator, in which the inpainting images are pulled closer to the ground truth images but pushed farther from the corrupted images. The proposed contrastive losses better guide the repaired images to move from the corrupted image data points to the real image data points in the feature representation space, resulting in more realistic completed images. We conduct extensive experiments on two benchmark datasets, demonstrating our model's effectiveness and superiority both qualitatively and quantitatively.
Zhiwen Zuo, Lei Zhao 0011, Ailin Li, Zhizhong Wang, Zhanjie Zhang, Jiafu Chen, Wei Xing 0001, Dongming Lu
AAAI6
2023 Rethinking Fast Fourier Convolution in Image Inpainting
abstract
Recently proposed LaMa [25] introduce Fast Fourier Convolution (FFC) [4] into image inpainting. FFC empowers the fully convolutional network to have a global receptive field in its early layers, and have the ability to produce robust repeating texture. However, LaMa has difficulty in generating clear and sharp complex content. In this paper, we analyze the fundamental flaws of using FFC in image inpainting, which are 1) spectrum shifting, 2) unexpected spatial activation, and 3) limited frequency receptive field. Such flaws make FFC-based inpainting framework difficult in generating complicated texture and performing faithful reconstruction. Based on the above analysis, we propose a novel Unbiased Fast Fourier Convolution (UFFC) module. UFFC is constructed by modifying the vanilla FFC module with 1) range transform and inverse transform, 2) absolute position embedding, 3) dynamic skip connection, and 4) adaptive clip, to overcome the above flaws. UFFC captures frequency information efficiently and realize reconstruction without introducing additional artifacts, achieving better inpainting results and more efficient training. In addition, we propose two novel perceptual losses for better generation quality and more robust training. Extensive experiments on several benchmark datasets demonstrate the effectiveness of our method, outperforming the state-of-the-art methods in both texture-capturing ability and expressiveness.
Tianyi Chu, Jiafu Chen, Jiakai Sun, Shuobin Lian, Zhizhong Wang, Zhiwen Zuo, Lei Zhao 0011, Wei Xing 0001, Dongming Lu
ICCV2
2023 Rethinking Multi-Contrast MRI Super-Resolution: Rectangle-Window Cross-Attention Transformer and Arbitrary-Scale Upsampling
abstract
Recently, several methods have explored the potential of multi-contrast magnetic resonance imaging (MRI) super-resolution (SR) and obtain results superior to single-contrast SR methods. However, existing approaches still have two shortcomings: (1) They can only address fixed integer upsampling scales, such as 2×, 3×, and 4×, which require training and storing the corresponding model separately for each upsampling scale in clinic. (2) They lack direct interaction among different windows as they adopt the square window (e.g., 8×8) transformer network architecture, which results in inadequate modelling of longer-range dependencies. Moreover, the relationship between reference images and target images is not fully mined. To address these issues, we develop a novel network for multi-contrast MRI arbitrary-scale SR, dubbed as McASSR. Specifically, we design a rectangle-window cross-attention transformer to establish longer-range dependencies in MR images without increasing computational complexity and fully use reference information. Besides, we propose the reference-aware implicit attention as an upsampling module, achieving arbitrary-scale super-resolution via implicit neural representation, further fusing supplementary information of the reference image. Extensive and comprehensive experiments on both public and clinical datasets show that our McASSR yields superior performance over SOTA methods, demonstrating its great potential to be applied in clinical practice. Code will be available at https://github.com/GuangYuanKK/McASSR.
Lei Zhao 0011, Jiakai Sun, Zehua Lan, Zhanjie Zhang, Jiafu Chen, Huaizhong Lin, Wei Xing 0001
ICCV6
2023 TeSTNeRF: Text-Driven 3D Style Transfer via Cross-Modal Learning
abstract
Text-driven 3D style transfer aims at stylizing a scene according to the text and generating arbitrary novel views with consistency. Simply combining image/video style transfer methods and novel view synthesis methods results in flickering when changing viewpoints, while existing 3D style transfer methods learn styles from images instead of texts. To address this problem, we for the first time design an efficient text-driven model for 3D style transfer, named TeSTNeRF, to stylize the scene using texts via cross-modal learning: we leverage an advanced text encoder to embed the texts in order to control 3D style transfer and align the input text and output stylized images in latent space. Furthermore, to obtain better visual results, we introduce style supervision, learning feature statistics from style images and utilizing 2D stylization results to rectify abrupt color spill. Extensive experiments demonstrate that TeSTNeRF significantly outperforms existing methods and provides a new way to guide 3D style transfer.
Jiafu Chen, Boyan Ji, Zhanjie Zhang, Tianyi Chu, Zhiwen Zuo, Lei Zhao 0011, Wei Xing 0001, Dongming Lu
IJCAI1
2023 VGOS: Voxel Grid Optimization for View Synthesis from Sparse Inputs
abstract
Neural Radiance Fields (NeRF) has shown great success in novel view synthesis due to its state-of-the-art quality and flexibility. However, NeRF requires dense input views (tens to hundreds) and a long training time (hours to days) for a single scene to generate high-fidelity images. Although using the voxel grids to represent the radiance field can significantly accelerate the optimization process, we observe that for sparse inputs, the voxel grids are more prone to overfitting to the training views and will have holes and floaters, which leads to artifacts. In this paper, we propose VGOS, an approach for fast (3-5 minutes) radiance field reconstruction from sparse inputs (3-10 views) to address these issues. To improve the performance of voxel-based radiance field in sparse input scenarios, we propose two methods: (a) We introduce an incremental voxel training strategy, which prevents overfitting by suppressing the optimization of peripheral voxels in the early stage of reconstruction. (b) We use several regularization techniques to smooth the voxels, which avoids degenerate solutions. Experiments demonstrate that VGOS achieves state-of-the-art performance for sparse inputs with super-fast convergence. Code will be available at https://github.com/SJoJoK/VGOS.
Jiakai Sun, Zhanjie Zhang, Jiafu Chen, Boyan Ji, Lei Zhao 0011, Wei Xing 0001
IJCAI3
2023 Caster: Cartoon style transfer via dynamic cartoon style casting
Zhanjie Zhang, Jiakai Sun, Jiafu Chen, Lei Zhao 0011, Boyan Ji, Zehua Lan, Wei Xing 0001, Duanqing Xu
Neurocomputing3