Yuwei Guo 0001

dblp:159/7655-1 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021
YearPublicationVenuePosition
2026 Causality-inspired learning semantic segmentation in unseen domain
Pei He, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Ronghua Shang, Yuwei Guo 0001, Puhua Chen, Shuyuan Yang 0001
Pattern Recognit.7
2026 Image Singularity Scattering Representation Learning Classification
abstract
The multi-scale geometric analysis is a great representation tool. It can be used to improve the feature representation and learning process of deep networks. In addition to extracting features, the multi-scale geometric prior knowledge can also be used for the structure improvement of deep networks. In this paper, we propose a multi-scale scattering representation learning network, abbreviated as MSRLN, for image classification tasks. The exploration of structure improvement can be made with multi-scale scattering operations. In this way, the better singularity representation learning process for networks can be achieved. Firstly, the filter banks and multi-scale scattering operator are introduced for non-linear and singularity representation. Secondly, the novel multi-scale scattering representation learning network structure is designed. The scaling- wise scattering process is deployed in the shallow layer as a non-linear layer. This structure essentially supplements deep networks with geometric prior knowledge. It can further improve the non-linear activation and singularity representation process. Thirdly, we put forward the multi-stage scattering representation strategy and the prior knowledge weakening mechanism. With flexible scaling factors and learning rates, the stepwise approximation and learning process of networks can be achieved. In sum, MSRLN is a kind of structural innovative, and the scattering singularity representation structure can be extended to other backbones or tasks. Extensive experimental results show that MSRLN can achieve better image classification accuracy. Finally, necessary convergence, insight, and adaptability analyses are provided in evaluation experiments.
Jie Gao 0013, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Puhua Chen, Yuwei Guo 0001, Fang Liu 0001, Shuyuan Yang 0001
IEEE Trans. Multim.6
2026 Edit-Your-Motion: Space-Time Diffusion Decoupling Learning for Video Motion Editing
abstract
Existing diffusion-based methods have achieved impressive results in human motion editing. However, these methods often exhibit significant ghosting and body distortion in unseen in-the-wild cases. In this paper, we introduce Edit-Your-Motion, a video motion editing method that tackles these challenges through one-shot fine-tuning on unseen cases. Specifically, firstly, we utilized DDIM inversion to initialize the noise, preserving the appearance of the source video and designed a lightweight motion attention adapter module to enhance motion fidelity. DDIM inversion aims to obtain the implicit representations by estimating the prediction noise from the source video, which serves as a starting point for the sampling process, ensuring the appearance consistency between the source and edited videos. The Motion Attention Module (MA) enhances the model's motion editing ability by resolving the conflict between the skeleton features and the appearance features. Secondly, to effectively decouple motion and appearance of source video, we design a spatio-temporal two-stage learning strategy (STL). In the first stage, we focus on learning temporal features of human motion and propose recurrent causal attention (RCA) to ensure consistency between video frames. In the second stage, we shift focus on learning the appearance features of the source video. With Edit-Your-Motion, users can edit the motion of humans in the source video, creating more engaging and diverse content. Extensive qualitative and quantitative experiments, along with user preference studies, show that Edit-Your-Motion outperforms other methods.
Yi Zuo 0003, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Wenping Ma 0001, Shuyuan Yang 0001, Yuwei Guo 0001
IEEE Trans. Multim.8
2025 Multidistribution Time-Series Prototype Learning for Crop Mapping With Sentinel-1 SAR Imagery
abstract
Time-series Synthetic Aperture Radar (SAR) offers significant potential for crop mapping due to its all-weather, weather-independent imaging capabilities. Existing crop mapping methods with time-series SAR data have achieved good performance. However, these methods often ignore the phenological diversity of the same crop in time-series data, and their limited performance significantly constrains their potential for application in large-scale crop mapping. To address these issues, this paper proposes a prototype-based time-series remote sensing crop mapping framework called Multi-Distribution time-series Prototype Learning (MDPL). The framework aims to learn multiple time-series prototypes for the same crop for various phenological distributions of crops, effectively capturing the complex and varied phenological characteristics of crops. Secondly, a phenological-invariant feature learning module is proposed to enhance the model’s generalization capability for large-scale crop mapping. Additionally, a new temporal metric is proposed to capture phenological differences. Experimental results on three benchmark datasets have demonstrated the effectiveness and superiority of MDPL compared with state-of-the-art time-series SAR crop mapping methods.
Yuwei Guo 0001, Licheng Jiao, Kairen Chen, Yujing Jia, Fang Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 CDFNet: Cross-Domain Feature Fusion Network for PolSAR Terrain Classification
abstract
The scarcity of labeled data and domain shift among polarimetric synthetic aperture radar (PolSAR) images degrades the performance of the supervised-learning-based algorithm. Some unsupervised domain adaptation (UDA) algorithms have been proposed to address this problem and achieve good performance. The existing UDA algorithms for PolSAR terrain classification focus on the feature distribution shift problem but ignore the label shift problem in UDA task. In addition, feature alignment-based algorithms generate pseudo labels for target domain which introduce label noise and compromising the UDA performance. To alleviate the problems above, we present a cross-domain feature fusion network (CDFNet) for PolSAR terrain classification. Specifically, a domain-balanced sampling (DBS) module is proposed to obtain a nearly balanced training dataset to alleviate the label shift problem. Then, a cross-domain feature fusion (CDF) module is presented to achieve class-wise feature alignment with no additional label noise introduction. Experimental results on four PolSAR datasets demonstrate that our algorithm outperforms state-of-the-art UDA algorithms in terms of target domain performance.
Shuang Wang 0001, Zhuangzhuang Sun, Tianquan Bian, Yuwei Guo 0001, Linwei Dai, Yanhe Guo, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2025 A 3D Self-Awareness Diffusion Network for Multimodal Classification
abstract
As imaging sensor technology in remote sensing has advanced quickly, multimodal fusion classification has become an important research direction in land cover and urban planning classification tasks. While generative models and image classification have greatly benefited from diffusion models, the present ones primarily concentrate on single-modality-driven diffusion processes. Therefore, this paper presents a 3D self-awareness diffusion network (3DSA-DiffNet) for multispectral (MS) and panchromatic (PAN) image fusion classification, which would make it easier to classify heterogeneous data from various sensors. First, in order to model the relationship between multi-channel spectra and multi-pixel spatial distributions as well as samples, respectively, a spatial-spectral joint denoising network (S$^{2}$JD-Net) is proposed. It can incorporate the diffusion process into the neural network to enhance the quality of diffusion features. Secondly, to imitate the brain's spatial-spectral coexistence learning mechanism, this work offers a 3D self-awareness module (3DSA-Module) that can learn the weight of each pixel in 3D space, resulting in extraordinarily high feature representation capabilities. Finally, experimental verification demonstrates that the 3D self-awareness diffusion fusion network driven by brain inspiration outperforms more sophisticated approaches on the Xi'an, Huhhot, and Muufl datasets.
Mengru Ma, Wenping Ma 0001, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Yuwei Guo 0001
IEEE Trans. Multim.8
2025 Heterogeneous Riemannian Few-Shot Learning Network
abstract
How to learn and accurately distinguish new concepts from few samples, as humans do, is a long-standing concern in artificial intelligence (AI). Studies in brain science and neuroscience have shown that human brain perception is based on nonlinear manifolds, and high-dimensional manifolds can facilitate concept learning in neural circuits. Based on this inspiration, in this paper, we propose a heterogeneous Riemannian few-shot learning network (HRFL-Net), which is the first few-shot learning method to perform end-to-end deep learning on heterogeneous Riemannian manifolds. Specifically, to enhance the geometric invariance of the image representation, the image features are projected into three heterogeneous Riemannian manifold spaces. Then, the implicit Riemannian kernel function maps the manifolds to the separable high-dimensional reproducing Hilbert space. It is assumed that the embedded kernel features of the complementary manifolds are mapped to the same common subspace. Thus, a novel neural network-based Riemannian metric learning method is designed to solve the subspace feature vectors by imposing orthogonal normalized projection, which overcomes the data extension limitation of the Riemannian metric. Finally, with the optimization objective of increasing the interclass distance and decreasing the intraclass distance in Hilbert space, the HRFL-Net is trained with end-to-end stochastic optimization, and the optimal aggregation subspace is learned during the gradient descent process. Thus, the proposed HRFL-Net can be easily generalized to challenging nonconvex data. The evaluation of four public datasets shows that the proposed HRFL-Net has significant superiority and also achieves competitive results compared with the state-of-the-art methods.
Jie Chen 0098, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Yuwei Guo 0001, Puhua Chen, Wenping Ma 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Multiplane Prior Guided Few-Shot Aerial Scene Rendering
abstract
Neural Radiance Fields (NeRF) have been successfully applied in various aerial scenes, yet they face challenges with sparse views due to limited supervision. The acquisition of dense aerial views is often prohibitive, as unmanned aerial vehicles (UAVs) may encounter constraints in perspective range and energy constraints. In this work, we introduce Multiplane Prior guided NeRF (MPNeRF), a novel approach tailored for few-shot aerial scene rendering-marking a pioneering effort in this domain. Our key insight is that the intrinsic geometric regularities specific to aerial imagery could be leveraged to enhance NeRF in sparse aerial scenes. By investigating NeRF's and Multiplane Image (MPI)'s behavior, we propose to guide the training process of NeRF with a Multiplane Prior. The proposed Multiplane Prior draws upon MPI's benefits and incorporates advanced image comprehension through a Swin V2 Transformer, pre-trained via SimMIM. Our extensive experiments demonstrate that MPN-eRF outperforms existing state-of-the-art methods applied in non-aerial contexts, by tripling the performance in SSIM and LPIPS even with three views available. We hope our work offers insights into the development of NeRF-based applications in aerial scenes with limited data.
Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Puhua Chen, Yuwei Guo 0001
CVPR7
2024 A Knowledge-Based Hierarchical Causal Inference Network for Video Action Recognition
abstract
Currently, existing action recognition methods mainly use a data-driven method to extract spatio-temporal representations of actions for recognition. However, this method may face performance bottlenecks. At the same time, existing action recognition methods are easily affected by the bias of scene information and object information in videos. In order to explore the essential causal relationship between factors and remove bias in action recognition, we introduce the theory of causal inference into the field of action recognition and propose a Knowledge-based Hierarchical Causal Inference Network (KHCIN) to help us step toward a new direction of inference in action recognition. First, we construct a Knowledge-based Hierarchical Causal Graph (KHCG) to structurally represent the scene, object and motion knowledge of a video. Then, in the model inference stage, we perform factual causal inference on a video on the constructed KHCG, and then deploy counterfactual inference on the Direct Content Hierarchy (DCH) and Indirect Interaction Hierarchy (IIH) in the KHCG. For DCH, we intervene in the model at the decision level to highlight bias errors in the model predictions. For the IIH, we focus on intervening in the feature modelling process. The biased interactions are revealed by interrupting the information communication in the feature space. By comparing the results of factual and counterfactual inference, we can easily expose the biased information in the original representations and eliminate them. Driven by counterfactual causal inference, our approach can significantly improve the performance of action recognition while improving model explainability. Extensive experiments demonstrate the effectiveness of this method. We hope that KHCIN can provide some new ideas for better introduction of causal inference theory in the action recognition community in the future.
Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Lingling Li 0002, Yuwei Guo 0001, Puhua Chen
IEEE Trans. Multim.6
2023 MinEnt: Minimum entropy for self-supervised representation learning
Shuo Li 0010, Fang Liu 0001, Zehua Hao, Licheng Jiao, Xu Liu 0006, Yuwei Guo 0001
Pattern Recognit.6
2022 Polarimetric Multipath Convolutional Neural Network for PolSAR Image Classification
abstract
Scatter targets of complex land covers in polarimetric synthetic aperture radar (PolSAR) images are often randomly oriented and cause randomly fluctuating echoes, which brings a challenge to PolSAR image classification. Therefore, many existing methods have alleviated this problem through orientation compensation. However, there are still two obstacles that limit the improvement of classification accuracy. On the one hand, generally, these methods process PolSAR images with fixed polarization rotation angles, which is experience-dependent and inflexible. On the other hand, for the different land covers of a PolSAR image, the existing methods do not consider these rotation angles separately. For the first obstacle, we design a group of convolution kernels called polarization rotation kernels (PRKs) and utilize them to build the polarimetric convolutional neural network (CNN) (PolCNN). The PolCNN is the base network of our final model, and it can learn polarization rotation angles adaptively. For the second obstacle, we extend the PolCNN into a multipath structure, the final model polarimetric multipath CNN (PolMPCNN). The polarization rotation angles of different land covers are directly related to the networks of different paths within the PolMPCNN. Furthermore, we also put forward the two-scale sampling and the stagewise training algorithm in order that our PolMPCNN can fit different scales of PolSAR targets and pays more attention to difficult training samples. Experiments on real PolSAR images show that the proposed model achieves the best classification results with an extremely low sampling rate of 0.1%.
Yuanhao Cui, Fang Liu 0001, Licheng Jiao, Yuwei Guo 0001, Xuefeng Liang, Lingling Li 0002, Shuyuan Yang 0001, Xiaoxue Qian
IEEE Trans. Geosci. Remote. Sens.4
2022 Adaptive Fuzzy Learning Superpixel Representation for PolSAR Image Classification
abstract
The increasing applications of polarimetric synthetic aperture radar (PolSAR) image classification demand for effective superpixels’ algorithms. Fuzzy superpixels’ algorithms reduce the misclassification rate by dividing pixels into superpixels, which are groups of pixels of homogenous appearance and undetermined pixels. However, two key issues remain to be addressed in designing a fuzzy superpixel algorithm for PolSAR image classification. First, the polarimetric scattering information, which is unique in PolSAR images, is not effectively used. Such information can be utilized to generate superpixels more suitable for PolSAR images. Second, the ratio of undetermined pixels is fixed for each image in the existing techniques, ignoring the fact that the difficulty of classifying different objects varies in an image. To address these two issues, we propose a polarimetric scattering information-based adaptive fuzzy superpixel (AFS) algorithm for PolSAR images classification. In AFS, the correlation between pixels’ polarimetric scattering information, for the first time, is considered through fuzzy rough set theory to generate superpixels. This correlation is further used to dynamically and adaptively update the ratio of undetermined pixels. AFS is evaluated extensively against different evaluation metrics and compared with the state-of-the-art superpixels’ algorithms on three PolSAR images. The experimental results demonstrate the superiority of AFS on PolSAR image classification problems.
Yuwei Guo 0001, Licheng Jiao, Rong Qu, Zhuangzhuang Sun, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Feature Split-Merge-Enhancement Network for Remote Sensing Object Detection
abstract
Recently, multicategory object detection in high-resolution remote sensing images is still a challenge. First, objects with significant scale differences exist in one scene simultaneously, so it is generally difficult for the detectors to balance the detection performance of large and small objects. Second, because of the complex background and the objects’ densely distributed characteristics in the remote sensing images, the extracted features usually have noise and blurred boundaries, which interfere with the detection performance of the object detectors. With this observation, we propose an end-to-end scale-aware network called feature split–merge–enhancement network (SME-Net) for remote sensing object detection, composed of the feature split-and-merge (FSM) module, the offset-error rectification (OER) module, and the object saliency enhancement (OSE) strategy. FSM eliminates salient information of large objects to highlight the features of small objects in the shallow feature maps. It also transmits the effective detailed features of large objects to the deep feature maps, alleviating feature confusion between multiscale objects. OER corrects the inconsistency of the features spatial layout among the multilayer feature maps by the proposed offset loss, so as to achieve supervised elimination and transmission in FSM. OSE enhances the features of interests and suppresses the background information by the proposed membership function, thus preventing false detection and missed detection caused by noise and blurred boundaries. The effectiveness of the proposed algorithm has been verified on multiple datasets. Our code is available at:https://github.com/Momuli/SMENet.git
Wenping Ma 0001, Hao Zhu 0009, Licheng Jiao, Xu Tang 0004, Yuwei Guo 0001, Biao Hou
IEEE Trans. Geosci. Remote. Sens.6
2022 Cluster Alignment With Target Knowledge Mining for Unsupervised Domain Adaptation Semantic Segmentation
abstract
Unsupervised domain adaptation (UDA) carries out knowledge transfer from the labeled source domain to the unlabeled target domain. Existing feature alignment methods in UDA semantic segmentation achieve this goal by aligning the feature distribution between domains. However, these feature alignment methods ignore the domain-specific knowledge of the target domain. In consequence, 1) the correlation among pixels of the target domain is not explored; and 2) the classifier is not explicitly designed for the target domain distribution. To conquer these obstacles, we propose a novel cluster alignment framework, which mines the domain-specific knowledge when performing the alignment. Specifically, we design a multi-prototype clustering strategy to make the pixel features within the same class tightly distributed for the target domain. Subsequently, a contrastive strategy is developed to align the distributions between domains, with the clustered structure maintained. After that, a novel affinity-based normalized cut loss is devised to learn task-specific decision boundaries. Our method enhances the model's adaptability in the target domain, and can be used as a pre-adaptation for self-training to boost its performance. Sufficient experiments prove the effectiveness of our method against existing state-of-the-art methods on representative UDA benchmarks.
Shuang Wang 0001, Dong Zhao 0007, Yuwei Guo 0001, Qi Zang, Yu Gu 0015, Yi Li 0054, Licheng Jiao
IEEE Trans. Image Process.4
2021 Graph Regular Loss for Semi-Supervised Polsar Terrain Classification
abstract
Amongst the utilizations of Polarimetric Synthetic Aperture Radar (PoISAR) data, semi-supervised terrain classification is much in demand. Samples of the same category are distributed in multiple regions of a PolSAR image, resulting in differences in the feature distributions of samples of the same category located in different regions. In addition, some samples of confusable categories, and samples of different categories located near the edges, have small feature differences in a PolSAR image. To address this problem, we introduce a graph regular loss constructed from pseudo-labels to constrain the intra-class similarity and inter-class similarity of features, and improve the discriminative property of features. In addition, since the quality of pseudo-labels affects the optimization of the model by the graph regular loss, we introduce the idea of clustering into the teacher-student model to improve the quality of pseudo-labels. Experiments on real PolSAR data show that our proposed method achieves an excellent performance.
Chunlei Han, Yuwei Guo 0001, Qi Zang, Baorui Duan, Dong Zhao 0007, Shuang Wang 0001
IGARSS4
2021 Deep multi-level fusion network for multi-source image pixel-wise classification
Xu Liu 0006, Licheng Jiao, Lingling Li 0002, Xu Tang 0004, Yuwei Guo 0001
Knowl. Based Syst.5
2021 Ridgelet-Nets With Speckle Reduction Regularization for SAR Image Scene Classification
abstract
With powerful feature representations, convolutional neural networks (CNNs) have produced tremendous achievements in image classification tasks and, typically, entail millions of labeled samples to train massive parameters. However, the sample labeling of synthetic aperture radar (SAR) images is extremely difficult, especially pixelwise labels, and has, sometimes, required field trips to accomplish labeling. Moreover, the inherent speckle noise may weaken the ability of networks to extract effective features from SAR images. In this article, we address these issues by labeling a few patchwise samples and propose Ridgelet-Nets with speckle reduction regularization for SAR image scene classification by combining deep learning with multiscale geometric analysis and statistical modeling of SAR images. First, we design Ridgelet-Nets with convolutional kernels constructed by ridgelet filters to reduce the training parameters and learn more discriminative features. Then, we embed speckle reduction regularization in the Ridgelet-Nets to restrain the influence of speckle noise and smooth the classification maps, in which the prior information of SAR image statistical modeling is introduced. Finally, we propose an adaptive SAR image scene classification framework based on an extended hierarchical visual semantic model, considering the differences in the structures and spatial relationships of different regions in the SAR images, particularly large-scale and complex scenes. Experimental results on real SAR images demonstrate that the proposed framework can achieve preferable classification performance using very limited labeled samples.
Xiaoxue Qian, Fang Liu 0001, Licheng Jiao, Xiangrong Zhang, Yuwei Guo 0001, Xu Liu 0006, Yuanhao Cui
IEEE Trans. Geosci. Remote. Sens.5
2020 Influence maximization based on the realistic independent cascade model
Jingyi Ding, Jianshe Wu, Yuwei Guo 0001
Knowl. Based Syst.4
2018 Fuzzy Sparse Autoencoder Framework for Single Image Per Person Face Recognition
abstract
The issue of single sample per person (SSPP) face recognition has attracted more and more attention in recent years. Patch/local-based algorithm is one of the most popular categories to address the issue, as patch/local features are robust to face image variations. However, the global discriminative information is ignored in patch/local-based algorithm, which is crucial to recognize the nondiscriminative region of face images. To make the best of the advantage of both local information and global information, a novel two-layer local-to-global feature learning framework is proposed to address SSPP face recognition. In the first layer, the objective-oriented local features are learned by a patch-based fuzzy rough set feature selection strategy. The obtained local features are not only robust to the image variations, but also usable to preserve the discrimination ability of original patches. Global structural information is extracted from local features by a sparse autoencoder in the second layer, which reduces the negative effect of nondiscriminative regions. Besides, the proposed framework is a shallow network, which avoids the over-fitting caused by using multilayer network to address SSPP problem. The experimental results have shown that the proposed local-to-global feature learning framework can achieve superior performance than other state-of-the-art feature learning algorithms for SSPP face recognition.
Yuwei Guo 0001, Licheng Jiao, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001
IEEE Trans. Cybern.1
2018 Fuzzy Superpixels for Polarimetric SAR Images Classification
abstract
Superpixels technique has drawn much attention in computer vision applications. Each superpixels algorithm has its own advantages. Selecting a more appropriate superpixels algorithm for a specific application can improve the performance of the application. In the last few years, superpixels are widely used in polarimetric synthetic aperture radar (PolSAR) image classification. However, no superpixel algorithm is especially designed for image classification. It is believed that both mixed superpixels and pure superpixels exist in an image. Nevertheless, mixed superpixels have negative effects on classification accuracy. Thus, it is necessary to generate superpixels containing as few mixed superpixels as possible for image classification. In this paper, first, a novel superpixels concept, named fuzzy superpixels, is proposed for reducing the generation of mixed superpixels. In fuzzy superpixels, not all pixels are assigned to a corresponding superpixel. We would rather ignore the pixels than assigning them to improper superpixels. Second, a new algorithm, named FuzzyS (FS), is proposed to generate fuzzy superpixels for PolSAR image classification. Three PolSAR images are used to verify the effect of the proposed FS algorithm. Experimental results demonstrate the superiority of the proposed FS algorithm over several state-of-the-art superpixels algorithms.
Yuwei Guo 0001, Licheng Jiao, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001, Wenqiang Hua
IEEE Trans. Fuzzy Syst.1
2015 A novel dynamic rough subspace based selective ensemble
Yuwei Guo 0001, Licheng Jiao, Shuang Wang 0001, Shuo Wang 0005, Fang Liu 0001, Kaixuan Rong
Pattern Recognit.1