Jianzhao Li

dblp:265/1973 · DBLP profile ↗
← Back
29ranked-venue papers
6as first author
29since 2021 · last 2026
0000-0002-1524-1363ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 1 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SVP: stratified vertical priors for LiDAR-based 3D object detection
Lei Ao, Wenkang Wan, Nan Ouyang, Jianzhao Li, Maoguo Gong
Neurocomputing4
2026 Federated Cross-Device Heterogeneous Few-Shot Adaptation for Edge IoT Systems
abstract
The deployment of federated learning in real-world IoT ecosystems presents intrinsic challenges stemming from hardware asymmetry and sample scarcity, the existing related approaches generally homogenize model architectures and assume abundant labeled data, resulting in an inability to achieve fast generalization on devices with varying computational capabilities and dynamic task conditions. To address the aforementioned challenges, we propose a novel federated cross-device heterogeneous few-shot adaptation (Fed-CHFSA) method for IoT systems. In Fed-CHFSA, collaborating with other devices, each edge device obtains a personalized model that can not only adapt well to the category distribution of respective local data but also recognize unseen categories without data leakage. Specifically, we designed a fine-grained personalized aggregation (FPA) module and an information entropy-driven adaptive feature constraint (EAFC) module for the devices possessing a small amount of labeled data in the model aggregation and training phases of Fed-CHFSA, respectively. In each round of global communication, the edge device performs a certain epoch of personalized training locally under the normalization of EAFC in the feature space. Subsequently, the central server follows the FPA to finely aggregate the received model updates parameter-wise, and redistribute the updated global model to participating devices. After multiple rounds of global communication, every edge device acquires an optimal model more adaptable to local data and more generalized to unseen categories. Compared with existing FL and PFL algorithms on three benchmark few-shot learning (FSL) datasets, the proposed Fed-CHFSA framework achieves the best performance. The effectiveness of FPA and EAFC is also demonstrated by extensive ablation experiments.
Jianzhao Li, Yiting Liu 0004, Boya Deng, Maoguo Gong, Zedong Tang, Mingyang Zhang 0002, Yourun Zhang, Zhuping Hu
IEEE Internet Things J.1
2026 Contrastive perception representation learning for image inpainting
Maoguo Gong, Jianzhao Li, Licheng Jiao, Xu Liu 0006, Fang Liu 0001
Knowl. Based Syst.3
2026 BCFNet: Bi-temporal collaborative fusion network for multi-modal humor detection
Boya Deng, Jianzhao Li, Maoguo Gong, Zedong Tang, Yourun Zhang, Kaiyuan Feng, Yue Wu 0004
Pattern Recognit.2
2026 PF2SMIS: Personalized federated few-shot learning for medical image segmentation
Shanfeng Wang, Wanrun Yu, Jianzhao Li, Zhao Wang 0011, Maoguo Gong
Pattern Recognit.4
2025 FedFSL-CFRD: Personalized Federated Few-Shot Learning with Collaborative Feature Representation Disentanglement
abstract
Federated few-shot learning (FedFSL) aims to enable the clients to obtain personalized generalization models for unseen categories with only a small number of referenceable samples in the distributed collaborative training paradigm. Most existing FedFSL-related algorithms suffer from domain bias and feature coupling in the presence of data heterogeneity and sample scarcity. In this work, we propose a collaborative feature representation disentanglement (CFRD) scheme for FedFSL to address these issues. After each client receives the global aggregation parameters, the original feature representation is decoupled into global communal features and local personality features with personalized bias representation, to maintain both global consistency and local relevance in the first feature representation disentanglement. On the few-shot metric space about the second feature representation disentanglement, category-independent information is encoded by class-specific and class-irrelevant reconstructions to separate the discriminative features. The proposed scheme collaboratively accomplishes global domain bias feature disentanglement and local category degradation feature disentanglement from client-wise and class-wise. Experiments on three few-shot benchmark datasets conforming to the FedFSL paradigm demonstrate that our proposed method outperforms state-of-the-art approaches in both global generality and local specificity.
Shanfeng Wang, Jianzhao Li, Zaitian Liu, Yourun Zhang, Maoguo Gong
AAAI2
2025 DT-FedSDC: A Dual-Target Federated Framework with Semantic Enhancement and Disentangled Contrastive Learning for Cross-Domain Recommendation
abstract
Federated cross-domain recommendation aims to alleviate the problem of data sparsity and enable collaborative modeling of user behavior data from different platforms or institutions while ensuring data privacy. Most existing federated cross-domain recommendation methods rely on item IDs for modeling, ignoring the mining and utilization of item semantic information. In addition, due to the heterogeneity of data between different domains, the model is prone to domain bias and feature coupling problems during the aggregation process, which negatively impacts the recommendation performance. This paper proposes a dual-target federated cross-domain recommendation framework with semantic enhancement and disentangled contrastive learning. First, to utilize semantic information of items, item IDs features and text semantic features are jointly fused to enhance the item embedding representations. Second, we propose a user representation decoupling mechanism to explicitly decouple users preferences into shared and domain-specific preferences, thereby alleviating domain bias and feature coupling problems. Furthermore, we design a cross-domain contrastive learning module on the server side to enhance the consistency and transferability of shared representations between user representations across different domains. Experimental results show that the proposed algorithm performs significantly better than existing optimal methods on multiple real-world datasets, demonstrating its excellent performance in federated cross-domain recommendations.
Shanyang Gao, Shanfeng Wang, Lanyu Yao, Jianzhao Li, Zhao Wang 0011, Maoguo Gong, Ke Pan 0001
CIKM4
2025 Heterogeneity-aware pruning framework for personalized federated learning in remote sensing scene classification
Zhuping Hu, Maoguo Gong, Zhuowei Dong, Yiheng Lu, Jianzhao Li, Yue Zhao 0024
Knowl. Based Syst.5
2025 Personalized Federated Contrastive Learning for Recommendation
abstract
Recommender systems play crucial roles in addressing the issue of information overload, but traditional centralized storage in recommendation poses significant privacy concerns. In recent years, federated learning has been successfully introduced into a recommendation, while these algorithms still encounter several challenges. First, real-world recommendation scenarios often suffer from sparse data, making it difficult for models to learn reliable representations. Second, data heterogeneity necessitates the design of personalized models to enhance recommendation performance. To address these challenges, we propose a federated recommendation approach based on graph neural networks, named federated personalized contrastive learning for recommendation. On the client side, we propose a contrastive learning approach to enhance the embedding quality of nodes (users or items) by maximizing positive similarities. Specifically, we formulate the concept of structural neighbors based on the graph structure and devise a contrastive learning objective. We treat nodes and their structural neighbors as positive pairs to better learn node representations. On the server side, we group users based on the learned representations and compute cluster-level federated models and a global model. Each user learns a personalized model by combining these two models. Extensive experiments on five real-world datasets demonstrate that the proposed algorithm outperforms existing methods in terms of performance.
Shanfeng Wang, Xiaolong Fan, Jianzhao Li, Zexuan Lei, Maoguo Gong
IEEE Trans. Comput. Soc. Syst.4
2025 Scale-Aware Pruning Framework for Remote Sensing Object Detection via Multifeature Representation
abstract
With the rapid advancements in computer vision, high-resolution remote sensing imagery has become a crucial data source for object detection. Nevertheless, effectively utilizing limited computational resources and reducing the burden on satellite edge devices remains a significant challenge. To effectively reduce model complexity while maintaining its representational capacity, this article proposes a scale-aware pruning framework (SAPF) to enhance remote sensing object detection ability. First, this article classifies the convolutional layers in object detection models into two categories: layers with a single-scale feature representation and layers with a multiscale feature representation. For convolutional layers with single-scale features, we utilize singular value decomposition (SVD) to quantify feature importance and assess filter redundancy to enhance model efficiency. By removing less critical filters, this pruning criteria aims to reduce the model size and computational load without compromising performance. However, convolutional layers with multiscale features are crucial for optimizing feature extraction and balancing information capture across various scales. To address this, this article evaluates the similarity between convolutional layers with different scales to determine the contribution of various scale features in multiscale fusion. Surprisingly, the SAPF can reduce the FLOPs and parameters, as well as ensure the representational ability obviously when the YOLO v5s and Faster-RCNN are adopted to classify the NWPU VHR-10, RSOD, and SIMD datasets. This means we can save the training computation resources for the model. Additionally, SAPF can significantly improve the efficiency of the model in object detection to ensure its real-time performance.
Zhuping Hu, Maoguo Gong, Yue Zhao 0024, Mingyang Zhang 0002, Yiheng Lu, Jianzhao Li, Yan Pu, Zhao Wang 0011
IEEE Trans. Geosci. Remote. Sens.6
2025 Toward Federated Customized Neural Architecture Search for Remote Sensing Scene Classification
abstract
Remote sensing (RS) scenarios usually involve sensitive geographic information on national security and regional development. In the commonly used centralized machine-learning paradigm, data dispersed in various locations are concentrated and processed on a single server, which is prone to privacy leakage and data security concerns. Besides, it is difficult to solve the high heterogeneity of RS images by simply applying federated learning (FL) algorithms to scene classification. In this article, we formulate a federated remote sensing scene classification (FedSC) framework, and design a customized neural architecture search (CNAS) to achieve both global generality for multiparty collaborative distributed training and local specificity for personalized RS scene customization. The proposed FedSC is generalizable to be implemented in any manually designed networks, network pruning strategies, or NAS methods related to remote sensing scene classification (RSSC). While the designed CNAS not only achieves collaborative distributed training in protecting participant data privacy to obtain a generalized global model, but also provides a customized local model for each participant that is more in line with the characteristics of private RS scenarios. Overall, the proposed FedSC$_{\textrm {CNAS}}$provides a novel federated collaborative training paradigm for RSSC in terms of data privacy, data heterogeneity, and personalized customization. Extensive analytical and comparative experiments on three benchmark RSSC datasets validate the versatility and effectiveness of our methods, and the proposed FedSC$_{\textrm {CNAS}}$exhibits superior competitiveness compared to state-of-the-art methods.
Jianzhao Li, Shanfeng Wang, Maoguo Gong, Zhuping Hu, Yu Zhou 0051
IEEE Trans. Geosci. Remote. Sens.1
2025 Collaborative Frequency-Aware Transformer for Unsupervised Multimodal Change Detection in Heterogeneous Remote Sensing Images
abstract
Multimodal change detection (MCD), as an emerging task, aims at recognizing change regions from bi-temporal remote sensing images (RSI) of different modalities. Inspired by the success of the self-attention mechanism in transformer, attempts have been made to solve MCD through the transformer variants. However, transformer-based network optimization requires high-quality training samples. In addition, due to the significant differences in the data distribution, semantic information, and feature representation of multimodal data, transformer-based methods have obvious deficiencies in local feature representation and spatial consistency, especially when dealing with heterogeneous images. To address the above challenges, we propose a collaborative frequency-aware transformer for MCD (CFAT-MCD). As an unsupervised framework, CFAT-MCD is capable of learning more fine-grained patterns of land cover change through a few pseudo-labels. The CFAT is designed to enhance spatial consistency and align the features on a multi-scale basis, which can effectively mitigate the effects of modal differences. In addition, we propose a window-based spatial-frequency collaborative representation (SFCR) module to introduce frequency information into the spatial domain and improve the discriminability of spatial features. Extensive experiments on public datasets and quantitative analyses have validated the superior detection performance of our approach and the effectiveness of each module.
Yan Pu, Maoguo Gong, Tongfei Liu, Mingyang Zhang 0002, Jianzhao Li, Hanhong Zheng, Yue Zhao 0024
IEEE Trans. Geosci. Remote. Sens.5
2025 Few-Shot Learning With Enhancements to Data Augmentation and Feature Extraction
abstract
The few-shot image classification task is to enable a model to identify novel classes by using only a few labeled samples as references. In general, the more knowledge a model has, the more robust it is when facing novel situations. Although directly introducing large amounts of new training data to acquire more knowledge is an attractive solution, it violates the purpose of few-shot learning with respect to reducing dependence on big data. Another viable option is to enable the model to accumulate knowledge more effectively from existing data, i.e., improve the utilization of existing data. In this article, we propose a new data augmentation method called self-mixup (SM) to assemble different augmented instances of the same image, which facilitates the model to more effectively accumulate knowledge from limited training data. In addition to the utilization of data, few-shot learning faces another challenge related to feature extraction. Specifically, existing metric-based few-shot classification methods rely on comparing the extracted features of the novel classes, but the widely adopted downsampling structures in various networks can lead to feature degradation due to the violation of the sampling theorem, and the degraded features are not conducive to robust classification. To alleviate this problem, we propose a calibration-adaptive downsampling (CADS) that calibrates and utilizes the characteristics of different features, which can facilitate robust feature extraction and benefit classification. By improving data utilization and feature extraction, our method shows superior performance on four widely adopted few-shot classification datasets.
Yourun Zhang, Maoguo Gong, Jianzhao Li, Kaiyuan Feng, Mingyang Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.3
2025 SPCNet: Deep Self-Paced Curriculum Network Incorporated With Inductive Bias
abstract
The vulnerability to poor local optimum and the memorization of noise data limit the generalizability and reliability of massively parameterized convolutional neural networks (CNNs) on complex real-world data. Self-paced curriculum learning (SPCL), which models the easy-to-hard learning progression from human beings, is considered as a potential savior. In spite of the fact that numerous SPCL solutions have been explored, it still confronts two main challenges exactly in solving deep networks. By virtue of various designed regularizers, existing weighting schemes independent of the learning objective heavily rely on the prior knowledge. In addition, alternative optimization strategy (AOS) enables the tedious iterative training procedure, thus there is still not an efficient framework that integrates the SPCL paradigm well with networks. This article delivers a novel insight that attention mechanism allows for adaptive enhancement in the contribution of diverse instance information to the gradient propagation. Accordingly, we propose a general-purpose deep SPCL paradigm that incorporates the preferences of implicit regularizer for different samples into the network structure with inductive bias, which in turn is formalized as the self-paced curriculum network (SPCNet). Our proposal allows simultaneous online difficulty estimation, adaptive sample selection, and model updating in an end-to-end manner, which significantly facilitates the collaboration of SPCL to deep networks. Experiments on image classification and scene classification tasks demonstrate that our approach surpasses the state-of-the-art schemes and obtains superior performance.
Yue Zhao 0024, Maoguo Gong, Mingyang Zhang 0002, A. K. Qin 0001, Fenlong Jiang, Jianzhao Li
IEEE Trans. Neural Networks Learn. Syst.6
2024 Evolutionary Multitasking Collaborative Neural Architecture Search for Scene Classification
abstract
With the acquisition of large-scale remote sensing data and the development of deep learning, convolutional neural networks have achieved great progress in scene clas-sification tasks. However, the current networks greatly rely on the experience design of experts, and a single network cannot cope with multiple complex scene categories. In this paper, we design a novel evolutionary multitasking collaborative neural architecture search (EMCNAS) for remote sensing scene classification. EMCNAS mainly explores the similar features of different remote sensing scene classification tasks, and utilizes the uniformly encoded population to achieve implicit collaborative transfer. EM CNAS is able to adaptively determine the degree of exchange of genetic material based on the similarity of tasks in different scenarios, enabling more effective positive transfer. Compared with excellent manually designed neural networks and NAS peers, the proposed EMCNAS achieved competitive results on two remote sensing scene classification benchmark datasets UC Merced LandUse and NWPU-RESISC45 datasets. Ablation experiments also demonstrate the effectiveness of the adaptive collaborative transfer designed in EMCNAS.
Shanfeng Wang, Zaitian Liu, Jianzhao Li, Maoguo Gong
CEC3
2024 Evolutionary multitasking cooperative transfer for multiobjective hyperspectral sparse unmixing
Jianzhao Li, Maoguo Gong, Jinxin Wei, Yourun Zhang, Yue Zhao 0024, Shanfeng Wang, Xiangming Jiang
Knowl. Based Syst.1
2024 Towards fair and personalized federated recommendation
Shanfeng Wang, Hao Tao, Jianzhao Li, Xinyuan Ji, Yuan Gao 0019, Maoguo Gong
Pattern Recognit.3
2024 Data Customization-Based Multiobjective Optimization Pruning Framework for Remote Sensing Scene Classification
abstract
Pruning techniques have been utilized widely for convolutional neural networks (CNNs) to reduce the computation resources in remote sensing scene image classification. However, conventional pruning techniques are weight-based, which can not balance the pruning ratio and representation ability appropriately. In this paper, we propose a Data Customization-based Multiobjective Optimization Pruning (DCMOP) framework for the pruning in remote sensing scene image classification, which can not only trade-off between pruning ratio and capability for CNNs, but also speed up the evolutionary process for the pruning. We adopt the multiobjective evolutionary algorithms (MOEAs) to search for a trade-off between the pruning ratio and capability for CNNs. However, a big concern of pruning for networks via MOEAs is that the evaluation of sub-networks is time-costing. This originates that the slimmed sub-networks require a lot of retraining operation, which will burden the hardware. In order to alleviate this limitation, we design a Data Customization-based Proxy Mechanism (DCPM) to reduce the size of the input dataset in terms of the structure of the slimmed sub-network to accelerate significantly the evolutionary process for the pruning. According to this, our proposed DCMOP achieves the pruning with higher efficiency and performance by cooperating with MOEAs and DCPM. Experimental results based on four datasets of AID, NWPURESISC45, PatternNet, and WHU-RS19 show that the proposed DCMOP can achieve a balance between model performance and pruning rate, while obviously reducing the time cost of the pruning.
Zhuping Hu, Maoguo Gong, Yiheng Lu, Jianzhao Li, Yue Zhao 0024, Mingyang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.4
2024 Toward Multiparty Personalized Collaborative Learning in Remote Sensing
abstract
The powerful deep learning models in remote sensing are inseparable from the support of massive data. However, the privacy and sensitivity of remote sensing data (RSD) restrict the possibility of each party to collaboratively train and share a large general model. Although multi-party learning (MPL) is a feasible solution, it is difficult for the existing MPL methods to uniformly process different remote sensing tasks (RSTs), and the data held by each party is non-independent and identically distributed, heterogeneous and multi-sources. Therefore, it is urgent to explore a solution for the personalized processing of different RSTs. In this paper, we formulate a novel multi-party personalized collaborative learning (MPCL) framework in terms of models and tasks. Specifically, in each iteration of the communication round, we aim to decouple personalized model optimization from global model learning. Different participants are allowed to explore their personalized local models at a certain distance from the global aggregation models according to the characteristics of their local data. In terms of task personalization, MPCL provides different personalized global models to handle the corresponding RSTs. For participants with different RSTs, it can be implemented in the multi-task collaborative training strategy to explore the connection between different tasks. To demonstrate the feasibility of MPCL, we take remote sensing image classification as a case study and provide a detailed feasibility scheme. We constructed four benchmark datasets compliant with MPL and personalized MPL, including single-source and multi-source about SAR, hyperspectral and optical RSD. The experimental results demonstrate that our MPCL is superior in these four RSD, which ranked first in the competition with the classic or state-of-the-art MPL and personalized MPL algorithms. In addition, the scalability of MPCL is also verified on image segmentation RSTs of building and road extraction.
Jianzhao Li, Maoguo Gong, Zaitian Liu, Shanfeng Wang, Yourun Zhang, Yu Zhou 0051, Yuan Gao 0019
IEEE Trans. Geosci. Remote. Sens.1
2024 MSANet: Multiscale Self-Attention Aggregation Network for Few-Shot Aerial Imagery Segmentation
abstract
Few-shot aerial imagery segmentation refers to the task of segmenting specific objects in scenes that have not been encountered during training with a small amount of annotated data for reference. However, most existing few-shot segmentation algorithms are primarily designed for natural images, and there is still a lack of exploration in the context of remote sensing aerial imagery. In this article, we propose a novel multiscale self-attention aggregation network (MS2A2Net), dubbed MS2A2Net, to address the challenge of few-shot aerial image segmentation in terms of scarce data and network architecture. Specifically, we first incorporate the designed asymmetric momentum contrastive learning (AMCL) into the pre-training stage, to improve the representation capability of the backbone without the expensive labeled data. Then the frozen encoder is transferred to the downstream few-shot segmentation task as the feature embedding. In terms of network architecture, we design self-attention aggregation in multiscale feature fusion, to construct the dual correlation of foreground and background between support and query features at the pixel level. Besides, the coordinate attention is designed to rearrange the distribution of feature importance in both horizontal and vertical spatial order perspectives, which facilitates adaptive fusion with the multiscale features. To verify the availability of the proposed MS2A2Net, we also reconstructed two novel datasets dedicated to few-shot aerial image segmentation, called DLRSD-$4^{i}$and iSAID-$4^{i}$. The experimental results show that our approach MS2A2Net is superior in three few-shot benchmark aerial imagery segmentation datasets, which achieves competitive segmentation performance. Extensive ablation experiments also reflect the effectiveness and scalability of the proposed components and overall network architecture.
Jianzhao Li, Maoguo Gong, Mingyang Zhang 0002, Yourun Zhang, Shanfeng Wang, Yue Wu 0004
IEEE Trans. Geosci. Remote. Sens.1
2024 Collaborative Self-Supervised Evolution for Few-Shot Remote Sensing Scene Classification
abstract
Self-supervised learning, which leverages unlabeled data to learn useful feature representations by constructing auxiliary tasks, has been widely explored in few-shot scene classification to improve the feature representation and generalization capabilities of deep models in scarce data. However, most of the current related work adopts specific self-supervised auxiliary tasks (SSATs) for combinatorial improvement, and does not explore the intrinsic connection between different pretext tasks. In practice, the linkage of SSATs is complex, and the optimization of task-sharing parameters by minimizing linear combinations of losses can be conflicting. In addition, although a single combination of SSAT can improve certain performance on the baseline, it is not the personalized optimal solution on various remote sensing datasets with diverse properties. In this article, we propose a collaborative self-supervised evolution (so-called CSENet) framework for few-shot remote sensing scene classification to automatically search for appropriate weights in balancing the task conflicts. In contrast to most existing methods, which consider all SSATs to be equally efficacious or fixed-weighted for the few-shot main task, CSENet achieves autonomous co-evolutionary optimization by encoding arbitrary self-supervised weights. Specifically, the complex self-supervised combinations for different remote sensing data are transformed into an evolutionary optimization problem, where chromosomes with weighting variables obtain the optimal combination with genetic operators. Based on the transfer learning few-shot training paradigm, CSENet first efficiently searches for optimal self-supervised combinations with potential by the proposed automatic collaborative evolution strategy and automatically adjusts the weights without manual settings. Importantly, CSENet provides both inductive and transductive inference, and supports the embedding of arbitrary SSATs. The effectiveness of the proposed framework is demonstrated by state-of-the-art (SOTA) results on three benchmark datasets.
Yiting Liu 0004, Jianzhao Li, Maoguo Gong, Huilin Liu, Yourun Zhang, Zedong Tang, Yu Zhou 0051
IEEE Trans. Geosci. Remote. Sens.2
2024 Personalized Multiparty Few-Shot Learning for Remote Sensing Scene Classification
abstract
The existing few-shot scene classification (FSSC) algorithms have achieved satisfactory results, but they are limited by the paradigm of centralized machine learning, i.e., private remote sensing data need to be centralized on a certain server for training. However, remote sensing images generally contain sensitive information such as national security and company privacy, so it is realistically difficult to collect remote sensing data from all the parties. Therefore, there is a pressing requirement in FSSC for a novel paradigm to achieve multi-party collaborative learning without compromising remote sensing data privacy. In this paper, we formulate a novel personalized multi-party few-shot learning (PMPFSL) paradigm for remote sensing scene classification. In PMPFSL, different participants can achieve multi-party collaborative learning without sacrificing the privacy of their local data, and their respective local models are able to recognize the unseen remote sensing scene categories with a small number of labeled samples. Importantly, the proposed PMPFSL is applicable to various multi-party learning algorithms and few-shot scene classification networks. Moreover, to address the problems of local model overfitting and poor discriminability of few-shot metrics, we propose the personalized adaptive distillation (PAD) scheme and multi-scale feature matching network (MSFMNet) on PMPFSL, respectively. Specifically, each participant obtains the MSFMNet with initialization parameters, and implements a certain number of local training on their respective private machines. Global aggregation is subsequently achieved by uploading only the local models to the central server. In a new round of local training, the participants realize personalized data adaptation to the global model based on the PAD. Overall, the proposed PMPFSL customizes a personalized few-shot model for each participant that is more tailored to their respective remote sensing scenarios. The experimental results demonstrate that our PMPFSL is superior in three benchmark FSSC datasets. We also extensively studied and analyzed the contributions of PAD and MSFMNet in the proposed PMPFSL framework.
Shanfeng Wang, Jianzhao Li, Zaitian Liu, Maoguo Gong, Yourun Zhang, Yue Zhao 0024, Boya Deng, Yu Zhou 0051
IEEE Trans. Geosci. Remote. Sens.2
2024 Self-Supervised Monocular Depth Estimation With Self-Perceptual Anomaly Handling
abstract
It is attractive to extract plausible 3-D information from a single 2-D image, and self-supervised learning has shown impressive potential in this field. However, when only monocular videos are available as training data, moving objects at similar speeds to the camera can disturb the reprojection process during training. Existing methods filter out some moving pixels by comparing pixelwise photometric error, but the illumination inconsistency between frames leads to incomplete filtering. In addition, existing methods calculate photometric error within local windows, which leads to the fact that even if an anomalous pixel is masked out, it can still implicitly disturb the reprojection process, as long as it is in the local neighborhood of a nonanomalous pixel. Moreover, the ill-posed nature of monocular depth estimation makes the same scene correspond to multiple plausible depth maps, which damages the robustness of the model. In order to alleviate the above problems, we propose: 1) a self-reprojection mask to further filter out moving objects while avoiding illumination inconsistency; 2) a self-statistical mask method to prevent the filtered anomalous pixels from implicitly disturbing the reprojection; and 3) a self-distillation augmentation consistency loss to reduce the impact of ill-posed nature of monocular depth estimation. Our method shows superior performance on the KITTI dataset, especially when evaluating only the depth of potential moving objects.
Yourun Zhang, Maoguo Gong, Mingyang Zhang 0002, Jianzhao Li
IEEE Trans. Neural Networks Learn. Syst.4
2023 Autonomous perception and adaptive standardization for few-shot learning
Yourun Zhang, Maoguo Gong, Jianzhao Li, Kaiyuan Feng, Mingyang Zhang 0002
Knowl. Based Syst.3
2023 Deep Fuzzy Variable C-Means Clustering Incorporated With Curriculum Learning
abstract
End-to-end deep clustering method utilizes deep neural networks to jointly learn representation features and clustering assignments. Although many k-means-friendly deep clustering models have been explored, the existing division-based methods tend to directly implement clustering with a specific number of clusters, which suffers from poor performance resulted from indistinguishable clusters, and contributes to bad local optimum. At the same time, the representation learning of fuzzy$c$-means clustering in the feature space still needs more research. In this article, a deep fuzzy curriculum clustering method with the learning strategy of clustering from easy to complex automatically is proposed to tackle the above issues. First, considering the soft flexible allocation of fuzzy$c$-means and the preservation of local structure of original data, the fuzzy clustering loss and the autoencoder's reconstruction loss are constructed to learn the embedded features and clustering centers simultaneously. Second, curriculum loss is introduced into the constraint to make clusters successively merge in line with implementing clustering from easy to complex, and realize the bottom-up deep aggregative clustering automatically. In addition, novel curriculum information is proposed as constraint to guide the merging of clusters belonging to the same class. Experimental results on four real-world datasets show the superiority of the proposal.
Maoguo Gong, Yue Zhao 0024, Hao Li 0009, A. K. Qin 0001, Lining Xing 0001, Jianzhao Li, Yiting Liu 0004
IEEE Trans. Fuzzy Syst.6
2023 Multiform Ensemble Self-Supervised Learning for Few-Shot Remote Sensing Scene Classification
abstract
Self-supervised learning is an effective way to solve model collapse for few-shot remote sensing scene classification (FSRSSC). However, most self-supervised contrastive learning auxiliary tasks perform poorly on the high interclass similarity problem in FSRSSC. Furthermore, it is time-consuming and computationally expensive to obtain the best combination among numerous self-supervised auxiliary tasks. In practical applications, we may encounter difficulties in remote sensing data acquisition and labeling, while most FSRSSC studies only focus on the former. To alleviate the above problems, we propose a multiform ensemble self-supervised learning (MES2L) framework for FSRSSC in this article. Based on the transfer learning-based few-shot scheme, we design a novel global–local contrastive learning auxiliary task to solve the low interclass separability problem. The self-attention mechanism is designed in the local contrast features to investigate the intrinsic associations between different remote sensing scene objectives. We also present a multiform ensemble enhancement (MEE) training method. Ensemble enhancement involves the concatenation of features extracted from different backbones trained by a combination of multiform self-supervised auxiliary tasks. MEE can not only be regarded as a more straightforward alternative to knowledge distillation but also can achieve an effective compromise between expensive computational cost and classification accuracy. In addition, we provide two scene classification schemes of inductive and transductive settings, corresponding to solving the difficulties of remote sensing data acquisition and labeling. The proposed network achieves state-of-the-art results on three benchmark FSRSSC datasets. The potential of the MES2L framework is also demonstrated in combination with classical metalearning-based and metric learning-based few-shot algorithms.
Jianzhao Li, Maoguo Gong, Huilin Liu, Yourun Zhang, Mingyang Zhang 0002, Yue Wu 0004
IEEE Trans. Geosci. Remote. Sens.1
2022 Financial credit risk assessment of online supply chain in construction industry with a hybrid model chain
abstract
Under the influence of COVID-19, although upstream small- and medium-sized enterprises (SMEs) in construction industry chain suffer from high operating costs and tight cash flow problems, their financing demands are even stronger. As an electronic and platform-based comprehensive service, online supply chain finance can ease the financing problems of SMEs in the construction industry. However, it has become an important issue for financial institutions to effectively assess the credit risk in the process of online supply chain financing. In this paper, an online supply chain risk assessment method, which is based on a hybrid model chain, including eXtreme Gradient Boosting (XGBoost), Synthetic Minority Oversampling TEchnique for Nominal and Continuous (SMOTENC), and Random Forest (RF), is proposed to identify and control the credit risk of financial institutions. Specifically, we establish the financial credit risk assessment system with respect to the characteristics of financing enterprises in the construction industry supply chain, including the status of financing enterprises, the status of core enterprises, the operating status of the supply chain, and the status of assets under financing. On the basis of the system, the best index number of the assessment system and the minority samples are obtained by the XGBoost algorithm and the SMOTENC algorithm, respectively. The classification method based on RF is applied to judge the credit risk of financing enterprises in the supply chain of construction industry. In the simulation stage, we take upstream SMEs in the supply chain of construction industry in China as an example for empirical analysis to validate the effectiveness of our proposed method. The credit risk assessment method proposed in this paper has better performance than the commonly used ones in the academic field with an average improvement on assessment accuracy for 6.39% and an average increase of Area Under Curve for 6.95%. Our study provides meaningful exploration on the fund monitoring system of the financing service platform to improve financing efficiency and risk management level.
Jia Liu 0069, Jianzhao Li
Int. J. Intell. Syst.4
2022 Two-Path Aggregation Attention Network With Quad-Patch Data Augmentation for Few-Shot Scene Classification
abstract
The few-shot scene classification is dedicated to identifying unseen remote sensing classes when only a very small number of labeled samples are available for reference. Most of the existing few-shot scene classification methods are based on meta-learning and employ the episodic learning for training, which lacks the consideration for the utilization of data efficiency. In this paper, instead of designing sophisticated meta-learning based algorithms, we are committed to training a feature extractor with good generalization performance and strong feature extraction capability. Specifically, we propose a novel two-path aggregation attention network with quad-patch data augmentation, called DANet, to solve the problem of few-shot scene classification from both data and architecture aspects. In terms of data, we design a new data augmentation strategy named quad-patch augmentation. We utilize the characteristics of remote sensing images to chunk and reassemble any existing data, thereby generating pseudo-new data to enrich the training set. In terms of architecture, we present a two-path aggregation attention module that makes it easier for the model to focus on the key clues in a targeted manner. The comparative experiments in natural image datasets and remote sensing image datasets demonstrate the effectiveness of our two innovations. In addition, DANet achieves competitive or state-of-the-art (SOTA) results on three benchmark scene classification datasets.
Maoguo Gong, Jianzhao Li, Yourun Zhang, Yue Wu 0004, Mingyang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.2
2022 Self-Supervised Monocular Depth Estimation With Multiscale Perception
abstract
Extracting 3D information from a single optical image is very attractive. Recently emerging self-supervised methods can learn depth representations without using ground truth depth maps as training data by transforming the depth prediction task into an image synthesis task. However, existing methods rely on a differentiable bilinear sampler for image synthesis, which results in each pixel in a synthetic image being derived from only four pixels in the source image and causes each pixel in the depth map to perceive only a few pixels in the source image. In addition, when calculating the photometric error between a synthetic image and its corresponding target image, existing methods only consider the photometric error within a small neighborhood of each single pixel and therefore ignore correlations between larger areas, which causes the model to tend to fall into the local optima for small patches. In order to extend the perceptual area of the depth map over the source image, we propose a novel multi-scale method that downsamples the predicted depth map and performs image synthesis at different resolutions, which enables each pixel in the depth map to perceive more pixels in the source image and improves the performance of the model. As for the locality of photometric error, we propose a structural similarity (SSIM) pyramid loss to allow the model to sense the difference between images in multiple areas of different sizes. Experimental results show that our method achieves superior performance on both outdoor and indoor benchmarks.
Yourun Zhang, Maoguo Gong, Jianzhao Li, Mingyang Zhang 0002, Fenlong Jiang, Hongyu Zhao 0007
IEEE Trans. Image Process.3