Yourun Zhang

dblp:319/3540 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-1086-9036ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Federated Cross-Device Heterogeneous Few-Shot Adaptation for Edge IoT Systems
abstract
The deployment of federated learning in real-world IoT ecosystems presents intrinsic challenges stemming from hardware asymmetry and sample scarcity, the existing related approaches generally homogenize model architectures and assume abundant labeled data, resulting in an inability to achieve fast generalization on devices with varying computational capabilities and dynamic task conditions. To address the aforementioned challenges, we propose a novel federated cross-device heterogeneous few-shot adaptation (Fed-CHFSA) method for IoT systems. In Fed-CHFSA, collaborating with other devices, each edge device obtains a personalized model that can not only adapt well to the category distribution of respective local data but also recognize unseen categories without data leakage. Specifically, we designed a fine-grained personalized aggregation (FPA) module and an information entropy-driven adaptive feature constraint (EAFC) module for the devices possessing a small amount of labeled data in the model aggregation and training phases of Fed-CHFSA, respectively. In each round of global communication, the edge device performs a certain epoch of personalized training locally under the normalization of EAFC in the feature space. Subsequently, the central server follows the FPA to finely aggregate the received model updates parameter-wise, and redistribute the updated global model to participating devices. After multiple rounds of global communication, every edge device acquires an optimal model more adaptable to local data and more generalized to unseen categories. Compared with existing FL and PFL algorithms on three benchmark few-shot learning (FSL) datasets, the proposed Fed-CHFSA framework achieves the best performance. The effectiveness of FPA and EAFC is also demonstrated by extensive ablation experiments.
Jianzhao Li, Yiting Liu 0004, Boya Deng, Maoguo Gong, Zedong Tang, Mingyang Zhang 0002, Yourun Zhang, Zhuping Hu
IEEE Internet Things J.7
2026 BCFNet: Bi-temporal collaborative fusion network for multi-modal humor detection
Boya Deng, Jianzhao Li, Maoguo Gong, Zedong Tang, Yourun Zhang, Kaiyuan Feng, Yue Wu 0004
Pattern Recognit.5
2025 FedFSL-CFRD: Personalized Federated Few-Shot Learning with Collaborative Feature Representation Disentanglement
abstract
Federated few-shot learning (FedFSL) aims to enable the clients to obtain personalized generalization models for unseen categories with only a small number of referenceable samples in the distributed collaborative training paradigm. Most existing FedFSL-related algorithms suffer from domain bias and feature coupling in the presence of data heterogeneity and sample scarcity. In this work, we propose a collaborative feature representation disentanglement (CFRD) scheme for FedFSL to address these issues. After each client receives the global aggregation parameters, the original feature representation is decoupled into global communal features and local personality features with personalized bias representation, to maintain both global consistency and local relevance in the first feature representation disentanglement. On the few-shot metric space about the second feature representation disentanglement, category-independent information is encoded by class-specific and class-irrelevant reconstructions to separate the discriminative features. The proposed scheme collaboratively accomplishes global domain bias feature disentanglement and local category degradation feature disentanglement from client-wise and class-wise. Experiments on three few-shot benchmark datasets conforming to the FedFSL paradigm demonstrate that our proposed method outperforms state-of-the-art approaches in both global generality and local specificity.
Shanfeng Wang, Jianzhao Li, Zaitian Liu, Yourun Zhang, Maoguo Gong
AAAI4
2025 Few-Shot Learning With Enhancements to Data Augmentation and Feature Extraction
abstract
The few-shot image classification task is to enable a model to identify novel classes by using only a few labeled samples as references. In general, the more knowledge a model has, the more robust it is when facing novel situations. Although directly introducing large amounts of new training data to acquire more knowledge is an attractive solution, it violates the purpose of few-shot learning with respect to reducing dependence on big data. Another viable option is to enable the model to accumulate knowledge more effectively from existing data, i.e., improve the utilization of existing data. In this article, we propose a new data augmentation method called self-mixup (SM) to assemble different augmented instances of the same image, which facilitates the model to more effectively accumulate knowledge from limited training data. In addition to the utilization of data, few-shot learning faces another challenge related to feature extraction. Specifically, existing metric-based few-shot classification methods rely on comparing the extracted features of the novel classes, but the widely adopted downsampling structures in various networks can lead to feature degradation due to the violation of the sampling theorem, and the degraded features are not conducive to robust classification. To alleviate this problem, we propose a calibration-adaptive downsampling (CADS) that calibrates and utilizes the characteristics of different features, which can facilitate robust feature extraction and benefit classification. By improving data utilization and feature extraction, our method shows superior performance on four widely adopted few-shot classification datasets.
Yourun Zhang, Maoguo Gong, Jianzhao Li, Kaiyuan Feng, Mingyang Zhang 0002
IEEE Trans. Neural Networks Learn. Syst.1
2024 Evolutionary multitasking cooperative transfer for multiobjective hyperspectral sparse unmixing
Jianzhao Li, Maoguo Gong, Jinxin Wei, Yourun Zhang, Yue Zhao 0024, Shanfeng Wang, Xiangming Jiang
Knowl. Based Syst.4
2024 Toward Multiparty Personalized Collaborative Learning in Remote Sensing
abstract
The powerful deep learning models in remote sensing are inseparable from the support of massive data. However, the privacy and sensitivity of remote sensing data (RSD) restrict the possibility of each party to collaboratively train and share a large general model. Although multi-party learning (MPL) is a feasible solution, it is difficult for the existing MPL methods to uniformly process different remote sensing tasks (RSTs), and the data held by each party is non-independent and identically distributed, heterogeneous and multi-sources. Therefore, it is urgent to explore a solution for the personalized processing of different RSTs. In this paper, we formulate a novel multi-party personalized collaborative learning (MPCL) framework in terms of models and tasks. Specifically, in each iteration of the communication round, we aim to decouple personalized model optimization from global model learning. Different participants are allowed to explore their personalized local models at a certain distance from the global aggregation models according to the characteristics of their local data. In terms of task personalization, MPCL provides different personalized global models to handle the corresponding RSTs. For participants with different RSTs, it can be implemented in the multi-task collaborative training strategy to explore the connection between different tasks. To demonstrate the feasibility of MPCL, we take remote sensing image classification as a case study and provide a detailed feasibility scheme. We constructed four benchmark datasets compliant with MPL and personalized MPL, including single-source and multi-source about SAR, hyperspectral and optical RSD. The experimental results demonstrate that our MPCL is superior in these four RSD, which ranked first in the competition with the classic or state-of-the-art MPL and personalized MPL algorithms. In addition, the scalability of MPCL is also verified on image segmentation RSTs of building and road extraction.
Jianzhao Li, Maoguo Gong, Zaitian Liu, Shanfeng Wang, Yourun Zhang, Yu Zhou 0051, Yuan Gao 0019
IEEE Trans. Geosci. Remote. Sens.5
2024 MSANet: Multiscale Self-Attention Aggregation Network for Few-Shot Aerial Imagery Segmentation
abstract
Few-shot aerial imagery segmentation refers to the task of segmenting specific objects in scenes that have not been encountered during training with a small amount of annotated data for reference. However, most existing few-shot segmentation algorithms are primarily designed for natural images, and there is still a lack of exploration in the context of remote sensing aerial imagery. In this article, we propose a novel multiscale self-attention aggregation network (MS2A2Net), dubbed MS2A2Net, to address the challenge of few-shot aerial image segmentation in terms of scarce data and network architecture. Specifically, we first incorporate the designed asymmetric momentum contrastive learning (AMCL) into the pre-training stage, to improve the representation capability of the backbone without the expensive labeled data. Then the frozen encoder is transferred to the downstream few-shot segmentation task as the feature embedding. In terms of network architecture, we design self-attention aggregation in multiscale feature fusion, to construct the dual correlation of foreground and background between support and query features at the pixel level. Besides, the coordinate attention is designed to rearrange the distribution of feature importance in both horizontal and vertical spatial order perspectives, which facilitates adaptive fusion with the multiscale features. To verify the availability of the proposed MS2A2Net, we also reconstructed two novel datasets dedicated to few-shot aerial image segmentation, called DLRSD-$4^{i}$and iSAID-$4^{i}$. The experimental results show that our approach MS2A2Net is superior in three few-shot benchmark aerial imagery segmentation datasets, which achieves competitive segmentation performance. Extensive ablation experiments also reflect the effectiveness and scalability of the proposed components and overall network architecture.
Jianzhao Li, Maoguo Gong, Mingyang Zhang 0002, Yourun Zhang, Shanfeng Wang, Yue Wu 0004
IEEE Trans. Geosci. Remote. Sens.5
2024 Collaborative Self-Supervised Evolution for Few-Shot Remote Sensing Scene Classification
abstract
Self-supervised learning, which leverages unlabeled data to learn useful feature representations by constructing auxiliary tasks, has been widely explored in few-shot scene classification to improve the feature representation and generalization capabilities of deep models in scarce data. However, most of the current related work adopts specific self-supervised auxiliary tasks (SSATs) for combinatorial improvement, and does not explore the intrinsic connection between different pretext tasks. In practice, the linkage of SSATs is complex, and the optimization of task-sharing parameters by minimizing linear combinations of losses can be conflicting. In addition, although a single combination of SSAT can improve certain performance on the baseline, it is not the personalized optimal solution on various remote sensing datasets with diverse properties. In this article, we propose a collaborative self-supervised evolution (so-called CSENet) framework for few-shot remote sensing scene classification to automatically search for appropriate weights in balancing the task conflicts. In contrast to most existing methods, which consider all SSATs to be equally efficacious or fixed-weighted for the few-shot main task, CSENet achieves autonomous co-evolutionary optimization by encoding arbitrary self-supervised weights. Specifically, the complex self-supervised combinations for different remote sensing data are transformed into an evolutionary optimization problem, where chromosomes with weighting variables obtain the optimal combination with genetic operators. Based on the transfer learning few-shot training paradigm, CSENet first efficiently searches for optimal self-supervised combinations with potential by the proposed automatic collaborative evolution strategy and automatically adjusts the weights without manual settings. Importantly, CSENet provides both inductive and transductive inference, and supports the embedding of arbitrary SSATs. The effectiveness of the proposed framework is demonstrated by state-of-the-art (SOTA) results on three benchmark datasets.
Yiting Liu 0004, Jianzhao Li, Maoguo Gong, Huilin Liu, Yourun Zhang, Zedong Tang, Yu Zhou 0051
IEEE Trans. Geosci. Remote. Sens.6
2024 Personalized Multiparty Few-Shot Learning for Remote Sensing Scene Classification
abstract
The existing few-shot scene classification (FSSC) algorithms have achieved satisfactory results, but they are limited by the paradigm of centralized machine learning, i.e., private remote sensing data need to be centralized on a certain server for training. However, remote sensing images generally contain sensitive information such as national security and company privacy, so it is realistically difficult to collect remote sensing data from all the parties. Therefore, there is a pressing requirement in FSSC for a novel paradigm to achieve multi-party collaborative learning without compromising remote sensing data privacy. In this paper, we formulate a novel personalized multi-party few-shot learning (PMPFSL) paradigm for remote sensing scene classification. In PMPFSL, different participants can achieve multi-party collaborative learning without sacrificing the privacy of their local data, and their respective local models are able to recognize the unseen remote sensing scene categories with a small number of labeled samples. Importantly, the proposed PMPFSL is applicable to various multi-party learning algorithms and few-shot scene classification networks. Moreover, to address the problems of local model overfitting and poor discriminability of few-shot metrics, we propose the personalized adaptive distillation (PAD) scheme and multi-scale feature matching network (MSFMNet) on PMPFSL, respectively. Specifically, each participant obtains the MSFMNet with initialization parameters, and implements a certain number of local training on their respective private machines. Global aggregation is subsequently achieved by uploading only the local models to the central server. In a new round of local training, the participants realize personalized data adaptation to the global model based on the PAD. Overall, the proposed PMPFSL customizes a personalized few-shot model for each participant that is more tailored to their respective remote sensing scenarios. The experimental results demonstrate that our PMPFSL is superior in three benchmark FSSC datasets. We also extensively studied and analyzed the contributions of PAD and MSFMNet in the proposed PMPFSL framework.
Shanfeng Wang, Jianzhao Li, Zaitian Liu, Maoguo Gong, Yourun Zhang, Yue Zhao 0024, Boya Deng, Yu Zhou 0051
IEEE Trans. Geosci. Remote. Sens.5
2024 Self-Supervised Monocular Depth Estimation With Self-Perceptual Anomaly Handling
abstract
It is attractive to extract plausible 3-D information from a single 2-D image, and self-supervised learning has shown impressive potential in this field. However, when only monocular videos are available as training data, moving objects at similar speeds to the camera can disturb the reprojection process during training. Existing methods filter out some moving pixels by comparing pixelwise photometric error, but the illumination inconsistency between frames leads to incomplete filtering. In addition, existing methods calculate photometric error within local windows, which leads to the fact that even if an anomalous pixel is masked out, it can still implicitly disturb the reprojection process, as long as it is in the local neighborhood of a nonanomalous pixel. Moreover, the ill-posed nature of monocular depth estimation makes the same scene correspond to multiple plausible depth maps, which damages the robustness of the model. In order to alleviate the above problems, we propose: 1) a self-reprojection mask to further filter out moving objects while avoiding illumination inconsistency; 2) a self-statistical mask method to prevent the filtered anomalous pixels from implicitly disturbing the reprojection; and 3) a self-distillation augmentation consistency loss to reduce the impact of ill-posed nature of monocular depth estimation. Our method shows superior performance on the KITTI dataset, especially when evaluating only the depth of potential moving objects.
Yourun Zhang, Maoguo Gong, Mingyang Zhang 0002, Jianzhao Li
IEEE Trans. Neural Networks Learn. Syst.1
2023 Autonomous perception and adaptive standardization for few-shot learning
Yourun Zhang, Maoguo Gong, Jianzhao Li, Kaiyuan Feng, Mingyang Zhang 0002
Knowl. Based Syst.1
2023 Multiform Ensemble Self-Supervised Learning for Few-Shot Remote Sensing Scene Classification
abstract
Self-supervised learning is an effective way to solve model collapse for few-shot remote sensing scene classification (FSRSSC). However, most self-supervised contrastive learning auxiliary tasks perform poorly on the high interclass similarity problem in FSRSSC. Furthermore, it is time-consuming and computationally expensive to obtain the best combination among numerous self-supervised auxiliary tasks. In practical applications, we may encounter difficulties in remote sensing data acquisition and labeling, while most FSRSSC studies only focus on the former. To alleviate the above problems, we propose a multiform ensemble self-supervised learning (MES2L) framework for FSRSSC in this article. Based on the transfer learning-based few-shot scheme, we design a novel global–local contrastive learning auxiliary task to solve the low interclass separability problem. The self-attention mechanism is designed in the local contrast features to investigate the intrinsic associations between different remote sensing scene objectives. We also present a multiform ensemble enhancement (MEE) training method. Ensemble enhancement involves the concatenation of features extracted from different backbones trained by a combination of multiform self-supervised auxiliary tasks. MEE can not only be regarded as a more straightforward alternative to knowledge distillation but also can achieve an effective compromise between expensive computational cost and classification accuracy. In addition, we provide two scene classification schemes of inductive and transductive settings, corresponding to solving the difficulties of remote sensing data acquisition and labeling. The proposed network achieves state-of-the-art results on three benchmark FSRSSC datasets. The potential of the MES2L framework is also demonstrated in combination with classical metalearning-based and metric learning-based few-shot algorithms.
Jianzhao Li, Maoguo Gong, Huilin Liu, Yourun Zhang, Mingyang Zhang 0002, Yue Wu 0004
IEEE Trans. Geosci. Remote. Sens.4
2022 Two-Path Aggregation Attention Network With Quad-Patch Data Augmentation for Few-Shot Scene Classification
abstract
The few-shot scene classification is dedicated to identifying unseen remote sensing classes when only a very small number of labeled samples are available for reference. Most of the existing few-shot scene classification methods are based on meta-learning and employ the episodic learning for training, which lacks the consideration for the utilization of data efficiency. In this paper, instead of designing sophisticated meta-learning based algorithms, we are committed to training a feature extractor with good generalization performance and strong feature extraction capability. Specifically, we propose a novel two-path aggregation attention network with quad-patch data augmentation, called DANet, to solve the problem of few-shot scene classification from both data and architecture aspects. In terms of data, we design a new data augmentation strategy named quad-patch augmentation. We utilize the characteristics of remote sensing images to chunk and reassemble any existing data, thereby generating pseudo-new data to enrich the training set. In terms of architecture, we present a two-path aggregation attention module that makes it easier for the model to focus on the key clues in a targeted manner. The comparative experiments in natural image datasets and remote sensing image datasets demonstrate the effectiveness of our two innovations. In addition, DANet achieves competitive or state-of-the-art (SOTA) results on three benchmark scene classification datasets.
Maoguo Gong, Jianzhao Li, Yourun Zhang, Yue Wu 0004, Mingyang Zhang 0002
IEEE Trans. Geosci. Remote. Sens.3
2022 Self-Supervised Monocular Depth Estimation With Multiscale Perception
abstract
Extracting 3D information from a single optical image is very attractive. Recently emerging self-supervised methods can learn depth representations without using ground truth depth maps as training data by transforming the depth prediction task into an image synthesis task. However, existing methods rely on a differentiable bilinear sampler for image synthesis, which results in each pixel in a synthetic image being derived from only four pixels in the source image and causes each pixel in the depth map to perceive only a few pixels in the source image. In addition, when calculating the photometric error between a synthetic image and its corresponding target image, existing methods only consider the photometric error within a small neighborhood of each single pixel and therefore ignore correlations between larger areas, which causes the model to tend to fall into the local optima for small patches. In order to extend the perceptual area of the depth map over the source image, we propose a novel multi-scale method that downsamples the predicted depth map and performs image synthesis at different resolutions, which enables each pixel in the depth map to perceive more pixels in the source image and improves the performance of the model. As for the locality of photometric error, we propose a structural similarity (SSIM) pyramid loss to allow the model to sense the difference between images in multiple areas of different sizes. Experimental results show that our method achieves superior performance on both outdoor and indoor benchmarks.
Yourun Zhang, Maoguo Gong, Jianzhao Li, Mingyang Zhang 0002, Fenlong Jiang, Hongyu Zhao 0007
IEEE Trans. Image Process.1