VLDB 2026 Research / reviewers in the wild / expert
Juepeng Zheng
dblp:237/5258
· DBLP profile ↗
46ranked-venue papers
12as first author
44since 2021 · last 2026
0000-0002-4403-593XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 8 first-author · 16 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 16 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Continual test-time adaptation for object detection with adaptive monitoring and randomized restoration
Shilei Cao 0005, Juepeng Zheng, Baoquan Zhao, Runmin Dong, Haohuan Fu |
Expert Syst. Appl. | 2 |
| 2026 | GALA: A GlobAl-LocAl cluster approach for Multi-Source Active Domain Adaptation
Juepeng Zheng, Yibin Wen, Peifeng Zhang, Zurong Mai, Qingmei Li, Haohuan Fu |
Pattern Recognit. | 1 |
| 2026 | Evidential Graph Contrastive Alignment for Source-Free Blending-Target Domain AdaptationabstractIn this article, we first tackle a more realistic domain adaptation (DA) setting: source-free blending-target DA (SF-BTDA), where we cannot access to source-domain data while facing mixed multiple target domains without any domain labels in prior. Compared to existing DA scenarios, SF-BTDA generally faces the coexistence of different label shifts in different targets, along with noisy target pseudolabels generated from the source model. In this article, we propose a new method called evidential graph contrastive alignment (EGCA) to decouple the blending-target domain and alleviate the effect of noisy target pseudolabels. First, to improve the quality of pseudo target labels, we propose a calibrated evidential learning (CEL) module to iteratively improve both the accuracy and certainty of the resulting model and adaptively generate high-quality pseudo target labels. Second, we design a graph contrastive learning with the domain distance matrix and confidence-uncertainty criterion, to minimize the distribution gap of samples of the same class in the blending-target domain, which alleviates the coexistence of different label shifts in blended targets. We conduct a new benchmark based on three standard DA datasets, and EGCA outperforms other methods with considerable gains and achieves comparable results compared with those that have domain labels or source data in prior. Juepeng Zheng, Yibin Wen, Jinxiao Zhang, Runmin Dong, Haohuan Fu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | TTA-FedDG: Leveraging Test-Time Adaptation to Address Federated Domain GeneralizationabstractIn recent years, Federated Domain Generalization (FedDG) has succeeded in generalizing to unknown clients (domains). However, current methods only utilize training data, and when there is a significant difference between the unknown client and source client domains (domain shift), these methods cannot ensure model performance. This limitation appears to have caused research in FedDG to reach a bottleneck. On the other hand, test data is a resource that can help models adapt while previous FedDG approaches have not taken this into account. In this paper, we introduce a new framework TTA-FedDG to address the FedDG problem, which leverages test-time adaptation (TTA) to adapt across different domains, thereby enhancing the generalization of the model. We propose the method Federated domain generalization based on select Strong Pseudo Label (FedSPL), which combines fast feature matching and knowledge distillation. Our method consists of two parts. Firstly, we use fast feature reordering for feature mixing during local updates on the client side, improving the robustness of the global model and enhancing its generalization ability to mitigate domain shift. Secondly, we employ a teacher-student model with contrastive learning and label selection during the testing phase, enabling the global model to better adapt to the distribution of the target client,thereby alleviating domain shift. Extensive experiments havedemonstrated the effectiveness of FedSPL in handling domain shift, outperforming existing FedDG methods across multiple datasets and model architectures. Haoyuan Liang, Shilei Cao 0005, Juepeng Zheng |
AAAI | 5 |
| 2025 | ADU: Adaptive Detection of Unknown Categories in Black-Box Domain AdaptationabstractBlack-box Domain Adaptation (BDA) utilizes a black-box predictor of the source domain to label target domain data, addressing privacy concerns in Unsupervised Domain Adaptation (UDA). However, BDA assumes identical label sets across domains, which is unrealistic. To overcome this limitation, we propose a study on BDA with unknown classes in the target domain. It uses a black-box predictor to label target data and identify "unknown" categories, without requiring access to source domain data or predictor parameters, thus addressing both data privacy and category shift issues in traditional UDA. Existing methods face two main challenges: (i) Noisy pseudo-labels in knowledge distillation (KD) accumulate prediction errors, and (ii) relying on a preset threshold fails to adapt to varying category shifts. To address these, we propose ADU, a framework that allows the target domain to autonomously learn pseudo-labels guided by quality and use an adaptive threshold to identify "unknown" categories. Specifically, ADU consists of Selective Amplification Knowledge Distillation (SAKD) and Entropy-Driven Label Differentiation (EDLD). SAKD improves KD by focusing on high-quality pseudo-labels, mitigating the impact of noisy labels. EDLD categorizes pseudo-labels by quality and applies tailored training strategies to distinguish "unknown" categories, improving detection accuracy and adaptability. Extensive experiments show that ADU achieves state-of-the-art results, outperforming the best existing method by 3.1% on VisDA in the OPBDA scenario. Yushan Lai, Haoyuan Liang, Juepeng Zheng, Zhiyu Ye |
CVPR | 4 |
| 2025 | Shift-Driven Learning for Unsupervised Domain AdaptationabstractSelf-training is widely used in unsupervised domain adaptation (UDA) by assigning pseudo labels to unlabeled samples. However, existing self-training strategies bring bias, while potentially inaccurate pseudo labels may accumulate errors during self-training (self-training shift) and the inability to accurately distinguish features may bring prediction bias (class shift). To address these issues, we propose Shift-Driven Learning (SDL). First, we decouple the generation and utilization of pseudo labels to mitigate the direct error accumulation. Second, we measure the maximum training shift of data, where the classifier achieves high accuracy on labeled data while making as many mistakes as possible on unlabeled data. Then we adversarially optimize the feature representations generation to indirectly decrease the self-training shift. Third, we minimize the class shift by data rearrangement strategy and joint contrastive learning, which find class-level discriminative feature representations. Extensive experiments justify that SDL outperforms SOTA methods on three UDA datasets with considerable gains. Wentang Chen, Yibin Wen, Juepeng Zheng |
ICME | 3 |
| 2025 | Instance-Distance Active Learning for Source-Free Cross-Domain Object DetectionabstractSource-Free Domain Adaptive Object Detection (SFDA-OD) aims to apply a pre-trained detector from the source domain to unlabeled data in the target domain. Existing methods typically utilize techniques such as domain alignment and fine-tuning model parameters to achieve satisfactory performance in the target domain. However, even using SFDA approaches, the accuracy of the model on the target domain remains significantly accuracy gap compared to the fully-supervised model. In this paper, we explore a new scenario Source-Free Active Domain Adaptive Object Detection (SFADA-OD) to better enhance model adaptability and performance within limited annotation, and propose a novel approach named Learning Domain Distance for Active Pick (LDDAP). Firstly, we design graph-aware distance learning to calculate the distance between the encoded images and the predicted instances, which better evaluate the distance between domains in SFDA scenario. Secondly, we propose instance-aware active sampling to facilitate instance-level distance associated with other indexes to select the most valuable images, thereby maximizing the effectiveness of the limited labeled information. The results demonstrate a significant enhancement in performance over baseline cross-domain models and other Active Learning related methods on several well-known public datasets.In addition, our model even surpasses fully-supervised results on the dataset Pascal → Watercolor. Code is avaliable at LDDAP Kangrui Du, Yujun Qian, Juepeng Zheng |
ICME | 3 |
| 2025 | MSPoint-Gait: Multi-Scale Point Cloud Analysis for 3D Gait Recognition via Cross-Modal LearningabstractRecent advances in LiDAR technology have enabled privacy-preserving gait recognition using 3D point cloud data. However, existing approaches struggle with the inherent challenges of point cloud processing and understanding such as spatial sparsity, irregular sampling, and complex temporal dynamics. In this paper, we present MSPoint-Gait, a novel framework that addresses these challenges through multi-scale analysis and cross-modal learning. At the core of our framework lies a Depth-Aware Attention Module (DAAM) that leverages rich 3D geometric information to generate attention-weighted depth representations, enabling fine-grained feature extraction from point cloud sequences. We further introduce a Multi-Scale Spatio-Temporal (MSST) network that hierarchically captures both local and global gait patterns through adaptive convolution kernels across multiple spatial and temporal scales. These components are unified through a novel cross-modal learning strategy that effectively bridges the semantic gap between raw point clouds and structured depth representations. The proposed frame-work achieves state-of-the-art performance on the challenging SUSTech1K dataset, with 91.9% Rank-1 and 98.0% Rank-5 accuracy, demonstrating significant improvements over existing methods across various walking conditions and viewpoints. Xinzhu Li, Yikun Chen, Guanghui Yue 0001, Wei Zhou 0021, Ruomei Wang 0001, Xudong Mao, Juepeng Zheng, Fan Zhou 0001, Ziqi Qiu, Baoquan Zhao |
ICME | 8 |
| 2025 | Federated Open-Set Domain Generalization with Adaptive Adjustment Boundary and WeightsabstractConcerns about privacy and the centralized collection of sensitive data have led to the development of Federated Learning, a paradigm enabling collaborative model training without the need to aggregate raw data centrally. However, variations in data distributions between source and target clients, a phenomenon known as domain shift, often lead to degraded model performance. While recent advancements in Federated Domain Generalization address this challenge, they typically operate under a closed-set assumption, disregarding scenarios where target domains introduce entirely new classes, referred to as category shift. This oversight can result in critical misclassifications in real-world applications. To overcome these limitations, we explore Federated Open-Set Domain Generalization (FedOSDG) setting for the first time, which not only preserves data privacy but also identifies new, unseen classes in the unseen target domains. Specifically, we propose the Adaptive Adjustment Boundary and Weights (AABAW) framework, comprising Stronger Classification Boundary (SCB) and Adaptive Adjustment of Weights (AAW). The SCB module reinforces the decision boundaries of the binary classifiers to handle category shift while the AAW module leverages local model diversity to increase the variance of global model, thereby enhancing the model’s generalization under domain shifts. Experimental results show that our proposed AABAW achieves state-of-the-art performance in recognizing unknown classes and H-scores in the FedOSDG task with considerable gains and maintains competitive performance in all classes. Haoyuan Liang, Shilei Cao 0005, Yushan Lai, Juepeng Zheng |
ICME | 4 |
| 2025 | Evidential Graph Contrastive Alignment for Source-Free Blending-Target Domain AdaptationabstractIn this paper, we firstly tackle a more realistic Domain Adaptation (DA) setting: Source-Free Blending-Target Domain Adaptation (SF-BTDA), where we can not access to source domain data while facing mixed multiple target domains without any domain labels in prior. Compared to existing DA scenarios, SF-BTDA generally faces the co-existence of different label shifts in different targets, along with noisy target pseudo labels generated from the source model. In this paper, we propose a new method called Evidential Graph Contrastive Alignment (EGCA) to decouple the blending target domain and alleviate the effect from noisy target pseudo labels. First, to improve the quality of pseudo target labels, we propose a Calibrated Evidential Learning module to iteratively improve both the accuracy and certainty of the resulting model and adaptively generate high-quality pseudo target labels. Second, we design a Graph Contrastive Learning with the domain distance matrix and confidence-uncertainty criterion, to minimize the distribution gap of samples of a same class in the blended target domains, which alleviates the co-existence of different label shifts in blended targets. We conduct a new benchmark based on three standard DA datasets and EGCA outperforms other methods with considerable gains and achieves comparable results compared with those have domain labels or source data in prior. Juepeng Zheng, Yibin Wen |
ICME | 1 |
| 2025 | DepthGait: Multi-Scale Cross-Level Feature Fusion of RGB-Derived Depth and Silhouette Sequences for Robust Gait RecognitionabstractRobust gait recognition requires highly discriminative representations, which are closely tied to input modalities. While binary silhouettes and skeletons have dominated recent literature, these 2D representations fall short of capturing sufficient cues that can be exploited to handle viewpoint variations, and capture finer and meaningful details of gait. In this paper, we introduce a novel framework, termed DepthGait, that incorporates RGB-derived depth maps and silhouettes for enhanced gait recognition. Specifically, apart from the 2D silhouette representation of the human body, the proposed pipeline explicitly estimates depth maps from a given RGB image sequence and uses them as a new modality to capture discriminative features inherent in human locomotion. In addition, a novel multi-scale and cross-level fusion scheme has also been developed to bridge the modality gap between depth maps and silhouettes. Extensive experiments on standard benchmarks demonstrate that the proposed DepthGait achieves state-of-the-art performance compared to peer methods and attains an impressive mean rank-1 accuracy on the challenging datasets. Xinzhu Li, Juepeng Zheng, Yikun Chen, Xudong Mao, Guanghui Yue 0001, Wei Zhou 0021, Chenlei Lv, Ruomei Wang 0001, Fan Zhou 0001, Baoquan Zhao |
ACM Multimedia | 2 |
| 2025 | Quantifying Samples with Invariance for Source-Free Class Incremental Domain AdaptationabstractIn response to the growing demands of real-world applications, models must be capable of learning continuously under inconsistent data distribution. However, existing Class-Incremental (CI) methods fail to alleviate domain shifts, while traditional Unsupervised Domain Adaptation (UDA) techniques suffer from catastrophic forgetting and privacy concerns. To address these limitations, we explore Source-Free Class Incremental Domain Adaptation (SFCIDA) and propose a novel approach, Quantifying Samples with Invariance (QSI), for this scenario. Our proposed method involves two main strategies: (1) Semantic Restructuring. We identify confusing source category pairs and restructure images to create a negative dataset that is semantically similar to the source features, refining accurate decision boundary among source categories. (2) Invariance Quantification. The sample's confidence is then quantified by its spatial location under the special data distribution, reflecting the trade-off between invariant features and domain shifts. Guided by such strategy, samples' confidence is accumulated for the target model to prioritize reliable categories, not only mitigating the poor performance of experience replay in unsupervised scenarios, but alleviating distribution discrepancies simultaneously. Experiments demonstrate that our approach outperforms previous methods, establishing new state-of-the-art performance on the Office-31, Office-Home and DomainNet-126 datasets, with average accuracy improvements of over 7.3%, 4.9% and 10.2% respectively. Zhiyu Ye, Haoyuan Liang, Shilei Cao 0005, Yushan Lai, Juepeng Zheng |
ACM Multimedia | 7 |
| 2025 | SeasonBench-EA: A Multi-Source Benchmark for Seasonal Prediction and Numerical Model Post-Processing in East AsiaabstractSeasonal-scale climate prediction plays a critical role in supporting agricultural planning, disaster prevention, and long-term decision making. In particular, reliable forecasts issued 1-6 months in advance are essential for early warning of flood and drought risks associated with precipitation during the East Asian summer monsoon season. However, while the use of machine learning techniques has advanced rapidly in weather and subseasonal-to-seasonal forecasting, partly driven by the availability of benchmark datasets, their application to seasonal-scale prediction remains limited. Existing seasonal prediction primarily relies on ensemble forecasts from numerical models, which, while physically grounded, are subject to biases and uncertainties at long lead times. Motivated by these challenges, we propose SeasonBench-EA, a benchmark dataset for seasonal prediction in East Asia region. It features multi-resolution, multi-source data with both regional and global coverage, integrating ERA5 reanalysis data and ensemble forecasts from multiple leading forecast centers. Beyond key atmospheric fields, the dataset also includes boundary-related variables, such as ocean state, soil and solar radiation, that are essential for capturing seasonal-scale atmospheric variability. Two tasks are defined and evaluated: 1) machine learning-based seasonal prediction using ERA5 reanalysis, and 2) post-processing of seasonal forecasts from numerical model ensembles. A suite of deterministic and probabilistic metrics is provided for tasks evaluation, along with a hindcast assessment focused on precipitation during the East Asian summer monsoon, aligned with model evaluation protocols used in operations. By offering a unified data and evaluation framework, SeasonBench-EA aims to promote the development and application of data-driven methods for seasonal prediction, a challenging yet highly impactful task with board implications for society and public well-being. Our benchmark is available at https://github.com/SauryChen/SeasonBench-EA. Mengxuan Chen, Ziheng Zou, Jinxiao Zhang, Runmin Dong, Juepeng Zheng, Haohuan Fu |
NeurIPS | 7 |
| 2025 | DiffLiG: Diffusion-enhanced Liquid Graph with Attention Propagation for Grid-to-Station Precipitation CorrectionabstractModern precipitation forecasting systems, including reanalysis datasets, numerical models, and AI-based approaches, typically produce coarse-resolution gridded outputs. The process of converting these outputs to station-level predictions often introduces substantial spatial biases relative to station-level observations, especially in complex terrains or under extreme conditions. These biases stem from two core challenges: (i) $\textbf{station-level heterogeneity}$, with site-specific temporal and spatial dynamics; and (ii) $\textbf{oversmoothing}$, which blurs fine-scale variability in graph-based models. To address these issues, we propose $\textbf{DiffLiG}$ ($\underline{Diff}$usion-enhanced $\underline{Li}$quid $\underline{G}$raph with Attention Propagation), a graph neural network designed for precise spatial correction from gridded forecasts to station observations. DiffLiG integrates a GeoLiquidNet that adapts temporal encoding via site-aware OU dynamics, a graph neural network with a dynamic edge modulator that learns spatially adaptive connectivity, and a Probabilistic Diffusion Selector that generates and refines ensemble forecasts to mitigate oversmoothing. Experiments across multiple datasets show that DiffLiG consistently outperforms other methods, delivering more accurate and robust corrections across diverse geographic and climatic settings. Moreover, it achieves notable gains on other key meteorological variables, underscoring its generalizability and practical utility. Mengxuan Chen, Haohuan Fu, Juepeng Zheng |
NeurIPS | 8 |
| 2025 | Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMindabstractLarge Multimodal Models (LMMs) has demonstrated capabilities across various domains, but comprehensive benchmarks for agricultural remote sensing (RS) remain scarce. Existing benchmarks designed for agricultural RS scenarios exhibit notable limitations, primarily in terms of insufficient scene diversity in the dataset and oversimplified task design. To bridge this gap, we introduce AgroMind, a comprehensive agricultural remote sensing benchmark covering four task dimensions: spatial perception, object understanding, scene understanding, and scene reasoning, with a total of 13 task types, ranging from crop identification and health monitoring to environmental analysis. We curate a high-quality evaluation set by integrating nine public datasets and one private global parcel dataset, containing 28,482 QA pairs and 20,850 images. The pipeline begins with multi-source data pre-processing, including collection, format standardization, and annotation refinement. We then generate a diverse set of agriculturally relevant questions through the systematic definition of tasks. Finally, we employ LMMs for inference, generating responses, and performing detailed examinations. We evaluated 20 open-source LMMs and 4 closed-source models on AgroMind. Experiments reveal significant performance gaps, particularly in spatial reasoning and fine-grained recognition, it is notable that human performance lags behind several leading LMMs. By establishing a standardized evaluation framework for agricultural RS, AgroMind reveals the limitations of LMMs in domain knowledge and highlights critical challenges for future work. Data and code can be accessed at https://rssysu.github.io/AgroMind/. Qingmei Li, Zurong Mai, Shuohong Lou, Henglian Huang, Jiarui Zhang 0008, Yibin Wen, Haohuan Fu, Jianxi Huang, Juepeng Zheng |
NeurIPS | 13 |
| 2025 | SPFL: Sequential updates with Parallel aggregation for Enhanced Federated Learning under Category and Domain ShiftsabstractFederated learning (FL) has recently emerged as the primary approach to overcoming data silos,
enabling collaborative model training without sharing sensitive or proprietary data.
Parallel federated learning (PFL) aggregates models trained independently on each client’s local data, which can lead to suboptimal convergence due to limited data exposure.
In contrast, Sequential Federated Learning (SFL) allows models to traverse client datasets sequentially, enhancing data utilization.
However, SFL effectiveness is limited in real-world non-IID scenarios characterized by category shift (inconsistent class distributions) and domain shift (distribution discrepancies).
These shifts cause two critical issues: update order sensitivity, where model performance varies significantly with the sequence of client updates, and catastrophic forgetting, where the model forgets previously learned features when trained on new client data.
We propose SPFL, a novel updating method that can be integrated into existing FL methods, integrating sequential updates with parallel aggregation to enhance data utilization and ease update order sensitivity. At the same time, we give the convergence analysis of SPFL under strong convex, general convex, and non-convex conditions, proving that this update scheme is significantly better than PFL and SFL.
Additionally, we introduce the Global-Local Alignment Module to mitigate catastrophic forgetting by aligning the predictions of the global model with those of the local and previous models during training.
Our extensive experiments demonstrate that integrating SPFL into existing PFL methods significantly improves performance under category and domain shifts. Haoyuan Liang, Shilei Cao 0005, Zhiyu Ye, Haohuan Fu, Juepeng Zheng |
NeurIPS | 6 |
| 2025 | GTPBD: A Fine-Grained Global Terraced Parcel and Boundary DatasetabstractAgricultural parcels serve as basic units for conducting agricultural practices and applications, which is vital for land ownership registration, food security assessment, soil erosion monitoring, etc. However, existing agriculture parcel extraction studies only focus on mid-resolution mapping or regular plain farmlands while lacking representation of complex terraced terrains due to the demands of precision agriculture. In this paper, we introduce a more fine-grained terraced parcel dataset named GTPBD (Global Terraced Parcel and Boundary Dataset), which is the first fine-grained dataset covering major worldwide terraced regions with more than 200,000 complex terraced parcels with manually annotation. GTPBD comprises 47,537 high-resolution images with three-level labels, including pixel-level boundary labels, mask labels, and parcel labels. It covers seven major geographic zones in China and transcontinental climatic regions around the world. Compared to the existing datasets, the GTPBD dataset brings considerable challenges due to the: (1) terrain diversity; (2) complex and irregular parcel objects; and (3) multiple domain styles. Our proposed GTPBD dataset is suitable for four different tasks, including semantic segmentation, edge detection, terraced parcel extraction and unsupervised domain adaptation (UDA) tasks. Accordingly, we benchmark the GTPBD dataset on eight semantic segmentation methods, four edge extraction methods, three parcel extraction methods and five UDA methods, along with a multi-dimensional evaluation framework integrating pixel-level and object-level metrics. GTPBD fills a critical gap in terraced remote sensing research, providing a basic infrastructure for fine-grained agricultural terrain analysis and cross-scenario knowledge transfer. The code and data are available at https://github.com/Z-ZW-WXQ/GTPBD/. Yibin Wen, Shuai Yuan 0005, Haohuan Fu, Jianxi Huang, Juepeng Zheng |
NeurIPS | 7 |
| 2025 | BAN: A Universal Paradigm for Cross-Scene Classification Under Noisy Annotations From RGB and Hyperspectral Remote Sensing ImagesabstractWhile domain adaptation (DA) methods have made significant strides in remote sensing community, most current works assume that the source domain labels are accurate. However, limited emphasis has been placed on the scenario where source data are mislabeled with noisy annotations, which is more common in real applications and referred to as noisy DA (NDA). This article formulates remote sensing cross-scene classification on NDA scenarios and proposes a novel network called bilateral adaptation network (BAN), which consists of two parts: 1) forward learning (FL), which utilizes a model learning from the noisy source domain and transfers knowledge to target domain; and 2) backward learning (BL), which utilizes a dual model to acquire knowledge from the target domain and transfer it to source domain. We conduct two parts alternately and adopt a symmetrical Kullback-Leibler (KL) loss to align predictions of the model and its dual model in the same domain. This interactive strategy is able to explore bilateral relationships between domains, implicitly reducing label noise in the source domain. In addition, BAN could serve as a universal paradigm to not only improve the existing NDA methods but also enhance recent DA approaches. Comprehensive evaluations on three publicly available RGB-band remote sensing datasets and two hyperspectral datasets validate the superior effectiveness of our proposed BAN. BAN improves the average accuracy by 6.70%–15.70% on RGB datasets and overall accuracy (OA) by 1.36%–3.14% on hyperspectral datasets with flip-20% noise compared to other state-of-the-art DA and NDA approaches. Promising results indicate the potential of our approach in tackling more general and practical problems with noisy source domain. Wentang Chen, Yibin Wen, Juepeng Zheng, Jianxi Huang, Haohuan Fu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Boosting Universal Domain Adaptation in Remote Sensing With Dual-Classifiers Consistency Discrimination and Cross-Domain Feature MixupabstractIn the field of remote sensing image classification, domain adaptation (DA) methods have been extensively utilized to overcome the challenges posed by data discrepancies between source and target domains that arise from varying imaging conditions, sensor differences, or geographical variations. Stemming from the existence of unseen classes in both the source and target domains, universal DA (UniDA) poses the greatest challenge that demands innovative solutions. Existing UniDA methods often overlook intra-domain variations within the target domain and face difficulties in distinguishing between similar known and unknown classes, which significantly hinder cross-domain transfer. To overcome these challenges, we propose a dual-classifier network tailored for cross-domain classification of remote sensing images, namedDCmix. DCmix introduces a dual-classifiers network that utilizes both closed-set and open-set classifiers to improve the accuracy of identifying unknown sample classes. To our knowledge, this is the first attempt to introduce dual classifiers into the UniDA remote sensing image classification task. We further enhance the feature generalization capability of the target domain based on sample neighborhood relations, resulting in a more adaptable and robust feature representation. A cross-domain feature mixup scheme is also designed based on the consistency discrimination of the dual classifiers, achieving smoother decision boundaries and simpler hidden layer representations. Extensive experiments conducted on four hyperspectral image datasets and three RGB datasets prove that the introduced approach attains state-of-the-art performance in remote sensing image classification under the UniDA scenario. Qingmei Li, Juepeng Zheng, Jianxi Huang, Haohuan Fu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Building Bridges Across Spatial and Temporal Resolutions: Reference-Based Super-Resolution via Change Priors and Conditional Diffusion ModelabstractReference-based super-resolution (RefSR) has the potential to build bridges across spatial and temporal resolutions of remote sensing images. However, existing RefSR methods are limited by the faithfulness of content reconstruction and the effectiveness of texture transfer in large scaling factors. Conditional diffusion models have opened up new opportunities for generating realistic high-resolution images, but effectively utilizing reference images within these models remains an area for further exploration. Furthermore, content fidelity is difficult to guarantee in areas without relevant reference information. To solve these issues, we propose a change-aware diffusion model named Ref-Diff for RefSR, using the land cover change priors to guide the denoising process explicitly. Specifically, we inject the priors into the denoising model to improve the utilization of reference information in unchanged areas and regulate the reconstruction of semantically relevant content in changed areas. With this powerful guidance, we decouple the semantics-guided denoising and reference texture-guided denoising processes to improve the model performance. Extensive experiments demonstrate the superior effectiveness and robustness of the proposed method compared with state-of-the-art RefSR methods in both quantitative and qualitative evaluations. The code and data are available at https://github.com/dongrunmin/RefDiff. Runmin Dong, Shuai Yuan 0005, Mengxuan Chen, Jinxiao Zhang, Lixian Zhang 0002, Juepeng Zheng, Haohuan Fu |
CVPR | 8 |
| 2024 | 3D Building Reconstruction from Monocular Remote Sensing Images with Multi-level Supervisionsabstract3D building reconstruction from monocular remote sensing images is an important and challenging research problem that has received increasing attention in recent years, owing to its low cost of data acquisition and availability for large-scale applications. However, existing methods rely on expensive 3D-annotated samples for fully-supervised training, restricting their application to large-scale cross-city scenarios. In this work, we propose MLS-BRN, a multi-level supervised building reconstruction network that can flexibly utilize training samples with different annotation levels to achieve better reconstruction results in an end-to-end manner. To alleviate the demand on full 3D supervision, we design two new modules, Pseudo Building Bbox Calculator and Roof-Offset guided Footprint Extractor, as well as new tasks and training strategies for different types of samples. Experimental results on several public and new datasets demonstrate that our proposed MLS-BRN achieves competitive performance using much fewer 3D-annotated samples, and significantly improves the footprint extraction and 3D reconstruction performance compared with current state-of-the-art. The code and datasets of this work will be released at https://github.com/opendatalab/MLS-BRN.git. Haote Yang, Zhenghao Hu, Juepeng Zheng, Gui-Song Xia, Conghui He |
CVPR | 4 |
| 2024 | Single Free-Hand Sketch Guided Free-Form Deformation For 3D Shape GenerationabstractSketch-guided point cloud reconstruction aims to provide an efficient and flexible pathway to generate plausible 3D shapes automatically from a free-hand sketch shaping the modeling intentions of end users. However, such a task is still in its infancy due to the complex, challenging, and highly variable patterns of sketches by nature. In this paper, we present a novel sketch-guided framework based on free-form deformation (FFD) for 3D point cloud generation. To capture sufficient meaningful features from a sketch, a dual-branch encoding architecture is devised to extract complementary semantic and geometric clues by formulating the input as a binary image and a 2D point cloud, respectively. The proposed encoder also learns useful features to guide content generation from a template point cloud before decoding the resultant global features into a set of control points for the use of FFD. We have also developed a large and diverse manually collected dataset, Sketch-3DPC, in which there are a total of 13,754 sketch and 3D point cloud pairs categorized into 11 classes. Both qualitative and quantitative experiment results demonstrate the superiority of the proposed methodology and dataset. Fei Wang 0056, Jianqiang Sheng, Zhineng Zhang, Juepeng Zheng, Baoquan Zhao |
ICME | 5 |
| 2024 | Universal Domain Adaptation for Hyperspectral Image ClassificationabstractAlthough enormous Domain Adaptation (DA) approaches have been proposed for cross-scene hyperspectral image (HSI) classification, most of them strongly rely on prior knowledge of the relationship between the label sets of source and target domains (including closed-set, partial and open-set DA), which significantly limits their applications. In a real-world application scenario, we often transfer knowledge between domains without any constraints on the label sets, which is called Universal Domain Adaptation (UniDA). In this paper, we propose HyUniDA, which is the first attempt to address UniDA scenario from HSIs. HyUniDA contains two major parts: the Shared Semantic Pairing (SSP) and Domain Similarity Score (DSS). The SSP identifies pairs of clusters that have coincident semantic features as the common classes. By examining the consistency level of samples across source and target domains, DSS can estimate the number of target clusters and generate distinct clusters without prior knowledge. We evaluate our proposed method on two transfer learning tasks for four typical HSI datasets, it turns out that our proposed method yields 6.41%∼34.71% improvements compared to other state-of-the-art DA methods. Qingmei Li, Yibin Wen, Juepeng Zheng, Haohuan Fu |
IGARSS | 3 |
| 2024 | Unveiling Annual Dynamics in Large-Scale Road Networks Through a Connectivity-Aware Approach Utilizing Sentinel-2 Multi-Spectral ImageryabstractEfficient and timely assessment of road network dynamic changes is crucial for the comprehensive evaluation of urban development, transportation accessibility, and environmental impacts. While existing methods mainly focus on optimizing performance with very-high-resolution remote sensing images on public road datasets, their practical applicability remains to unlock when confronted with large-scale real-world applications utilizing multi-spectral remote sensing images. The limitations manifest in unsatisfactory model generalization and fragmented segmentation, reducing the effectiveness of road extraction outcomes. To address these challenges, this study introduces a novel connectivity-aware approach tailored to address road extraction challenges in real-world scenarios. Leveraging Sentinel-2 multi-spectral imagery, this study conducts a 6-year road change mapping over an expansive 10,097 square kilometers in Xi’an, China. Experimental and evaluation results underscore the efficacy of the proposed methodology for widespread applications in urban planning and environmental management, offering a robust solution for practical and efficient road extraction in diverse and extensive large-scale urban investigation. Lixian Zhang 0002, Kangrui Du, Shuai Yuan 0005, Runmin Dong, Juepeng Zheng, Haohuan Fu |
IGARSS | 5 |
| 2024 | Single Domain Generalization For Scene Classification Using Style-Oriented Data AugmentationabstractDomain generalization (DG), which tries to improve the performance of models trained with known domains but applied to unknown domains, is an important step towards practical solutions in real-world scenarios. In this paper, we tackle a much more difficult scenario called single domain generation in scene classification problem, where only one source domain is available during training. Existing DG methods usually focus on extracting invariant features from different known domains and often suffer from overfitting issues. Therefore, to tackle the above challenge, we propose a randomly-stylized data augmentation method, which enables randomized style perturbation of the training data, to alleviate the overfitting problem and to improve the robustness of the resulting model. On a multidomain scene classification benchmark, our method achieves an accuracy improvement of 0.4%-2.5% compared to other DG methods. Yi Zhao 0024, Guancong Lin, Juepeng Zheng, Yang You 0001, Haohuan Fu |
IGARSS | 3 |
| 2024 | DeepLight: Reconstructing High-Resolution Observations of Nighttime Light With Multi-Modal Remote Sensing Data
Lixian Zhang 0002, Runmin Dong, Shuai Yuan 0005, Jinxiao Zhang, Mengxuan Chen, Juepeng Zheng, Haohuan Fu |
IJCAI | 6 |
| 2024 | Circular Reconfigurable Parallel Processor for Edge Computing : Industrial Product ✶abstractGraphics Processing Units (GPUs) have emerged as the predominant hardware platforms for massively parallel computing. However, their inherent von-Neumann architecture still suffers performance inefficiency stemming from the sequential instruction execution and frequent data transfer overheads within the memory system. These intrinsic architectural flaws lead to heavy overhead on the latency, area, and energy efficiency, rendering GPUs suboptimal for edge computing applications. To tackle these challenges, this paper introduces a novel circular Reconfigurable Parallel Processor (RPP) to enable massively parallel applications in edge computing with high efficiency. RPP features a novel circular array of reconfigurable compute engines, enabling efficient streaming dataflow processing. In contrast to traditional Coarse Grained Reconfigurable Architecture (CGRA), the circular network topology of RPP is formed by linear switch networks with an innovative gasket memory, which reduces complicated network routing overheads while allowing versatile datapath mapping and optimized data reuse. A dedicated hierarchical memory system is proposed to support different memory access patterns and address mapping strategies, enabling flexible data access with high memory efficiency. Several hardware optimizations are further introduced to improve hardware utilization and performance such as concurrent kernel execution, register split&refill and heterogeneous scalar&vector computing. To fully utilize the hardware capability of RPP, we develop an end-to-end software stack consisting of a compiler, runtime environment, and different RPP libraries. This software stack is designed to be compatible with the GPGPU computing paradigm, enhancing its potential for broader adoption. Fabricated in a 14nm process, RPP occupies an area of 119 mm2and operates at a maximum power of 15W with a 1GHz clock frequency. From the runtime measurement of various workloads, RPP achieves up to 27.5 × higher energy efficiency than Nvidia edge GPUs in deep learning inference and up to 14062 × lower latency than AMD Ryzen 5 CPU in linear algebra operations. Jianbin Zhu, Toshio Nagata, Ryan Braidwood, Haohuan Fu, Juepeng Zheng, Wayne Luk, Hongxiang Fan |
ISCA | 8 |
| 2024 | Spatial-Temporal Context Model for Remote Sensing Imagery CompressionabstractWith the increasing spatial and temporal resolutions of obtained remote sensing (RS) images, effective compression becomes critical for storage, transmission, and large-scale in-memory processing. Although image compression methods achieve a series of breakthroughs for daily images, a straightforward application of these methods to RS domain underutilizes the properties of the RS images, such as content duplication, homogeneity, and temporal redundancy. This paper proposes a Spatial-Temporal Context model (STCM) for RS image compression, jointly leveraging context from a broader spatial scope and across different temporal images. Specifically, we propose a stacked diagonal masked module to expand the contextual reference scope, which is stackable and maintains its parallel capability. Furthermore, we propose spatial-temporal contextual adaptive coding to enable the entropy estimation to reference context across different temporal RS images at the same geographic location. Experiments show that our method outperforms previous state-of-the-art compression methods on rate-distortion (RD) performance. For downstream tasks validation, our method reduces the bitrate by 52 times for single temporal images in the scene classification task while maintaining accuracy. Jinxiao Zhang, Runmin Dong, Juepeng Zheng, Mengxuan Chen, Lixian Zhang 0002, Yi Zhao 0024, Haohuan Fu |
ACM Multimedia | 3 |
| 2024 | FUSU: A Multi-temporal-source Land Use Change Segmentation Dataset for Fine-grained Urban Semantic UnderstandingabstractFine urban change segmentation using multi-temporal remote sensing images is essential for understanding human-environment interactions in urban areas. Although there have been advances in high-quality land cover datasets that reveal the physical features of urban landscapes, the lack of fine-grained land use datasets hinders a deeper understanding of how human activities are distributed across landscapes and the impact of these activities on the environment, thus constraining proper technique development. To address this, we introduce FUSU, the first fine-grained land use change segmentation dataset for Fine-grained Urban Semantic Understanding. FUSU features the most detailed land use classification system to date, with 17 classes and 30 billion pixels of annotations. It includes bi-temporal high-resolution satellite images with 0.2-0.5 m ground sample distance and monthly optical and radar satellite time series, covering 847 km^2 across five urban areas in the southern and northern of China with different geographical features. The fine-grained land use pixel-wise annotations and high spatial-temporal resolution data provide a robust foundation for developing proper deep learning models to provide contextual insights on human activities and urbanization. To fully leverage FUSU, we propose a unified time-series architecture for both change detection and segmentation. We benchmark FUSU on various methods for several tasks. Dataset and code are available at: https://github.com/yuanshuai0914/FUSU. Shuai Yuan 0005, Guancong Lin, Lixian Zhang 0002, Runmin Dong, Jinxiao Zhang, Juepeng Zheng, Jie Wang 0036, Haohuan Fu |
NeurIPS | 7 |
| 2024 | C³DA: A Universal Domain Adaptation Method for Scene Classification From Remote Sensing ImageryabstractVarious remote sensing applications have widely used domain adaptation (DA) methods. Since it does not need to add human interpretation in the target domain, it can be used in cross-region, multi-temporal, and multi-sensor application scenarios. In order to further optimize the design of the loss function and better address the challenges of DA in remote sensing, in this paper, we propose a new universal DA method named C3DA for scene recognition of remote sensing images. It has a comprehensive C3criterion for recognizing the "unknown" classes by innovatively fusing confidence, consistency, and certainty of samples to make our network training more efficient. We evaluate the performance of our proposed method based on six transfer tasks on three remote sensing datasets. The evaluation results show that our proposed method achieves an average H-score of 58.44%, significantly higher than other SOTA universal DA methods with an average improvement of 2.32~29.43%. Compared to the baseline ResNet-50, it achieves up to 19.92% improvement, demonstrating that the proposed method outperforms in the universal DA scenario. In the future, we also plan to expand the application of this method to more scenarios. Jiaxu Guo, Yushan Lai, Jinxiao Zhang, Juepeng Zheng, Haohuan Fu, Lin Gan 0008, Liang Hu 0001, Gaochao Xu, Xilong Che |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Relational Part-Aware Learning for Complex Composite Object Detection in High-Resolution Remote Sensing ImagesabstractIn high-resolution remote sensing images (RSIs), complex composite object detection (e.g., coal-fired power plant detection and harbor detection) is challenging due to multiple discrete parts with variable layouts leading to complex weak inter-relationship and blurred boundaries, instead of a clearly defined single object. To address this issue, this article proposes an end-to-end framework, i.e., relational part-aware network (REPAN), to explore the semantic correlation and extract discriminative features among multiple parts. Specifically, we first design a part region proposal network (P-RPN) to locate discriminative yet subtle regions. With butterfly units (BFUs) embedded, feature-scale confusion problems stemming from aliasing effects can be largely alleviated. Second, a feature relation Transformer (FRT) plumbs the depths of the spatial relationships by part-and-global joint learning, exploring correlations between various parts to enhance significant part representation. Finally, a contextual detector (CD) classifies and detects parts and the whole composite object through multirelation-aware features, where part information guides to locate the whole object. We collect three remote sensing object detection datasets with four categories to evaluate our method. Consistently surpassing the performance of state-of-the-art methods, the results of extensive experiments underscore the effectiveness and superiority of our proposed method. Shuai Yuan 0005, Lixian Zhang 0002, Runmin Dong, Juepeng Zheng, Haohuan Fu, Peng Gong 0002 |
IEEE Trans. Cybern. | 5 |
| 2024 | Weakly Supervised 3-D Building Reconstruction From Monocular Remote Sensing Imagesabstract3D building reconstruction from monocular remote sensing imagery is an important research problem that has been extensively studied for several decades. Although monocular remote sensing imagery is a more economic data source compared with the LiDAR data and multi-view imagery, its limited information results in great challenges and restricts the performance of existing monocular reconstruction methods. Moreover, the expensive cost and the limited quantity of 3D annotations also restrict the application scenes of existing methods, which are mostly based on fully-supervised learning. In our previous work, we have proposed MTBR-Net, a monocular building reconstruction method that consists of a fully-supervised multi-task network and a post-processing module for optimizing the reconstruction results. In this work, we further propose WS-MTBR-Net, a weakly-supervised building reconstruction network that uses fewer 3D annotations and achieves better performance in an end-to-end manner. Specifically, our WS-MTBR-Net fully leverages the relation between different components of a 3D building instance and the property of off-nadir images to improve the footprint segmentation boundary, based on six modified tasks and a new network structure with an improved feature warping module to support weakly-supervised learning. We also design a new training strategy via a hybrid loss function that enables utilizing the training samples with different annotation levels, i.e., complete 3D annotations, 2D footprint annotations, and image-level angle annotations. Results on BONAI Shanghai and Xi’an test datasets demonstrate that our method achieves competitive performance when using 50% fewer 3D-annotated samples, and improves the footprint segmentation F1-score by around 4% compared with current state-of-the-art. Zhenghao Hu, Lingxuan Meng, Jinwang Wang, Juepeng Zheng, Runmin Dong, Conghui He, Gui-Song Xia, Haohuan Fu, Dahua Lin |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | HyUniDA: Breaking Label Set Constraints for Universal Domain Adaptation in Cross-Scene Hyperspectral Image ClassificationabstractAlthough enormous Domain Adaptation (DA) approaches have been proposed for cross-scene hyperspectral image (HSI) classification, majority DA methods strongly depend on much prior knowledge of the association among the label sets of source and target domains (encompassing closed set, partial and open set DA), thereby significantly hindering their applications. Realistic application scenarios often require knowledge transfer between domains without restrictions on the label space, which is called Universal Domain Adaptation (UniDA). In this paper, we propose HyUniDA, which is the first attempt to address UniDA scenario from HSIs. HyUniDA contains two major parts: the Shared Semantic Pairing (SSP) and Domain Similarity Score (DSS). We group both source and target domains to form discriminative clusters. The SSP identifies pairs of clusters that have coincident semantic features as the common classes. By examining the consistency level of samples across source and target domains, DSS can estimate the quantity of target clusters and generate distinct clusters without prior knowledge. Meanwhile, we apply the contrastive domain discrepancy to alleviate the offset of samples distribution, with a representative regularizer to assist distinguish target domain clusters. We evaluate our proposed method on three transfer learning tasks for six typical HSI datasets, it turns out that our proposed method yields 3.83%~37.57% improvements compared to other state-of-the-art DA methods. Qingmei Li, Yibin Wen, Juepeng Zheng, Haohuan Fu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Large-Scale Land Cover Mapping with Fine-Grained Classes via Class-Aware Semi-Supervised Semantic SegmentationabstractSemi-supervised learning has attracted increasing attention in the large-scale land cover mapping task. However, existing methods overlook the potential to alleviate the class imbalance problem by selecting a suitable set of unlabeled data. Besides, in class-imbalanced scenarios, existing pseudo-labeling methods mostly only pick confident samples, failing to exploit the hard samples during training. To tackle these issues, we propose a unified Class-Aware Semi-Supervised Semantic Segmentation framework. The proposed framework consists of three key components. To construct a better semi-supervised learning dataset, we propose a class-aware unlabeled data selection method that is more balanced towards the minority classes. Based on the built dataset with improved class balance, we propose a Class-Balanced Cross Entropy loss, jointly considering the annotation bias and the class bias to re-weight the loss in both sample and class levels to alleviate the class imbalance problem. Moreover, we propose the Class Center Contrast method to jointly utilize the labeled and unlabeled data. Specifically, we decompose the feature embedding space using the ground truth and pseudo-labels, and employ the embedding centers for hard and easy samples of each class per image in the contrast loss to exploit the hard samples during training. Compared with state-of-the-art class-balanced pseudo-labeling methods, the proposed method improves the mean accuracy and mIoU by 4.28% and 1.70%, respectively, on the large-scale Sentinel-2 dataset with 24 land cover classes. Runmin Dong, Lichao Mou, Mengxuan Chen, Xin-Yi Tong 0003, Shuai Yuan 0005, Lixian Zhang 0002, Juepeng Zheng, Xiao Xiang Zhu 0001, Haohuan Fu |
ICCV | 8 |
| 2023 | CO-Detector: Towards Complex Object Detection with Cross-Part Feature Learning in Remote SensingabstractObject detection in remote sensing imagery builds the essential foundation of aerial and satellite image understanding, being an important role in many common real-world tasks and attracting world-wide attention. In recent years, despite the great progress of common object detection in remote sensing and the proven success of deep learning in this field, yet complex object detection which consists of multiple objects with variable layouts in remote sensing (e.g., coal-fired power plant, airport, sewage treatment plant, etc.) is still challenging for complex composite spatial relationship, non-rigid boundaries, and complicated surrounding textures. These challenges necessitate developing specific complex object detection methods to learn inter-relationship and distinctive and discriminative features in complex objects. To address this problem, in this paper, we propose a method, i.e., CO-Detector, in an end-to-end manner, to achieve various complex composite object detection in remote sensing images with high accuracy and efficiency. The effectiveness of CO-Detector is built on three main parts: (a) First, as surrounding contexts are normally complicated and similar to complex objects, we propose a Tandem Attention Network (TAN), including a channel enhanced network and a spatial enhanced network, with a K-global max/average pooling, to restrain noise disturbance and highlight complex object features and boundaries. (b) Second, we design a Part Region Proposal Network (P-RPN) to learn the interrelationship between parts in one object, generating part proposals and locating discriminative and distinctive object parts finely. (c) Third, to detect the whole complex object as well as the parts, we propose a Part Detection Network (PDN) to detect the individual parts, and detect the whole object through multi-level fused features. We train our CO-Detector model with three selected categories (i.e., coal-fired power plant, airport, oil storage tank) in three datasets, and conduct comparative experiments to evaluate and verify the performance. The comprehensive experiment results show that our CO-Detector achieves a mAP of 80.23%, outperforming 4.17%-17.83% against other cutting-edge deep learning-based detection methods. The experiment results indicate our CO-Detector has promising performance and potential in various complex object detection in highresolution remote sensing images, pending to be utilized in real large-scale applications. Shuai Yuan 0005, Juepeng Zheng, Jierui Liu, Haohuan Fu, Ray C. C. Cheung |
IGARSS | 2 |
| 2023 | Achieving 10m China Land Cover Mapping within Three Minutes Using a New Sunway SupercomputerabstractLand Cover Mapping (LCM) is an important task to detect and understand the change of the earth surface. However, most LCM methods adopt supervised classifiers, and suffer from a lack of labels at a large scale. In this paper, we propose Fast-LCM, a scalable and weakly-supervised LCM method on a new Sunway supercomputer to achieve large-scale land cover mapping, requiring no manual annotations. Fast-LCM includes two major parts: (1) a distance-guided k-means module that combines textural, spectral, and temporal features, and (2) an automatic voting-based merging strategy to give each cluster a real meaning of classification system. Through careful parallelization, our Fast-LCM method scales to over 38 million cores, and provides a sustained performance for the task of China LCM. We produce a 10m resolution land cover map of China within only 3 minutes, including 1.2 minutes for IO and only 55 seconds to finish the computation. Fast-LCM achieves an accuracy of 68.85% (25-class), with 3.64% to 7.05% higher than best existing products. Juepeng Zheng, Yi Zhao 0024, Jinxiao Zhang, Wenzhao Wu, Shuai Yuan 0005, Haohuan Fu |
IGARSS | 1 |
| 2023 | SW-LCM: A Scalable and Weakly-supervised Land Cover Mapping Method on a New Sunway SupercomputerabstractHigh-resolution land cover mapping (LCM) is an important application for studying and understanding the change of the earth surface. While deep learning (DL) methods demonstrate great potential in analyzing satellite images, they largely depend on massive high-quality labels. This paper proposes SW-LCM, a Scalable and Weakly-supervised two-stage Land Cover Mapping method on a new Sunway Supercomputer. Our method consists of a k-means clustering module as a first stage, and an iterative deep learning module as a second stage. With the k-means module providing a good enough starting point (taking inaccurate results as noisy labels), the deep learning module improves the classification results in an iterative way, without any labelling efforts required for processing large scenarios. To achieve efficiency for country-level land cover mapping, we design a customized data partition scheme and an on-the-fly assembly for k-means. Through careful parallelization and optimization, our k-means module scales to 98,304 computing nodes (over 38 million cores), and provides a sustained performance of 437.56 PFLOPS, in a real LCM task of the entire region of China; the iterative updating part scales to 24,576 nodes, with a performance of 11 PFLOPS. We produce a 10-m resolution land cover map of China, with an accuracy of 83.5% (10-class) or 73.2% (25-class), 7% to 8% higher than best existing products, paving ways for finer land surveys to support sustainability-related applications. Yi Zhao 0024, Juepeng Zheng, Haohuan Fu, Wenzhao Wu, Mengxuan Chen, Jinxiao Zhang, Lixian Zhang 0002, Runmin Dong, Zhenrong Du, Xin Liu 0081, Shaoqing Zhang, Le Yu 0001 |
IPDPS | 2 |
| 2023 | Partial Domain Adaptation for Scene Classification From Remote Sensing ImageryabstractAlthough domain adaptation approaches have been proposed to tackle cross-regional, multitemporal, and multisensor remote sensing applications since they do not require any human interpretation in the target domain, most current works assume identical label space across the source and the target domains. However, in real-world applications, we often transfer knowledge from a large-scale dataset with rich annotations to a small-scale target dataset with scarcity of labels. In most cases, the label space of the source domain is usually large enough to subsume that of the target domain, which is termed partial domain adaptation. In this article, we propose a new partial domain adaptation algorithm for remote sensing scene classification and our proposed method contains three major parts. First, we employ a progressive auxiliary domain module to alleviate the negative transfer effect caused by outlier classes. Second, we adopt an improved domain adversarial neural network (DANN) with multiweights to better encourage domain confusion. Last but not least, we design an attentive complement entropy regularization to improve the prediction confidence for samples and avoid untransferable samples (such as the samples belonging to outlier classes in the source domain) being mistakenly classified. We collect three common remote sensing datasets to evaluate our proposed method. Our method achieves an average accuracy of 79.36%, which considerably outperforms other state-of-the-art partial domain adaptation methods with an average accuracy improvement of 1.90%–12.45% and attaining a 13.67% gain compared to the straightforward deep learning model (ResNet-50). The experiment results indicate that our approach shows promising prospects for solving more general and practical domain adaptation problems where the label space of the source domain subsumes that of the target domain. Juepeng Zheng, Yi Zhao 0024, Wenzhao Wu, Mengxuan Chen, Haohuan Fu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Melting Glacier: A 37-Year (1984-2020) High-Resolution Glacier-Cover Record of MT. KilimanjaroabstractCommonly recognized as an important symbol of the tropics and global warming, the glacier loss on Mt. Kilimanjaro has received worldwide attention for decades. In this paper, we propose a high-resolution glacier-cover (GC) record of Mt. Kilimanjaro over the period from 1984 to 2020, using a novel deep learning-based semantic segmentation method and Google Earth images, as well as digital elevation model (DEM) and ERA5-Land (ERA5) for snowline and temperature variations analysis. Our method achieves an accuracy of 94.37%, which proves the model's capability to record the GC areas precisely. The results show that (1) the GC area dramatically decreases from 19.2 km2to 3.6 km2during 37 years, which decreases about 4% and 2% per year from 1984 to 2000 and from 2000 to 2020 respectively, (2) the snowline altitude rises from$4,651 m$to$5,088 m$by about$437 m$, and (3) the average$5,000 m$air temperature on Mt. Kilimanjaro increases from −2.1 °C to −1.1 °C by about 1 °C. This study indicates that there will be no GC within a few decades if the current loss continues. Shuai Yuan 0005, Juepeng Zheng, Lixian Zhang 0002, Runmin Dong, Yile Xing, Yuhan She, Haohuan Fu, Ray C. C. Cheung |
IGARSS | 2 |
| 2022 | A Parallel Approach for Oil Palm Tree Detection on a SW26010 Many-Core ProcessorabstractCounting and detecting oil palm trees from high-resolution remotely sensed images is a significant work for improving economy of several countries such as Malaysia, Indonesia, etc. However, rare attention have been paid on accelerating tree crown detection algorithms on high performance platforms. In this paper, we design a parallel approach for oil palm tree detection on a SW26010 many-core processor, which is used in a world-leading supercomputer, Sunway TaihuLight. Our parallel framework contains three steps, local maximum filtering, oil palm tree crown center reassignment and oil palm tree crown center merging. Experimental results indicates that our parallel approach of oil palm tree detection obtains the speedup of 32.30 times and 1.74 times for a QuickBird image with a size of$12,188\times 12,576$pixels compared with the well-optimized software implementation of the original algorithm on an Intel 12-core CPU and FPGAs. Juepeng Zheng, Wenzhao Wu, Yi Zhao 0024, Shuai Yuan 0005, Runmin Dong, Lixian Zhang 0002, Haohuan Fu |
IGARSS | 1 |
| 2022 | Multisource-Domain Generalization-Based Oil Palm Tree Detection Using Very-High-Resolution (VHR) Satellite ImagesabstractProviding accurate and timely oil palm information on a large scale is essential for both economic development and ecological significance. However, owing to different sensors, photograph acquisition conditions, and environmental heterogeneity, the large volume and the variety of the data make it extremely challenging for large-scale and cross-regional oil palm tree detection. It is computationally expensive to train a model from images covering large heterogeneous regions and all environmental conditions for continuously accumulated multisource remote sensing data. In this letter, we propose a new multisource domain generalization (DG) method, Maximum Mean Discrepancy Deep Reconstruction Classification Network (MMD-DRCN). It learns representations from multiple source domains and obtains inspiring performance in an unknown and “unseen” target domain. Besides classification loss, our MMD-DRCN distills more representative features through reconstruction loss and aligns multisource latent features by MMD loss, both of which effectively enhance the capacity of generalization. MMD-DRCN achieves an average F1-score of 82.70% in all transfer tasks, attaining a 5.83% gain compared to Baseline (a straightforward convolutional neural network (CNN) model). Experimental results demonstrate DG poses a promising potential for large-scale and cross-regional oil palm tree detection without any information of the target domain. Juepeng Zheng, Wenzhao Wu, Shuai Yuan 0005, Haohuan Fu, Le Yu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | A Two-Stage Adaptation Network (TSAN) for Remote Sensing Scene Classification in Single-Source-Mixed-Multiple-Target Domain Adaptation (S²M²T DA) ScenariosabstractOver the past decade, domain adaptation (DA) algorithms have been proposed to address domain gap problems as they do not need any interpretation in the target domain. However, most existing efforts focus on scenarios with only one source domain and one target domain. In this article, we explore the scenario with one source domain and mixed multiple target domains for remote sensing applications and propose a new algorithm, named the two-stage adaptation network (TSAN). First, we utilize the adversarial learning approach to confuse the classifier to discriminate between the source domain and the whole mixed-multiple-target domain. Second, we adopt self-supervised learning to divide the mixed-multiple-target domain with automated generation of “pseudo”-domain labels, which guides our network to learn intrinsic features of multiple target domains. Finally, these two steps are combined as an iterative procedure. We integrate a test dataset that includes five remote sensing datasets and ten classes. Our method achieves an average accuracy of 63.25% and 73.68% with two typical backbones, considerably outperforming other DA methods with an average accuracy improvement of 4.84%–20.19% and 9.06%–17.04%, respectively. Furthermore, we identify the negative transfer effect in existing mainstream DA methods in remote sensing image classification with multiple different domains. Juepeng Zheng, Wenzhao Wu, Shuai Yuan 0005, Yi Zhao 0024, Lixian Zhang 0002, Runmin Dong, Haohuan Fu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Transresnet: Transferable Resnet For Domain AdaptationabstractAlthough Deep Convolutional Neural Network (DCNN) has been admittedly witnessed as an enormous success in a wide range of applications, most of them require sufficient annotations with time-consuming and labor-exhausting efforts. Existing domain adaptation (DA) approaches delve into designing an effective loss module to minimize the distribution gap between the source and target domains. However, few studies pay attention to improve the backbone or network architecture for DA issues. In this paper, we propose a new backbone for DA specially, i.e., Transferable ResNet (TransResNet). TransResNet remedies the residual block in ResNet, separating source and target input features and highlighting more transferable channels in each block. It can be easily applied to all kinds of DA methods, without adding any extra learning parameters. We conduct substantial experiments on two general DA datasets and embed TransResNet into two seminal DA methods, including DANN and CDAN. Experimental results demonstrate TransResNet improves the transferability of the architecture, indicating that it is a great substitute for ResNet as a network backbone in DA issues. Juepeng Zheng, Wenzhao Wu, Yi Zhao 0024, Haohuan Fu |
ICIP | 1 |
| 2021 | Coconut Trees Detection on the Tenarunga Using High-Resolution Satellite Images and Deep LearningabstractThe Coconut tree is of great importance in economic values and ecological impacts for many tropical developing countries and lots of islands in the Pacific Ocean. Detecting and counting coconut is a meaningful and valuable research. In this paper, we present a coconut tree crown detection method to detect and count the coconut trees in the Tenarunga from high-resolution satellite images acquired by Google Earth. Our coconut tree detection method contains three major procedures: feature extraction, a multi-level Region Proposal Network (RPN) and a large-scale coconut tree detection workflow. We manually annotate all coconut trees for our study regions in the Tenarunga. Eventually, we achieve a higher average F1-score of 77.14% in our four test regions than pure Faster R-CNN. Experiment results demonstrate the potential for large-scale individual coconut tree detection and counting from high-resolution satellite images using deep learning. Juepeng Zheng, Wenzhao Wu, Le Yu 0001, Haohuan Fu |
IGARSS | 1 |
| 2020 | Unsupervised Mixed Multi-Target Domain Adaptation for Remote Sensing Images ClassificationabstractAlthough deep learning has been successfully applied in the field of remote sensing image classification, it still requires time-consuming and costly annotations. In recent years, domain adaptation has been witnessed to address this problem as they do not need any human interpreted in the target domain dataset. However, most of the existing works dedicate effort on the circumstance where there is only one source domain and only one target domain. In this paper, we firstly explore one source and multiple target domains issue for remote sensing application and build a challenging mixed multi-target dataset to contribute to the community. Our method constitutes three parts. Firstly, as we are blind for the multitarget domain, we adopt meta learning to divide the mixed multi-target dataset and insert sub-target domain loss as the part of the loss function. Secondly, we apply the adversarial learning to confuse the classifier to discriminate between the source domain images and the whole mixed multi-target domain images. Finally, the meta learning and the adversarial learning are dynamically iterative procedures and the labels for domain classification in mixed multi-target dataset will be updated for a particular iteration. Our method is well-performed in the four common remote sensing dataset (AID, NWPU-RESISC45, UC Merced and WHU-RS19), including five classes (agriculture, forest, river, residential and parking). Our method achieved an average accuracy of 81.59% and outperformed other domain adaptation method. The experiment results indicate our method is promising for large-scale, multi-regional and multi-temporal remote sensing applications. Juepeng Zheng, Wenzhao Wu, Haohuan Fu, Runmin Dong, Lixian Zhang 0002, Shuai Yuan 0005 |
IGARSS | 1 |
| 2019 | Large-Scale Oil Palm Tree Detection from High-Resolution Remote Sensing Images Using Faster-RCNNabstractOil palm is of great importance in agricultural productivity for many tropic developing countries and accordingly investigating as well as counting oil palms is a meaningful and valuable research. In this paper, we firstly apply Faster-RCNN, one of the most popular object detection algorithms, to detect tree crowns from satellite images. Although Faster-RCNN has an excellent performance in well-known datasets of general object detection, it does not have obvious advantages in oil palm tree detection in this study compared with other classical machine learning based methods. We argue two reasons accounting for the drawbacks of Faster-RCNN: (1) the size of each oil palm tree is too small (only 17 × 17 pixels on average) in 0.6m-resolution QuickBird satellite images; (2) there are lots of other similar trees around the oil palm trees that make it difficult to detect them correctly. In order to reach a satisfying accuracy, we tailored the Region Proposal Network (RPN) and proposed a simple but practical post-processing strategy based on empirical planting rules, filtering out the wrongly detected trees (False Positives) effectively. Eventually we achieved a higher average F1-score of 94.99% (using IOU based evaluation matrices) in our six study regions compared wtih state-of-the-art oil palm detection methods. In addition, we proposed a workflow of large-scale oil palm tree detection using high-resolution remotely sensed images based deep learning methods. Juepeng Zheng, Maocai Xia, Runmin Dong, Haohuan Fu, Shuai Yuan 0005 |
IGARSS | 1 |