VLDB 2026 Research / reviewers in the wild / expert
Wanxuan Lu
dblp:55/8420
· DBLP profile ↗
23ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0003-4612-508XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 1 first-author · 18 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Novel Driving Risk Assessment Method Based on Traffic Scene Understanding and Causal ReasoningabstractAccurate understanding of complex driving scenarios and the prediction of potential risks have always been critical prerequisites for autonomous vehicles to generate safe and reliable behavioral decisions. However, the uncertain states of multiple objects and their complex interactions in intricate scenes make it difficult for autonomous vehicles to extract sufficient information from representation learning to comprehend road scenarios. To address this, this paper proposes a novel method of understanding the road scene aimed at identifying potential driving risks. Firstly, we integrate the driving risk field and individual behavior features into the scene graph to enhance the dynamic characteristics of perceived targets, thereby improving the system’s sensitivity to identifying potential risks in complex scenarios. Subsequently, a causal reasoning completion module is introduced to infer and reconstruct the state of traffic entities when they are occluded or undetected, effectively enhancing the completeness of complex scene perception. Finally, we incorporate a learning framework that combines graph convolution and temporal transformers to capture node features and relationships, outputting risk prediction results. Experimental results demonstrate that the proposed method outperforms baseline approaches on both public datasets and self-constructed occlusion datasets, validating its effectiveness and superiority. Peng Ping, Qida Yao, Yeqin Shao, Yingyan Hou, Wanxuan Lu, Weiping Ding 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Positive2Negative: Breaking the Information-Lossy Barrier in Self-Supervised Single Image DenoisingabstractImage denoising enhances image quality, serving as a foundational technique across various computational photography applications. The obstacle to clean image acquisition in real scenarios necessitates the development of self-supervised image denoising methods only depending on noisy images, especially a single noisy image. Existing self-supervised image denoising paradigms (Noise2Noise and Noise2Void) rely heavily on information-lossy operations, such as downsampling and masking, culminating in low-quality denoising performance. In this paper, we propose a novel self-supervised single image denoising paradigm, Positive2Negative, to break the information-lossy barrier. Our paradigm involves two key steps: Renoised Data Construction (RDC) and Denoised Consistency Supervision (DCS). RDC renoises the predicted denoised image by the predicted noise to construct multiple noisy images, preserving all the information of the original image. DCS ensures consistency across the multiple denoised images, supervising the network to learn robust denoising. Our Positive2Negative paradigm achieves state-of-the-art performance in self-supervised single image denoising with significant speed improvements. The code is released to the public at https://github.com/Li-Tong-621/P2N. Tong Li 0016, Lizhi Wang 0001, Lin Zhu 0012, Wanxuan Lu, Hua Huang 0001 |
CVPR | 5 |
| 2025 | Complementary Advantages: Exploiting Cross-Field Frequency Correlation for NIR-Assisted Image DenoisingabstractExisting single-image denoising algorithms often struggle to restore details when dealing with complex noisy images. The introduction of near-infrared (NIR) images offers new possibilities for RGB image denoising. However, due to the inconsistency between NIR and RGB images, the existing works still struggle to balance the contributions of two fields in the process of image fusion. In response to this, in this paper, we develop a cross-field Frequency Correlation Exploiting Network (FCENet) for NIR-assisted image denoising. We first propose the frequency correlation prior based on an in-depth statistical frequency analysis of NIR-RGB image pairs. The prior reveals the complementary correlation of NIR and RGB images in the frequency domain. Leveraging frequency correlation prior, we then establish a frequency learning framework composed of Frequency Dynamic Selection Mechanism (FDSM) and Frequency Exhaustive Fusion Mechanism (FEFM). FDSM dynamically selects complementary information from NIR and RGB images in the frequency domain, and FEFM strengthens the control of common and differential features during the fusion process of NIR and RGB features. Extensive experiments on simulated and real data validate that the proposed method outperforms other state-of-the-art methods. The code will be released at https://github.com/yuchenwang815/FCENet. Lizhi Wang 0001, Lin Zhu 0012, Wanxuan Lu, Hua Huang 0001 |
CVPR | 6 |
| 2025 | Scene-Specific Multiprototype Network for Remote Sensing Scene Graph GenerationabstractRemote sensing scene graph generation aims to capture both objects and their semantic relationships, offering a comprehensive understanding of complex scenes. However, two major challenges hinder the performance of existing methods. First, remote sensing images often contain a large number of objects, many of which are unrelated. Performing global feature interactions across all objects introduces noise from irrelevant pairs, degrading feature quality. Second, relationship categories in remote sensing scenes exhibit significant intra-class variation across different contexts, and long-tailed distribution further complicates learning due to limited samples for tail classes. To address these issues, we propose the Scene-specific Multi-Prototype Network (SSMP). Our method performs contextual interactions selectively based on object and relationship categories, reducing interference from irrelevant features. Moreover, we introduce a scene-specific multi-prototype classification framework that better captures the diverse visual manifestations of each relationship class, while also improving discrimination under long-tailed distributions. Experimental results demonstrate that the proposed model achieves state-of-the-art (SOTA) performance, with a minimum improvement of 4.5% and a maximum improvement of 21.3% in mR@20 on the PredCls task over baseline models. Zhongyan Hou, Chubo Deng, Qiwei Yan, Tong Ling, Wanxuan Lu, Yingyan Hou, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Multibranch Mutual-Guiding Learning for Infrared Small Target DetectionabstractAt present, many infrared target detection approaches focus on designing modules that address the two key characteristics of targets: their weak signals and small size. However, these approaches often fail to fully leverage guided learning for weak and small target content, resulting in sub-optimal detection performance, particularly in terms of shape preservation and target positioning. To tackle this challenge, this paper proposes a multi-branch mutual-guiding learning network (MMLNet) that enhances the accuracy of infrared target detection, even in the absence of clear morphological and textural features in images. The method consists of three branches: edge, positioning, and detection, each of which is designed with a specialized module from a unique perspective. In the detection branch, we introduce a multi-dimensional lossless encoder optimized through a downsampling strategy and multi-level feature fusion to mitigate feature loss in small targets. In the positioning branch, a target positioning strategy is proposed to explicitly identify candidate targets from the image by means of a learnable multi-kernel pattern. In the edge branch, a simple architecture is adopted to enhance the ability of the model to preserve the target shape. To effectively utilize the knowledge of different branches, a mutual-guiding fusion module is developed to adjust information within and between branches. The manner adaptively utilizes the specific knowledge from each input branch. Experiment results demonstrate that the proposed method achieves comparable performance, and the visualization results show the advantages of our method in shape preservation and positioning of the targets. Our code is publicly available at https://github.com/qianngli/MMLNet. Qiang Li 0042, Wei Zhang 0250, Wanxuan Lu, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Ringmo-SenseV2: Remote Sensing Foundation Model for Spatiotemporal Prediction Based on Multisource Heterogeneous Time-Series DataabstractThe rapid development of Remote Sensing (RS) technology has generated a vast amount of heterogeneous time series data from various sources, including drone videos, satellite time-series images, and multi-object trajectories. Effectively processing and analyzing this multi-source heterogeneous data for accurate spatiotemporal prediction is crucial in fields such as environmental protection and disaster response. In this paper, we propose a universal predictive foundation model named Ringmo-SenseV2 to learn the general evolutionary patterns of RS elements from massive heterogeneous data. Ringmo- SenseV2 features a Mixture-of-Heterogeneous-Experts (MoHE) Transformer, which unifies the modeling of multi-source heterogeneous time-series data. Additionally, to better capture the complex dependencies across different spatiotemporal locations, we introduce a hypergraph translator, treating embeddings of different spatiotemporal locations as nodes and employing hypergraph convolution for information propagation. Furthermore, to enhance the model’s adaptability to different evolution speeds during pre-training, we implement the Adaptive tube Masking (AM) strategy, which controls prediction difficulty by adaptively setting mask proportions for sequences with varying evolution speeds. Extensive experiments demonstrate that Ringmo-SenseV2 exhibits outstanding performance across various RS prediction tasks. Further tests on scene graph generation for RS images showcase the model’s ability to extract image features, thereby enhancing image perception tasks. Liangyu Xu, Wanxuan Lu, Leiyi Hu, Heming Yang 0003, Chubo Deng, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | ReCon1M: A Large-Scale Benchmark Dataset for Relation Comprehension in Remote Sensing ImageryabstractScene graph generation (SGG) is a high-level visual understanding and reasoning task aimed at extracting entities (such as objects) and their interrelationships from images. Significant progress has been made in the study of SGG in natural images in recent years, but its exploration in the domain of remote sensing images remains very limited. The complex characteristics of remote sensing images necessitate higher time and manual interpretation costs for annotation compared to natural images. The lack of a large-scale public SGG benchmark is a major impediment to the advancement of SGG-related research in aerial imagery. In this article, we introduce the first publicly available large-scale, million-level relation dataset in the field of remote sensing images, which is named ReCon1M. Specifically, our dataset is built upon FAIR1M and comprises 22 262 images. It includes annotations for 873 761 object bounding boxes across 60 categories and 1 052 223 relation triplets across 59 categories based on these bounding boxes. We provide a detailed description of the dataset’s characteristics and statistical information. In addition, an efficient global context-aware network (EGCAN) is proposed to improve inference efficiency in dense relation prediction through an object-pair pre-screening mechanism. By integrating visual, spatial, and semantic features, EGCAN captures fine-grained pairwise features and object-level contextual information to enhance its ability to discriminate relation. We conduct two object detection tasks and three subtasks within SGG on this dataset, assessing the performance of mainstream methods on these tasks. The experimental results show that the proposed EGCAN achieves state-of-the-art (SOTA) performance in 17 out of 24 accuracy metrics across three tasks and delivers the best performance in frames per second (FPS) for model inference. The ReCon1M dataset and related resources are available athttps://recon1m-dataset.github.io/ Qiwei Yan, Chubo Deng, Zhongyan Hou, Wanxuan Lu, Fanglong Yao, Lingxiang Hao, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | SoPerModel: Leveraging Social Perception for Multi-Agent Trajectory PredictionabstractTrajectory prediction is an essential task within various automation systems. Recent studies have highlighted that the social interactions among multiple agents are crucial for accurate predictions, relying on empirically derived human-imposed constraints to model these interactions. However, from a sociological perspective, agents’ interactions exhibit significant inherent randomness. Dependence on a priori knowledge may lead to biased estimations of data distributions across different scenarios, failing to account for this randomness. Consequently, such methodologies often do not comprehensively capture the full spectrum of social influences, thus limiting the models’ predictive efficacy. To address these issues, we propose a novel multi-agent trajectory prediction framework, SoPerModel, which incorporates a freeform social evolution module (FSEM) and a local perception attention mechanism (LPA). The FSEM enables SoPerModel to naturally capture representative social interactions among agents without the reliance on additional human-derived priors. Through LPA, the model integrates both local and global social interaction information and leverages them to enhance trajectory prediction performance. Our framework is empirically evaluated on real-world trajectory prediction datasets, and the results demonstrate that our approach achieves a highly competitive performance compared with state-of-the-art models. Heming Yang 0003, Changyuan Tian 0001, Wanxuan Lu, Chubo Deng, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | TEA: A Training-Efficient Adapting Framework for Tuning Foundation Models in Remote SensingabstractWith well-pretrained foundation models (FMs), the performance of almost every remote sensing interpretation task has been boosted. The parameter volume of FMs increases with their continuously enhanced capabilities, leading to increased costs of fine-tuning. To apply FMs more effectively and efficiently, there are already some arts that introduce the parameter-efficient fine-tuning (PEFT) concept into remote sensing and achieve competitive performance with much lower parameter cost. However, the training efficiency of most PEFT frameworks may be not satisfactory. To make tuning FMs for remote sensing applications more efficient, we propose a training-efficient adapting (TEA) framework. Specifically, we attach a SIDE adapter network (SIDEAN) to the frozen powerful FMs and only update the SIDEAN to perform the downstream tasks. Moreover, to make TEA perceive remote sensing scenes from a macroscopic perspective and boost the performance, we propose a top-down guidance mechanism to inject macro scene information into the SIDEAN during adapting. TEA is also parameter-efficient, as SIDEAN is designed to be lightweight. We conduct extensive experiments to demonstrate the effectiveness and efficiency of TEA on ten widely adopted datasets covering four primary remote sensing tasks, e.g., object detection, orientated object detection, semantic segmentation, and scene classification. By training only 5.43% of the frozen FM parameters, TEA can save more than 57% of training memory footprint and up to 15% of time cost on average while achieving competitive performance on all datasets. Furthermore, TEA can surpass full fine-tuning on several datasets. Leiyi Hu, Wanxuan Lu, Dongshuo Yin, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | AiRs: Adapter in Remote Sensing for Parameter-Efficient Transfer LearningabstractRemote sensing is stepping into the era of the foundation model, where the fine-tuning paradigm is widely adopted to transfer the profound knowledge of pretrained foundation models to downstream tasks. However, the full fine-tuning method would become inefficient in terms of training and storage, as the foundation models are getting larger and larger. Recently, a lot of deep learning research has proposed various parameter-efficient fine-tuning (PEFT) methods that perform well with a few trainable parameters. However, most of them focus on fine-tuning general foundation models without considering the special properties of remote sensing. In this article, we propose an adapter in remote sensing (AiRs) to fine-tune large foundation models for remote sensing downstream tasks by introducing the adapter-tuning framework. Specifically, we construct AiRs from two aspects: more expressive adaptation modules and a more efficient integration strategy. Specialized adaptation modules are applied to different functional layers in AiRs, which encode the inductive bias of remote sensing images and enhance the semantic concepts of geography. Moreover, AiRs establishes pathways between trainable modules with residual connections, which reduces training difficulty and improves performance. We conduct extensive experiments on object detection, semantic segmentation, and scene classification tasks. By training only 4.4% parameters of the pretrained backbone, AiRs surpasses the previous state-of-the-art (SOTA) PEFT competitors on all experimental datasets and outperforms the full fine-tuning on six out of ten datasets. Leiyi Hu, Wanxuan Lu, Dongshuo Yin, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Dynamic and Adaptive Self-Training for Semi-Supervised Remote Sensing Image Semantic SegmentationabstractRemote sensing technology has made remarkable progress, providing a wealth of data for various applications, such as ecological conservation and urban planning. However, the meticulous annotation of this data is labor-intensive, leading to a shortage of labeled data, particularly in tasks like semantic segmentation. Semi-supervised methods, combining consistency regularization with self-training, offer a solution to efficiently utilize labeled and unlabeled data. However, these methods encounter challenges due to imbalanced data ratios. To tackle these challenges, we introduce a self-training approach namedDAST(Dynamic andAdaptiveSelf-Training), which is combined with dynamic pseudo-label sampling, distribution matching, and adaptive threshold updating. Dynamic pseudo-label sampling is tailored to address the issue of class distribution imbalance by giving priority to classes with fewer samples. Meanwhile, distribution matching and adaptive threshold updating aim to reduce distribution disparities by adjusting model predictions across augmented images within the framework of consistency regularization, ensuring they align with the actual data distribution. Experiment results on the Potsdam and iSAID datasets demonstrate that DAST effectively balances class distribution, aligns model predictions with data distribution, and stabilizes pseudo-labels, leading to state-of-the-art performance on both datasets. These findings highlight the potential of DAST in overcoming the challenges associated with significant disparities in labeled-to-unlabeled data ratios. Jidong Jin, Wanxuan Lu, Xuee Rong, Xian Sun 0001, Yirong Wu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Attention-Based Contrastive Learning for Few-Shot Remote Sensing Image ClassificationabstractFew-shot remote sensing image classification entails identifying images using a limited set of labeled data within remote sensing scenes, holding significant theoretical and practical implications. However, owing to the intricacy and variety of remote sensing images, traditional classification methods usually struggle to extract effective features and learn robust classifiers. To address this issue, an end-to-end metric learning framework named Attention-based Contrastive Learning Network is introduced in this paper. Specifically, the Attention-based Feature Optimization (ABFO) module is employed to align and enhance target image features, highlighting the target region and strengthening the network’s feature extraction capability. Additionally, the Dictionary-based Contrastive Loss (DBCL) module is assigned to optimize image feature vectors, improving category distinguishability and consequently enhancing classification accuracy. The experimental results on five publicly available Few-shot remote sensing classification datasets demonstrate the high competitiveness of our proposed method. Furthermore, it illustrates superior classification accuracy compared to other pertinent Few-shot learning algorithms in the 5-way 1-shot scenario. Hanbo Bi, Wanxuan Lu, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | TAFormer: A Unified Target-Aware Transformer for Video and Motion Joint Prediction in Aerial ScenesabstractAs drone technology advances, using unmanned aerial vehicles for aerial surveys has become the dominant trend in modern low-altitude remote sensing. The surge in aerial video data necessitates accurate prediction for future scenarios and motion states of the interested target, particularly in applications like traffic management and disaster response. Existing video prediction methods focus solely on predicting future scenes (video frames), suffering from the neglect of explicitly modeling target’s motion states, which is crucial for aerial video interpretation. To address this issue, we introduce a novel task called Target-Aware Aerial Video Prediction, aiming to simultaneously predict future scenes and motion states of the target. Further, we design a model specifically for this task, named TAFormer, which provides a unified modeling approach for both video and target motion states. Specifically, we introduce Spatiotemporal Attention (STA), which decouples the learning of video dynamics into spatial static attention and temporal dynamic attention, effectively modeling the scene appearance and motion. Additionally, we design an Information Sharing Mechanism (ISM), which elegantly unifies the modeling of video and target motion by facilitating information interaction through two sets of messenger tokens. Moreover, to alleviate the difficulty of distinguishing targets in blurry predictions, we introduce Target-Sensitive Gaussian Loss (TSGL), enhancing the model’s sensitivity to both target’s position and content. Extensive experiments on UAV123VP and VisDroneVP (derived from single-object tracking datasets) demonstrate the exceptional performance of TAFormer in target-aware video prediction, showcasing its adaptability to the additional requirements of aerial video interpretation for target awareness. Liangyu Xu, Wanxuan Lu, Yongqiang Mao, Hanbo Bi, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | SFTformer: A Spatial-Frequency-Temporal Correlation-Decoupling Transformer for Radar Echo ExtrapolationabstractExtrapolating future weather radar echoes from past observations is a complex task vital for precipitation nowcasting. The spatial morphology and temporal evolution of radar echoes exhibit a certain degree of correlation, yet they also possess independent characteristics. Existing methods learn unified spatial and temporal representations in a highly coupled feature space, emphasizing the correlation between spatial and temporal features but neglecting the explicit modeling of their independent characteristics, which may result in mutual interference between them. To effectively model the spatiotemporal dynamics of radar echoes, we propose a spatial-frequency-temporal correlation-decoupling transformer (SFTformer). The model leverages stacked multiple SFT-Blocks to not only mine the correlation of the spatiotemporal dynamics of echo cells but also avoid the mutual interference between the temporal modeling and the spatial morphology refinement by decoupling them. Furthermore, inspired by the practice that weather forecast experts effectively review historical echo evolution to make accurate predictions, SFTfomer incorporates a joint training paradigm for historical echo sequence reconstruction and future echo sequence prediction. Experimental results on the HKO-7 dataset and ChinaNorth-2021 dataset demonstrate the superior performance of SFTfomer in short-term (1 h), mid-term (2 h), and long-term (3 h) precipitation nowcasting. Liangyu Xu, Wanxuan Lu, Fanglong Yao, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Multi-Stage Semi-Supervised Transformer for Remote Sensing Semantic Segmentation with Various Data AugmentationabstractExisting research in semantic segmentation heavily relies on numerous manually annotated data, while the vast amount of unlabeled data still needs to be fully utilized. To address this challenge, this paper introduces a novel multi-stage semi-supervised method for remote sensing semantic segmentation, building upon a modified classical self-training scheme that leverages pseudo-labels. By dividing the data augmentation process into two stages, we employ various data augmentation strategies and balance the size of labels and pseudo-labels validated through rigorous experimentation, which can alleviate the student model from overfitting pseudo-labels. Furthermore, we also explore the efficacy of the Vision Transformer model in semi-supervised semantic segmentation, leading to further performance enhancements. The experimental results show that our semi-supervised remote sensing semantic segmentation method exhibits a more intuitive structure, easier deployment, and superior performance. Wanxuan Lu, Zhi Guo |
IGARSS | 2 |
| 2023 | Semi-Supervised Semantic Generative Networks For Remote Sensing Image SegmentationabstractSemi-supervised remote sensing semantic segmentation is an efficient way to increase the use of unlabeled data and cut labelling costs. The unlabeled-to-labeled data ratio is employed in more recent methods, which is very different from what is really used in practise. In this paper, we propose a semi-supervised semantic generative network for remote sensing images, introducing a self-supervised learning method to enhance the feature representation of the model when the data ratio is high. Specifically, we design a new branch for unlabeled data, which includes modules for both semantic reconstruction and appearance reconstruction. It can effectively alleviate the category confusion in complicated remote sensing image when there are few labeled data. Comprehensive experiments on the ISPRS POTSDAM dataset demonstrate that the proposed method achieves promising results. Wanxuan Lu, Jidong Jin, Xian Sun 0001, Kun Fu 0001 |
IGARSS | 1 |
| 2023 | From single- to multi-modal remote sensing imagery interpretation: a survey and taxonomy
Xian Sun 0001, Wanxuan Lu, Peijin Wang, Ruigang Niu, Kun Fu 0001 |
Sci. China Inf. Sci. | 3 |
| 2023 | RingMo: A Remote Sensing Foundation Model With Masked Image ModelingabstractDeep learning approaches have contributed to the rapid development of remote sensing (RS) image interpretation. The most widely used training paradigm is to use ImageNet pretrained models to process RS data for specified tasks. However, there are issues such as domain gap between natural and RS scenes and the poor generalization capacity of RS models. It makes sense to develop a foundation model with general RS feature representation. Since a large amount of unlabeled data is available, the self-supervised method has more development significance than the fully supervised method in RS. However, most of the current self-supervised methods use contrastive learning, whose performance is sensitive to data augmentation, additional information, and selection of positive and negative pairs. In this article, we leverage the benefits of generative self-supervised learning (SSL) for RS images and propose an RS foundationmodel framework called RingMo, which consists of two parts. First, a large-scale dataset is constructed by collecting two million RS images from satellite and aerial platforms, covering multiple scenes and objects around the world. Second, we propose an RS foundation model training method designed for dense and small objects in complicated RS scenes. We show that the foundation model trained on our dataset with RingMo method achieves state-of-the-art (SOTA) on eight datasets across four downstream tasks, demonstrating the effectiveness of the proposed framework. Through in-depth exploration, we believe it is time for RS researchers to embrace generative SSL and leverage its general representation capabilities to speed up the development of RS applications. Xian Sun 0001, Peijin Wang, Wanxuan Lu, Zicong Zhu, Qibin He 0001, Junxi Li, Xuee Rong, Zhujun Yang, Qinglin He, Ruiping Wang 0001, Jiwen Lu, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | RingMo-Sense: Remote Sensing Foundation Model for Spatiotemporal Prediction via Spatiotemporal Evolution DisentanglingabstractRemote sensing spatiotemporal prediction aims to infer future trends from historical spatiotemporal data, e.g., videos and time series images, has a broad application prospect in many fields. The foundation model is a promising research direction for spatiotemporal information mining because of its robust feature extraction capability, and has made rapid progress in natural scenes. Nevertheless, due to the spatially multi-scale and temporally multi-scale properties in remote sensing data, these methods still encounter bottlenecks when applied to remote sensing. Therefore, we propose a foundation model for remote sensing spatiotemporal prediction via spatiotemporal evolution decoupling, abbreviated as RingMo-Sense. Considering spatial affinity, temporal continuity, and spatiotemporal interaction, we construct spatial, temporal, and spatiotemporal triple-branch prediction networks. Specifically, we use parameter-sharing and progressive joint training strategies to achieve stable long-range prediction and parameter reduction simultaneously. In addition, we build a remote sensing spatiotemporal dataset by collecting various remote sensing videos and time series images. The experimental results on six downstream spatiotemporal tasks demonstrate that the proposed model yields competitive performance. Fanglong Yao, Wanxuan Lu, Heming Yang 0003, Liangyu Xu, Leiyi Hu, Nayu Liu, Chubo Deng, Deke Tang, Changshuo Chen, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Semisupervised Convolution Neural Network for Partial Unlabeled Remote-Sensing Image SegmentationabstractSemantic segmentation methods for remote-sensing images based on the deep learning framework have achieved significant performance improvements. However, most of the existing work is based on fully supervised methods, which rely on a large number of manually annotated pixel-level labels. However, for remote-sensing images, labeling the ground-truth takes time and effort. To solve the problem in existing methods of overly relying on manual labeling, in this study, we propose a semisupervised convolution neural network based on contrastive loss for partial unlabeled remote-sensing image segmentation. In the design of the contrastive loss function, to capture the semantic relationship of pixels and improve the separability between different categories, we propose pixel-level and region-level contrastive loss. The pixel-level contrastive loss is designed to learn the correlation between different images, while region-level contrastive loss is designed to improve the quality of generated pseudo-labels. In addition, we designed a propagated self-training method that further guarantees the quality of the pseudo-labels and improves the richness of the labeled data. Experiments on POTSDAM and Vaihingen datasets demonstrate that the proposed method achieves the highest Mean Intersection over Union (mIOU) and significantly outperforms previous methods. Wanxuan Lu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Boundarymix: Generating pseudo-training images for improving segmentation with scribble annotations
Wanxuan Lu, Dong Gong, Kun Fu 0001, Xian Sun 0001, Wenhui Diao, Lingqiao Liu |
Pattern Recognit. | 1 |
| 2015 | Inferring User Preference in Good Abandonment from Eye Movements
Wanxuan Lu, Yunde Jia |
WAIM | 1 |
| 2014 | An Eye-Tracking Study of User Behavior in Web Image Search
Wanxuan Lu, Yunde Jia |
PRICAI | 1 |