Yansheng Li 0001

dblp:133/4525-1 · DBLP profile ↗
← Back
72ranked-venue papers
24as first author
50since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 33 · 14 first-author · 19 since 2021Artificial intelligence and machine learning · 27 · 9 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 15 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Predicting risks of hemorrhagic fever with renal syndrome using a Bayesian spatiotemporal hierarchical model
abstract
Hemorrhagic fever with renal syndrome (HFRS), a severe infectious disease primarily hosted by rodents, poses significant public health challenges in endemic regions. However, the intricate interplay of factors driving HFRS transmission, characterized by spatiotemporal heterogeneity and nonlinear relationships across different urbanization contexts, remains insufficiently explored. Utilizing historical HFRS cases, remote sensing imagery, and census data, this study constructed a Bayesian spatiotemporal hierarchical model to reveal the multiple risk drivers of HFRS transmission in the Guanzhong Plain, China. By integrating structured space-time components, nonlinear smoothing functions, and spatiotemporally varying coefficients, this model effectively captures the dynamic, non-stationary, and regionally heterogeneous nature of HFRS transmission. In addition, we constructed three integrated indices, the Human Activity Intensity Index (HAI), Eco-environmental Quality Index (EQI), and Habitat Suitability Index (HSI) to quantify the synergistic effects of human-environment systems. The model demonstrated strong performance (R2 = 0.814), identifying host habitat suitability as the dominant risk driver (RR = 2.283). Results further elucidated how meteorological conditions, ecological quality, and human activities influence HFRS risk across urbanization gradients. This study provides a generalized Bayesian framework for investigating geospatial health issues characterized by instability and complex nonlinear relationships.
Mengna Wei, Tiezhi Jin, Zhenfan Xu, Yuetong Chen, Mengyan Ye, Yaqin Su, Yansheng Li 0001, Pengbo Yu
Int. J. Geogr. Inf. Sci.10
2026 REST: Holistic Learning for End-to-End Semantic Segmentation of Whole-Scene Remote Sensing Imagery
abstract
Semantic segmentation of remote sensing imagery (RSI) is a fundamental task that aims at assigning a category label to each pixel. To pursue precise segmentation with one or more fine-grained categories, semantic segmentation often requires holistic segmentation of whole-scene RSI (WRI), which is normally characterized by a large size. However, conventional deep learning methods struggle to handle holistic segmentation of WRI due to the memory limitations of the graphics processing unit (GPU), thus requiring to adopt suboptimal strategies such as cropping or fusion, which result in performance degradation. Here, we introduce the Robust End-to-end semantic Segmentation architecture for whole-scene remoTe sensing imagery (REST). REST is the first intrinsically endtoend framework for truly holistic segmentation of WRI, supporting a wide range of encoders and decoders in a plugandplay fashion. It enables seamless integration with mainstream semantic segmentation methods, and even more advanced foundation models. Specifically, we propose a novel spatial parallel interaction mechanism (SPIM) within REST to overcome GPU memory constraints and achieve global context awareness. Unlike traditional parallel methods, SPIM enables REST to process a WRI effectively and efficiently by combining parallel computation with a divideandconquer strategy. Both theoretical analysis and experiments demonstrate that REST attains nearlinear throughput scalability as additional GPUs are employed. Extensive experiments demonstrate that REST consistently outperforms existing cropping-based and fusion-based methods across a variety of scenarios, ranging from single-class to multi-class segmentation, from multispectral to hyperspectral imagery, and from satellite to drone platforms. The robustness and versatility of REST are expected to offer a promising solution for the holistic segmentation of WRI, with the potential for further extension to large-size medical imagery segmentation.
Wei Chen 0089, Lorenzo Bruzzone, Bo Dang 0002, Yuan Gao 0015, Youming Deng, Jin-Gang Yu, Liangqi Yuan, Yansheng Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2026 Full-Scope Vectorization of Geographical Elements from Large-Size Remote Sensing Imagery
abstract
Large-size very-high-resolution (VHR) remote sensing imagery has emerged as a critical data source for high-precision vector mapping of multi-scale geographical elements such as building, water, road and etc. When dealing with the large-size image, due to the limited memory of GPU, the deep learning-based vector mapping methods often employ the sliding block strategy. This inevitably leads to the degenerated performance because of the stitching difficulty of the sliding blocks' vector mapping results. Therefore, it is necessary to conduct full-scope vector mapping via mining the consistent cue in large-size remote sensing imagery. To this end, this paper presents a novel global context-aware local point optimization method. To leverage the global context, this paper proposes a novel pyramid fusion network (PFNet) to conduct semantic segmentation of the large-size image in an end-to-end manner. Under the constraint of the global semantic segmentation result, a new inflection-point perception network (IPNet) is proposed to generate a set of stable points to depict the boundary of each element. Extensive experiments on building, water and road datasets, where each image has over 100 million pixels, show that our method obviously outperforms the existing methods.
Yansheng Li 0001, Wanchun Li, Bo Dang 0002, Yu Wang 0222, Wei Chen 0089, Bingnan Yang, Yongjun Zhang 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning
Xue Yang 0005, Qi Zhu 0010, Jingdong Chen, Yansheng Li 0001
ICCV8
2025 Towards Privacy-preserved Pre-training of Remote Sensing Foundation Models with Federated Mutual-Guidance Learning
abstract
Traditional Remote Sensing Foundation models (RSFMs) are pre-trained with a data-centralized paradigm, through self-supervision on large-scale curated remote sensing data. For each institution, however, pre-training RSFMs with limited data in a standalone manner may lead to suboptimal performance, while aggregating remote sensing data from multiple institutions for centralized pre-training raises privacy concerns. Seeking for collaboration is a promising solution to resolve this dilemma, where multiple institutions can collaboratively train RSFMs without sharing private data. In this paper, we propose a novel privacy-preserved pre-training framework (FedSense), which enables multiple institutions to collaboratively train RSFMs without sharing private data. However, it is a non-trivial task hindered by a vicious cycle, which results from model drift by remote sensing data heterogeneity and high communication overhead. To break this vicious cycle, we introduce Federated Mutual-guidance Learning. Specifically, we propose a Server-to-Clients Guidance (SCG) mechanism to guide clients updates towards global-flatness optimal solutions. Additionally, we propose a Clients-to-Server Guidance (CSG) mechanism to inject local knowledge into the server by low-bit communication. Extensive experiments on four downstream tasks demonstrate the effectiveness of our FedSense in both full-precision and communication-reduced scenarios, showcasing remarkable communication efficiency and performance gains.
Jieyi Tan, Bo Dang 0002, Yansheng Li 0001
ICCV4
2025 HoliTracer: Holistic Vectorization of Geographic Objects from Large-Size Remote Sensing Imagery
abstract
With the increasing resolution of remote sensing imagery (RSI), large-size RSI has emerged as a vital data source for high-precision vector mapping of geographic objects. Existing methods are typically constrained to processing small image patches, which often leads to the loss of contextual information and produces fragmented vector outputs. To address these, this paper introduces HoliTracer, the first framework designed to holistically extract vectorized geographic objects from large-size RSI. In HoliTracer, we enhance segmentation of large-size RSI using the Context Attention Net (CAN), which employs a local-to-global attention mechanism to capture contextual dependencies. Furthermore, we achieve holistic vectorization through a robust pipeline that leverages the Mask Contour Reformer (MCR) to reconstruct polygons and the Polygon Sequence Tracer (PST) to trace vertices. Extensive experiments on large-size RSI datasets, including buildings, water bodies, and roads, demonstrate that HoliTracer outperforms state-of-the-art methods. Our code and data are available in https://github.com/vvangfaye/HoliTracer.
Yu Wang 0222, Bo Dang 0002, Wanchun Li, Wei Chen 0089, Yansheng Li 0001
ICCV5
2025 SkySense V2: A Unified Foundation Model for Multi-Modal Remote Sensing
abstract
The multi-modal remote sensing foundation model (MM-RSFM) has significantly advanced various Earth observation tasks, such as urban planning, environmental monitoring, and natural disaster management. However, most existing approaches generally require the training of separate backbone networks for each data modality, leading to redundancy and inefficient parameter utilization. Moreover, prevalent pre-training methods typically apply self-supervised learning (SSL) techniques from natural images without adequately accommodating the characteristics of remote sensing (RS) images, such as the complicated semantic distribution within a single RS image. In this work, we present SkySense V2, a unified MM-RSFM that employs a single transformer backbone to handle multiple modalities. This backbone is pre-trained with a novel SSL strategy tailored to the distinct traits of RS data. In particular, SkySense V2 incorporates an innovative adaptive patch merging module and learnable modality prompt tokens to address challenges related to varying resolutions and limited feature diversity across modalities. In additional, we incorporate the mixture of experts (MoE) module to further enhance the performance of the foundation model. SkySense V2 demonstrates impressive generalization abilities through an extensive evaluation involving 16 datasets over 7 tasks, outperforming SkySense by an average of 1.8 points.
Lixiang Ru, Lei Yu 0005, Yansheng Li 0001, Jingdong Chen
ICCV6
2025 PointOBB-v3: Expanding Performance Boundaries of Single Point-Supervised Oriented Object Detection
Peiyuan Zhang, Xue Yang 0005, Yi Yu 0010, Qingyun Li, Yue Zhou 0005, Xiaosong Jia, Jingdong Chen, Xiang Li 0041, Junchi Yan, Yansheng Li 0001
Int. J. Comput. Vis.12
2025 Knowledge Graph-Guided Deep Network for Hyperspectral Remote Sensing Image Classification
abstract
For the classification of hyperspectral images (HSIs), most deep learning networks are data-driven and lack the usage of prior knowledge. In this letter, we propose a knowledge graph-guided classification network (KGNet), attempting to utilize the prior knowledge of land cover categories to enhance the classification performance. We first construct a knowledge graph on several hyperspectral scenes, which can characterize not only the attributes of land cover categories but also the rich connections between categories. Semantic features are then derived to represent the knowledge in the graph. Knowledge-guided learning is achieved by performing feature alignment between semantic and visual features. Finally, classification is performed on visual features that have contained the knowledge from semantic features. Experiments on three datasets demonstrate the effectiveness of applying the knowledge graph for the classification of hyperspectral remote sensing images.
Li Ma 0005, Yansheng Li 0001, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.3
2025 STAR: A First-Ever Dataset and a Large-Scale Benchmark for Scene Graph Generation in Large-Size Satellite Imagery
abstract
Scene graph generation (SGG) in satellite imagery (SAI) benefits promoting understanding of geospatial scenarios from perception to cognition. In SAI, objects exhibit great variations in scales and aspect ratios, and there exist rich relationships between objects (even between spatially disjoint objects), which makes it attractive to holistically conduct SGG in large-size very-high-resolution (VHR) SAI. However, there lack such SGG datasets. Due to the complexity of large-size SAI, mining triplets subject, relationship, object heavily relies on long-range contextual reasoning. Consequently, SGG models designed for small-size natural imagery are not directly applicable to large-size SAI. This paper constructs a large-scale dataset for SGG in large-size VHR SAI with image sizes ranging from 512 × 768 to 27,860 × 31,096 pixels, named STAR (Scene graph generaTion in lArge-size satellite imageRy), encompassing over 210K objects and over 400K triplets. To realize SGG in large-size SAI, we propose a context-aware cascade cognition (CAC) framework to understand SAI regarding object detection (OBD), pair pruning and relationship prediction for SGG. We also release a SAI-oriented SGG toolkit with about 30 OBD and 10 SGG methods which need further adaptation by our devised modules on our challenging STAR dataset. The dataset and toolkit are available at: https://linlin-dev.github.io/project/STAR.
Yansheng Li 0001, Tingzhu Wang, Xue Yang 0005, Qi Wang 0009, Youming Deng, Xian Sun 0001, Haifeng Li 0007, Bo Dang 0002, Yongjun Zhang 0002, Yi Yu 0010, Junchi Yan
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Wholly-WOOD: Wholly Leveraging Diversified-Quality Labels for Weakly-Supervised Oriented Object Detection
abstract
Accurately estimating the orientation of visual objects with compact rotated bounding boxes (RBoxes) has become a prominent demand, which challenges existing object detection paradigms that only use horizontal bounding boxes (HBoxes). To equip the detectors with orientation awareness, supervised regression/classification modules have been introduced at the high cost of rotation annotation. Meanwhile, some existing datasets with oriented objects are already annotated with horizontal boxes or even single points. It becomes attractive yet remains open for effectively utilizing weaker single point and horizontal annotations to train an oriented object detector (OOD). We develop Wholly-WOOD, a weakly-supervised OOD framework, capable of wholly leveraging various labeling forms (Points, HBoxes, RBoxes, and their combination) in a unified fashion. By only using HBox for training, our Wholly-WOOD achieves performance very close to that of the RBox-trained counterpart on remote sensing and other areas, significantly reducing the tedious efforts on labor-intensive annotation for oriented objects.
Yi Yu 0010, Xue Yang 0005, Yansheng Li 0001, Zhenjun Han, Feipeng Da, Junchi Yan
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 T2EA: Target-Aware Taylor Expansion Approximation Network for Infrared and Visible Image Fusion
abstract
In the image fusion mission, the crucial task is to generate high-quality images for highlighting the key objects while enhancing the scenes to be understood. To complete this task and provide a powerful interpretability as well as a strong generalization ability in producing enjoyable fusion results which are comfortable for vision tasks (such as objects detection and their segmentation), we present a novel interpretable decomposition scheme and develop a target-aware Taylor expansion approximation (T2EA) network for infrared and visible image fusion, where our T2EA includes the following key procedures: Firstly, visible and infrared images are both decomposed into feature maps through a designed Taylor expansion approximation (TEA) network. Then, the Taylor feature maps are hierarchically fused by a dual-branch feature fusion (DBFF) network. Next, the fused map of each layer is contributed to synthesize an enjoyable fusion result by the inverse Taylor expansion. Finally, a segmentation network is jointed to refine the fusion network parameters which can promote the pleasing fusion results to be more suitable for segmenting the objects. To validate the effectiveness of our reported T2EA network, we first discuss the selection of Taylor expansion layers and fusion strategies. Then, both quantitatively and qualitatively experimental results generated by the selected SOTA approaches on three datasets (MSRS, TNO, andLLVIP) are compared in testing, generalization, and target detection and segmentation, demonstrating that our T2EA can produce more competitive fusion results for vision tasks and is more powerful for image adaption. The code will be available at https://github.com/MysterYxby/T2EA.
Zhenghua Huang, Biyun Xu, Menghan Xia, Qian Li 0019, Yansheng Li 0001, Nong Sang
IEEE Trans. Circuits Syst. Video Technol.6
2025 Toward Generalizable Physical Attacks Against Remote Sensing Building Change Detection
abstract
Deep networks have been widely used in remote sensing building change detection (BCD) and achieved good performances. However, the BCD networks are vulnerable to adversarial attacks. Previous physical adversarial attacks in remote sensing focused on image-level tasks and instance-level tasks, we propose a generalizable physical attacks (GPA) framework for the pixel-level BCD task. The GPA framework is in the joint task-instance space to generate the universal physical adversarial patch which can cheat unseen models. It involves white-box attacks and black-box attacks. To solve the problem of the singular adversarial patch position and the singular target task in existing works, we propose an adaptive adversarial patch position module and present a feature loss for bi-temporal remote sensing images in white-box attacks. To enhance the generalization of the patch, we propose a method in the joint task-instance space to learn the black-box adversarial patch. Extensive experiments on the widely used remote sensing image BCD dataset shows the proposed adversarial attacks against BCD achieve better performance compared with state-of-the-art methods. In particular, our GPA elevatesASRpandCD_Ratefor all buildings by an average of 29.31%, 13.92% in the white-box scenario, and 3.92%, 3.59% in the strict black-box scenario. Real-world experiments with printed patches on physical building models further demonstrate that our GPA can attack the BCD networks. The source code and associated datasets are available at https://github.com/hhhh420/adversarial patch BCD.
Mengze Hao, Bo Dang 0002, Yansheng Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 AllSpark: A Multimodal Spatiotemporal General Intelligence Model With Ten Modalities via Language as a Reference Framework
abstract
RGB, multispectral, point and other spatio-temporal modal data fundamentally represent different observational approaches for the same geographic object. Therefore, leveraging multimodal data is an inherent requirement for comprehending geographic objects. However, due to the high heterogeneity in structure and semantics among various spatio-temporal modalities, the joint interpretation of multimodal spatio-temporal data has long been an extremely challenging problem. The primary challenge resides in striking a trade-off between the cohesion and autonomy of diverse modalities. This trade-off becomes progressively nonlinear as the number of modalities expands. Inspired by the human cognitive system and linguistic philosophy, where perceptual signals from the five senses converge into language, we introduce the Language as Reference Framework (LaRF), a fundamental principle for constructing a multimodal unified model. Building upon this, we propose AllSpark, a multimodal spatio-temporal general artificial intelligence model. Our model integrates ten different modalities into a unified framework, including one-dimensional (language, code, table), two-dimensional (RGB, SAR, multispectral, hyperspectral, graph, trajectory), and three-dimensional (point cloud) modalities. To achieve modal cohesion, AllSpark introduces a modal bridge and multimodal large language model (LLM) to map diverse modal features into the language feature space. To maintain modality autonomy, AllSpark uses modality-specific encoders to extract the tokens of various spatio-temporal modalities. Finally, observing a gap between the model’s interpretability and downstream tasks, we designed modality-specific prompts and task heads, enhancing the model’s generalization capability across specific tasks. Experiments indicate that the incorporation of language enables AllSpark to excel in few-shot classification tasks for RGB and point cloud modalities without additional training, surpassing baseline performance by up to 41.82%. Additionally, AllSpark, despite lacking expert knowledge in most spatio-temporal modalities and utilizing a unified structure, demonstrates strong adaptability across ten modalities. LaRF and AllSpark contribute to the shift in the research paradigm in spatio-temporal intelligence, transitioning from a modality-specific and task-specific paradigm to a general paradigm. The source code is available at https://github.com/GeoX-Lab/AllSpark.
Run Shao, Qiujun Li, Linrui Xu, Qing Zhu 0012, Yongjun Zhang 0002, Yansheng Li 0001, Yu Liu 0003, Shizhong Yang, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.10
2025 R2PLoc: A Region-to-Point UAV Visual Geo-Localization Framework Leveraging Hierarchical Semantic Representation
abstract
The challenges in UAV visual geo-localization primarily stem from discrepancies between satellite maps and aerial images, including scale variations, viewpoint deviations, and spatiotemporal mismatches. Current approaches adopt retrieval-based or keypoint-matching-based localization, and some studies employ a cascaded approach. However, these methods still exhibit limitations in addressing discrepancies. To address these challenges, we propose a region-to-point UAV visual geo-localization framework named R2PLoc. Specifically, we consider UAV visual geo-localization as the process of retrieving corresponding regions from satellite map databases using aerial images while establishing projective relationship between them. First, we employ a shared backbone network for semantic feature extraction to conserve computational resources. Then, the Hierarchical Semantic Aggregation Module (HSAM) is designed to address the feature distribution shifts by fusing multi-scale semantics that combine both global contexts and local structures. Additionally, the Semantic-Enhanced Hierarchical Refinement Matcher (SHRM) is constructed to improve the geometric consistency of keypoint matching by integrating high-level semantic information. Furthermore, the UAV-R2P dataset is constructed for the region-to-point geo-localization task. The qualitative and quantitative experimental results demonstrate that our method outperforms most state-of-the-art methods with similar model size on most available datasets.
Ruitao Lu, Yansheng Li 0001, Yunsong Li 0001, Dingwen Zhang
IEEE Trans. Geosci. Remote. Sens.4
2025 Hierarchical Prototype Learning via Aggregation-Decomposition for Fine-Grained Geospatial Scene Graph Generation
abstract
Geospatial scene graph generation (G-SGG) in remote sensing imagery (RSI) aims to identify tripletsthat reflect interactions among objects on the Earth’s surface. A key challenge in G-SGG is the large intra-class variation of geospatial relationships, which leads to frequent misclassification. To address this, we propose a hierarchical prototype learning network (HPL-Net) via aggregation-decomposition, designed to enhance fine-grained G-SGG by addressing intra-class variation in geospatial relationships. In HPL-Net, hierarchical constraint (HC) of prototype learning, employing a tree structure, is designed to learn discriminative representations of relationship prototypes. Additionally, a class-balanced constraint (CC) of instance learning with class-aware margins is introduced to mitigate misclassification caused by imbalanced relationship categories. Extensive experiments demonstrate that our method outperforms state-of-the-art G-SGG models on the satellite-based STAR and drone-based AUG datasets. Specifically, it outperforms the latest models, such as RPCM and LPG, by significant margins of 4.90%/5.32% and 4.52%/4.31% on TMR@1500/2000 and F@50/150, respectively, for the PredCls task on the STAR and AUG datasets. The code will be available at https://github.com/linlin-dev/HPL-Net.
Tingzhu Wang, Yansheng Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Topology-Aware Hierarchical Mamba for Salient Object Detection in Remote Sensing Imagery
Wei Yang 0043, Zhiqi Yi, Andong Huang, Ying Wang 0123, Yongxiang Yao, Yansheng Li 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 DDRL: Domain Distribution Reconstruction Learning for Binary Change Detection in Remote Sensing Images
abstract
Change detection (CD) aims to identify and locate changes in the same observed surface coverage area across bitemporal images. This technique has widespread applications in urban planning, land use, and disaster damage extraction. Deep learning-based CD methods typically use learnable encoders to map bitemporal images to a common domain distribution space, allowing for the discrimination and localization of change and invariant features. However, due to differences in imaging mechanisms, seasons, and shooting angles, a large number of pseudochanges may easily appear, affecting the accurate recognition of the domain distribution space. In addition, binary CD focuses solely on whether scene targets have changed, resulting in change labels that encompass a variety of different objects, thus increasing the significance of intraclass differences. To address the aforementioned issues, we propose a domain distribution reconstruction learning (DDRL) framework for binary CD, which effectively mitigates the problem of pseudochanges by detecting abnormal feature domain distributions. Specifically, DDRL first extracts multiscale features from bitemporal images using a Siamese cross-window self-attention module, achieving feature domain transformation from the original space. Subsequently, it employs a graph attention enhanced (GAE) module to improve the low-level domain distribution, enabling it to focus on change regions. In addition, DDRL utilizes a cross-domain feature contrastive learning (CFCL) module for reconstructive learning of high-level fused features. This process ensures that intraclass features are compact, while interclass features are dispersed within the high-level domain distribution, thereby significantly improving the domain distribution representation to discriminate pseudochanges. Experimental results show that the proposed DDRL performs excellently across multiple public datasets, surpassing mainstream methods and significantly improving CD performance. The source code will be made available athttps://github.com/yzygit1230/DDRL.
Wei Yang 0043, Zhaoyi Ye, Liye Mei, Yongxiang Yao, Yansheng Li 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 Multimodal Remote Sensing Image Robust Matching Based on Second-Order Tensor Orientation Feature Transformation
abstract
Nonrigid deformation (NRD) and image noise in multimodal remote sensing images (MRSI) lead to abrupt changes in feature directions, resulting in sensitivity to rotational variation, sparse correct matches, and high false match rates. In order to address these challenges, this article proposes a second-order tensor orientation feature transformation (SOFT) method to improve the rotational invariance of MRSI matching and increase the number of correct matches (NCMs). The SOFT method has two main contributions: 1) a novel second-order tensor orientation descriptor is constructed by generating a tensor orientation feature map using a designed second-order tensor function, which is then combined with a gradient location and orientation histogram (GLOH)-like descriptor framework to achieve robust rotational invariance in multimodal image matching and 2) an error-removal global-local iterative optimization (EGIO) is introduced, employing a skewness of mixed pixel intensity (SMPI) function to automatically select matching seed points, followed by an iterative partition optimization strategy for refining corresponding points. Experiments on 744 groups of typical MRSIs demonstrate that the SOFT method significantly outperforms nine state-of-the-art methods, achieving an average 97% improvement in the NCMs, an average 25.51% improvement in the rate of correct matches (RCMs), and an average reduction in RMSE of 2.69 pixels. The proposed SOFT method, thus, offers robust MRSI matching with strong rotational invariance and precise identification of corresponding points, proving its effectiveness for complex remote sensing scenarios. Access to experiment-related data and codes will be provided athttps://skyearth.org/research/.
Yongjun Zhang 0002, Peihao Wu, Yongxiang Yao, Yi Wan 0001, Wenfei Zhang, Yansheng Li 0001, Xiaohu Yan
IEEE Trans. Geosci. Remote. Sens.6
2024 ResMatch: Residual Attention Learning for Feature Matching
abstract
Attention-based graph neural networks have made great progress in feature matching. However, the literature lacks a comprehensive understanding of how the attention mechanism operates for feature matching. In this paper, we rethink cross- and self-attention from the viewpoint of traditional feature matching and filtering. To facilitate the learning of matching and filtering, we incorporate the similarity of descriptors into cross-attention and relative positions into self-attention. In this way, the attention can concentrate on learning residual matching and filtering functions with reference to the basic functions of measuring visual and spatial correlation. Moreover, we leverage descriptor similarity and relative positions to extract inter- and intra-neighbors. Then sparse attention for each point can be performed only within its neighborhoods to acquire higher computation efficiency. Extensive experiments, including feature matching, pose estimation and visual localization, confirm the superiority of the proposed method. Our codes are available at https://github.com/ACuOoOoO/ResMatch.
Yuxin Deng 0002, Kaining Zhang, Yansheng Li 0001, Jiayi Ma 0001
AAAI4
2024 GLH-Water: A Large-Scale Dataset for Global Surface Water Detection in Large-Size Very-High-Resolution Satellite Imagery
abstract
Global surface water detection in very-high-resolution (VHR) satellite imagery can directly serve major applications such as refined flood mapping and water resource assessment. Although achievements have been made in detecting surface water in small-size satellite images corresponding to local geographic scales, datasets and methods suitable for mapping and analyzing global surface water have yet to be explored. To encourage the development of this task and facilitate the implementation of relevant applications, we propose the GLH-water dataset that consists of 250 satellite images and 40.96 billion pixels labeled surface water annotations that are distributed globally and contain water bodies exhibiting a wide variety of types (e.g. , rivers, lakes, and ponds in forests, irrigated fields, bare areas, and urban areas). Each image is of the size 12,800 × 12,800 pixels at 0.3 meter spatial resolution. To build a benchmark for GLH-water, we perform extensive experiments employing representative surface water detection models, popular semantic segmentation models, and ultra-high resolution segmentation models. Furthermore, we also design a strong baseline with the novel pyramid consistency loss (PCL) to initially explore this challenge, increasing IoU by 2.4% over the next best baseline. Finally, we implement the cross-dataset generalization and pilot area application experiments, and the superior performance illustrates the strong generalization and practical application value of GLH-water dataset. Project page: https://jack-bo1220.github.io/project/GLH-water.html
Yansheng Li 0001, Bo Dang 0002, Wanchun Li, Yongjun Zhang 0002
AAAI1
2024 SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery
abstract
Prior studies on Remote Sensing Foundation Model (RSFM) reveal immense potential towards a generic model for Earth Observation. Nevertheless, these works primar-ily focus on a single modality without temporal and geo-context modeling, hampering their capabilities for diverse tasks. In this study, we present SkySense, a generic billion-scale model, pretrained on a curated multimodal Remote Sensing Imagery (RSI) dataset with 21.5 million temporal sequences. SkySense incorporates a factorized multimodal spatiotemporal encoder taking temporal sequences of opti-cal and Synthetic Aperture Radar (SAR) data as input. This encoder is pretrained by our proposed Multi-Granularity Contrastive Learning to learn representations across different modal and spatial granularities. To further enhance the RSI representations by the geo-context clue, we introduce Geo-Context Prototype Learning to learn region-aware prototypes upon RSI's multimodal spatiotemporal features. To our best knowledge, SkySense is the largest Multi-Modal RSFM to date, whose modules can be flexibly combined or used individually to accommodate various tasks. It demonstrates remarkable generalization capabilities on a thor-ough evaluation encompassing 16 datasets over 7 tasks, from single- to multimodal, static to temporal, and classification to localization. SkySense surpasses 18 recent RSFMs in all test scenarios. Specifically, it outperforms the latest models such as GFM, SatLas and Scale-MAE by a large margin, i.e., 2.76%, 3.67% and 3.61% on average respectively. We will release the pretrained weights to facilitate future research and Earth Observation applications.
Xin Guo 0010, Jiangwei Lao, Bo Dang 0002, Lei Yu 0005, Lixiang Ru, Liheng Zhong, Dingxiang Hu, Huimei He, Jian Wang 0108, Jingdong Chen, Ming Yang 0007, Yongjun Zhang 0002, Yansheng Li 0001
CVPR16
2024 PointOBB: Learning Oriented Object Detection via Single Point Supervision
abstract
Single point-supervised object detection is gaining attention due to its cost-effectiveness. However, existing approaches focus on generating horizontal bounding boxes (HBBs) while ignoring oriented bounding boxes (OBBs) commonly used for objects in aerial images. This paper proposes PointOBB, the first single Point-based OBB generation method, for oriented object detection. PointOBB operates through the collaborative utilization of three distinctive views: an original view, a resized view, and a ro-tatedlflipped (rot/flp) view. Upon the original view, we leverage the resized and rot/flp views to build a scale augmentation module and an angle acquisition module, respectively. In the former module, a Scale-Sensitive Consistency (SSC) loss is designed to enhance the deep network's ability to perceive the object scale. For accurate object angle predictions, the latter module incorporates self-supervised learning to predict angles, which is associated with a scale-guided Dense-to-Sparse (DS) matching strategy for aggre-gating dense angles corresponding to sparse objects. The resized and rot/flp views are switched using a progressive multi- view switching strategy during training to achieve coupled optimization of scale and angle. Experimental re-sults on the DIOR-R and DOTA-v1.0 datasets demonstrate that PointOBB achieves promising performance, and significantly outperforms potential point-supervised baselines.
Xue Yang 0005, Yi Yu 0010, Qingyun Li, Junchi Yan, Yansheng Li 0001
CVPR6
2024 Fine-Grained Scene Graph Generation via Sample-Level Bias Prediction
Yansheng Li 0001, Tingzhu Wang, Xin Guo 0010
ECCV (26)1
2024 Domain Knowledge-Aware Remote Sensing Foundation Model for Flood Detection in Multi-Spectral Imagery
abstract
Obtaining accurate and timely flood information is crucial for effective disaster management and response. To address the limitations of existing methods in terms of accuracy and model robustness, this research significantly improves the accuracy and stability of flood detection by integrating domain knowledge into the Remote Sensing Foundation Model (RSFM). Specifically, we employ advanced RSFM to focus on extracting spatial texture features from images after super-resolution. The Automatic Water Extraction Index (AWEI) is introduced to leverage spectral information from multi-spectral imagery, while model fusion techniques further enhance the accuracy of segmentation results. Moreover, we incorporate prior knowledge such as land use products and Digital Elevation Models (DEM) for knowledge rules-driven post-processing, refining the final flood detection results. Experimental results demonstrate that our approach achieve the second-place ranking in the 2024 IEEE GRSS Data Fusion Contest (DFC) Track 2 test phase (F1: 88.25%), highlighting the effectiveness and competitiveness of our method.
Yansheng Li 0001, Bo Dang 0002, Fanyi Wei, Jieyi Tan, Yangjie Lin
IGARSS1
2024 Parameter-Efficient Complementary Expert Learning for Long-Tailed Visual Recognition
abstract
Long-tailed recognition (LTR) aims to learn balanced models from extremely unbalanced training data. Fine-tuning pretrained foundation models has recently emerged as a promising research direction for LTR. However, we observe that the fine-tuning process tends to degrade the intrinsic representation capability of pretrained models and lead to model bias towards certain classes, thereby hindering the overall recognition performance. To unleash the intrinsic representation capability of pretrained foundation models, in this work, we propose a new Parameter-Efficient Complementary Expert Learning (PECEL) for LTR. Specifically, PECEL consists of multiple experts, where individual experts are trained via Parameter-Efficient Fine-Tuning (PEFT) and encouraged to learn different expertise on complementary sub-categories via the proposed sample-aware logit adjustment loss. By aggregating the predictions of different experts, PECEL effectively achieves a balanced performance on long-tailed classes. Nevertheless, learning multiple experts generally introduces extra trainable parameters. To ensure parameter efficiency, we further propose a parameter sharing strategy which decomposes and shares the parameters in each expert. Extensive experiments on 4 LTR benchmarks show that the proposed PECEL can effectively learn multiple complementary experts without increasing the trainable parameters and achieve new state-of-the-art performance.
Lixiang Ru, Xin Guo 0010, Lei Yu 0005, Jiangwei Lao, Jian Wang 0108, Jingdong Chen, Yansheng Li 0001, Ming Yang 0007
ACM Multimedia8
2024 MMKDGAT: Multi-modal Knowledge graph-aware Deep Graph Attention Network for remote sensing image recommendation
Fei Wang 0094, Xianzhang Zhu, Yongjun Zhang 0002, Yansheng Li 0001
Expert Syst. Appl.5
2024 Learning to Holistically Detect Bridges From Large-Size VHR Remote Sensing Imagery
abstract
Bridge detection in remote sensing images (RSIs) plays a crucial role in various applications, but it poses unique challenges compared to the detection of other objects. In RSIs, bridges exhibit considerable variations in terms of their spatial scales and aspect ratios. Therefore, to ensure the visibility and integrity of bridges, it is essential to perform holistic bridge detection in large-size very-high-resolution (VHR) RSIs. However, the lack of datasets with large-size VHR RSIs limits the deep learning algorithms' performance on bridge detection. Due to the limitation of GPU memory in tackling large-size images, deep learning-based object detection methods commonly adopt the cropping strategy, which inevitably results in label fragmentation and discontinuous prediction. To ameliorate the scarcity of datasets, this paper proposes a large-scale dataset named GLH-Bridge comprising 6,000 VHR RSIs sampled from diverse geographic locations across the globe. These images encompass a wide range of sizes, varying from 2,048 × 2,048 to 16,384 × 16,384 pixels, and collectively feature 59,737 bridges. These bridges span diverse backgrounds, and each of them has been manually annotated, using both an oriented bounding box (OBB) and a horizontal bounding box (HBB). Furthermore, we present an efficient network for holistic bridge detection (HBD-Net) in large-size RSIs. The HBD-Net presents a separate detector-based feature fusion (SDFF) architecture and is optimized via a shape-sensitive sample re-weighting (SSRW) strategy. The SDFF architecture performs inter-layer feature fusion (IFF) to incorporate multi-scale context in the dynamic image pyramid (DIP) of the large-size image, and the SSRW strategy is employed to ensure an equitable balance in the regression weight of bridges with various aspect ratios. Based on the proposed GLH-Bridge dataset, we establish a bridge detection benchmark including the OBB and HBB tasks, and validate the effectiveness of the proposed HBD-Net. Additionally, cross-dataset generalization experiments on two publicly available datasets illustrate the strong generalization capability of the GLH-Bridge dataset.
Yansheng Li 0001, Yongjun Zhang 0002, Yihua Tan, Jin-Gang Yu, Song Bai 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Trusted Multimodal Socio-Cyber Sentiment Analysis Based on Disentangled Hierarchical Representation Learning
abstract
The rapid development of the digital age has led to a qualitative leap in social media. To meet the cognitive needs of users, social media platforms have been mining users’ private information and disseminating information through various means. However, these platforms lack effective management of information release and various forms of emotional expressions make public propaganda increasingly diverse and complex. Therefore, accurately identifying the relationships between multimodal data poses a challenge. An effective modal representation must consider both the consistency of multimodal data and the complementarity of single-modal data. However, existing methods focus on fusing different modal features into a unified feature representation, while neglecting to evaluate the reliability of prediction results. In this article, we disentangle the consistency and complementarity in the fused representation problem of multimodal data. We construct the modal private task (unique) by using the Dirichlet distribution and evidence theory to solve the uncertainty of each modal prediction. The model can output the uncertainty of prediction and learn complementary information through the fusion of decision layers. At the same time, we construct the modal common task using a low-rank tensor fusion model to learn consistent features. Finally, we compare the model with the current mainstream methods on three public datasets, and the experimental results show that the performance of our method reaches the level of current advanced algorithms.
Guoxia Xu, Lizhen Deng, Yansheng Li 0001, Yantao Wei, Xiaokang Zhou, Hu Zhu
IEEE Trans. Comput. Soc. Syst.3
2024 SCD-SAM: Adapting Segment Anything Model for Semantic Change Detection in Remote Sensing Imagery
abstract
Semantic change detection (SCD) has gradually emerged as a prominent research focus in remote sensing image processing due to its critical role in earth observation applications. In view of its powerful semantic-driven feature extraction capability, the Segment Anything Model (SAM) has demonstrated its suitability across various visual scenes. However, it suffers from significant performance degradation when confronted with remote sensing images, especially those containing various ground objects that possess significant inter-class similarity and substantial intra-class variations. To address the above issues, we propose SCD-SAM, aiming to leverage the potent visual recognition capabilities of SAM for enhanced accuracy and robustness in SCD. Specifically, we introduce a contextual semantic change-aware dual encoder that combines MobileSAM and CNN to extract progressive semantic change features in parallel, and inject local features into the MobileSAM encoder through depth feature interaction to compensate for the Transformer’s limitations in perceiving local semantic details. Besides, in order to utilize the strong visual feature extraction capability of MobileSAM in remote sensing images, we propose a semantic adaptor that aggregates semantic-oriented information about changing objects. To better integrate the extracted contextual semantic information, we devise a progressive feature aggregation dual decoder that aggregates binary change features and semantic change features respectively, alleviating the semantic gap across different scales. The quantitative and visual results show that SCD-SAM outperforms the state-of-the-art SCD methods on publicly open SCD datasets (e.g., SECOND-CD and Landsat-CD). The code will be made available at https://github.com/yzygit1230/SCD-SAM.
Liye Mei, Zhaoyi Ye, Hongzhu Wang, Ying Wang 0123, Wei Yang 0043, Yansheng Li 0001
IEEE Trans. Geosci. Remote. Sens.8
2024 Scene Graph-Aware Hierarchical Fusion Network for Remote Sensing Image Retrieval With Text Feedback
abstract
In the realm of image retrieval with text feedback, existing studies have predominantly concentrated on the intrinsic attribute of target objects, neglecting extrinsic information essential for remote sensing (RS) images, such as spatial relationships. This research addresses this gap by incorporating RS image scene graphs as side information, given their capacity to encapsulate internal object attributes, external structural features between objects, and the relationships among images. To fully leverage the features from the reference RS image, scene graph, and modifier sentence, we propose a Scene graph-aware Hierarchical Fusion Network (SHF), which optimally integrates the multimodal features in a two-stage fusion process. Initially, image and scene graph features are fused hierarchically, followed by transforming content information with a proposed Multi-modal Global Content block (MGC), ultimately transforming style information. To validate the superiority of SHF, we constructed three datasets with images from several popular RS datasets, named Airplane (3461 image+text-image pairs), Tennis (1924 image+text-image pairs), and WHIRT (3344 image+text-image pairs). Extensive experiments conducted on these datasets show that SHF significantly outperforms state-of-the-art methods.
Fei Wang 0094, Xianzhang Zhu, Xiaojian Liu 0004, Yongjun Zhang 0002, Yansheng Li 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Progressive Learning With Cross-Window Consistency for Semi-Supervised Semantic Segmentation
abstract
Semi-supervised semantic segmentation focuses on the exploration of a small amount of labeled data and a large amount of unlabeled data, which is more in line with the demands of real-world image understanding applications. However, it is still hindered by the inability to fully and effectively leverage unlabeled images. In this paper, we reveal that cross-window consistency (CWC) is helpful in comprehensively extracting auxiliary supervision from unlabeled data. Additionally, we propose a novel CWC-driven progressive learning framework to optimize the deep network by mining weak-to-strong constraints from massive unlabeled data. More specifically, this paper presents a biased cross-window consistency (BCC) loss with an importance factor, which helps the deep network explicitly constrain confidence maps from overlapping regions in different windows to maintain semantic consistency with larger contexts. In addition, we propose a dynamic pseudo-label memory bank (DPM) to provide high-consistency and high-reliability pseudo-labels to further optimize the network. Extensive experiments on three representative datasets of urban views, medical scenarios, and satellite scenes with consistent performance gain demonstrate the superiority of our framework. Our code is released at https://jack-bo1220.github.io/project/CWC.html.
Bo Dang 0002, Yansheng Li 0001, Yongjun Zhang 0002, Jiayi Ma 0001
IEEE Trans. Image Process.2
2023 MFVNet: a deep adaptive fusion network with multiple field-of-views for remote sensing image semantic segmentation
Yansheng Li 0001, Wei Chen 0089, Xin Huang 0002, Zhi Gao 0005, Tao He 0002, Yongjun Zhang 0002
Sci. China Inf. Sci.1
2023 To see further: Knowledge graph-aware deep graph convolutional network for recommender systems
Fei Wang 0094, Yongjun Zhang 0002, Yansheng Li 0001, Chenming Zhu
Inf. Sci.4
2023 CloudViT: A Lightweight Vision Transformer Network for Remote Sensing Cloud Detection
abstract
Clouds inevitably exist in satellite images, which limit the processing and application of satellite images to a certain extent. Therefore, cloud detection is a preprocessing task in satellite image extraction and analysis processing. However, the existing methods are difficult to mine robust features, and the number of parameters and computation are large, which is not conducive to the deployment of the model. In this letter, cloud vision transformer (CloudViT), a lightweight vision transformer network for cloud detection from satellite imagery, is proposed. In detail, to utilize dark channel priors in multispectral imagery to guide the network to learn features, a multiscale dark channel extractor is used to first predict dark channels, and then, the dark channel features and image features are input to the attention mechanism-based dark channel-guided context aggregation module to enhance image features, which in turn makes cloud detection results more accurate. At the same time, to enhance the transfer ability of the network between different satellite sensors, a plug-and-play channel adaptive module is proposed to deal with the inconsistency of the number of different satellite sensor bands. The experimental results on the Landsat7 dataset show that our network CloudViT outperforms the state-of-the-art methods while keeping the number of parameters and computation small. At the same time, the experimental results on transfer to three other datasets show that using the channel adaptation module can greatly improve the transfer ability of the model.
Bin Zhang 0046, Yongjun Zhang 0002, Yansheng Li 0001, Yi Wan 0001, Yongxiang Yao
IEEE Geosci. Remote. Sens. Lett.3
2023 MSINet: Mining scale information from digital surface models for semantic segmentation of aerial images
Chengli Peng, Haifeng Li 0007, Chao Tao 0001, Yansheng Li 0001, Jiayi Ma 0001
Pattern Recognit.4
2023 Semantic-Aware Attack and Defense on Deep Hashing Networks for Remote-Sensing Image Retrieval
abstract
Deep hashing networks have been successful in retrieving interesting images from massive remote sensing images. There is no doubt that security and reliability are critical in remote sensing image retrieval. Recent studies about natural image retrieval have shown the vulnerability of deep hashing networks to adversarial examples, but there do not exist any researches about the attack and defense on deep hashing networks in remote sensing image retrieval. Due to the large intra-class difference and high inter-class similarity of remote sensing images, the attack and defense methods on deep hashing networks for natural images cannot be directly applied to the remote sensing images. Different from the widely adopted instance-aware hash codes which often present the suboptimum performance of the attack and defense on deep hashing networks, this paper recommends the usage of semantic-aware hash codes, which take into account multiple samples in the given semantic categories, in both attack and defense. To pursue the strongest attack on remote sensing image retrieval, a novel semantic-aware attack with weights via multiple random initialization (RWC) is proposed. To alleviate the retrieval degradation caused by adversarial attacks, a new adversarial training defense method on deep hashing networks with the adversarial semantic-aware consistency constraint (ACN) is proposed. Extensive experiments on three typical open remote sensing image datasets (i.e., UCM, AID, NWPU-RESISC45) show the proposed attack and defense methods on various deep hashing networks achieve better performance compared with the state-of-the-art methods. The source code will be made publicly available along with this paper.
Yansheng Li 0001, Mengze Hao, Hu Zhu, Yongjun Zhang 0002
IEEE Trans. Geosci. Remote. Sens.1
2023 SPGAN-DA: Semantic-Preserved Generative Adversarial Network for Domain Adaptive Remote Sensing Image Semantic Segmentation
abstract
Unsupervised domain adaptation for remote sensing semantic segmentation seeks to adapt a model trained on the labeled source domain to the unlabeled target domain. One of the most promising ways is to translate images from the source domain to the target domain to align the spectral information or imaging mode by the generative adversarial network (GAN). However, source-to-target translation often brings bias in the translated images causing limited performance, as semantic information is not well considered in the translation procedure. To overcome this limitation, we present an innovative semantic-preserved generative adversarial network (SPGAN), designed to mitigate the image translation bias and then leverage the translated images as well as unlabeled target images by class distribution alignment (CDA) module to train a domain adaptive semantic segmentation model. The above two stages are coupled together to form a unified framework called SPGAN-DA. Specifically, we first conduct semantic invariant translation from source to target domain, which is achieved by introducing representation-invariant and semantic-preserved constraints to the GAN model. To further narrow the landscape layout gap between the translated and target images, CDA semantic segmentation is proposed. CDA semantic segmentation consists of two aspects. At the model input level, object discrepancy is eliminated by introducing the ClassMix operation. At the model output level, boundary enhancement is proposed to refine the performance of object boundaries. Extensive experiments on three typical remote sensing cross-domain semantic segmentation benchmarks demonstrate the effectiveness and generality of our proposed method, which competes favorably against existing state-of-the-art methods.
Yansheng Li 0001, Te Shi 0001, Yongjun Zhang 0002, Jiayi Ma 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Learning Relative Feature Displacement for Few-Shot Open-Set Recognition
abstract
Few-shot learning (FSL) usually assumes that the query is drawn from the same label space as the support set, while queries from unknown classes may emerge unexpectedly in many open-world application scenarios. Such an open-set issue will limit the practical deployment of FSL systems, which remains largely unexplored. In this paper, we investigate the problem of few-shot open-set recognition (FSOR) and propose a novel solution, called Relative Feature Displacement Network (RFDNet), which empowers FSL systems to reject queries from unknown classes while accurately classifying those from known classes. First, we suggest a different relative feature displacement learning (RFDL) paradigm for FSOR, i.e., meta-learning a feature displacement relative to a pretrained reference feature embedding, based on our insightful observations on the randomness drift issue of previous meta-learning based for FSOR methods, as well as the generalization ability of the feature embedding pretrained for general classification. Second, we design the RFDNet framework to implement the RFDL paradigm, which is mainly featured by a task-aware RFD generator and a marginal open-set loss. Comprehensive experiments on three public datasets, i.e., miniImageNet, CIFAR-FS and tieredImageNet, demonstrate that RFDNet can consistently outperform the state-of-the-art methods, achieving improvement of 5.2%, 2.0% and 1.7% respectively, in terms of AUROC for unknown-class rejection under the 5-way 5-shot setting.
Shule Deng, Jin-Gang Yu, Zihao Wu 0004, Hongxia Gao, Yansheng Li 0001, Yang Yang 0066
IEEE Trans. Multim.5
2022 Learning Soft Estimator of Keypoint Scale and Orientation with Probabilistic Covariant Loss
abstract
Estimating keypoint scale and orientation is crucial to extracting invariant features under significant geometric changes. Recently, the estimators based on self-supervised learning have been designed to adapt to complex imaging conditions. Such learning-based estimators generally predict a single scalar for the keypoint scale or orientation, called hard estimators. However, hard estimators are difficult to handle the local patches containing structures of different objects or multiple edges. In this paper, a Soft Self-Supervised Estimator (S3Esti) is proposed to overcome this problem by learning to predict multiple scales and orientations. S3Esti involves three core factors. First, the estimator is constructed to predict the discrete distributions of scales and orientations. The elements with high confidence will be kept as the final scales and orientations. Second, a probabilistic covariant loss is proposed to improve the consistency of the scale and orientation distributions under different transformations. Third, an optimization algorithm is designed to minimize the loss function, whose convergence is proved in theory. When combined with different keypoint extraction models, S3Esti generally improves over 50% accuracy in image matching tasks under significant viewpoint changes. In the 3D reconstruction task, S3Esti decreases more than 10% reprojection error and improves the number of registered images. [code release]
Pei Yan, Yihua Tan, Shengzhou Xiong, Yuan Tai, Yansheng Li 0001
CVPR5
2022 Hierarchical Memory Learning for Fine-Grained Scene Graph Generation
Youming Deng, Yansheng Li 0001, Yongjun Zhang 0002, Xiang Xiang 0001, Jian Wang 0108, Jingdong Chen, Jiayi Ma 0001
ECCV (27)2
2022 RLPath: a knowledge graph link prediction method using reinforcement learning based attentive relation path searching and representation learning
Ling Chen 0001, Jun Cui 0003, Xing Tang 0006, Yuntao Qian, Yansheng Li 0001, Yongjun Zhang 0002
Appl. Intell.5
2022 KLGCN: Knowledge graph-aware Light Graph Convolutional Network for recommender systems
Fei Wang 0094, Yansheng Li 0001, Yongjun Zhang 0002
Expert Syst. Appl.2
2022 Combining deep learning and ontology reasoning for remote sensing image semantic segmentation
Yansheng Li 0001, Song Ouyang, Yongjun Zhang 0002
Knowl. Based Syst.1
2022 DACHA: A Dual Graph Convolution Based Temporal Knowledge Graph Representation Learning Method Using Historical Relation
abstract
Temporal knowledge graph (TKG) representation learning embeds relations and entities into a continuous low-dimensional vector space by incorporating temporal information. Latest studies mainly aim at learning entity representations by modeling entity interactions from the neighbor structure of the graph. However, the interactions of relations from the neighbor structure of the graph are neglected, which are also of significance for learning informative representations. In addition, there still lacks an effective historical relation encoder to model the multi-range temporal dependencies. In this article, we propose a d ual gr a ph c onvolution network based TKG representation learning method using h istorical rel a tions (DACHA). Specifically, we first construct the primal graph according to historical relations, as well as the edge graph by regarding historical relations as nodes. Then, we employ the dual graph convolution network to capture the interactions of both entities and historical relations from the neighbor structure of the graph. In addition, the temporal self-attentive historical relation encoder is proposed to explicitly model both local and global temporal dependencies. Extensive experiments on two event based TKG datasets demonstrate that DACHA achieves the state-of-the-art results.
Ling Chen 0001, Xing Tang 0006, Yuntao Qian, Yansheng Li 0001, Yongjun Zhang 0002
ACM Trans. Knowl. Discov. Data5
2021 Representation Learning of Remote Sensing Knowledge Graph for Zero-Shot Remote Sensing Image Scene Classification
abstract
Although deep learning has revolutionized remote sensing image scene classification, current deep learning-based approaches highly depend on the massive supervision of the predetermined scene categories and have disappointingly poor performance on new categories which go beyond the predetermined scene categories. In reality, the classification task often has to be extended along with the emergence of new applications that inevitably involve new categories of remote sensing image scenes, so how to make the deep learning model own the inference ability to recognize the remote sensing image scenes from unseen categories becomes incredibly important. By fully exploiting the remote sensing domain characteristic, this paper proposes a novel remote sensing knowledge graph-guided deep alignment network to address zero-shot remote sensing image scene classification. To improve the semantic representation ability of remote sensing-oriented scene categories, this paper, for the first time, tries to generate the semantic representations of remote sensing scene categories by representation learning of remote sensing knowledge graph (SR-RSKG). In addition, this paper proposes a novel deep alignment network with a series of constraints (DAN) to conduct robust cross-modal alignment between visual features and semantic representations. Extensive experiments on one merged remote sensing image scene dataset, which is the integration of multiple publicly open remote sensing image scene datasets, show that the presented SR-RSKG obviously outperforms the existing semantic representation methods (e.g., the natural language processing models and manually annotated attribute vectors), and our proposed DAN shows better performance compared with the state-of- the-art methods under different kinds of semantic representations.
Yansheng Li 0001, Yongjun Zhang 0002, Ruixian Chen, Jingdong Chen
IGARSS1
2021 Rotation Consistency-Preserved Generative Adversarial Networks for Cross-Domain Aerial Image Semantic Segmentation
abstract
Due to its wide applications, aerial image semantic segmentation attracts increasing research interest in recent years. As well known, deep semantic segmentation network (DSSN) has been widely used to deal with aerial image segmentation and achieves spectacular success. However, when applying the DSSN trained with the labeled aerial images (i.e., the source domain) to predict the aerial images acquired with different acquisition conditions (i.e., the target domain), the performance often dramatically degrades. To alleviate the negative influence of cross-domain data shift, this paper proposes a domain adaptation approach to deal with cross-domain aerial image semantic segmentation. More precisely, this paper proposes a novel rotation consistency-preserved generative adversarial network (RCP-GAN) to carry out domain adaptation for mapping aerial images in the source domain to the target domain. Furthermore, the mapped aerial imageries with labels are used to train DSSN, which is further used to classify aerial imagery in the target domain. To verify the validity of the presented approach, we give two cross-domain experimental settings including: (I) variation of geographic location; (II) variation of both geographic location and imaging mode. Extensive experiments under two typical cross-domain settings show that our proposed method can effectively address the domain shift problem and outperform the state-of-the-art methods with a large margin.
Te Shi 0001, Yansheng Li 0001, Yongjun Zhang 0002
IGARSS2
2021 AG3line: Active grouping and geometry-gradient combined validation for fast line segment extraction
Yongjun Zhang 0002, Yansheng Li 0001
Pattern Recognit.3
2021 Error-Tolerant Deep Learning for Remote Sensing Image Scene Classification
abstract
Due to its various application potentials, the remote sensing image scene classification (RSSC) has attracted a broad range of interests. While the deep convolutional neural network (CNN) has recently achieved tremendous success in RSSC, its superior performances highly depend on a large number of accurately labeled samples which require lots of time and manpower to generate for a large-scale remote sensing image scene dataset. In contrast, it is not only relatively easy to collect coarse and noisy labels but also inevitable to introduce label noise when collecting large-scale annotated data in the remote sensing scenario. Therefore, it is of great practical importance to robustly learn a superior CNN-based classification model from the remote sensing image scene dataset containing non-negligible or even significant error labels. To this end, this article proposes a new RSSC-oriented error-tolerant deep learning (RSSC-ETDL) approach to mitigate the adverse effect of incorrect labels of the remote sensing image scene dataset. In our proposed RSSC-ETDL method, learning multiview CNNs and correcting error labels are alternatively conducted in an iterative manner. It is noted that to make the alternative scheme work effectively, we propose a novel adaptive multifeature collaborative representation classifier (AMF-CRC) that benefits from adaptively combining multiple features of CNNs to correct the labels of uncertain samples. To quantitatively evaluate the performance of error-tolerant methods in the remote sensing domain, we construct remote sensing image scene datasets with: 1) simulated noisy labels by corrupting the open datasets with varying error rates and 2) real noisy labels by deploying the greedy annotation strategies that are practically used to accelerate the process of annotating remote sensing image scene datasets. Extensive experiments on these datasets demonstrate that our proposed RSSC-ETDL approach outperforms the state-of-the-art approaches.
Yansheng Li 0001, Yongjun Zhang 0002, Zhihui Zhu
IEEE Trans. Cybern.1
2021 Learning Deep Cross-Modal Embedding Networks for Zero-Shot Remote Sensing Image Scene Classification
abstract
Due to its wide applications, remote sensing (RS) image scene classification has attracted increasing research interest. When each category has a sufficient number of labeled samples, RS image scene classification can be well addressed by deep learning. However, in the RS big data era, it is extremely difficult or even impossible to annotate RS scene samples for all the categories in one time as the RS scene classification often needs to be extended along with the emergence of new applications that inevitably involve a new class of RS images. Hence, the RS big data era fairly requires a zero-shot RS scene classification (ZSRSSC) paradigm in which the classification model learned from training RS scene categories obeys the inference ability to recognize the RS image scenes from unseen categories, in common with the humans’ evolutionary perception ability. Unfortunately, zero-shot classification is largely unexploited in the RS field. This article proposes a novel ZSRSSC method based on locality-preservation deep cross-modal embedding networks (LPDCMENs). The proposed LPDCMENs, which can fully assimilate the pairwise intramodal and intermodal supervision in an end-to-end manner, aim to alleviate the problem of class structure inconsistency between two hybrid spaces (i.e., the visual image space and the semantic space). To pursue a stable and generalization ability, which is highly desired for ZSRSSC, a set of explainable constraints is specially designed to optimize LPDCMENs. To fully verify the effectiveness of the proposed LPDCMENs, we collect a new large-scale RS scene data set, including the instance-level visual images and class-level semantic representations (RSSDIVCS), where the general and domain knowledge is exploited to construct the class-level semantic representations. Extensive experiments show that the proposed ZSRSSC method based on LPDCMENs can obviously outperform the state-of-the-art methods, and the domain knowledge further improves the performance of ZSRSSC compared with the general knowledge. The collected RSSDIVCS will be made publicly available along with this article.
Yansheng Li 0001, Zhihui Zhu, Jin-Gang Yu, Yongjun Zhang 0002
IEEE Trans. Geosci. Remote. Sens.1
2020 Deep Networks Under Block-Level Supervision for Pixel-Level Cloud Detection in Multi-Spectral Satellite Imagery
abstract
Cloud cover hinders the usability of optical remote sensing imagery. Existing cloud detection methods either require hand-crafted features or utilize deep networks. Generally, deep networks perform better than hand-crafted features. However, deep networks for cloud detection need massive and expensive pixel-level annotation labels. To alleviate that, this paper proposes a weakly supervised deep learning-based cloud detection method using only block-level labels, with a new global convolutional pooling operation and a local pooling pruning strategy to improve the performance. For evaluating, we collect a training dataset containing over 160,000 image blocks with block-level labels and a testing dataset including ten large image scenes with pixel-level labels. Even under extremely weak supervision, our method performed well with the average overall accuracy reached 97.2 %. Experiments demonstrate that our proposed method obviously outperforms the state-of-the-art methods.
Wei Chen 0089, Yansheng Li 0001, Yongjun Zhang 0002, Xiaolong Hao
IGARSS2
2020 A CNN-GCN Framework for Multi-Label Aerial Image Scene Classification
abstract
As one of the fundamental tasks in aerial image understanding, multi-label aerial image scene classification attracts increasing research interest. In general, the semantic category of a scene is reflected by the object information and the topological relations among objects. Most of existing deep learning-based aerial image scene classification methods (e.g., convolutional neural network (CNN)) classify the image scene by perceiving object information, while how to learn spatial relationships from image scene is still a challenging problem. In literature, graph convolutional network (GCN) has been successfully used for learning spatial characteristics of topological data, but it is rarely adopted in aerial image scene classification. To simultaneously mine both the object visual information and spatial relationships among multiple objects, this paper proposes a novel framework combining CNN and GCN to address multi-label aerial image scene classification. Extensive experimental results on two public datasets show that our proposed method can achieve better performance than the state-of-the-art methods.
Yansheng Li 0001, Ruixian Chen, Yongjun Zhang 0002
IGARSS1
2020 Unsupervised Style Transfer via Dualgan for Cross-Domain Aerial Image Classification
abstract
Due to its wide applications, aerial image classification, which is also called semantic segmentation of aerial imagery, attracts increasing research interest in recent years. Until now, deep semantic segmentation network (DSSN) has been widely adopted to address aerial image classification and achieves tremendous success. However, the superior performance of DSSN highly depends on massive targeted data with labels. When DSSN is trained on data from the source domain but tested on data from the target domain, the performance of DSSN is often very limited due to the data shift between source and target domains. To alleviate the disadvantage influence of data shift, this paper proposes a domain adaptation approach via unsupervised style transfer to cope with cross-domain aerial image classification. More specifically, this paper innovatively recommends DualGAN to conduct unsupervised style transfer for mapping aerial images in the source domain to the target domain. The mapped aerial imagery with labels is adopted to train DSSN, which is further used to classify aerial imagery in the target domain. To verify the validity of the presented approach, we give two cross-domain experimental settings including: (I) variation of geographic location; (II) variation of both geographic location and imaging mode. Extensive experiments under two typical cross-domain settings show that our proposed method can obviously outperform the state-of-the-art methods.
Yansheng Li 0001, Te Shi 0001, Wei Chen 0089, Yongjun Zhang 0002, Zhibin Wang 0004, Hao Li 0030
IGARSS1
2020 Infrared Small Target Detection via Low-Rank Tensor Completion With Top-Hat Regularization
abstract
Infrared small target detection technology is one of the key technologies in the field of computer vision. In recent years, several methods have been proposed for detecting small infrared targets. However, the existing methods are highly sensitive to challenging heterogeneous backgrounds, which are mainly due to: 1) infrared images containing mostly heavy clouds and chaotic sea backgrounds and 2) the inefficiency of utilizing the structural prior knowledge of the target. In this article, we propose a novel approach for infrared small target detection in order to take both the structural prior knowledge of the target and the self-correlation of the background into account. First, we construct a tensor model for the high-dimensional structural characteristics of multiframe infrared images. Second, inspired by the low-rank background and morphological operator, a novel method based on low-rank tensor completion with top-hat regularization is proposed, which integrates low-rank tensor completion and a ring top-hat regularization into our model. Third, a closed solution to the optimization algorithm is given to solve the proposed tensor model. Furthermore, the experimental results from seven real infrared sequences demonstrate the superiority of the proposed small target detection method. Compared with traditional baseline methods, the proposed method can not only achieve an improvement in the signal-to-clutter ratio gain and background suppression factor but also provide a more robust detection model in situations with low false-positive rates.
Hu Zhu, Shiming Liu, Lizhen Deng, Yansheng Li 0001, Fu Xiao 0001
IEEE Trans. Geosci. Remote. Sens.4
2020 Exemplar-Based Recursive Instance Segmentation With Application to Plant Image Analysis
abstract
Instance segmentation is a challenging computer vision problem which lies at the intersection of object detection and semantic segmentation. Motivated by plant image analysis in the context of plant phenotyping, a recently emerging application field of computer vision, this paper presents the Exemplar-Based Recursive Instance Segmentation (ERIS) framework. A three-layer probabilistic model is firstly introduced to jointly represent hypotheses, voting elements, instance labels and their connections. Afterwards, a recursive optimization algorithm is developed to infer the maximum a posteriori (MAP) solution, which handles one instance at a time by alternating among the three steps of detection, segmentation and update. The proposed ERIS framework departs from previous works mainly in two respects. First, it is exemplar-based and model-free, which can achieve instance-level segmentation of a specific object class given only a handful of (typically less than 10) annotated exemplars. Such a merit enables its use in case that no massive manually-labeled data is available for training strong classification models, as required by most existing methods. Second, instead of attempting to infer the solution in a single shot, which suffers from extremely high computational complexity, our recursive optimization strategy allows for reasonably efficient MAP-inference in full hypothesis space. The ERIS framework is substantialized for the specific application of plant leaf segmentation in this work. Experiments are conducted on public benchmarks to demonstrate the superiority of our method in both effectiveness and efficiency in comparison with the state-of-the-art.
Jin-Gang Yu, Yansheng Li 0001, Changxin Gao, Hongxia Gao, Gui-Song Xia, Zhu Liang Yu, Yuanqing Li 0001
IEEE Trans. Image Process.2
2019 Learning Deep Networks under Noisy Labels for Remote Sensing Image Scene Classification
abstract
The deep convolutional neural network (DCNN) has been successfully applied in remote sensing (RS) image scene classification, and the superior performances of DCNNs highly depend on not only the large number of samples, but also the high accuracy of labels. In practice, collecting definitely accurate labels for a large-scale RS image scene dataset needs large amounts of manual intervention, but collecting roughly noisy labels for a dataset would be significantly simplified in the RS scenario. Therefore, how to robustly learn the superior DCNN-based classification model from one RS image scene dataset containing some error labels is a practical problem of great importance. To this end, this paper proposes a simple but effective RS-oriented error-tolerant deep learning (RS-ETDL) approach to mitigate the adverse effect of incorrect labels of the corrupted RS image scene dataset. In our proposed RS-ETDL method, learning multi-view DCNN models and correcting error labels are jointly conducted in an iteratively alternative manner. Extensive experiments on the noisy RS image scene dataset demonstrate that our proposed method outperforms the state-of-the-art approaches with a large margin.
Yansheng Li 0001, Yongjun Zhang 0002, Zhihui Zhu
IGARSS1
2019 Scene Context-Driven Vehicle Detection in High-Resolution Aerial Images
abstract
As the spatial resolution of remote sensing images is improving gradually, it is feasible to realize “scene-object” collaborative image interpretation. Unfortunately, this idea is not fully utilized in vehicle detection from high-resolution aerial images, and most of the existing methods may be promoted by considering the variability of vehicle spatial distribution in different image scenes and treating vehicle detection tasks scene-specific. With this motivation, a scene context-driven vehicle detection method is proposed in this paper. At first, we perform scene classification using the deep learning method and, then, detect vehicles in roads and parking lots separately through different vehicle detectors. Afterward, we further optimize the detection results using different postprocessing rules according to different scene types. Experimental results show that the proposed approach outperforms the state-of-the-art algorithms in terms of higher detection accuracy rate and lower false alarm rate.
Chao Tao 0001, Li Mi, Yansheng Li 0001, Ji Qi 0001
IEEE Trans. Geosci. Remote. Sens.3
2018 Salient Object Detection Via Double Sparse Representations Under Visual Attention Guidance
abstract
This paper introduces a novel method for salient object detection from the perspective of sparse representation under visual attention guidance. After pretreatment and regional analysis with eye fixation detection and multi scale segmentation, regions that are used to make up the foreground and background dictionaries are respectively selected by sorting the visual attraction level of all image regions. For saliency measurement, the reconstruction errors instead of common local and global contrasts are used as the saliency indicator, which is expected to improve the object integrity. In addition, the multi scale workflow is conductive to enhance the robustness for objects of different sizes. The proposed method was compared to six state-of-the-art saliency detection methods using three benchmark datasets, and it was confirmed to have more favorable performance in the detection of multiple objects as well as maintaining the integrity of the object area.
Yongjun Zhang 0002, Xunwei Xie, Yansheng Li 0001
IGARSS4
2018 Adaptive top-hat filter based on quantum genetic algorithm for infrared small target detection
Lizhen Deng, Hu Zhu, Quan Zhou 0004, Yansheng Li 0001
Multim. Tools Appl.4
2018 Robust infrared small target detection using local steering kernel reconstruction
Yansheng Li 0001, Yongjun Zhang 0002
Pattern Recognit.1
2018 Learning Source-Invariant Deep Hashing Convolutional Neural Networks for Cross-Source Remote Sensing Image Retrieval
abstract
Due to the urgent demand for remote sensing big data analysis, large-scale remote sensing image retrieval (LSRSIR) attracts increasing attention from researchers. Generally, LSRSIR can be divided into two categories as follows: uni-source LSRSIR (US-LSRSIR) and cross-source LSRSIR (CS-LSRSIR). More specifically, US-LSRSIR means the inquiry remote sensing image and images in the searching data set come from the same remote sensing data source, whereas CS-LSRSIR is designed to retrieve remote sensing images with a similar content to the inquiry remote sensing image that are from a different remote sensing data source. In the literature, US-LSRSIR has been widely exploited, but CS-LSRSIR is rarely discussed. In practical situations, remote sensing images from different kinds of remote sensing data sources are continually increasing, so there is a great motivation to exploit CS-LSRSIR. Therefore, this paper focuses on CS-LSRSIR. To cope with CS-LSRSIR, this paper proposes source-invariant deep hashing convolutional neural networks (SIDHCNNs), which can be optimized in an end-to-end manner using a series of well-designed optimization constraints. To quantitatively evaluate the proposed SIDHCNNs, we construct a dual-source remote sensing image data set that contains eight typical land-cover categories and 10 000 dual samples in each category. Extensive experiments show that the proposed SIDHCNNs can yield substantial improvements over several baselines involving the most recent techniques.
Yansheng Li 0001, Yongjun Zhang 0002, Xin Huang 0002, Jiayi Ma 0001
IEEE Trans. Geosci. Remote. Sens.1
2018 Large-Scale Remote Sensing Image Retrieval by Deep Hashing Neural Networks
abstract
As one of the most challenging tasks of remote sensing big data mining, large-scale remote sensing image retrieval has attracted increasing attention from researchers. Existing large-scale remote sensing image retrieval approaches are generally implemented by using hashing learning methods, which take handcrafted features as inputs and map the high-dimensional feature vector to the low-dimensional binary feature vector to reduce feature-searching complexity levels. As a means of applying the merits of deep learning, this paper proposes a novel large-scale remote sensing image retrieval approach based on deep hashing neural networks (DHNNs). More specifically, DHNNs are composed of deep feature learning neural networks and hashing learning neural networks and can be optimized in an end-to-end manner. Rather than requiring to dedicate expertise and effort to the design of feature descriptors, we can automatically learn good feature extraction operations and feature hashing mapping under the supervision of labeled samples. To broaden the application field, DHNNs are evaluated under two representative remote sensing cases: scarce and sufficient labeled samples. To make up for a lack of labeled samples, DHNNs can be trained via transfer learning for the former case. For the latter case, DHNNs can be trained via supervised learning from scratch with the aid of a vast number of labeled samples. Extensive experiments on one public remote sensing image data set with a limited number of labeled samples and on another public data set with plenty of labeled samples show that the proposed remote sensing image retrieval approach based on DHNNs can remarkably outperform state-of-the-art methods under both of the examined conditions.
Yansheng Li 0001, Yongjun Zhang 0002, Xin Huang 0002, Hu Zhu, Jiayi Ma 0001
IEEE Trans. Geosci. Remote. Sens.1
2017 Feature guided Gaussian mixture model with semi-supervised EM and local geometric constraint for retinal image registration
Jiayi Ma 0001, Junjun Jiang, Chengyin Liu, Yansheng Li 0001
Inf. Sci.4
2017 Object-Based Visual Saliency via Laplacian Regularized Kernel Regression
abstract
Saliency object detection has been a very active research topic recently, due to its extensive applications in image compression, scene understanding, image retrieval, and so forth. The overwhelming majority of existing computational models are designed based on computer vision techniques by using a lot of image cues and priors. In fact, salient object detection is derived from the biological perceptual mechanism, and biological evidence shows that the object-based saliency stems from the spread of the spatial attention. Inspired by this, we attempt to utilize the emerging spread mechanism of object attention to construct a new computational model. A novel Laplacian regularized kernel regression diffusion model is proposed to fulfill the spread process. The proposed diffusion model, which is able to fully capture both global and local structures of the image, thereby allows for effective propagation of spatial attention with visual grouping cues, yielding a well-structured object-based saliency map. Experimental results demonstrate that our method can achieve encouraging performance in comparison with the state-of-the-art methods.
Hao Dou, Delie Ming, Zhihong Pan 0002, Yansheng Li 0001, Jinwen Tian
IEEE Trans. Multim.5
2016 A novel spatio-temporal saliency approach for robust dim moving target detection from airborne infrared image sequences
Yansheng Li 0001, Yongjun Zhang 0002, Jin-Gang Yu, Yihua Tan, Jinwen Tian, Jiayi Ma 0001
Inf. Sci.1
2016 Unsupervised Multilayer Feature Learning for Satellite Image Scene Classification
abstract
This letter proposes a simple but effective approach to automatically learn a multilayer image feature for satellite image scene classification. Different from the hand-crafted features which are empirically designed but lack high generalization ability, the proposed approach can autonomously extract the data-dependent feature. The presented feature extraction algorithm is composed of two layers, and the bases of these two layers are uniformly learned by a plain $K$-means clustering algorithm. Coincidentally, the feature extraction performance of the aforementioned two layers is consistent with visual processing of human visual cortex. More specifically, the first layer can generate edgelike bases, which are analogous to the neuron responses of primary visual cortex (V1), and the second layer can produce cornerlike bases, which resemble the neuron responses of visual extrastriate cortical area two (V2). The proposed feature extraction approach can automatically extract not only simple structure features (e.g., edges) but also complex structure features (e.g., corners and junctions). The learned feature is further discriminated by the linear support vector machine classifier for scene classification. In order to fairly demonstrate the validity of the proposed feature extraction approach, its satellite image scene classification performance is evaluated on the public UCM-21 data set. Experimental results show that the proposed approach can outperform several recent state-of-the-art approaches.
Yansheng Li 0001, Chao Tao 0001, Yihua Tan, Jinwen Tian
IEEE Geosci. Remote. Sens. Lett.1
2015 Kernel regression in mixed feature spaces for spatio-temporal saliency detection
Yansheng Li 0001, Yihua Tan, Jin-Gang Yu, Shengxiang Qi, Jinwen Tian
Comput. Vis. Image Underst.1
2015 Salient object detection via contrast information and object vision organization cues
Shengxiang Qi, Jin-Gang Yu, Jie Ma 0003, Yansheng Li 0001, Jinwen Tian
Neurocomputing4
2015 Built-Up Area Detection From Satellite Images Using Multikernel Learning, Multifield Integrating, and Multihypothesis Voting
abstract
This letter proposes a novel supervised approach for accurate built-up area detection from high-resolution remote sensing images. In existing supervised built-up area detection approaches based on block-based image interpretation, the determination of the block size and the pursuit of the pixel-level result are not well addressed. Concerning these issues, this letter proposes a complete and systematic approach. It first utilizes multikernel learning to incorporate multiple features to implement the block-level image interpretation. Then, multifield integrating (i.e., the image interpretation results using different block sizes are fused) is proposed to obtain the block-level result. On the basis of the achieved result of the second step, multihypothesis voting is finally presented for working toward the pixel-level built-up area detection result through multihypothesis superpixel representation and graph smoothing. The proposed approach has been validated in the ZY-3 and GF-1 satellite images, and experimental results show that the proposed approach can outperform the state-of-the-art approaches.
Yansheng Li 0001, Yihua Tan, Shengxiang Qi, Jinwen Tian
IEEE Geosci. Remote. Sens. Lett.1
2015 Unsupervised Ship Detection Based on Saliency and S-HOG Descriptor From Optical Satellite Images
abstract
With the development of high-resolution imagery, ship detection in optical satellite images has attracted a lot of research interest because of the broad applications in fishery management, vessel salvage, etc. Major challenges for this task include cloud, wave, and wake clutters, and even the variability of ship sizes. In this letter, we propose an unsupervised ship detection method toward overcoming these existing issues. Visual saliency, which focuses on highlighting salient signals from scenes, is applied to extract candidate regions followed by a homogeneous filter presented to confirm suspected ship targets with complete profiles. Then, a novel descriptor, ship histogram of oriented gradient, which characterizes the gradient symmetry of ship sides, is provided to discriminate real ships. Experimental results on numerous panchromatic satellite images demonstrate the good performance of our method compared to state-of-the-art methods.
Shengxiang Qi, Jie Ma 0003, Yansheng Li 0001, Jinwen Tian
IEEE Geosci. Remote. Sens. Lett.4
2015 Unsupervised Spectral-Spatial Feature Learning With Stacked Sparse Autoencoder for Hyperspectral Imagery Classification
abstract
In this letter, different from traditional methods using original spectral features or handcraft spectral-spatial features, we propose to adaptively learn a suitable feature representation from unlabeled data. This is achieved by learning a feature mapping function based on stacked sparse autoencoder. Considering that hyperspectral imagery (HSI) is intrinsically defined in both the spectral and spatial domains, we further establish two variants of feature learning procedures for sparse spectral feature learning and multiscale spatial feature learning. Finally, we embed the learned spectral-spatial feature into a linear support vector machine for classification. Experiments on two hyperspectral images indicate the following: 1) the learned spectral-spatial feature representation is more discriminative for HSI classification compared to previously hand-engineered spectral-spatial features, especially when the training data are limited and 2) the learned features appear not to be specific to a particular image but general in that they are applicable to multiple related images (e.g., images acquired by the same sensor but varying with location or time).
Chao Tao 0001, Yansheng Li 0001, Zhengrou Zou
IEEE Geosci. Remote. Sens. Lett.3
2014 Urban building extraction via visual graphical topic model
abstract
This paper addresses the automatic building extraction problem from high-resolution remote sensing images. The buildings in remote sensing images generally represent different shapes (i.e., simple rectangular or complex hybrid shape), it is intractable to extract all the buildings with different shapes once. Therefore, we adopt a hierarchical extraction style with a visual graphical topic model embedded, which includes two stages: the first stage detects the simple rectangular buildings and the second stage extracts the complex hybrid buildings. More specifically, the first stage is mainly responsible for regular buildings detection, and unsupervised visual graphical topic model (i.e., replicated softmax restricted boltzmann machine) and supervised discriminative model learning, and the second stage is mainly in charge of complex buildings extraction using the learned semantic feature mapping and discriminative model. Experimental results show that the second stage can obviously improve the building detection rate with slightly increasing the false alarm rate.
Yansheng Li 0001, Yihua Tan, Jinwen Tian
IGARSS1