VLDB 2026 Research / reviewers in the wild / expert
Shengyang Li
dblp:125/9522
· DBLP profile ↗
46ranked-venue papers
5as first author
36since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 30 · 2 first-author · 22 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PoFEL: Energy-Efficient Consensus for Blockchain-Based Hierarchical Federated LearningabstractFacilitated by mobile edge computing, client-edge-cloud hierarchical federated learning (HFL) enables communication-efficient model training in a widespread area but also incurs additional security and privacy challenges from intermediate model aggregations and remains vulnerable to the single point of failure issue. To tackle these challenges, we propose a blockchain-based HFL (BHFL) system that operates a permissioned blockchain among edge servers for model aggregation without the need for a centralized cloud server. The employment of blockchain, however, introduces additional overhead. To enable a compact and efficient workflow, we design a novel lightweight consensus algorithm, named Proof of Federated Edge Learning (PoFEL), to reuse computational work performed for local model training. Specifically, the leader node is selected by evaluating the intermediate FEL models from all edge servers instead of other additional mechanisms used solely for leader elections. This design thus improves the system efficiency compared with traditional BHFL frameworks. To prevent model plagiarism and bribery voting during the consensus process, we propose Hash-based Commitment and Digital Signature (HCDS) and Bayesian Truth Serum-based Voting (BTSV) schemes. Finally, we devise an incentive mechanism to motivate continuous contributions from clients to the learning task. Experimental results demonstrate that our proposed BHFL system with the corresponding consensus protocol and incentive mechanism achieves effectiveness, low computational cost, and fairness. Shengyang Li, Qin Hu 0001, Zhilin Wang, Minghui Xu 0001, Zhipeng Cai 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2025 | Zs-Drosophila: Learning Transferable Representations for Drosophila Behavior Analysis Via Visual-Language Hyperbolic AlignmentabstractAutomatic behavioral analysis of Drosophila has drawn consistent research interest in the field, as understanding quantitative behavior plays a crucial role in neuroscience, genetics, and space biology. Existing machine learning approaches often rely heavily on expert domain knowledge and manual annotations to build supervised models, limiting their scalability and adaptability (e.g., to novel actions and unseen domains). Meanwhile, recent advances in vision-language models offer a more flexible, expressive, and interpretable medium for behavior representation, enabling broader semantic understanding and generalization. To this end, we propose ZS-Drosophila, the first framework that introduces language-guided multimodal alignment for Drosophila behavior analysis. Our foundation model ZS-Drosophila learns transferable behavior representations that can generalize to unseen behaviors and domains. Furthermore, we construct SpaceAnimal-Drosophila, a benchmark dataset of Drosophila video recordings collected both on Earth and in space, comprising annotated skeletal sequences and behaviorlanguage pairs as ground truths. It has a series of evaluation protocols to showcase the strong transferability of ZS-Drosophila, achieved with only prompt tuning at test time. We demonstrate the capabilities of our method in the following scenarios: (1) conventional supervised action recognition on on-Earth data, (2) zero-shot recognition of unseen behaviors, and (3) crossdomain generalization to in-orbit microgravity data. Notably, our proposed model enables zero-shot recognition of novel behaviors potentially induced by microgravity without requiring additional annotations, which, to our knowledge, is the first attempt in the field. Kang Liu 0020, Han Wang 0049, Yixuan Lv, Shengyang Li, Jianing You |
BIBM | 6 |
| 2025 | Enhance Document-Level Event Argument Extraction Through Text Diffusion and Event Argument ConstraintsabstractEvent Extraction (EE) is a critical task in space science, pivotal in supporting task analysis, knowledge discovery, and decision-making processes. The rapid development of China's manned space program has resulted in the emergence of a large volume of complex events and associated parameter information. While prompt tuning has demonstrated robust performance in sentence-level event extraction tasks within this domain, its effectiveness diminishes when applied to long-span event extraction across multiple documents. To address this challenge, we propose an enhanced prompt tuning method to improve the model's ability to fuse information and increase extraction accuracy. This approach is composed of two key components: 1) Text Diffusion Module: This module reconstructs noisy text based on encoded prompt statements, thereby augmenting the model's capacity for information integration and mitigating the issue of incomplete or noisy data in such environments. 2) Argument Constraint Module: Utilizing a Cross-Attention Transformer, this module integrates information from other event arguments, thereby enhancing extraction accuracy by addressing the problem of incomplete argument integration. Experimental results demonstrate the effectiveness of our proposed method, with improvements of 3.5% in the average F1 score for the Bart-base model and 3.0% for the Bart-large model across three distinct datasets. Additionally, ablation experiments confirm that these enhancements are both necessary and effective. Yunfei Liu 0003, Shiyi Hao, Shengyang Li |
CSCWD | 4 |
| 2025 | Cross-Modal Ship Re-Identification via Optical and SAR Imagery: A Novel Dataset and MethodabstractDetecting and tracking ground objects using earth observation imagery remains a significant challenge in the field of remote sensing. Continuous maritime ship tracking is crucial for applications such as maritime search and rescue, law enforcement, and shipping analysis. However, most current ship tracking methods rely on geostationary satellites or video satellites. The former offer low resolution and are susceptible to weather conditions, while the latter have short filming durations and limited coverage areas, making them less suitable for the real-world requirements of ship tracking. To address these limitations, we present the Hybrid Optical and Synthetic Aperture Radar (SAR) Ship Re-Identification Dataset (HOSS ReID dataset), designed to evaluate the effectiveness of ship tracking using low-Earth orbit constellations of optical and SAR sensors. This approach ensures shorter re-imaging cycles and enables all-weather tracking. HOSS ReID dataset includes images of the same ship captured over extended periods under diverse conditions, using different satellites of different modalities at varying times and angles. Furthermore, we propose a baseline method for cross-modal ship re-identification, TransOSS, which is built on the Vision Transformer architecture. It refines the patch embedding structure to better accommodate cross-modal tasks, incorporates additional embeddings to introduce more reference information, and employs contrastive learning to pre-train on large-scale optical-SAR image pairs, ensuring the model's ability to extract modality-invariant features. Our dataset and baseline method are publicly available on https://github.com/Alioth2000/Hoss-ReID. Shengyang Li, Yixuan Lv |
ICCV | 2 |
| 2025 | MR4SseC: A Multimodal Representation Learning Framework for Space Science Experiment of China's Space StationabstractExisting vision-language contrastive learning frameworks, such as CLIP, focus on aligning the paired image and caption embeddings while pushing non-matching pairs apart, enhancing the models' ability to generalize and facilitating zero-shot learning. However, space science experiments images and their text descriptions exhibit a significant domain difference from the internet images and captions. In this work, we introduce a novel framework designed to learn a cohesive representation for space science experiment images and their textual descriptions. Our approach is based on a novel alignment mechanism that integrates both local and global perspectives. Furthermore, rather than relying solely on image-text contrastive supervision, we leverage the full potential of the data through (1) self-contrastive learning within visual modality; (2) multi-view contrastive learning across modalities, to exponentially expand trainable data with minimal additional expense. We demonstrate that our Multimodal Representation learning framework for Space Science Experiments of China's space station (denoted as MR4SseC) is simple yet effective for various space science image recognition tasks with limited labeled data: Our evaluations confirm its superior performance in classification and image-text retrieval tasks. Yunfei Liu 0003, Anqi Liu 0002, Yunziwei Deng, Shengyang Li |
ICMR | 6 |
| 2025 | SSCNet: Structure-Aware Segmentation Network for C.elegans in Scientific Experiments on China Space Station
Silei Liu, Kang Liu 0020, Han Wang 0049, Yuhan Sun 0004, Shengyang Li |
PRCV (2) | 7 |
| 2025 | Less is more: A semi-supervised fine-grained object detection for satellite video
Shengyang Li |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Fine-grained multi-modal prompt learning for vision-language models
Yunfei Liu 0003, Yunziwei Deng, Anqi Liu 0002, Shengyang Li |
Neurocomputing | 5 |
| 2025 | The Time and Frequency Distribution Characteristics of Interference Signals Based on Artificial Intelligence TechnologyabstractIn this paper, we propose a method with artificial intelligence to optimize manual astronomical observation works. For the large amount of data generated by radio astronomy monitoring, we compile 4 algorithms including VTD, WSV, MAD, and MAS in the procedure of data analysis. Then the platform can recognize the radio interference signals from radio astronomy monitoring data and analyze the spatiotemporal distribution characteristics, and generate reports automatically. Through this method, the amount of work would greatly improve work efficiency and accuracy. The distribution patterns and changes of radio frequency interference signals in the area can be grasped and analyzed efficiently and quickly by astronomy researchers. Shengyang Li, Zhixiang Zhao, Junwen Tang, Ning Fu |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2025 | Fine-Grained Object Detection of Satellite Video in the Frequency DomainabstractSatellite video objects often have small scales and the occlusion of their distinguishable regions. Existing fine-grained object detection methods estimate object locations and categories by enhancing object features and increasing feature differences between categories. However, they fail to account for the impact of limited feature information on fine-grained detection, specifically reflected in two aspects: 1) limited pixels lead to limited features. The small pixel coverage of satellite video objects results in a low upper bound on the available feature information, hindering significant improvements in fine-grained detection accuracy and 2) limited differences exacerbate limitations. Occlusion of distinguishable regions in small-scale objects exacerbates the indistinctness of features between different fine-grained categories. It prevents the network from accurately learning unique features for certain object classes, thereby degrades detector performance. To address these challenges, we propose a frequency auxiliary network (FANet), which integrates frequency domain feature learning into fine-grained object detection networks. Specifically, we propose the spectral augmented module (SAM) to extract multispectral features from various frequency components of satellite video frames to complement spatial-domain features, enabling the network to leverage hidden semantic information from the frequency domain. In addition, to better distinguish small-scale objects from the background and emphasize fine-grained category features, we design the frequency domain attention (FDA) mechanism. FDA assigns dynamic weights to spatial and frequency domain features, suppressing background information and enhancing feature differences between fine-grained categories. Extensive experiments on the SAT-MTB dataset demonstrate that FANet achieves superior performance compared to existing methods. Yuhan Sun 0004, Shengyang Li |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | DFDNet: Deep Feature Decoupling for Oriented Object DetectionabstractObjects in remote sensing images exhibit diverse orientations. Current oriented object detection (OOD) methods estimate object angle by designing different loss functions and bounding box representations. However, these approaches do not account for the effects of coupling between rotation-equivariant and rotation-invariant features on the regression of oriented bounding box (OBB) parameters. We manifest the problem in two aspects, (1)The coupling of parameters with different attributes. Current OOD methods overlook the inherent differences among features representing an object’s location, scale, and angle, making it challenging to accurately predict OBB parameters with different attributes. (2)The coupling of object and background features. Conventional OOD methods apply convolution kernels uniformly across objects and background regions, leading to feature entanglement and degradation in detection performance. To address the above issues, we propose Deep Feature Decoupling Network (DFDNet) to decouple the extracted features. Specifically, we propose Parameter Regression Decoupling (PRD) to separate feature maps based on their attributes, subsequently assigning them to distinct branches for OBB parameter regression. This approach ensures the decoupling of features related to an object’s location, shape, angle, and category. Additionally, to enhance the ability of OOD networks to differentiate between object and background features, we design the Mask Reinforcement Module (MRM), which is integrated into the PRD branches. The MRM dynamically adjusts the weights of object features, suppressing background interference and enhancing the distinction between object and background features. Extensive experiments conducted on the DOTA, HRSC2016, and UCAS-AOD datasets validate the effectiveness of DFDNet, demonstrating that it achieves state-of-the-art performance. Yuhan Sun 0004, Shengyang Li |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | Semantic Affinity-Driven Spatiotemporal Transformer Network for Satellite Video Moving-Object SegmentationabstractSatellite video intelligent processing plays a critical role in Earth observation applications such as traffic monitoring and environmental surveillance. However, moving-object segmentation in satellite videos faces several challenges. First, spatiotemporal redundancy makes it difficult to model long-range dependencies because large-scale scenes with slow background changes lead to fragmented segmentation. Second, semantic ambiguity arises when stationary objects like parked aircraft share category-level similarities with moving targets, which causes false positives. Besides, insufficient feature discrimination occurs as small, rigid objects such as ships exhibit weak texture and edge details under low-resolution imaging. To overcome these issues, we introduce a semantic affinity-driven spatiotemporal Transformer network that leverages a Transformer-based architecture to capture pixel-level dependencies across spatial and temporal dimensions. Furthermore, our network employs a contextual affinity-constrained decoder to suppress category-level interference and integrates a triple-branch feature extractor with edge priors for enhanced contour delineation. Our framework operates in an end-to-end manner without requiring fine-tuning during inference, which ensures deployment efficiency. Extensive experiments on a dataset built upon SAT-MTB demonstrate state-of-the-art performance with a J&F Mean of 71.7%. The proposed method outperforms the baseline by 3.7% with improvements of 4.4% in J-Mean and 3.1% in F-Mean. In addition, it surpasses the optimized SAM2 with a 10.7% higher J-Mean while maintaining a significantly smaller parameter count (34.6 M versus 224 M). Both qualitative and quantitative evaluations confirm the method’s superiority and temporal stability. This work offers a robust and efficient solution for accurate moving-object segmentation in satellite videos. Yixuan Lv, Kang Liu 0020, Han Wang 0049, Shengyang Li, Jianing You, Kailun Zhang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Language-Empowered Conversion for Remote Sensing Image Retrieval With Text FeedbackabstractRemote sensing image retrieval with text feedback (RSIR-TF) presents a challenging multi-modal retrieval task that leverages a reference image, modification text, and scene graph to retrieve the relevant target image from a gallery. Existing approaches rely on cross-modal combiners to integrate multi-modal features extracted separately from modality-specific encoders. However, the modality-specific encoders often suffer from insufficient representational capacity and limited cross-modal alignment due to the lack of effective pre-training. Recently, vision language models (VLMs) pre-trained on large-scale image-text pairs have demonstrated exceptional representation and alignment capabilities in the remote sensing domain. However, these VLMs struggle with processing structured scene graphs, limiting their applicability to tasks like RSIR-TF that require composite reasoning over structured and unstructured modalities. To address these limitations, we propose a novel pipeline, Language-Empowered Conversion (LEmpo), which effectively migrates the large language model (LLM) and VLM to the RSIR-TF task. Firstly, we perform Pseudo Caption Generation and Scene Graph Interpretation powered by LLM to convert structured scene graphs into natural language captions. This conversion bridges the gap between structured scene graphs and unstructured text captions, enabling unified feature extraction and alignment. Subsequently, we employ the pre-trained VLM to extract robust visual and textual features within a joint visual-textual feature space. To fully utilizing the complementary information from visual, textual, and structured data, we introduce a Hybrid Similarity Tuning strategy, which aggregates triplet similarity, language similarity, and pseudo caption similarity into a unified hybrid similarity. The hybrid similarity is optimized during training through vision-fixed tuning, which anchors visual features while refining textual features to enhance alignment with target images. Comprehensive experiments conducted on the Airplane, Tennis and WHIRT datasets demonstrate that LEmpo significantly outperforms all comparison methods, achieving a substantial improvement in recall performance. Shengyang Li, Yuhan Sun 0004, Han Wang 0049 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Parameter-Efficient Reparameterization Tuning for Remote Sensing Image-Text RetrievalabstractVision language models have been gradually adapted to various tasks of remote sensing domain with full fine-tuning paradigm,e.g., remote sensing image-text retrieval (RSITR). Superior performance enhancements have proven the powerful generalization and robustness of vision language models. However, full fine-tuning vision language models is resource-intensive and poses risks of overfitting. Moreover, existing RSITR methods usually assume that remote sensing images correspond to text captions one by one and utilize the bidirectional matching training objective, which is not aligned with evaluation benchmarks and real-world applications. To tackle the mentioned problems, we propose a novel Parameter-Efficient Reparameterization Tuning with Ranking and Matching (PERT-RaMa) framework, which effectively migrates the vision language model (i.e.CLIP) to RSITR task. To overcome the overfitting issue, we build a lightweight, plug-and-play module called Kronecker product for low-rank adaptation (KPLoRA). KPLoRA obtains higher intrinsic rank with fewer parameters. Furthermore, we design the Ranking and Matching (RaMa) training method that converts RSITR task into one-to-one matching and one-to-many ranking, which is aligned with current RSITR benchmarks and accelerates training speed through removal of unessential computations. Comprehensive experiments on three public RSITR benchmarks demonstrate that the effectiveness and efficiency of the proposed retrieval model. Our method outperforms full fine-tuning methods without CLIP by nearly 5-10%, and achieves comparable or superior retrieval capability than CLIP with full fine-tuning and parameter-efficient fine-tuning. Furthermore, our RaMa training method increases the training speed by 5x compared to current training method. Shengyang Li, Manqi Zhao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Video process detection for space electrostatic suspension material experiment in China's Space Station
Manqi Zhao, Shengyang Li |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Target-Aware Transformer for Satellite Video Object TrackingabstractRecent years have witnessed the astonishing development of transformer-based paradigm in single object tracking (SOT) in generic videos. However, due to the fact that the targets of interest in satellite videos are small in size and weak in visual appearance, the advancements of transformer-based paradigm in satellite video object tracking are impeded. To alleviate this issue, a novel transformer-based recipe is proposed, which consists of a bi-direction propagation and fusion (Bi-PF) strategy and a target-aware enhancement (TAE) module. Concretely, we first adopt the Bi-PF strategy to make full use of multiscale information to generate discriminative representations of tracking targets. Then, the TAE module is employed to decouple an object query into content-aware embedding and spatial-aware embedding and produce a target prototype to help get high-quality content-aware embedding. It is worth mentioning that, different from the previous methods in satellite video tracking most of which evaluate their performance using only several videos, we conduct extensive experiments on the SatSOT dataset which consists of 105 videos. In particular, the proposed method achieves the success score of 45.6% and the precision score of 57.6%, surpassing the baseline method by 5.0% and 9.5%, respectively. The code will be released athttps://github.com/laybebe/TATrans_SVOT. Pujian Lai, Meili Zhang, Gong Cheng 0003, Shengyang Li, Xiankai Huang, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | ABFL: Angular Boundary Discontinuity Free Loss for Arbitrary Oriented Object Detection in Aerial ImagesabstractArbitrary oriented object detection (AOOD) in aerial images is a widely concerned and highly challenging task, and plays an important role in many scenarios. The core of AOOD involves the representation, encoding, and feature augmentation of oriented bounding-boxes (Bboxes). Existing methods lack intuitive modeling of angle difference measurement in oriented Bbox representations. Oriented Bboxes under different representations exhibit rotational symmetry with varying periods due to angle periodicity. The angular boundary discontinuity (ABD) problem at periodic boundary positions is caused by rotational symmetry in measuring angular differences. In addition, existing methods also use additional encoding-decoding structures for oriented Bboxes. In this paper, we design an angular boundary free loss (ABFL) based on the von Mises distribution. The ABFL aims to solve the ABD problem when detecting oriented objects. Specifically, ABFL proposes to treat angles as circular data rather than linear data when measuring angle differences, aiming to introduce angle periodicity to alleviate the ABD problem and improve the accuracy of angle difference measurement. In addition, ABFL provides a simple and effective solution for various periodic boundary discontinuities caused by rotational symmetry in AOOD tasks, as it does not require additional encoding-decoding structures for oriented Bboxes. Extensive experiments on the DOTA and HRSC2016 datasets show that the proposed ABFL loss outperforms some state-of-the-art methods focused on addressing the ABD problem. Zifei Zhao, Shengyang Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | MP2Net: Mask Propagation and Motion Prediction Network for Multiobject Tracking in Satellite VideosabstractMainstream multi-object tracking (MOT) algorithms employ global object detection and association methods. However, when dealing with scenarios involving crowded tiny objects in satellite videos, existing global trackers often yield numerous missed detections and unstable trajectories. To address this issue, we propose a novel joint-detection-and-tracking framework, MP2Net, which integrates local detection enhancements for tiny targets and bridges the gap between detection and association. Specifically, our approach incorporates a mask propagation network that enhances feature representation for tiny targets by matching frame-by-frame to capture local details. Additionally, we utilize an implicit and explicit motion prediction strategy that merges tracking information into detection at both feature and instance levels, thereby improving tracking robustness. Experimental results on two large-scale datasets demonstrate the effectiveness and robustness of MP2Net, achieving state-of-the-art performance on typical moving objects in satellite videos, such as 66.7% MOTA and 75.9% IDF1 on the SatVideoDT challenge dataset. The code will be available at https://github.com/DonDominic/MP2Net. Manqi Zhao, Shengyang Li, Han Wang 0049, Yuhan Sun 0004, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Transformer Tracking for Satellite Video: Matching, Propagation, and PredictionabstractRecently, transformer-based trackers have brought overwhelming advantages in general video. However, their performance in satellite video has been hindered by insufficient satellite-specific training and a lack of designs tailored to satellite targets and scene characteristics. To tackle these challenges, we propose a novel transformer-based tracking framework for satellite video object tracking: Transformer Matching, Propagation, and Prediction (TransMPP). TransMPP combines three stages: static matching, dynamic propagation, and prediction, to ensure accurate tracking in satellite videos. Specifically, the Matching model uses a one-stream pipeline for simultaneous feature extraction and relationship modeling across extensive search and template areas, thereby improving foreground and background discrimination capabilities. In addition, the Propagation and Prediction models enhance temporal modeling capabilities through local long-term and short-term feature propagation and global sequence prediction, respectively, boosting tracking robustness. Moreover, to ensure a fair comparison and evaluation, we also developed SatSOT-train, a large-scale training dataset for the SatSOT benchmark. After comprehensive training, TransMPP demonstrates state-of-the-art (SOTA) performance on the SatSOT dataset, achieving an area under the curve (AUC) score of 59.9% and a precision score of 71.5%, bringing improvements of 6.3% and 5.3%, respectively. The code will be available athttps://github.com/DonDominic/TransMPP. Manqi Zhao, Shengyang Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Satellite Video Multi-Label Scene Classification With Spatial and Temporal Feature Cooperative Encoding: A Benchmark Dataset and MethodabstractSatellite video multi-label scene classification predicts semantic labels of multiple ground contents to describe a given satellite observation video, which plays an important role in applications like ocean observation, smart cities, et al. However, the lack of a high-quality and large-scale dataset prevents further improvement of the task. And existing methods on general videos have the difficulty to represent the local details of ground contents when directly applied to the satellite videos. In this paper, our contributions include (1) we develop the first publicly available and large-scale satellite video multi-label scene classification dataset. It consists of 18 classes of static and dynamic ground contents, 3549 videos, and 141960 frames. (2) we propose a baseline method with the novel Spatial and Temporal Feature Cooperative Encoding (STFCE). It exploits the relations between local spatial and temporal features, and models long-term motion information hidden in inter-frame variations. In this way, it can enhance features of local details and obtain the powerful video-scene-level feature representation, which raises the classification performance effectively. Experimental results show that our proposed STFCE outperforms 13 state-of-the-art methods with a global average precision (GAP) of 0.8106 and the careful fusion and joint learning of the spatial, temporal, and motion features are beneficial to achieve a more robust and accurate model. Moreover, benchmarking results show that the proposed dataset is very challenging and we hope it could promote further development of the satellite video multi-label scene classification task. Weilong Guo, Shengyang Li, Feixiang Chen, Yuhan Sun 0004, Yanfeng Gu |
IEEE Trans. Image Process. | 2 |
| 2023 | Interferometric Calibration of Wide Swath Altimeters with Inclined Baseline Using Inland LakesabstractWide swath altimeters are a new generation of radar altimeters, which can provide elevations of both the ocean and the inland surface waters with unprecedented swath and spatial resolution. Interferometric calibration that gives estimates of the unknown variations of the baseline length, the baseline roll angle, and the phase offset, are essential to deriving topography accurately. Current methods are mainly developed on a cross-track height error model for systems with a horizontal baseline. However, for altimeters with a deliberately-inclined baseline, the height error model turns into a complicated form. This paper proposes a calibration method for wide swath altimeters with an inclined baseline using phase measurements from large inland lakes. The validity of the method has been confirmed by simulated buoy samples and together with real data acquired by Interferometric Imaging Radar Altimeter (InIRA), which was equipped on the Tiangong-2 space laboratory of China. Hong Tan, Shengyang Li |
IGARSS | 2 |
| 2023 | SSUIE 1.0: A Dataset for Chinese Space Science and Utilization Information Extraction
Yunfei Liu 0003, Shengyang Li, Chen Wang 0091, Yifeng Zheng 0002, Shiyi Hao |
NLPCC (2) | 2 |
| 2023 | A multi-frame sparse self-learning PWC-Net for motion estimation in satellite video scenes
Tengfei Wang 0001, Yanfeng Gu, Shengyang Li |
Sci. China Inf. Sci. | 3 |
| 2023 | Siamese Graph Attention Networks for robust visual object trackingabstractSiamese-based trackers usually convert the object tracking task into a similarity matching problem between the target template and the search region. Since fixed or manually updated templates are not robust when tracking moving objects with dramatically changing appearance, this paper proposes an improved siamese graph attention network with adaptive template update called SiamGT. By establishing spatiotemporal and context dependencies between historical images and search regions, a frame selection mechanism is added to improve the richness of information. In addition, a graph attention network with residual connections is used in the template update mechanism which enables the propagation and aggregation of information to generate robust templates. Extensive experimental results on challenging benchmarks such as UAV123, OTB100, and VOT2019 demonstrate that the proposed SiamGT has achieved state-of-the-art performance in visual object tracking. Shengyang Li, Weilong Guo, Manqi Zhao, Yunfei Liu 0003 |
Comput. Vis. Image Underst. | 2 |
| 2023 | Improved 2-D Inland Water Surface Elevations of Wide Swath Altimeters Using Surrounding Lakeshores and RiverbanksabstractTwo-dimensional (2-D) water surface elevations (WSEs) can be measured for the first time using a new generation of radar altimeter, which uses the interferometric synthetic aperture radar (InSAR) technique on near-nadir swath to measure elevations of both the inland surface waters and the ocean with unprecedented swath and spatial resolution. The interferometric imaging radar altimeter (InIRA) in the Tiangong-2 space laboratory of China is the first spaceborne interferometric radar altimeter, which was launched in 2016. Calibration is essential to a guarantee of the measurement accuracy. Not only the cross-track errors, but also the along-track errors have to be estimated and corrected. This letter proposes a correction method to remove the azimuth time shift of the radar system using geographic positions of the surrounding lakeshores and riverbanks. Real acquisitions of InIRA over the largest inland lake of China, i.e., the Qinghai Lake, have been selected to validate the method. Both azimuth and range system time errors, the media delays, and the cross-track phase errors have been estimated and corrected on the level-1 complex products. The derived 2-D WSEs over an area of ~400 square kilometers have been compared with the high-resolution data of Sentinel-3A synthetic aperture radar altimeter (SRAL). The validation results have shown good agreements in both along-track and cross-track directions with a root mean square error of 28.1 cm and a mean absolute error of 22.7 cm. Hong Tan, Shengyang Li |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Scattering Information Fusion Network for Oriented Ship Detection in SAR ImagesabstractSynthetic aperture radar (SAR) image ship detection is a popular area of ocean remote sensing, which has broad application prospects in ocean monitoring, maritime rescue and other tasks. Recently, deep learning has been used in this field, but convolutional neural network (CNN) based SAR ship detection still faces some challenges. First, due to the characteristic of CNN’s local convolution, the global information of the ship is not sufficiently learned and the detection is vulnerable to complex background interference. Second, SAR ship imaging varies greatly under different imaging conditions and postures, so CNN is difficult to adapt to scattering change imaging. To solve these problems, we propose a Scattering Information Fusion Network (SIFNet) for oriented ship detection in SAR Images consisting of a multi-scale contextual semantic information fusion (MCSIF) module and a scattering points information learning (SPIL) module. The MCSIF module enhances the acquisition of global information, enabling the network to extract more efficient feature maps. The SPIL module takes advantage of the fact that scattering points can stably represent the key features of the ship under different imaging conditions to make detection more robust through scattering information learning. Experiments show that our method achieves the highest F1-score and AP50 on both HRSID and RSDD-SAR datasets. Han Wang 0049, Silei Liu, Yixuan Lv, Shengyang Li |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | A Multitask Benchmark Dataset for Satellite Video: Object Detection, Tracking, and SegmentationabstractVideo satellites can continuously image large areas and provide dynamic, real-time monitoring of hotspots and objects. The intelligent processing and analysis of satellite video have become a research hotspot in the field of remote sensing. However, the lack of high-quality satellite video datasets limits the development of relevant object detection, object tracking, and object segmentation. In this paper, we build the largest scale satellite video dataset with the most task types supported and object categories, named Satellite Video Multi-Mission Benchmark (SAT-MTB). First, multi-task annotation of aircraft, ships, cars, trains, and their corresponding 14 categories of fine-grained objects in 249 satellite videos is performed based on horizontal bounding boxes (HBB), oriented bounding boxes (OBB), masks, which cover more than 50,000 frames and 1,033,511 annotated object instances. Then, we review the tasks of object detection, object tracking, and object segmentation based on satellite videos, providing a comprehensive overview of progress in related datasets and algorithm research. Finally, we establish the first public benchmark of multi-task algorithms for satellite video object detection, object tracking, and object segmentation, evaluating and analyzing the performance of a total of 47 representative algorithms under different tasks on the constructed dataset. The proposed SAT-MTB will significantly advance research in intelligent processing and analysis of satellite video and related applications. Shengyang Li, Manqi Zhao, Weilong Guo, Yixuan Lv, Longxuan Kou, Han Wang 0049, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | On Improving Bounding Box Representations for Oriented Object DetectionabstractDetecting objects in remote sensing images (RSIs) using oriented bounding boxes (OBBs) is flourishing but challenging, wherein the design of OBB representations is the key to achieving accurate detection. In this article, we focus on two issues that hinder the performance of the two-stage oriented detectors: 1) the notorious boundary discontinuity problem, which would result in significant loss increases in boundary conditions, and 2) the inconsistency in regression schemes between the two stages. We propose a simple and effective bounding box representation by drawing inspiration from the polar coordinate system and integrate it into two detection stages to circumvent the two issues. The first stage specifically initializes four quadrant points as the starting points of the regression for producing high-quality oriented candidates without any postprocessing. In the second stage, the final localization results are refined using the proposed novel bounding box representation, which can fully release the capabilities of the oriented detectors. Such consistency brings a good trade-off between accuracy and speed. With only flipping augmentation and single-scale training and testing, our approach with ResNet-50-FPN harvests 76.25% mAP on the DOTA dataset with a speed of up to 16.5 frames/s, achieving the best accuracy and the fastest speed among the mainstream two-stage oriented detectors. Additional results on the DIOR-R and HRSC2016 datasets also demonstrate the effectiveness and robustness of our method. The source code is publicly available athttps://github.com/yanqingyao1994/QPDet. Gong Cheng 0003, Guangxing Wang 0001, Shengyang Li, Peicheng Zhou, Xingxing Xie, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | A Mutual Information Domain Adaptation Network for Remotely Sensed Semantic SegmentationabstractAlthough deep learning has made semantic segmentation of very-high-resolution (VHR) remote sensing (RS) images practical and efficient, its large-scale application is still limited. Given the diversity of imaging sensors, acquisition conditions, and regional styles, a deep learning network well-trained on one source domain dataset often suffers from drastic performance drops when applied to other target domain datasets. Thus, we propose a novel end-to-end mutual information domain adaptation network (MIDANet) that can shift between semantic segmentation domains by integrating multitask learning in the convolutional neural networks within an entropy adversarial learning (EAL) framework. Through the joint learning of semantic segmentation and elevation estimation, the features extracted by MIDANet can concentrate more on the elevation clues while dropping the domain-variant information (i.e., texture, spectral information). First, one encoder is applied to excavate general semantic features. Two decoders that share the same architecture are used to perform pixel-level classification and digital surface model (DSM) regression. Second, feature interaction modules (FIMs) and a mutual information attention unit (MIAU) are designed to mine the latent relationships between the two tasks and enhance their feature representations. Finally, a final MIDANet is obtained for semantic segmentation that does not require any semantic segmentation labels in the target domain after the adversarial learning of the classification entropy at the output level. Extensive comparative experiments and ablation studies were conducted on the International Society for Photogrammetry and Remote Sensing (ISPRS) Potsdam and Vaihingen test datasets. The results show that MIDANet outperforms other state-of-the-art domain adaptation (DA) methods in both evaluation metrics and visual assessment. Hongyu Chen 0003, Hongyan Zhang 0001, Shengyang Li, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Dual-Aligned Oriented DetectorabstractIn the past few years, object detection in remote sensing images has achieved remarkable progress. However, the detection of oriented and densely packed objects are still unsatisfactory due to the following spatial and feature misalignments. 1) Most two-stage oriented detectors only introduce an orientation regression branch in the detection head, while still leverage horizontal proposals for classification and regression. This inevitably results in the spatial misalignment problem between horizontal proposals and oriented objects. 2) The features used for classification are in fact extracted from the region proposals which have shifted to the final predictions via the regression branch. This leads to the feature misalignment problem between the classification and the localization tasks. In this article, we present a two-stage oriented object detection method, termed dual-aligned oriented detector (DODet), toward evading the aforementioned problems of spatial and feature misalignments. In DODet, the first stage is an oriented proposal network (OPN), which generates high-quality oriented proposals via a novel representation scheme of oriented objects. The second stage is a localization-guided detection head (LDH) that aims at alleviating the feature misalignment between classification and localization. Comprehensive and extensive evaluations on three benchmarks, including DIOR-R, DOTA, and HRSC2016, indicate that our method could obtain consistent and substantial gains compared with the baseline method. The source code is publicly available athttps://github.com/yanqingyao1994/DODet. Gong Cheng 0003, Shengyang Li, Ke Li 0005, Xingxing Xie, Jiabao Wang 0005, Xiwen Yao, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Satellite Video Intrinsic DecompositionabstractExisting satellite video processing methods are mainly based on original video, ignoring the use of invariant background characteristics of staring satellites, and easy to be disturbed by rapid light changes. In order to improve application capability of satellite video, this paper establishes the satellite video intrinsic decomposition (SVID) model, including satellite video signal composition model, decomposition constraint with time-spatial unity similarity constraint, static and dynamic components separation by improving TRPCA, and decomposition acceleration based on reflectance transfer. With SVID, intrinsic decomposition and dynamic and static component separation are realized. Five Jilin-1 satellite videos are used to verify the validity, superiority and the potential applications of the proposed algorithm. By comparing with state-of-the-art intrinsic image decomposition method and intrinsic video decomposition method, the experimental results prove the superiority of the SVID method in extracting reflectance component. In addition, the experimental results also prove SVID has excellent application ability in scene background analysis and moving target tracking. Guoming Gao, Yanfeng Gu, Shengyang Li |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Intrinsic Satellite Video Decomposition With Motion Target Energy ConstraintabstractSatellite videos dynamically monitor the Earth’s surface by using staring imaging, which has gained increased attention and enabled target tracking applications. While various tracking methods are processed on the original video, the rapid light changes due to staring imaging are not considered and have a negative effect on tracking. To reduce the effects of illumination and improve the performance of satellite video target tracking, an intrinsic satellite video decomposition model with motion target energy constraint, called MTE-ISVD, is proposed in this paper. The proposed algorithm introduces two main constraints: The first is a temporal constraint of reflectance, which can solve the flicker problem by preserving reflectance coherence in the time domain with the property that the background pixels in satellite videos are nearly consistent between adjacent frames. The second is a motion target energy constraint, which can concentrate the signal energy of the motion targets in the reflectance by representing them with the surrounding background in the shading. The decomposition problem is reformulated as a quadratic function minimization, which can be addressed using the standard conjugate gradient in closed form. For visual and quantitative comparisons, we perform experiments on five Jilin-1 satellite videos and analyze the results in terms of visual comparison, target tracking improvement, stability evaluation and processing time comparison. The experimental results demonstrate that our proposed method outperforms the other representative intrinsic decomposition methods in terms of processing speed, stability, and motion target representation. Jialei Pan, Yanfeng Gu, Shengyang Li, Guoming Gao, Shaochuan Wu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Terrain Aided Planetary UAV Localization Based on Geo-referencingabstractThe autonomous real-time optical navigation of planetary unmanned aerial vehicle (UAV) is of the key technologies to ensure the success of the exploration. In such a GPS-denied environment, vision-based localization is an optimal approach. In this article, we proposed a terrain aided simultaneous localisation and mapping (SLAM) algorithm, which simultaneously reconstructs the 3-D map point of environment and estimates the location of a planet UAV based on preexisting digital elevation model (DEM). To directly georeference the onboard UAV images to the digital terrain model, a theoretical model is proposed to prove that topographic features of UAV image and DEM can be correlated in the frequency domain via cross power spectrum. To provide the six-DOF of the UAV, we developed an optimization approach, which fuses the geo-referencing result into an SLAM system via local bundle adjustment (LBA) to achieve robust and accurate vision-based navigation even in featureless planetary areas. To test the robustness and effectiveness of the proposed localization algorithm, a new dataset for planetary drone navigation is proposed based on simulation engine. The proposed dataset includes 40 200 synthetic drone images taken from nine planetary scenes with related DEM query images. Comparison experiments are carried out to demonstrate that over the flight distance of 33.8 km, the proposed method achieved an average localization error of 0.45 m, compared to 1.32 m by ORB-SLAM2 and 0.75 m by ORB-SLAM3, with the processing speed of 12 Hz, which will ensure real-time performance. We will make our datasets available to encourage further work on this topic. Xue Wan, Yuanbin Shao, Shengyang Zhang, Shengyang Li |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | SatSOT: A Benchmark Dataset for Satellite Video Single Object TrackingabstractBy imaging a specific area continuously, satellite video shows excellent capability in various applications such as surveillance and traffic management. Although object tracking has made significant progress in recent years, development in satellite object tracking is limited by the lack of open-source satellite datasets. It is thus essential to establish a satellite video object-tracking benchmark to fill the gap and advance the research. In this work, we present SatSOT, the first densely annotated satellite video single object-tracking benchmark dataset. SatSOT consists of 105 sequences with 27664 frames, 11 attributes, and four categories of typical moving targets in satellite videos: car, plane, ship, and train. Based on the proposed dataset and the significant challenges in satellite video object tracking, such as small targets, background interference, and severe occlusion, detailed evaluation and analysis are performed on 15 among the best and most representative tracking algorithms, which provides a basis for further research on satellite video object tracking. Manqi Zhao, Shengyang Li, Shiyu Xuan, Longxuan Kou, Shuai Gong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Rotation adaptive correlation filter for moving object tracking in satellite videos
Shiyu Xuan, Shengyang Li, Zifei Zhao, Wanfeng Zhang, Hong Tan, Gui-Song Xia, Yanfeng Gu |
Neurocomputing | 2 |
| 2021 | Siamese networks with distractor-reduction method for long-term visual object trackingabstractMany trackers which divide the tracking process into two stages have recently been proposed to solve the problem of long-term tracking. Their outstanding performance makes them become one of the mainstream algorithms of long-term tracking. To further improve the performance of two-stage tracking algorithms, some improvements are proposed in this paper. (a) A hard negative mining method is proposed. It can optimize the training process of the verification network and bridge the gap between the two sub-networks. (b) The architecture of the verification network is designed as a Siamese structure; therefore, the semantic ambiguity in classification can be alleviated. Extensive experiments performed on benchmarks demonstrate that the proposed approach significantly outperforms the state-of-the-art methods, yielding 7% relative gain in the VOT2018-LT dataset and 14.2% relative gain in the OxUvA dataset. Shiyu Xuan, Shengyang Li, Zifei Zhao, Longxuan Kou, Gui-Song Xia |
Pattern Recognit. | 2 |
| 2020 | Deep feature extraction and motion representation for satellite video scene classification
Yanfeng Gu, Tengfei Wang 0001, Shengyang Li, Guoming Gao |
Sci. China Inf. Sci. | 4 |
| 2020 | Satellite Video Super-Resolution Based on Adaptively Spatiotemporal Neighbors and Nonlocal Similarity RegularizationabstractRecently, super-resolution (SR) of satellite videos has received increasing attention as it can overcome the limitation of spatial resolution in applications of satellite videos to dynamic analysis. The low quality of satellite videos presents big challenges to the development of the spatial SR techniques, e.g., accurate motion estimation and motion compensation for multiframe SR. Therefore, reasonable image priors in maximum a posteriori (MAP) framework, where motion information among adjacent frames is involved, are needed to regularize the solution space and generate the corresponding high-resolution frames. In this article, an effective satellite video SR framework based on locally spatiotemporal neighbors and nonlocal similarity modeling is proposed. Firstly, local prior knowledge is represented by means of adaptively exploiting spatiotemporal neighbors. In this way, implicitly local motion information can be captured without explicit motion estimation. Secondly, the nonlocal spatial similarity is integrated into the proposed SR framework to enhance texture details. Finally, the locally spatiotemporal regularization and nonlocal similarity modeling bring out a complex optimization problem, which is solved via the iterated reweighted least squares in the proposed SR framework. The videos from the Jilin-1 satellite and the OVS-1A satellite are used for evaluating the proposed method. Experimental results show that the proposed method demonstrates better SR performance in preserving edges and texture details compared with the-state-of-art video SR methods. Yanfeng Gu, Tengfei Wang 0001, Shengyang Li |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Object Tracking in Satellite Videos by Improved Correlation Filters With Motion EstimationsabstractAs a new method of Earth observation, video satellite is capable of monitoring specific events on the Earth's surface continuously by providing high-temporal resolution remote sensing images. The video observations enable a variety of new satellite applications such as object tracking and road traffic monitoring. In this article, we address the problem of fast object tracking in satellite videos, by developing a novel tracking algorithm based on correlation filters embedded with motion estimations. Based on the kernelized correlation filter (KCF), the proposed algorithm provides the following improvements: 1) proposing a novel motion estimation (ME) algorithm by combining the Kalman filter and motion trajectory averaging and mitigating the boundary effects of KCF by using this ME algorithm and 2) solving the problem of tracking failure when a moving object is partially or completely occluded. The experimental results demonstrate that our algorithm can track the moving object in satellite videos with 95% accuracy. Shiyu Xuan, Shengyang Li, Mingfei Han 0002, Xue Wan, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | The Multi-task Fully Convolutional Siamese Network with Correlation Filter Layer for Real-Time Visual Tracking
Shiyu Xuan, Shengyang Li, Zifei Zhao, Mingfei Han 0002 |
PRCV (1) | 2 |
| 2019 | An Illumination-Invariant Change Detection Method Based on Disparity Saliency Map for Multitemporal Optical Remotely Sensed ImagesabstractMultitemporal airborne and satellite imagery data with frequent repeat coverage provide great capability for change detection (CD). When comparing two images taken at different times of day or in different seasons for CD, the variation of topographic shades and shadows caused by the change of sunlight angle can be so significant that it overwhelms the real object and environmental changes, making automatic detection unreliable. An effective CD algorithm, therefore, has to be robust to the illumination variation. In this paper, the robustness of phase correlation (PC) to shadow effects is proven via mathematical analysis, and then, an illumination-invariant change detection (IICD) metric is proposed based on pixel-wise dense PC matching. In the proposed IICD method, a graph-based visual saliency map is introduced for the initial CD followed by an active contour-based segmentation to precisely quantize the change region. Compared to the state-of-the-art CD algorithms, experiments using daily images of a landscape model and Landsat satellite images demonstrate that only the proposed method can effectively detect and precisely segment appearance changes under daily and seasonal sunlight changes. Xue Wan, Jian Guo Liu 0005, Shengyang Li, John Dawson 0002, Hongshi Yan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Phase Correlation Decomposition: The Impact of Illumination Variation for Robust Subpixel Remotely Sensed Image MatchingabstractIllumination variation is one of the major problems in multitemporal earth observation (EO) image matching. Despite much research has been focused on illumination invariant image matching, subpixel image matching under large illumination variation without prior knowledge is still a challenge. This paper proposes a phase correlation decomposition (PCD) theory model in order to analyze the joint effects of zenith and azimuth angles of the lighting source. A novel stepwise least-squares fitting-based PC (SLSF-PC) is proposed to accurately calculate the subpixel image shift by a stepwise function in the frequency domain. Our mathematical investigation is validated by image alignment and stereo dense matching experiments using simulated terrain shading images representing four different landscapes and a multi-illumination remotely sensed image data set containing eight different scenes under seasonal and daily illumination variation. Image matching experiments demonstrate the superior performance of the proposed SLSF-PC compared to the state-of-the-art image matching algorithms, such as speeded up robust features (SURF), mutual information (MI), and normalized cross correlation (NCC). Even under great illumination angle change, the proposed SLSF-PC is able to achieve 0.1 subpixel matching accuracy on average while other methods fail to find the correspondence. Xue Wan, Jian Guo Liu 0005, Shengyang Li, Hongshi Yan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | How Agile is the Adaptive Data Rate Mechanism of LoRaWAN?abstractThe LoRaWAN based Low Power Wide Area networks aim to provide long-range connectivity to a large number of devices by exploiting limited radio resources. The Adaptive Data Rate (ADR) mechanism controls the assignment of these resources to individual end-devices by a runtime adaptation of their communication parameters when the quality of links inevitably changes over time. This paper provides a detailed performance analysis of the ADR technique presented in the recently released LoRaWan Specifications (v1.1). We show that the ADR technique lacks the agility to adapt to the changing link conditions, requiring a number of hours to days to converge to a reliable and energy-efficient communication state. As a vital step towards improving this situation, we then change different control knobs or parameters in the ADR technique to observe their effects on the convergence time. Shengyang Li, Usman Raza, Aftab Khan 0001 |
GLOBECOM | 1 |
| 2018 | Crops Classification from Sentinel-2A Multi-spectral Remote Sensing Images Based on Convolutional Neural NetworksabstractDeep learning technology such as convolutional neural networks (CNN) can extract the distinguishable and representative features of different land cover from remote sensing images in a hierarchical way to classify. However, in the field of agriculture, there are few application of crops classification from multi-spectral remote sensing images based on deep learning. In this context, we compared the classification methods of CNN and support vector machines (SVM) in extracting the spatial distribution of crops planting area from Sentineal-2A multi-spectral remote sensing images in Yuanyang county, China. For the region of study, both methods obtained reasonable spatial distribution of different crops, the verification results show that the overall accuracy of CNN is 95.6% which is superior to SVM. Shengyang Li, Yuyang Shao |
IGARSS | 2 |
| 2005 | Spatial resolution improvement by iterative resampling and restoration based on MTF computation
Shengyang Li, Chongguang Zhu, Shukui Bo, Famao Ye |
IGARSS | 1 |
| 2005 | Extraction of complex object contour by particle filteringabstractContour extraction plays an important role in the field of image processing, and currently the hybrid approach combining the manual and automatic methods is widely used in practice. One of the techniques is called JetStream that is based on particle filtering and is a considerable advance on tracing contour, but it tends to give undesirable results when it deals with complex object. This paper handles this problem by choosing a good importance density and using more information to calculate the likelihood. By those measures, the method is capable of locating the target object contour of sharp tips accurately using a few user interactions. Famao Ye, Shengyang Li |
IGARSS | 4 |