Yanxing Liu

dblp:154/6015 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Locate Anything on Earth: Advancing Open-Vocabulary Object Detection for Remote Sensing Community
abstract
Object detection, particularly open-vocabulary object detection, plays a crucial role in Earth sciences, such as environmental monitoring, natural disaster assessment, and land-use planning. However, existing open-vocabulary detectors, primarily trained on natural-world images, struggle to generalize to remote sensing images due to a significant data domain gap. Thus, this paper aims to advance the development of open-vocabulary object detection in remote sensing community. To achieve this, we first reformulate the task as Locate Anything on Earth (LAE) with the goal of detecting any novel concepts on Earth. We then developed the LAE-Label Engine which collects, auto-annotates, and unifies up to 10 remote sensing datasets creating the LAE-1M — the first large-scale remote sensing object detection dataset with broad category coverage. Using the LAE-1M, we further propose and train the novel LAE-DINO Model, the first open-vocabulary foundation object detector for the LAE task, featuring Dynamic Vocabulary Construction (DVC) and Visual-Guided Text Prompt Learning (VisGT) modules. DVC dynamically constructs vocabulary for each training batch, while VisGT maps visual features to semantic space, enhancing text features. We comprehensively conduct experiments on established remote sensing benchmark DIOR, DOTAv2.0, as well as our newly introduced 80-class LAE-80C benchmark. Results demonstrate the advantages of the LAE-1M dataset and the effectiveness of the LAE-DINO method.
Jiancheng Pan, Yanxing Liu, Yuqian Fu, Muyuan Ma, Jiahao Li 0005, Danda Pani Paudel, Luc Van Gool, Xiaomeng Huang
AAAI2
2025 Diverse Instance Generation via Diffusion Models for Enhanced Few-Shot Object Detection in Remote Sensing Images
abstract
Few-shot object detection (FSOD) aims to detect novel instances with only a limited number of labeled training samples, presenting a challenge that is particularly prominent in numerous remote sensing applications such as endangered species monitoring and disaster assessment. Existing FSOD methods for remote sensing images (RSIs) have achieved promising progress but remain constrained by the limited diversity of instances. To address this issue, we propose a novel framework that can leverage a diffusion model pretrained on large-scale natural images to synthesize diverse remote sensing instances, thereby improving the performance of few-shot object detectors. Instead of directly synthesizing complete remote sensing images, we first generate instance-level slices via a specialized slice-to-slice module, and then embed these slices into full-scale imagery for enhanced data augmentation. To further adapt diffusion models for remote sensing scenarios, we develop a class-agnostic image inversion module that can invert remote sensing instance slices into semantic space. Additionally, we introduce contrastive loss to semantically align the synthesized images with their corresponding classes. Experimental results show that our method has achieved an average performance improvement of 4.4% across multiple datasets and various approaches. Ablation experiments indicate that the elaborately designed inversion module can effectively enhance the performance of FSOD methods, and the semantic contrastive loss can further boost the performance.
Yanxing Liu, Jiancheng Pan, Tiancheng Chen, Peiling Zhou, Bingchen Zhang
IEEE Geosci. Remote. Sens. Lett.1
2025 Sublook Contrastive Learning for SAR Representation Learning and Image Classification
abstract
Deep learning networks, such as convolutional neural networks (CNNs), are increasingly applied to synthetic aperture radar (SAR) feature representation and image classification. However, the performance of most deep learning methods relies on sufficient labeled data, which is difficult to collect for SAR images. As a result, self-supervised contrastive learning (CL) methods have attracted enthusiasm in recent studies to learn SAR representations with unlabeled data. Most existing CL-based methods for SAR representation learning simply use the amplitude information as the network input but neglect the particular features of complex-valued SAR images such as spectral information, leading to insufficient feature extraction. To address this issue, we propose a self-supervised sublook CL (SCL) method to learn spectral representations from unlabeled data. First, we extract sublooks from SAR images to establish a spectral representation branch (SRB) to discover the spectral information. Second, a novel SCL method with a sublook-matching task is proposed based on a contrastive network to learn spectral representations. The branch is applied to SAR image classification tasks by integrating it with a typical amplitude network through feature concatenation. Experimental results have validated the effectiveness and generalization ability of the proposed method with a different experimental protocol that distinguishes the pretrain and downstream datasets.
Peiling Zhou, Zongxu Pan, Yanxing Liu, Ben Niu 0008
IEEE Geosci. Remote. Sens. Lett.3
2024 A privacy-preserving vehicle trajectory clustering framework
abstract
As one of the essential tools for spatio–temporal traffic data mining, vehicle trajectory clustering is widely used to mine the behavior patterns of vehicles. However, uploading original vehicle trajectory data to the server and clustering carry the risk of privacy leakage. Therefore, one of the current challenges is determining how to perform vehicle trajectory clustering while protecting user privacy. We propose a privacy-preserving vehicle trajectory clustering framework and construct a vehicle trajectory clustering model (IKV) based on the variational autoencoder (VAE) and an improved K -means algorithm. In the framework, the client calculates the hidden variables of the vehicle trajectory and uploads the variables to the server; the server uses the hidden variables for clustering analysis and delivers the analysis results to the client. The IKV’ workflow is as follows: first, we train the VAE with historical vehicle trajectory data (when VAE’s decoder can approximate the original data, the encoder is deployed to the edge computing device); second, the edge device transmits the hidden variables to the server; finally, clustering is performed using improved K -means, which prevents the leakage of the vehicle trajectory. IKV is compared to numerous clustering methods on three datasets. In the nine performance comparison experiments, IKV achieves optimal or sub-optimal performance in six of the experiments. Furthermore, in the nine sensitivity analysis experiments, IKV not only demonstrates significant stability in seven experiments but also shows good robustness to hyperparameter variations. These results validate that the framework proposed in this paper is not only suitable for privacy-conscious production environments, such as carpooling tasks, but also adapts to clustering tasks of different magnitudes due to the low sensitivity to the number of cluster centers.
Pulun Gao, Yanxing Liu
Frontiers Inf. Technol. Electron. Eng.3
2024 Few-Shot Object Detection in Remote-Sensing Images via Label-Consistent Classifier and Gradual Regression
abstract
With the abomination of time-consuming or even impractical large-scale labeling, few-shot object detection (FSOD) based on natural scenes has attracted extensive attention. However, directly migrating FSOD methods designed for natural images to large-size remote sensing images (RSIs) still remains challenges. 1) Labels of novel instances within the base dataset are inconsistently assigned between the base training and the few-shot fine-tuning stage, which confuses the detector and leads to significant performance degradation over novel classes. 2) The region proposal network (RPN) of detectors cannot provide sufficient high-quality proposals for remote sensing objects with various aspect ratios and irregular shapes, leading to decreased detection performance. To tackle these issues, we specify a novel few-shot object detector for RSIs, to avoid the significant performance degradation caused by inconsistent labeling assignments, as well as efficiently leveraging the novel instances that existed in the base dataset. Furthermore, the proposed detector utilizes a coarse-to-fine regression method with an enhanced feature extractor called Gradual RPN to improve the recall of RPN. Experiments on a newly constructed few-shot detection benchmark show that our approach improves the mAP of novel classes by up to 8.4% and the average recall of RPN by up to 12.3%. The source code is available at here.
Yanxing Liu, Zongxu Pan, Bingchen Zhang, Qixiang Ye
IEEE Trans. Geosci. Remote. Sens.1
2023 LDformer: a parallel neural network model for long-term power forecasting
abstract
Accurate long-term power forecasting is important in the decision-making operation of the power grid and power consumption management of customers to ensure the power system’s reliable power supply and the grid economy’s reliable operation. However, most time-series forecasting models do not perform well in dealing with long-time-series prediction tasks with a large amount of data. To address this challenge, we propose a parallel time-series prediction model called LDformer. First, we combine Informer with long short-term memory (LSTM) to obtain deep representation abilities in the time series. Then, we propose a parallel encoder module to improve the robustness of the model and combine convolutional layers with an attention mechanism to avoid value redundancy in the attention mechanism. Finally, we propose a probabilistic sparse (ProbSparse) self-attention mechanism combined with UniDrop to reduce the computational overhead and mitigate the risk of losing some key connections in the sequence. Experimental results on five datasets show that LDformer outperforms the state-of-the-art methods for most of the cases when handling the different long-time-series prediction tasks.
Xinmei Li, Zhongyu Ma, Yanxing Liu, Jingxia Wang
Frontiers Inf. Technol. Electron. Eng.4
2018 A deep belief network to predict the hot deformation behavior of a Ni-based superalloy
Yongcheng Lin, Mingsong Chen 0004, Yanxing Liu, Yingjie Liang
Neural Comput. Appl.4