Yixuan Lv

dblp:174/0749 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Zs-Drosophila: Learning Transferable Representations for Drosophila Behavior Analysis Via Visual-Language Hyperbolic Alignment
abstract
Automatic behavioral analysis of Drosophila has drawn consistent research interest in the field, as understanding quantitative behavior plays a crucial role in neuroscience, genetics, and space biology. Existing machine learning approaches often rely heavily on expert domain knowledge and manual annotations to build supervised models, limiting their scalability and adaptability (e.g., to novel actions and unseen domains). Meanwhile, recent advances in vision-language models offer a more flexible, expressive, and interpretable medium for behavior representation, enabling broader semantic understanding and generalization. To this end, we propose ZS-Drosophila, the first framework that introduces language-guided multimodal alignment for Drosophila behavior analysis. Our foundation model ZS-Drosophila learns transferable behavior representations that can generalize to unseen behaviors and domains. Furthermore, we construct SpaceAnimal-Drosophila, a benchmark dataset of Drosophila video recordings collected both on Earth and in space, comprising annotated skeletal sequences and behaviorlanguage pairs as ground truths. It has a series of evaluation protocols to showcase the strong transferability of ZS-Drosophila, achieved with only prompt tuning at test time. We demonstrate the capabilities of our method in the following scenarios: (1) conventional supervised action recognition on on-Earth data, (2) zero-shot recognition of unseen behaviors, and (3) crossdomain generalization to in-orbit microgravity data. Notably, our proposed model enables zero-shot recognition of novel behaviors potentially induced by microgravity without requiring additional annotations, which, to our knowledge, is the first attempt in the field.
Kang Liu 0020, Han Wang 0049, Yixuan Lv, Shengyang Li, Jianing You
BIBM5
2025 Cross-Modal Ship Re-Identification via Optical and SAR Imagery: A Novel Dataset and Method
abstract
Detecting and tracking ground objects using earth observation imagery remains a significant challenge in the field of remote sensing. Continuous maritime ship tracking is crucial for applications such as maritime search and rescue, law enforcement, and shipping analysis. However, most current ship tracking methods rely on geostationary satellites or video satellites. The former offer low resolution and are susceptible to weather conditions, while the latter have short filming durations and limited coverage areas, making them less suitable for the real-world requirements of ship tracking. To address these limitations, we present the Hybrid Optical and Synthetic Aperture Radar (SAR) Ship Re-Identification Dataset (HOSS ReID dataset), designed to evaluate the effectiveness of ship tracking using low-Earth orbit constellations of optical and SAR sensors. This approach ensures shorter re-imaging cycles and enables all-weather tracking. HOSS ReID dataset includes images of the same ship captured over extended periods under diverse conditions, using different satellites of different modalities at varying times and angles. Furthermore, we propose a baseline method for cross-modal ship re-identification, TransOSS, which is built on the Vision Transformer architecture. It refines the patch embedding structure to better accommodate cross-modal tasks, incorporates additional embeddings to introduce more reference information, and employs contrastive learning to pre-train on large-scale optical-SAR image pairs, ensuring the model's ability to extract modality-invariant features. Our dataset and baseline method are publicly available on https://github.com/Alioth2000/Hoss-ReID.
Shengyang Li, Yixuan Lv
ICCV5
2025 Semantic Affinity-Driven Spatiotemporal Transformer Network for Satellite Video Moving-Object Segmentation
abstract
Satellite video intelligent processing plays a critical role in Earth observation applications such as traffic monitoring and environmental surveillance. However, moving-object segmentation in satellite videos faces several challenges. First, spatiotemporal redundancy makes it difficult to model long-range dependencies because large-scale scenes with slow background changes lead to fragmented segmentation. Second, semantic ambiguity arises when stationary objects like parked aircraft share category-level similarities with moving targets, which causes false positives. Besides, insufficient feature discrimination occurs as small, rigid objects such as ships exhibit weak texture and edge details under low-resolution imaging. To overcome these issues, we introduce a semantic affinity-driven spatiotemporal Transformer network that leverages a Transformer-based architecture to capture pixel-level dependencies across spatial and temporal dimensions. Furthermore, our network employs a contextual affinity-constrained decoder to suppress category-level interference and integrates a triple-branch feature extractor with edge priors for enhanced contour delineation. Our framework operates in an end-to-end manner without requiring fine-tuning during inference, which ensures deployment efficiency. Extensive experiments on a dataset built upon SAT-MTB demonstrate state-of-the-art performance with a J&F Mean of 71.7%. The proposed method outperforms the baseline by 3.7% with improvements of 4.4% in J-Mean and 3.1% in F-Mean. In addition, it surpasses the optimized SAM2 with a 10.7% higher J-Mean while maintaining a significantly smaller parameter count (34.6 M versus 224 M). Both qualitative and quantitative evaluations confirm the method’s superiority and temporal stability. This work offers a robust and efficient solution for accurate moving-object segmentation in satellite videos.
Yixuan Lv, Kang Liu 0020, Han Wang 0049, Shengyang Li, Jianing You, Kailun Zhang
IEEE Trans. Geosci. Remote. Sens.1
2024 Distributed Nash equilibrium searching for multi-agent games under false data injection attacks
Yixuan Lv, Yan-Jun Liu 0003, Lei Liu 0006, Dengxiu Yu, Yang Chen 0027
Neurocomputing1
2023 Scattering Information Fusion Network for Oriented Ship Detection in SAR Images
abstract
Synthetic aperture radar (SAR) image ship detection is a popular area of ocean remote sensing, which has broad application prospects in ocean monitoring, maritime rescue and other tasks. Recently, deep learning has been used in this field, but convolutional neural network (CNN) based SAR ship detection still faces some challenges. First, due to the characteristic of CNN’s local convolution, the global information of the ship is not sufficiently learned and the detection is vulnerable to complex background interference. Second, SAR ship imaging varies greatly under different imaging conditions and postures, so CNN is difficult to adapt to scattering change imaging. To solve these problems, we propose a Scattering Information Fusion Network (SIFNet) for oriented ship detection in SAR Images consisting of a multi-scale contextual semantic information fusion (MCSIF) module and a scattering points information learning (SPIL) module. The MCSIF module enhances the acquisition of global information, enabling the network to extract more efficient feature maps. The SPIL module takes advantage of the fact that scattering points can stably represent the key features of the ship under different imaging conditions to make detection more robust through scattering information learning. Experiments show that our method achieves the highest F1-score and AP50 on both HRSID and RSDD-SAR datasets.
Han Wang 0049, Silei Liu, Yixuan Lv, Shengyang Li
IEEE Geosci. Remote. Sens. Lett.3
2023 A Multitask Benchmark Dataset for Satellite Video: Object Detection, Tracking, and Segmentation
abstract
Video satellites can continuously image large areas and provide dynamic, real-time monitoring of hotspots and objects. The intelligent processing and analysis of satellite video have become a research hotspot in the field of remote sensing. However, the lack of high-quality satellite video datasets limits the development of relevant object detection, object tracking, and object segmentation. In this paper, we build the largest scale satellite video dataset with the most task types supported and object categories, named Satellite Video Multi-Mission Benchmark (SAT-MTB). First, multi-task annotation of aircraft, ships, cars, trains, and their corresponding 14 categories of fine-grained objects in 249 satellite videos is performed based on horizontal bounding boxes (HBB), oriented bounding boxes (OBB), masks, which cover more than 50,000 frames and 1,033,511 annotated object instances. Then, we review the tasks of object detection, object tracking, and object segmentation based on satellite videos, providing a comprehensive overview of progress in related datasets and algorithm research. Finally, we establish the first public benchmark of multi-task algorithms for satellite video object detection, object tracking, and object segmentation, evaluating and analyzing the performance of a total of 47 representative algorithms under different tasks on the constructed dataset. The proposed SAT-MTB will significantly advance research in intelligent processing and analysis of satellite video and related applications.
Shengyang Li, Manqi Zhao, Weilong Guo, Yixuan Lv, Longxuan Kou, Han Wang 0049, Yanfeng Gu
IEEE Trans. Geosci. Remote. Sens.6
2022 SCAN: Scattering Characteristics Analysis Network for Few-Shot Aircraft Classification in High-Resolution SAR Images
abstract
Recently, deep learning in synthetic aperture radar (SAR) automatic target recognition (ATR) has made significant progress, but the sample limitation problem in the SAR field is still obvious. Compared with the optical remote sensing images, the SAR images are insufficient, especially those containing the geospatial targets with certain target attitude angles (TAAs). To solve these problems, a novel few-shot learning framework named scattering characteristics analysis network (SCAN) is proposed in this article. First, a scattering extraction module (SEM) is designed to combine the target imaging mechanism with the network, which learns the number and distribution of the scattering points for each target type via explicit supervision. Besides, considering the imaging variability of SAR targets, a TAA-guided metalearning network consisting of an angle self-adaption classifier (ASC) and a frequency embedded module (FEM) is designed. ASC guides the network to focus on the positive sample pairs with different TAAs. FEM combines pulse cosine transform (PCT) with the network training process effectively to enrich frequency-domain information. In addition, a new dataset named SAR aircraft category dataset is constructed for the experiments. Compared with other few-shot SAR target classification approaches, our model efficiently integrates the scattering characteristics with the learning process, and the test accuracy for 5-way 1-shot has been improved by 4.74%. Finally, the experimental results are provided to demonstrate the validity of the proposed method.
Xian Sun 0001, Yixuan Lv, Zhirui Wang 0003, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.2
2016 Direct heuristic dynamic programming design with extreme learning machine
abstract
Extreme learning machine (ELM) as a learning algorithm for neural networks (NN) could provide the best generalization performance at extremely fast leaning speed. Through the use of ELM, it is thus possible to improve the existing schemes especially the ones whose learning speed is not fast enough while addressing control problems. As a popular NN-based approach for control applications, direct heuristic dynamic programming (DHDP) with a good capability of adaptive learning has been successfully applied to solve control problems. But limited by slow learning algorithms in NN, it imposes very challenging obstacles to the real-time controller design of DHDP, which keeps it from widely applied. In this paper, driven by the interest of improving learning speed of DHDP while maintaining its good approximation performance, we employ ELM as a learning algorithm in DHDP. The proposed ELM-based DHDP learning scheme is tested on a cart-pole balancing control problem. The simulation results show the proposed scheme has better learning performance than traditional DHDP. Furthermore, this paper provides a novel idea of applying ELM in control problems.
Xiong Luo, Yixuan Lv, Weiping Wang 0007, Wenbing Zhao 0001
IJCNN2
2016 A laguerre neural network-based ADP learning scheme with its application to tracking control in the Internet of Things
Xiong Luo, Yixuan Lv, Weiping Wang 0007, Wenbing Zhao 0001
Pers. Ubiquitous Comput.2