EDBT 2026 Demo / reviewers in the wild / expert
Han Wang 0049
dblp:67/1771-49
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0003-0924-5138ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Zs-Drosophila: Learning Transferable Representations for Drosophila Behavior Analysis Via Visual-Language Hyperbolic AlignmentabstractAutomatic behavioral analysis of Drosophila has drawn consistent research interest in the field, as understanding quantitative behavior plays a crucial role in neuroscience, genetics, and space biology. Existing machine learning approaches often rely heavily on expert domain knowledge and manual annotations to build supervised models, limiting their scalability and adaptability (e.g., to novel actions and unseen domains). Meanwhile, recent advances in vision-language models offer a more flexible, expressive, and interpretable medium for behavior representation, enabling broader semantic understanding and generalization. To this end, we propose ZS-Drosophila, the first framework that introduces language-guided multimodal alignment for Drosophila behavior analysis. Our foundation model ZS-Drosophila learns transferable behavior representations that can generalize to unseen behaviors and domains. Furthermore, we construct SpaceAnimal-Drosophila, a benchmark dataset of Drosophila video recordings collected both on Earth and in space, comprising annotated skeletal sequences and behaviorlanguage pairs as ground truths. It has a series of evaluation protocols to showcase the strong transferability of ZS-Drosophila, achieved with only prompt tuning at test time. We demonstrate the capabilities of our method in the following scenarios: (1) conventional supervised action recognition on on-Earth data, (2) zero-shot recognition of unseen behaviors, and (3) crossdomain generalization to in-orbit microgravity data. Notably, our proposed model enables zero-shot recognition of novel behaviors potentially induced by microgravity without requiring additional annotations, which, to our knowledge, is the first attempt in the field. Kang Liu 0020, Han Wang 0049, Yixuan Lv, Shengyang Li, Jianing You |
BIBM | 2 |
| 2025 | SSCNet: Structure-Aware Segmentation Network for C.elegans in Scientific Experiments on China Space Station
Silei Liu, Kang Liu 0020, Han Wang 0049, Yuhan Sun 0004, Shengyang Li |
PRCV (2) | 5 |
| 2025 | Semantic Affinity-Driven Spatiotemporal Transformer Network for Satellite Video Moving-Object SegmentationabstractSatellite video intelligent processing plays a critical role in Earth observation applications such as traffic monitoring and environmental surveillance. However, moving-object segmentation in satellite videos faces several challenges. First, spatiotemporal redundancy makes it difficult to model long-range dependencies because large-scale scenes with slow background changes lead to fragmented segmentation. Second, semantic ambiguity arises when stationary objects like parked aircraft share category-level similarities with moving targets, which causes false positives. Besides, insufficient feature discrimination occurs as small, rigid objects such as ships exhibit weak texture and edge details under low-resolution imaging. To overcome these issues, we introduce a semantic affinity-driven spatiotemporal Transformer network that leverages a Transformer-based architecture to capture pixel-level dependencies across spatial and temporal dimensions. Furthermore, our network employs a contextual affinity-constrained decoder to suppress category-level interference and integrates a triple-branch feature extractor with edge priors for enhanced contour delineation. Our framework operates in an end-to-end manner without requiring fine-tuning during inference, which ensures deployment efficiency. Extensive experiments on a dataset built upon SAT-MTB demonstrate state-of-the-art performance with a J&F Mean of 71.7%. The proposed method outperforms the baseline by 3.7% with improvements of 4.4% in J-Mean and 3.1% in F-Mean. In addition, it surpasses the optimized SAM2 with a 10.7% higher J-Mean while maintaining a significantly smaller parameter count (34.6 M versus 224 M). Both qualitative and quantitative evaluations confirm the method’s superiority and temporal stability. This work offers a robust and efficient solution for accurate moving-object segmentation in satellite videos. Yixuan Lv, Kang Liu 0020, Han Wang 0049, Shengyang Li, Jianing You, Kailun Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Language-Empowered Conversion for Remote Sensing Image Retrieval With Text FeedbackabstractRemote sensing image retrieval with text feedback (RSIR-TF) presents a challenging multi-modal retrieval task that leverages a reference image, modification text, and scene graph to retrieve the relevant target image from a gallery. Existing approaches rely on cross-modal combiners to integrate multi-modal features extracted separately from modality-specific encoders. However, the modality-specific encoders often suffer from insufficient representational capacity and limited cross-modal alignment due to the lack of effective pre-training. Recently, vision language models (VLMs) pre-trained on large-scale image-text pairs have demonstrated exceptional representation and alignment capabilities in the remote sensing domain. However, these VLMs struggle with processing structured scene graphs, limiting their applicability to tasks like RSIR-TF that require composite reasoning over structured and unstructured modalities. To address these limitations, we propose a novel pipeline, Language-Empowered Conversion (LEmpo), which effectively migrates the large language model (LLM) and VLM to the RSIR-TF task. Firstly, we perform Pseudo Caption Generation and Scene Graph Interpretation powered by LLM to convert structured scene graphs into natural language captions. This conversion bridges the gap between structured scene graphs and unstructured text captions, enabling unified feature extraction and alignment. Subsequently, we employ the pre-trained VLM to extract robust visual and textual features within a joint visual-textual feature space. To fully utilizing the complementary information from visual, textual, and structured data, we introduce a Hybrid Similarity Tuning strategy, which aggregates triplet similarity, language similarity, and pseudo caption similarity into a unified hybrid similarity. The hybrid similarity is optimized during training through vision-fixed tuning, which anchors visual features while refining textual features to enhance alignment with target images. Comprehensive experiments conducted on the Airplane, Tennis and WHIRT datasets demonstrate that LEmpo significantly outperforms all comparison methods, achieving a substantial improvement in recall performance. Shengyang Li, Yuhan Sun 0004, Han Wang 0049 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | ICPR 2024 Competition on Moving Object Detection and Tracking in Satellite Videos: Methods and Results
Yulan Guo, Qingyong Hu, Feng Zhang 0046, Ye Zhang 0037, Hanyun Wang, Han Wang 0049, Furui Chen, Silei Liu, Xiaomin Huang, Shining Wang, Ying Li 0017, Peng Wang 0015, Shiyong Peng, Xiaokai Bi, Renbin Zou, Wenjing Deng, Zhen Cui 0001 |
ICPR (34) | 9 |
| 2024 | MP2Net: Mask Propagation and Motion Prediction Network for Multiobject Tracking in Satellite VideosabstractMainstream multi-object tracking (MOT) algorithms employ global object detection and association methods. However, when dealing with scenarios involving crowded tiny objects in satellite videos, existing global trackers often yield numerous missed detections and unstable trajectories. To address this issue, we propose a novel joint-detection-and-tracking framework, MP2Net, which integrates local detection enhancements for tiny targets and bridges the gap between detection and association. Specifically, our approach incorporates a mask propagation network that enhances feature representation for tiny targets by matching frame-by-frame to capture local details. Additionally, we utilize an implicit and explicit motion prediction strategy that merges tracking information into detection at both feature and instance levels, thereby improving tracking robustness. Experimental results on two large-scale datasets demonstrate the effectiveness and robustness of MP2Net, achieving state-of-the-art performance on typical moving objects in satellite videos, such as 66.7% MOTA and 75.9% IDF1 on the SatVideoDT challenge dataset. The code will be available at https://github.com/DonDominic/MP2Net. Manqi Zhao, Shengyang Li, Han Wang 0049, Yuhan Sun 0004, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Frequency and Spatial Domain Filter Network for Visual Object Tracking
Manqi Zhao, Shenyang Li, Han Wang 0049 |
PRCV (6) | 3 |
| 2023 | Scattering Information Fusion Network for Oriented Ship Detection in SAR ImagesabstractSynthetic aperture radar (SAR) image ship detection is a popular area of ocean remote sensing, which has broad application prospects in ocean monitoring, maritime rescue and other tasks. Recently, deep learning has been used in this field, but convolutional neural network (CNN) based SAR ship detection still faces some challenges. First, due to the characteristic of CNN’s local convolution, the global information of the ship is not sufficiently learned and the detection is vulnerable to complex background interference. Second, SAR ship imaging varies greatly under different imaging conditions and postures, so CNN is difficult to adapt to scattering change imaging. To solve these problems, we propose a Scattering Information Fusion Network (SIFNet) for oriented ship detection in SAR Images consisting of a multi-scale contextual semantic information fusion (MCSIF) module and a scattering points information learning (SPIL) module. The MCSIF module enhances the acquisition of global information, enabling the network to extract more efficient feature maps. The SPIL module takes advantage of the fact that scattering points can stably represent the key features of the ship under different imaging conditions to make detection more robust through scattering information learning. Experiments show that our method achieves the highest F1-score and AP50 on both HRSID and RSDD-SAR datasets. Han Wang 0049, Silei Liu, Yixuan Lv, Shengyang Li |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | A Multitask Benchmark Dataset for Satellite Video: Object Detection, Tracking, and SegmentationabstractVideo satellites can continuously image large areas and provide dynamic, real-time monitoring of hotspots and objects. The intelligent processing and analysis of satellite video have become a research hotspot in the field of remote sensing. However, the lack of high-quality satellite video datasets limits the development of relevant object detection, object tracking, and object segmentation. In this paper, we build the largest scale satellite video dataset with the most task types supported and object categories, named Satellite Video Multi-Mission Benchmark (SAT-MTB). First, multi-task annotation of aircraft, ships, cars, trains, and their corresponding 14 categories of fine-grained objects in 249 satellite videos is performed based on horizontal bounding boxes (HBB), oriented bounding boxes (OBB), masks, which cover more than 50,000 frames and 1,033,511 annotated object instances. Then, we review the tasks of object detection, object tracking, and object segmentation based on satellite videos, providing a comprehensive overview of progress in related datasets and algorithm research. Finally, we establish the first public benchmark of multi-task algorithms for satellite video object detection, object tracking, and object segmentation, evaluating and analyzing the performance of a total of 47 representative algorithms under different tasks on the constructed dataset. The proposed SAT-MTB will significantly advance research in intelligent processing and analysis of satellite video and related applications. Shengyang Li, Manqi Zhao, Weilong Guo, Yixuan Lv, Longxuan Kou, Han Wang 0049, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 8 |