Haolin Qin

dblp:362/5864 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0001-8569-7430ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 HSOD-BIT-V2: A Challenging Benchmark for Hyperspectral Salient Object Detection
abstract
Salient Object Detection (SOD) is crucial in computer vision, yet RGB-based methods face limitations in challenging scenes, such as small objects and similar color features. Hyperspectral images provide a promising solution for more accurate Hyperspectral Salient Object Detection (HSOD) by abundant spectral information, while HSOD methods are hindered by the lack of extensive and available datasets. In this context, we introduce HSOD-BIT-V2, the largest and most challenging HSOD benchmark dataset to date. Five distinct challenges focusing on small objects and foreground-background similarity are designed to emphasize spectral advantages and real-world complexity. To tackle these challenges, we propose Hyper-HRNet, a high-resolution HSOD network. Hyper-HRNet effectively extracts, integrates, and preserves effective spectral information while reducing dimensionality by capturing the self-similar spectral features. Additionally, it conveys fine details and precisely locates object contours by incorporating comprehensive global information and detailed object saliency representations. Experimental analysis demonstrates that Hyper-HRNet outperforms existing models, especially in challenging scenarios.
Yuhao Qiu, Shuyan Bai, Tingfa Xu, Peifu Liu, Haolin Qin, Jianan Li 0001
AAAI5
2025 MUST: The First Dataset and Unified Framework for Multispectral UAV Single Object Tracking
abstract
UAV tracking faces significant challenges in real-world scenarios, such as small-size targets and occlusions, which limit the performance of RGB-based trackers. Multispectral images (MSI), which capture additional spectral information, offer a promising solution to these challenges. However, progress in this field has been hindered by the lack of relevant datasets. To address this gap, we introduce the first large-scale Multispectral UAV Single Object Tracking dataset (MUST), which includes 250 video sequences spanning diverse environments and challenges, providing a comprehensive data foundation for multispectral UAV tracking. We also propose a novel tracking framework, UNTrack, which encodes unified spectral, spatial, and temporal features from spectrum prompts, initial templates, and sequential searches. UNTrack employs an asymmetric transformer with a spectral background eliminate mechanism for optimal relationship modeling and an encoder that continuously updates the spectrum prompt to refine tracking, improving both accuracy and efficiency. Extensive experiments show that our proposed UNTrack outperforms state-of-the-art UAV trackers. We believe our dataset and framework will drive future research in this area. The dataset is available on https://github.com/q2479036243/MUST-Multispectral-UAV-Single-Object-Tracking.
Haolin Qin, Tingfa Xu, Jianan Li 0001
CVPR1
2025 MSITrack: A Challenging Benchmark for Multispectral Single Object Tracking
abstract
Visual object tracking in real-world scenarios presents numerous challenges including occlusion, interference from similar objects and complex backgrounds - all of which limit the effectiveness of RGB-based trackers. Multispectral imagery, which captures pixel-level spectral reflectance, enhances target discriminability. However, the availability of multispectral tracking datasets remains limited. To bridge this gap, we introduce MSITrack, the largest and most diverse multispectral single object tracking dataset to date. MSITrack offers the following key features: (i) More Challenging Attributes - including interference from similar objects and similarity in color and texture between targets and backgrounds in natural scenarios, along with a wide range of real-world tracking challenges; (ii) Richer and More Natural Scenes - spanning 55 object categories and 300 distinct natural scenes, MSITrack far exceeds the scope of existing benchmarks. Many of these scenes and categories are introduced to the multispectral tracking domain for the first time; (iii) Larger Scale - 300 videos comprising over 129k frames of multispectral imagery. To ensure annotation precision, each frame has undergone meticulous processing, manual labeling and multi-stage verification. Extensive evaluations using representative trackers demonstrate that the multispectral data in MSITrack significantly improves performance over RGB-only baselines, highlighting its potential to drive future advancements in the field. The MSITrack dataset is publicly available at: https://github.com/Fengtao191/MSITrack.
Tingfa Xu, Haolin Qin, Shuaihao Han, Xuyang Zou, Zhan Lv, Jianan Li 0001
ACM Multimedia3
2025 MMOT: The First Challenging Benchmark for Drone-based Multispectral Multi-Object Tracking
abstract
Drone-based multi-object tracking is essential yet highly challenging due to small targets, severe occlusions, and cluttered backgrounds. Existing RGB-based multi-object tracking algorithms heavily depend on spatial appearance cues such as color and texture, which often degrade in aerial views, compromising tracking reliability. Multispectral imagery, capturing pixel-level spectral reflectance, provides crucial spectral cues that significantly enhance object discriminability under degraded spatial conditions. However, the lack of dedicated multispectral UAV datasets has hindered progress in this domain. To bridge this gap, we introduce MMOT, the first challenging benchmark for drone-based multispectral multi-object tracking dataset. It features three key characteristics: (i) Large Scale — 125 video sequences with over 488.8K annotations across eight object categories; (ii) Comprehensive Challenges — covering diverse real-world challenges such as extreme small targets, high-density scenarios, severe occlusions and complex platform motion; and (iii) Precise Oriented Annotations — enabling accurate localization and reduced object ambiguity under aerial perspectives. To better extract spectral features and leverage oriented annotations, we further present a multispectral and orientation-aware MOT scheme adapting existing MOT methods, featuring: (i) a lightweight Spectral 3D-Stem integrating spectral features while preserving compatibility with RGB pretraining; (ii) a orientation-aware Kalman filter for precise state estimation; and (iii) an end-to-end orientation-adaptive transformer architecture. Extensive experiments across representative trackers consistently show that multispectral input markedly improves tracking performance over RGB baselines, particularly for small and densely packed objects. We believe our work will benefit the community for advancing drone-based multispectral multi-object tracking research. Our MMOT, code and benchmarks are publicly available at https://github.com/Annzstbl/MMOT.
Tingfa Xu, Ying Wang 0064, Haolin Qin, Jianan Li 0001
NeurIPS4
2025 IRSTD-YOLO: An Improved YOLO Framework for Infrared Small Target Detection
abstract
Detecting small targets in infrared images, especially in low-contrast and complex backgrounds, remains challenging. To tackle this, we propose infrared small target detection YOLO (IRSTD-YOLO), a novel detection network. The Edge and Feature Extraction (EFE) module enhances feature representation by integrating a SobelConv branch and a 2DConv branch. The SobelConv branch applies Sobel operators to extract gradient information, enhancing edge contrast and making small targets more distinguishable from the background. Unlike standard convolutions, which process all features uniformly, this edge-aware operation emphasizes structural information crucial for detecting small infrared targets. The 2DConv branch captures spatial context, complementing the edge features to create a more comprehensive representation. To further refine detection, we introduce the Infrared Small Target Enhancement (IRSTE) module, addressing the limitations of conventional feature pyramid networks. Instead of merely adding a shallow detection head, IRSTE processes and enhances shallow-layer features, which are rich in small target information, and fuses them with deeper features. By leveraging a multi-branch strategy that integrates local, global, and large-scale contexts, IRSTE enhances small target representation and detection robustness, particularly in low-contrast environments where traditional networks often fail. Experimental results show that IRSTD-YOLO achieves an [email protected]:0.95 of 36.7% on the InfraredUAV dataset and 51.6% on the AntiUAV310 dataset, outperforming YOLOv11-s by 4.4% and 4.2%, respectively.
Tingfa Xu, Haolin Qin, Jianan Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2025 CVT-Track: Concentrating on Valid Tokens for One-Stream Tracking
abstract
In the domain of single object tracking, the Ground Truth bounding box is intentionally sized larger than the minimum dimensions required to enclose the target in the initial video frame, inadvertently including extraneous elements and interferences in the template image. Moreover, significant appearance changes of the target during movement present substantial challenges for maintaining robust tracking. To address these issues, this study introduces a novel one-stream tracking framework named CVT-Track. CVT-Track comprises two main components: the Target Valid Token Collection (TaVTC) and the Temporal Valid Token Collection (TeVTC) modules. The TaVTC module effectively mitigates background noise and interference from similar targets, thereby sharpening the focus on the target’s unique features and enhancing tracking accuracy. Conversely, the TeVTC module skillfully extracts target information from historical frames, capturing the target’s dynamic appearance changes throughout the tracking process and thereby improving tracking robustness. The synergistic operation of these modules markedly enhances both the accuracy and robustness of tracking. Empirical evaluations demonstrate that CVT-Track achieves state-of-the-art performance across multiple datasets and maintains superior inference speeds.
Jianan Li 0001, Xiaoying Yuan, Haolin Qin, Ying Wang 0064, Xincong Liu, Tingfa Xu
IEEE Trans. Circuits Syst. Video Technol.3
2025 OSFormer: One-Step Transformer for Infrared Video Small Object Detection
abstract
Infrared video small object detection is pivotal in numerous security and surveillance applications. However, existing deep learning-based methods, which typically rely on a two-step paradigm of frame-by-frame detection followed by temporal refinement, struggle to effectively utilize temporal information. This is particularly challenging when detecting small objects against complex backgrounds. To address these issues, we introduce the One-Step Transformer (OSFormer), a novel method that pioneeringly integrates a small-object-friendly transformer with a one-step detection paradigm. Unlike traditional methods, OSFormer processes the video sequence only through a single inference, encoding the sequence into cube format data and tracking object motion trajectories. Additionally, we propose the Varied-Size Patch Attention (VPA) module, which generates patches of varying sizes to capture adaptive attention features, bridging the gap between transformer architectures and small object detection. To further enhance detection accuracy, OSFormer incorporates a Doppler Adaptive Filter, which integrates traditional filtering techniques into an end-to-end neural network to suppress background noise and accentuate small objects. OSFormer outperforms YOLOv8-s on both the AntiUAV dataset (+ $3.1\%~\text {mAP}_{50}$ , - $35.1\%~\text {Params}$ ) and the InfraredUAV dataset (+ $4.0\%~\text {mAP}_{50-95}$ , - $51.0\%~\text {FLOPs}$ ), demonstrating superior efficiency and effectiveness in small object detection. The code is available on https://github.com/q2479036243/OSFormer.
Haolin Qin, Tingfa Xu, Fengxiang Xu, Jianan Li 0001
IEEE Trans. Image Process.1
2025 Factorization Vision Transformer: Modeling Long-Range Dependency With Local Window Cost
abstract
Transformers have astounding representational power but typically consume considerable computation which is quadratic with image resolution. The prevailing Swin transformer reduces computational costs through a local window strategy. However, this strategy inevitably causes two drawbacks: 1) the local window-based self-attention (WSA) hinders global dependency modeling capability and 2) recent studies point out that local windows impair robustness. To overcome these challenges, we pursue a preferable trade-off between computational cost and performance. Accordingly, we propose a novel factorization self-attention (FaSA) mechanism that enjoys both the advantages of local window cost and long-range dependency modeling capability. By factorizing the conventional attention matrix into sparse subattention matrices, FaSA captures long-range dependencies, while aggregating mixed-grained information at a computational cost equivalent to the local WSA. Leveraging FaSA, we present the factorization vision transformer (FaViT) with a hierarchical structure. FaViT achieves high performance and robustness, with linear computational complexity concerning input image spatial resolution. Extensive experiments have shown FaViT's advanced performance in classification and downstream tasks. Furthermore, it also exhibits strong model robustness to corrupted and biased data and hence demonstrates benefits in favor of practical applications. In comparison to the baseline model Swin-T, our FaViT-B2 significantly improves classification accuracy by 1% and robustness by 7%, while reducing model parameters by 14%. Our code will soon be publicly available: at https://github.com/q2479036243/FaViT.
Haolin Qin, Daquan Zhou, Tingfa Xu, Ziyang Bian, Jianan Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 BACTrack: Building Appearance Collection for Aerial Tracking
abstract
Siamese network-based trackers have shown remarkable success in aerial tracking. Most previous works, however, usually perform template matching only between the initial template and the search region and thus fail to deal with rapidly changing targets that often appear in aerial tracking. As a remedy, this work presents Building Appearance Collection Tracking (BACTrack). This simple yet effective tracking framework builds a dynamic collection of target templates online and performs efficient multi-template matching to achieve robust tracking. Specifically, BACTrack mainly comprises a Mixed-Temporal Transformer (MTT) and an appearance discriminator. The former is responsible for efficiently building relationships between the search region and multiple target templates in parallel through a mixed-temporal attention mechanism. At the same time, the appearance discriminator employs an online adaptive template-update strategy to ensure that the collected multiple templates remain reliable and diverse, allowing them to closely follow rapid changes in the target’s appearance and suppress background interference during tracking. Extensive experiments show that our BACTrack achieves top performance on four challenging aerial tracking benchmarks while maintaining an impressive speed of over 87 FPS on a single GPU. Speed tests on embedded platforms also validate our potential suitability for deployment on UAV platforms.
Xincong Liu, Tingfa Xu, Ying Wang 0064, Zhinong Yu, Xiaoying Yuan, Haolin Qin, Jianan Li 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Multi-Step Temporal Modeling for UAV Tracking
abstract
In the realm of unmanned aerial vehicle (UAV) tracking, Siamese-based approaches have gained traction due to their optimal balance between efficiency and precision. However, UAV scenarios often present challenges such as insufficient sampling resolution, fast motion and small objects with limited feature information. As a result, temporal context in UAV tracking tasks plays a pivotal role in target location, overshadowing the target’s precise features. In this paper, we introduce MT-Track, a streamlined and efficient multi-step temporal modeling framework designed to harness the temporal context from historical frames for enhanced UAV tracking. This temporal integration occurs in two steps: correlation map generation and correlation map refinement. Specifically, we unveil a unique temporal correlation module that dynamically assesses the interplay between the template and search region features. This module leverages temporal information to refresh the template feature, yielding a more precise correlation map. Subsequently, we propose a mutual transformer module to refine the correlation maps of historical and current frames by modeling the temporal knowledge in the tracking sequence. This method significantly trims computational demands compared to the raw transformer. The compact yet potent nature of our tracking framework ensures commendable tracking outcomes, particularly in extended tracking scenarios. Comprehensive tests across four renowned UAV benchmarks substantiate the superior efficacy of our approach, delivering real-time performance at 84.7 FPS on a single GPU. Real-world test on the NVIDIA AGX hardware platform achieves a speed exceeding 30 FPS, validating the practicality of our method.
Xiaoying Yuan, Tingfa Xu, Xincong Liu, Ying Wang 0064, Haolin Qin, Yuqiang Fang, Jianan Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 DMSSN: Distilled Mixed Spectral-Spatial Network for Hyperspectral Salient Object Detection
abstract
Hyperspectral salient object detection (HSOD) has exhibited remarkable promise across various applications, particularly in intricate scenarios where conventional RGB-based approaches fall short. Despite the considerable progress in HSOD method advancements, two critical challenges require immediate attention. Firstly, existing hyperspectral data dimension reduction techniques incur a loss of spectral information, which adversely affects detection accuracy. Secondly, previous methods insufficiently harness the inherent distinctive attributes of hyperspectral images (HSIs) during the feature extraction process. To address these challenges, we propose a novel approach termed the Distilled Mixed Spectral-Spatial Network (DMSSN), comprising a Distilled Spectral Encoding process and a Mixed Spectral-Spatial Transformer (MSST) feature extraction network. The encoding process utilizes knowledge distillation to construct a lightweight autoencoder for dimension reduction, striking a balance between robust encoding capabilities and low computational costs. The MSST extracts spectral-spatial features through multiple attention head groups, collaboratively enhancing its resistance to intricate scenarios. Moreover, we have created a large-scale HSOD dataset, HSOD-BIT, to tackle the issue of data scarcity in this field and meet the fundamental data requirements of deep network training. Extensive experiments demonstrate that our proposed DMSSN achieves state-of-the-art performance on multiple datasets. We will soon make the code and dataset publicly available on https://github.com/anonymous0519/HSOD-BIT.
Haolin Qin, Tingfa Xu, Peifu Liu, Jingxuan Xu, Jianan Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Spectrum-Driven Mixed-Frequency Network for Hyperspectral Salient Object Detection
abstract
Hyperspectral salient object detection (HSOD) aims to detect spectrally salient objects in hyperspectral images (HSIs). However, existing methods inadequately utilize spectral information by either converting HSIs into false-color images or converging neural networks with clustering. We propose a novel approach that fully leverages the spectral characteristics by extracting two distinct frequency components from the spectrum: low-frequency Spectral Saliency and high-frequency Spectral Edge. The Spectral Saliency approximates the region of salient objects, while the Spectral Edge captures edge information of salient objects. These two complementary components, crucial for HSOD, are derived by computing from the inter-layer spectral angular distance of the Gaussian pyramid and the intra-neighborhood spectral angular gradients, respectively. To effectively utilize this dual-frequency information, we introduce a novel lightweight Spectrum-driven Mixed-frequency Network (SMN). SMN incorporates two parameter-free plug-and-play operators, namely Spectral Saliency Generator and Spectral Edge Operator, to extract the Spectral Saliency and Spectral Edge components from the input HSI independently. Subsequently, the Mixed-frequency Attention module, comprised of two frequency-dependent heads, intelligently combines the embedded features of edge and saliency information, resulting in a mixed-frequency feature representation. Furthermore, a saliency-edge-aware decoder progressively scales up the mixed-frequency feature while preserving rich detail and saliency information for accurate salient object prediction. Extensive experiments conducted on the HS-SOD benchmark and our custom dataset HSOD-BIT demonstrate that our SMN outperforms state-of-the-art methods regarding HSOD performance. Code and dataset will be available athttps://github.com/laprf/SMN.
Peifu Liu, Tingfa Xu, Huan Chen 0018, Shiyun Zhou, Haolin Qin, Jianan Li 0001
IEEE Trans. Multim.5