Tingfa Xu

dblp:93/1709 · DBLP profile ↗
← Back
87ranked-venue papers
1as first author
69since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 29 since 2021Artificial intelligence and machine learning · 35 · 1 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 20 since 2021Computer networks · 6 · 6 since 2021
YearPublicationVenuePosition
2026 HyperCOD: The First Challenging Benchmark and Baseline for Hyperspectral Camouflaged Object Detection
abstract
RGB-based camouflaged object detection struggles in real-world scenarios where color and texture cues are ambiguous. While hyperspectral image offers a powerful alternative by capturing fine-grained spectral signatures, progress in hyperspectral camouflaged object detection (HCOD) has been critically hampered by the absence of a dedicated, large-scale benchmark. To spur innovation, we introduce HyperCOD, the first challenging benchmark for HCOD. Comprising 350 high-resolution hyperspectral images, It features complex real-world scenarios with minimal objects, intricate shapes, severe occlusions, and dynamic lighting to challenge current models.The advent of foundation models like the Segment Anything Model (SAM) presents a compelling opportunity. To adapt the Segment Anything Model (SAM) for HCOD, we propose HyperSpectral Camouflage-aware SAM (HSC-SAM). HSC-SAM ingeniously reformulates the hyperspectral image by decoupling it into a spatial map fed to SAM's image encoder and a spectral saliency map that serves as an adaptive prompt. This translation effectively bridges the modality gap. Extensive experiments show that HSC-SAM sets a new state-of-the-art on HyperCOD and generalizes robustly to other public HSI datasets. The HyperCOD dataset and our HSC-SAM baseline provide a robust foundation to foster future research in this emerging area.
Shuyan Bai, Tingfa Xu, Peifu Liu, Yuhao Qiu, Huiyan Bai, Huan Chen 0018, Yanyan Peng, Jianan Li 0001
AAAI2
2026 MODA: The First Challenging Benchmark for Multispectral Object Detection in Aerial Images
abstract
Aerial object detection faces significant challenges in real-world scenarios, such as small objects and extensive background interference, which limit the performance of RGB-based detectors with insufficient discriminative information. Multispectral images (MSIs) capture additional spectral cues across multiple bands, offering a promising alternative. However, the lack of training data has been the primary bottleneck to exploiting the potential of MSIs. To address this gap, we introduce the first large-scale dataset for Multispectral Object Detection in Aerial images (MODA), which comprises 14,041 MSIs and 330,191 annotations across diverse, challenging scenarios, providing a comprehensive data foundation for this field. Furthermore, to overcome challenges inherent to aerial object detection using MSIs, we propose OSSDet, a framework that integrates spectral and spatial information with object-aware cues. OSSDet employs a cascaded spectral-spatial modulation structure to optimize target perception, aggregates spectrally related features by exploiting spectral similarities to reinforce intra-object correlations, and suppresses irrelevant background via object-aware masking. Moreover, cross-spectral attention further refines object-related representations under explicit object-aware guidance. Extensive experiments demonstrate that OSSDet outperforms existing methods with comparable parameters and efficiency.
Shuaihao Han, Tingfa Xu, Peifu Liu, Jianan Li 0001
AAAI2
2026 Local grid rendering networks for 3D object detection in point clouds
Jianan Li 0001, Lihe Ding, Jie Wang 0097, Tingfa Xu
Pattern Recognit.4
2026 COXNet: Cross-Layer Fusion With Adaptive Alignment and Scale Integration for RGBT Tiny Object Detection
abstract
Detecting tiny objects in multimodal Red-Green-Blue-Thermal (RGBT) imagery is a critical challenge in computer vision, particularly in surveillance, search and rescue, and autonomous navigation. Drone-based scenarios exacerbate these challenges due to spatial misalignment, low-light conditions, occlusion, and cluttered backgrounds. Current methods struggle to leverage the complementary information between visible and thermal modalities effectively. We propose COXNet, a novel framework for RGBT tiny object detection, addressing these issues through three core innovations: i) the Cross-Layer Fusion Module, fusing high-level visible and low-level thermal features for enhanced semantic and spatial accuracy; ii) the Dynamic Alignment and Scale Refinement module, correcting cross-modal spatial misalignments and preserving multi-scale features; and iii) an optimized label assignment strategy using the GeoShape Similarity Measure for better localization. COXNet achieves a 3.32% mAP50improvement on the RGBTDronePerson dataset over state-of-the-art methods, demonstrating its effectiveness for robust detection in complex environments.
Peiran Peng, Tingfa Xu, Liqiang Song, Mengqi Zhu, Yuqiang Fang, Jianan Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 PromptReg: Interactive Registration by "Corresponding Prompts" for Segment Anything Model (SAM)
abstract
Effectively establishing correspondence between two images is at the centre of image registration methods. Spatially omnipresent representations, including dense displacement fields (DDFs) and spatial (non-)rigid transformations, have been used to parameterise such correspondence. Alternatively, region-based representation uses paired regions of interest (ROIs) to represent region-level correspondence, while retaining its local and dense representation capability at pixel/voxel level if required. Thus, registration can be re-envisioned as a problem of segmenting corresponding paired ROIs in the to-be-registered images. In this work, we utilize models such as SAM, which are pre-trained on substantive datasets, to segment ROIs of the same class from two images, for a new training-free, non-iterative registration algorithm. First, a "corresponding prompt problem" is posed to find a corresponding Prompt Y on Image Y, given any vision Prompt X on Image X, such that the two respectively prompt-conditioned segmentations are a pair of corresponding ROIs from the two images. Second, we propose an "inverse prompt" solution to the corresponding prompt problem, by inverting Prompt X to the Image Y prompt space, where the Jacobian of prototypical features is used. Third, we propose a new registration algorithm that identifies multiple paired corresponding ROIs, by marginalizing the inverted Prompt X over both prompt and spatial spaces, random sampling Prompt X and spatial warping Image X. Comprehensive experiments were conducted on five applications of registering 3D prostate MR, 3D abdomen CT, 3D lung CT, 2D histopathology and, as a non-medical example, 2D aerial images. Based on metrics including Dice and target registration errors on anatomical structures, the proposed registration outperforms both intensity-based iterative algorithms and learning-based networks, even yielding competitive performance with weakly-supervised registration which requires fully-segmented training data.
Shiqi Huang 0001, Tingfa Xu, Jianan Li 0001, Shaheer U. Saeed, Ziyi Shen, Dean C. Barratt, Yipeng Hu
IEEE Trans. Image Process.2
2025 HSOD-BIT-V2: A Challenging Benchmark for Hyperspectral Salient Object Detection
abstract
Salient Object Detection (SOD) is crucial in computer vision, yet RGB-based methods face limitations in challenging scenes, such as small objects and similar color features. Hyperspectral images provide a promising solution for more accurate Hyperspectral Salient Object Detection (HSOD) by abundant spectral information, while HSOD methods are hindered by the lack of extensive and available datasets. In this context, we introduce HSOD-BIT-V2, the largest and most challenging HSOD benchmark dataset to date. Five distinct challenges focusing on small objects and foreground-background similarity are designed to emphasize spectral advantages and real-world complexity. To tackle these challenges, we propose Hyper-HRNet, a high-resolution HSOD network. Hyper-HRNet effectively extracts, integrates, and preserves effective spectral information while reducing dimensionality by capturing the self-similar spectral features. Additionally, it conveys fine details and precisely locates object contours by incorporating comprehensive global information and detailed object saliency representations. Experimental analysis demonstrates that Hyper-HRNet outperforms existing models, especially in challenging scenarios.
Yuhao Qiu, Shuyan Bai, Tingfa Xu, Peifu Liu, Haolin Qin, Jianan Li 0001
AAAI3
2025 FBRT-YOLO: Faster and Better for Real-Time Aerial Image Detection
abstract
Embedded flight devices with visual capabilities have become essential for a wide range of applications. In aerial image detection, while many existing methods have partially addressed the issue of small target detection, challenges remain in optimizing small target detection and balancing detection accuracy with efficiency. These issues are key obstacles to the advancement of real-time aerial image detection. In this paper, we propose a new family of real-time detectors for aerial image detection, named FBRT-YOLO, to address the imbalance between detection accuracy and efficiency. Our method comprises two lightweight modules: Feature Complementary Mapping Module (FCM) and Multi-Kernel Perception Unit (MKP), designed to enhance object perception for small targets in aerial images. FCM focuses on alleviating the problem of information imbalance caused by the loss of small target information in deep networks. It aims to integrate spatial positional information of targets more deeply into the network, better aligning with semantic information in the deeper layers to improve the localization of small targets. We introduce MKP, which leverages convolutions with kernels of different sizes to enhance the relationships between targets of various scales and improve the perception of targets at different scales. Extensive experimental results on three major aerial image datasets, including Visdrone, UAVDT, and AI-TOD, demonstrate that FBRT-YOLO outperforms various real-time detectors in terms of performance and speed.
Tingfa Xu, Jianan Li 0001
AAAI2
2025 MUST: The First Dataset and Unified Framework for Multispectral UAV Single Object Tracking
abstract
UAV tracking faces significant challenges in real-world scenarios, such as small-size targets and occlusions, which limit the performance of RGB-based trackers. Multispectral images (MSI), which capture additional spectral information, offer a promising solution to these challenges. However, progress in this field has been hindered by the lack of relevant datasets. To address this gap, we introduce the first large-scale Multispectral UAV Single Object Tracking dataset (MUST), which includes 250 video sequences spanning diverse environments and challenges, providing a comprehensive data foundation for multispectral UAV tracking. We also propose a novel tracking framework, UNTrack, which encodes unified spectral, spatial, and temporal features from spectrum prompts, initial templates, and sequential searches. UNTrack employs an asymmetric transformer with a spectral background eliminate mechanism for optimal relationship modeling and an encoder that continuously updates the spectrum prompt to refine tracking, improving both accuracy and efficiency. Extensive experiments show that our proposed UNTrack outperforms state-of-the-art UAV trackers. We believe our dataset and framework will drive future research in this area. The dataset is available on https://github.com/q2479036243/MUST-Multispectral-UAV-Single-Object-Tracking.
Haolin Qin, Tingfa Xu, Jianan Li 0001
CVPR2
2025 PvNeXt: Rethinking Network Design and Temporal Motion for Point Cloud Video Recognition
abstract
Point cloud video perception has become an essential task for the realm of 3D vision. Current 4D representation learning techniques typically engage in iterative processing coupled with dense query operations. Although effective in capturing temporal features, this approach leads to substantial computational redundancy. In this work, we propose a framework, named as PvNeXt, for effective yet efficient point cloud video recognition, via personalized one-shot query operation. Specially, PvNeXt consists of two key modules, the Motion Imitator and the Single-Step Motion Encoder. The former module, the Motion Imitator, is designed to capture the temporal dynamics inherent in sequences of point clouds, thus generating the virtual motion corresponding to each frame. The Single-Step Motion Encoder performs a one-step query operation, associating point cloud of each frame with its corresponding virtual motion frame, thereby extracting motion cues from point cloud sequences and capturing temporal dynamics across the entire sequence. Through the integration of these two modules, {PvNeXt} enables personalized one-shot queries for each frame, effectively eliminating the need for frame-specific looping and intensive query processes. Extensive experiments on multiple benchmarks demonstrate the effectiveness of our method.
Jie Wang 0097, Tingfa Xu, Lihe Ding, Long Bai 0008, Jianan Li 0001
ICLR2
2025 Register Anything: Estimating "Corresponding Prompts" for Segment Anything Model
Shiqi Huang 0001, Tingfa Xu, Wen Yan 0005, Dean C. Barratt, Yipeng Hu
MICCAI (4)2
2025 MSITrack: A Challenging Benchmark for Multispectral Single Object Tracking
abstract
Visual object tracking in real-world scenarios presents numerous challenges including occlusion, interference from similar objects and complex backgrounds - all of which limit the effectiveness of RGB-based trackers. Multispectral imagery, which captures pixel-level spectral reflectance, enhances target discriminability. However, the availability of multispectral tracking datasets remains limited. To bridge this gap, we introduce MSITrack, the largest and most diverse multispectral single object tracking dataset to date. MSITrack offers the following key features: (i) More Challenging Attributes - including interference from similar objects and similarity in color and texture between targets and backgrounds in natural scenarios, along with a wide range of real-world tracking challenges; (ii) Richer and More Natural Scenes - spanning 55 object categories and 300 distinct natural scenes, MSITrack far exceeds the scope of existing benchmarks. Many of these scenes and categories are introduced to the multispectral tracking domain for the first time; (iii) Larger Scale - 300 videos comprising over 129k frames of multispectral imagery. To ensure annotation precision, each frame has undergone meticulous processing, manual labeling and multi-stage verification. Extensive evaluations using representative trackers demonstrate that the multispectral data in MSITrack significantly improves performance over RGB-only baselines, highlighting its potential to drive future advancements in the field. The MSITrack dataset is publicly available at: https://github.com/Fengtao191/MSITrack.
Tingfa Xu, Haolin Qin, Shuaihao Han, Xuyang Zou, Zhan Lv, Jianan Li 0001
ACM Multimedia2
2025 MCOD: The First Challenging Benchmark for Multispectral Camouflaged Object Detection
abstract
Camouflaged Object Detection (COD) aims to identify objects that blend seamlessly into natural scenes. Although RGB-based methods have advanced, their performance remains limited under challenging conditions. Multispectral imagery, providing rich spectral information, offers a promising alternative for enhanced foreground-background discrimination. However, existing COD benchmark datasets are exclusively RGB-based, lacking essential support for multispectral approaches, which has impeded progress in this area. To address this gap, we introduce MCOD, the first challenging benchmark dataset specifically designed for multispectral camouflaged object detection. MCOD features three key advantages: (i) Comprehensive challenge attributes: It captures real-world difficulties such as small object sizes and extreme lighting conditions commonly encountered in COD tasks. (ii) Diverse real-world scenarios: The dataset spans a wide range of natural environments to better reflect practical applications. (iii) High-quality pixel-level annotations: Each image is manually annotated with precise object masks and corresponding challenge attribute labels. We benchmark eleven representative COD methods on MCOD, observing a consistent performance drop due to increased task difficulty. Notably, integrating multispectral modalities substantially alleviates this degradation, highlighting the value of spectral information in enhancing detection robustness. We anticipate MCOD will provide a strong foundation for future research in multispectral camouflaged object detection. The dataset is publicly accessible at https://github.com/yl2900260-bit/MCOD.
Tingfa Xu, Shuyan Bai, Peifu Liu, Jianan Li 0001
ACM Multimedia2
2025 MMOT: The First Challenging Benchmark for Drone-based Multispectral Multi-Object Tracking
abstract
Drone-based multi-object tracking is essential yet highly challenging due to small targets, severe occlusions, and cluttered backgrounds. Existing RGB-based multi-object tracking algorithms heavily depend on spatial appearance cues such as color and texture, which often degrade in aerial views, compromising tracking reliability. Multispectral imagery, capturing pixel-level spectral reflectance, provides crucial spectral cues that significantly enhance object discriminability under degraded spatial conditions. However, the lack of dedicated multispectral UAV datasets has hindered progress in this domain. To bridge this gap, we introduce MMOT, the first challenging benchmark for drone-based multispectral multi-object tracking dataset. It features three key characteristics: (i) Large Scale — 125 video sequences with over 488.8K annotations across eight object categories; (ii) Comprehensive Challenges — covering diverse real-world challenges such as extreme small targets, high-density scenarios, severe occlusions and complex platform motion; and (iii) Precise Oriented Annotations — enabling accurate localization and reduced object ambiguity under aerial perspectives. To better extract spectral features and leverage oriented annotations, we further present a multispectral and orientation-aware MOT scheme adapting existing MOT methods, featuring: (i) a lightweight Spectral 3D-Stem integrating spectral features while preserving compatibility with RGB pretraining; (ii) a orientation-aware Kalman filter for precise state estimation; and (iii) an end-to-end orientation-adaptive transformer architecture. Extensive experiments across representative trackers consistently show that multispectral input markedly improves tracking performance over RGB baselines, particularly for small and densely packed objects. We believe our work will benefit the community for advancing drone-based multispectral multi-object tracking research. Our MMOT, code and benchmarks are publicly available at https://github.com/Annzstbl/MMOT.
Tingfa Xu, Ying Wang 0064, Haolin Qin, Jianan Li 0001
NeurIPS2
2025 An information-Enhanced memory library for multi-modal continuous clustering
Tingfa Xu, Jianan Li 0001
Knowl. Based Syst.2
2025 IRSTD-YOLO: An Improved YOLO Framework for Infrared Small Target Detection
abstract
Detecting small targets in infrared images, especially in low-contrast and complex backgrounds, remains challenging. To tackle this, we propose infrared small target detection YOLO (IRSTD-YOLO), a novel detection network. The Edge and Feature Extraction (EFE) module enhances feature representation by integrating a SobelConv branch and a 2DConv branch. The SobelConv branch applies Sobel operators to extract gradient information, enhancing edge contrast and making small targets more distinguishable from the background. Unlike standard convolutions, which process all features uniformly, this edge-aware operation emphasizes structural information crucial for detecting small infrared targets. The 2DConv branch captures spatial context, complementing the edge features to create a more comprehensive representation. To further refine detection, we introduce the Infrared Small Target Enhancement (IRSTE) module, addressing the limitations of conventional feature pyramid networks. Instead of merely adding a shallow detection head, IRSTE processes and enhances shallow-layer features, which are rich in small target information, and fuses them with deeper features. By leveraging a multi-branch strategy that integrates local, global, and large-scale contexts, IRSTE enhances small target representation and detection robustness, particularly in low-contrast environments where traditional networks often fail. Experimental results show that IRSTD-YOLO achieves an [email protected]:0.95 of 36.7% on the InfraredUAV dataset and 51.6% on the AntiUAV310 dataset, outperforming YOLOv11-s by 4.4% and 4.2%, respectively.
Tingfa Xu, Haolin Qin, Jianan Li 0001
IEEE Geosci. Remote. Sens. Lett.2
2025 BiMAConv: Bimodal Adaptive Convolution for Multispectral Point Cloud Segmentation
abstract
Multispectral point cloud segmentation, leveraging both spatial and spectral information to classify individual points, is crucial for applications such as remote sensing, autonomous driving, and urban planning. However, existing methods primarily focus on spatial information and merge it with spectral data without fully considering their differences, limiting the effective use of spectral information. In this paper, we introduce a novel approach, Bi-Modal Adaptive Convolution (BiMAConv), which fully exploits information from different modalities, based on the divide-and-conquer philosophy. Specifically, BiMAConv leverages the spectral features provided by the Spectral Information Divergence (SID) and the weight information provided by the Modal-Weight Block (MW-Block) module. The SID highlights slight differences in spectral information, providing the detailed differential feature information. The MW-Block module utilizes attention mechanism to combine generated features with the original point cloud, thereby generating weights to maintain learning balance sharply. Additionally, we reconstruct a large-scale urban point cloud dataset GRSS DFC 2018 3D based on dataset GRSS DFC 2018 to advance the field of multispectral remote sensing point cloud, with a greater number of categories, more precise annotations, and registered multispectral channels. BiMAConv is fundamentally plug-and-play and supports different shared-MLP methods with almost no architectural changes. Extensive experiments on GRSS DFC 2018 3D and Toronto-3D benchmarks demonstrate that our method significantly boosts the performance of popular detectors.
Tingfa Xu, Peng Lou, Tiehong Tian, Jianan Li 0001
IEEE Geosci. Remote. Sens. Lett.2
2025 Towards Robust Point Cloud Recognition With Sample-Adaptive Auto-Augmentation
abstract
Robust 3D perception amidst corruption is a crucial task in the realm of 3D vision. Conventional data augmentation methods aimed at enhancing corruption robustness typically apply random transformations to all point cloud samples offline, neglecting sample structure, which often leads to over- or under-enhancement. In this study, we propose an alternative approach to address this issue by employing sample-adaptive transformations based on sample structure, through an auto-augmentation framework named AdaptPoint++. Central to this framework is an imitator, which initiates with Position-aware Feature Extraction to derive intrinsic structural information from the input sample. Subsequently, a Deformation Controller and a Mask Controller predict per-anchor deformation and per-point masking parameters, respectively, facilitating corruption simulations. In conjunction with the imitator, a discriminator is employed to curb the generation of excessive corruption that deviates from the original data distribution. Moreover, we integrate a perception-guidance feedback mechanism to steer the generation of samples towards an appropriate difficulty level. To effectively train the classifier using the generated augmented samples, we introduce a Structure Reconstruction-assisted learning mechanism, bolstering the classifier's robustness by prioritizing intrinsic structural characteristics over superficial discrepancies induced by corruption. Additionally, to alleviate the scarcity of real-world corrupted point cloud data, we introduce two novel datasets: ScanObjectNN-C and MVPNET-C, closely resembling actual data in real-world scenarios. Experimental results demonstrate that our method attains state-of-the-art performance on multiple corruption benchmarks.
Jianan Li 0001, Jie Wang 0097, Tingfa Xu
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 CVT-Track: Concentrating on Valid Tokens for One-Stream Tracking
abstract
In the domain of single object tracking, the Ground Truth bounding box is intentionally sized larger than the minimum dimensions required to enclose the target in the initial video frame, inadvertently including extraneous elements and interferences in the template image. Moreover, significant appearance changes of the target during movement present substantial challenges for maintaining robust tracking. To address these issues, this study introduces a novel one-stream tracking framework named CVT-Track. CVT-Track comprises two main components: the Target Valid Token Collection (TaVTC) and the Temporal Valid Token Collection (TeVTC) modules. The TaVTC module effectively mitigates background noise and interference from similar targets, thereby sharpening the focus on the target’s unique features and enhancing tracking accuracy. Conversely, the TeVTC module skillfully extracts target information from historical frames, capturing the target’s dynamic appearance changes throughout the tracking process and thereby improving tracking robustness. The synergistic operation of these modules markedly enhances both the accuracy and robustness of tracking. Empirical evaluations demonstrate that CVT-Track achieves state-of-the-art performance across multiple datasets and maintains superior inference speeds.
Jianan Li 0001, Xiaoying Yuan, Haolin Qin, Ying Wang 0064, Xincong Liu, Tingfa Xu
IEEE Trans. Circuits Syst. Video Technol.6
2025 Hyperspectral Remote Sensing Images Salient Object Detection: The First Benchmark Dataset and Baseline
abstract
The objective of hyperspectral remote sensing image salient object detection (HRSI-SOD) is to identify objects or regions that exhibit distinct spectrum contrasts with the background. This area holds significant promise for practical applications; however, progress has been limited by a notable scarcity of dedicated datasets and methodologies. To bridge this gap and stimulate further research, we introduce the first HRSI-SOD dataset, termed HRSSD, which includes 704 hyperspectral images and 5327 pixel-level annotated salient objects. The HRSSD dataset poses substantial challenges for salient object detection algorithms due to large scale variation, diverse foreground-background relations, and multi-salient objects. Additionally, we propose an innovative and efficient baseline model for HRSISOD, termed the Deep Spectral Saliency Network (DSSN). The core of DSSN is the Cross-level Saliency Assessment Block, which performs pixel-wise attention and evaluates the contributions of multi-scale similarity maps at each spatial location, effectively reducing erroneous responses in cluttered regions and emphasizes salient regions across scales. Additionally, the High-resolution Fusion Module combines bottom-up fusion strategy and learned spatial upsampling to leverage the strengths of multi-scale saliency maps, ensuring accurate localization of small objects. Experiments on the HRSSD dataset robustly validate the superiority of DSSN, underscoring the critical need for specialized datasets and methodologies in this domain. Further evaluations on the HSOD-BIT and HS-SOD datasets demonstrate the generalizability of the proposed method. The dataset and source code are publicly available at https://github.com/laprf/HRSSD.
Peifu Liu, Huiyan Bai, Tingfa Xu, Jihui Wang, Huan Chen 0018, Jianan Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 FocusTrack: A Self-Adaptive Local Sampling Algorithm for Efficient Anti-UAV Tracking
abstract
Anti-UAV tracking poses significant challenges, including small target sizes, abrupt camera motion, and cluttered infrared backgrounds. Existing tracking paradigms can be broadly categorized intoglobal-basedandlocal-basedmethods. Global-based trackers, such as SiamDT [1] and SiamSTA [2], achieve high accuracy by scanning the entire field of view but suffer from excessive computational overhead, limiting real-world deployment. In contrast, local-based methods, including OSTrack [3] and ROMTrack [4], efficiently restrict the search region but struggle when targets undergo significant displacements due to abrupt camera motion. Through preliminary experiments, it is evident that a local tracker, when paired with adaptive search region adjustment, can significantly enhance tracking accuracy, narrowing the gap between local and global trackers. To address this challenge, we propose FocusTrack, a novel framework that dynamically refines the search region and strengthens feature representations, achieving an optimal balance between computational efficiency and tracking accuracy. Specifically, our Search Region Adjustment (SRA) strategy estimates the target presence probability and adaptively adjusts the field of view, ensuring the target remains within focus. Furthermore, to counteract feature degradation caused by varying search regions, the Attention-to-Mask (ATM) module is proposed. This module integrates hierarchical information, enriching the target representations with fine-grained details. Experimental results demonstrate that FocusTrack achieves state-of-the-art performance, obtaining 67.7% AUC on AntiUAV [5] and 62.8% AUC on AntiUAV410 [1], outperforming the baseline tracker by 8.5% and 9.1% AUC, respectively. In terms of efficiency, FocusTrack surpasses global-based trackers, requiring only 30G MACs and achieving 143 fps with FocusTrack (SRA) and 44 fps with the full version, both enabling real-time tracking. Code and models are available at https://github.com/vero1925/FocusTrack.
Ying Wang 0064, Tingfa Xu, Jianan Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 OSFormer: One-Step Transformer for Infrared Video Small Object Detection
abstract
Infrared video small object detection is pivotal in numerous security and surveillance applications. However, existing deep learning-based methods, which typically rely on a two-step paradigm of frame-by-frame detection followed by temporal refinement, struggle to effectively utilize temporal information. This is particularly challenging when detecting small objects against complex backgrounds. To address these issues, we introduce the One-Step Transformer (OSFormer), a novel method that pioneeringly integrates a small-object-friendly transformer with a one-step detection paradigm. Unlike traditional methods, OSFormer processes the video sequence only through a single inference, encoding the sequence into cube format data and tracking object motion trajectories. Additionally, we propose the Varied-Size Patch Attention (VPA) module, which generates patches of varying sizes to capture adaptive attention features, bridging the gap between transformer architectures and small object detection. To further enhance detection accuracy, OSFormer incorporates a Doppler Adaptive Filter, which integrates traditional filtering techniques into an end-to-end neural network to suppress background noise and accentuate small objects. OSFormer outperforms YOLOv8-s on both the AntiUAV dataset (+ $3.1\%~\text {mAP}_{50}$ , - $35.1\%~\text {Params}$ ) and the InfraredUAV dataset (+ $4.0\%~\text {mAP}_{50-95}$ , - $51.0\%~\text {FLOPs}$ ), demonstrating superior efficiency and effectiveness in small object detection. The code is available on https://github.com/q2479036243/OSFormer.
Haolin Qin, Tingfa Xu, Fengxiang Xu, Jianan Li 0001
IEEE Trans. Image Process.2
2025 Dynamic Client Distillation for Semi-Supervised Federated Learning in a Realistic Scenario
abstract
Recent advancements in semi-supervised federated learning (SSFL) have significantly enhanced public health services by enabling medical institutions to share model updates via a central server. However, most SSFL approaches are based on conservative assumptions, such as labels-at-server and labels-at-client, which fail to fully capture the complex and diverse data distributions inherent in medical institutions. To address this limitation, we introduce a novel application of SSFL tailored to a realistic client data scenario, encompassing clients with fully-labeled, partially-labeled, and fully-unlabeled data. This approach effectively navigates varying levels of data annotation by maximizing the utility of unlabeled samples within the client federation. To tackle the challenges posed by such a complex scenario, we propose a new SSFL framework, FedCD. FedCD incorporates three client-distilled models, each corresponding to a distinct client data distribution, alongside server-client federation. First, each client-distilled model condenses the diverse parameters of the client federation into robust knowledge through distillation. The contribution of each client model is then dynamically adjusted based on its proximity to the client-distilled model, ensuring that the framework adapts to the heterogeneous characteristics of individual clients. By aggregating client-distilled models, FedCD implements model drift correction, effectively mitigating parameter drift across heterogeneous models. This dynamic federated approach not only harnesses unlabeled data efficiently but also accommodates diverse annotation levels while adapting to varying data distributions. Extensive experiments on two medical image segmentation tasks and one classification task demonstrate the superiority of our method, highlighting its ability to address realistic challenges in medical data scenarios.
Tingfa Xu, Shiqi Huang 0001, Jianan Li 0001
IEEE Trans. Medical Imaging2
2025 Mixed-Granularity Implicit Representation for Continuous Hyperspectral Compressive Reconstruction
abstract
Hyperspectral images (HSIs) are crucial across numerous fields but are hindered by the long acquisition times associated with traditional spectrometers. The coded aperture snapshot spectral imaging (CASSI) system mitigates this issue through a compression technique that accelerates the acquisition process. However, reconstructing HSIs from compressed data presents challenges due to fixed spatial and spectral resolution constraints. This study introduces a novel method using implicit neural representation (INR) for continuous HSI reconstruction. We propose the mixed-granularity implicit representation (MGIR) framework, which includes a hierarchical spectral-spatial implicit encoder (HSSIE) for efficient multiscale implicit feature extraction. This is complemented by a mixed-granularity local feature aggregator (MGLFA) that adaptively integrates local features across scales, combined with a decoder that merges coordinate information for precise reconstruction. By leveraging INRs, the MGIR framework enables reconstruction at any desired spatial-spectral resolution, significantly enhancing the flexibility and adaptability of the CASSI system. Extensive experimental evaluations confirm that our model produces reconstructed images at arbitrary resolutions and matches the state-of-the-art methods across varying spectral-spatial compression ratios (CRs). The code will be released at https://github.com/chh11/MGIR.
Jianan Li 0001, Huan Chen 0018, Wangcai Zhao, Tingfa Xu
IEEE Trans. Neural Networks Learn. Syst.5
2025 Factorization Vision Transformer: Modeling Long-Range Dependency With Local Window Cost
abstract
Transformers have astounding representational power but typically consume considerable computation which is quadratic with image resolution. The prevailing Swin transformer reduces computational costs through a local window strategy. However, this strategy inevitably causes two drawbacks: 1) the local window-based self-attention (WSA) hinders global dependency modeling capability and 2) recent studies point out that local windows impair robustness. To overcome these challenges, we pursue a preferable trade-off between computational cost and performance. Accordingly, we propose a novel factorization self-attention (FaSA) mechanism that enjoys both the advantages of local window cost and long-range dependency modeling capability. By factorizing the conventional attention matrix into sparse subattention matrices, FaSA captures long-range dependencies, while aggregating mixed-grained information at a computational cost equivalent to the local WSA. Leveraging FaSA, we present the factorization vision transformer (FaViT) with a hierarchical structure. FaViT achieves high performance and robustness, with linear computational complexity concerning input image spatial resolution. Extensive experiments have shown FaViT's advanced performance in classification and downstream tasks. Furthermore, it also exhibits strong model robustness to corrupted and biased data and hence demonstrates benefits in favor of practical applications. In comparison to the baseline model Swin-T, our FaViT-B2 significantly improves classification accuracy by 1% and robustness by 7%, while reducing model parameters by 14%. Our code will soon be publicly available: at https://github.com/q2479036243/FaViT.
Haolin Qin, Daquan Zhou, Tingfa Xu, Ziyang Bian, Jianan Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 PAPooling: Graph-based Position Adaptive Aggregation of Local Geometry in Point Clouds
abstract
Fine-grained geometry, obtained through the assimilation of localized point features, is crucial in the realms of object recognition and scene comprehension within point cloud contexts. Traditional point cloud backbones predominantly utilize max pooling for the amalgamation of local features, a process that tends to overlook spatial interrelations among points, consequently leading to the potential loss of fine-grained geometric details. To overcome this limitation, we introduce an innovative operation termed Position Adaptive Pooling (PAPooling), which is designed to amalgamate local features while sensitively considering the spatial positions of points. This is achieved by employing a graph-based representation to explicitly model the spatial relationships of points. PAPooling involves two principal components: first, the local graph construction , which establishes a local graph for a set of points by linking a central point with its adjacent points, thereby transforming pairwise relative positions into channel-specific attention weights; second, the attentive feature aggregation , which adeptly takes into account the contribution of each node and simulates the inter-node relationships within the local graph, effectively extracting representations of local features through a Graph Convolution Network (GCN). PAPooling’s simplicity and efficacy make it a versatile addition to widely used point-based backbones such as PointNet++ and DGCNN, offering a plug-and-play solution. Comprehensive experimental analysis demonstrates PAPooling’s enhanced capability in capturing local geometry, contributing significantly across a spectrum of applications including 3D shape classification, part segmentation, scene segmentation, and corruption defense, all with minimal computational increase. Code will be public at https://github.com/Roywangj/PAPooling/ .
Jie Wang 0097, Tingfa Xu, Liqiang Song, Lihe Ding, Peng Jiang 0013, Yuqi Han, Jianan Li 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Dual-Stage Hyperspectral Image Classification Model with Spectral Supertoken
Peifu Liu, Tingfa Xu, Jie Wang 0097, Huan Chen 0018, Huiyan Bai, Jianan Li 0001
ECCV (34)2
2024 One Registration is Worth Two Segmentations
Shiqi Huang 0001, Tingfa Xu, Ziyi Shen, Shaheer U. Saeed, Wen Yan 0005, Dean C. Barratt, Yipeng Hu
MICCAI (12)2
2024 Multi-scale Change-Aware Transformer for Remote Sensing Image Change Detection
abstract
Change detection identifies differences between images captured at different times. Real-world change detection faces challenges posed by the diverse and intricate nature of change areas, while current datasets and algorithms are often limited to simpler, consistent changes, reducing their effectiveness in practical applications. Existing dual-branch methods process images independently, risking the loss of change information due to insufficient early interaction. In contrast, single-stream approaches, though improving early integration, lack efficacy in capturing complex changes. To address these limitations, we introduce a novel single-stream framework, the Multi-scale Change-Aware Transformer (MCAT), which features the Dynamic Change-Aware Attention module and the Multi-scale Change-Enhanced Aggregator. The Dynamic Change-Aware Attention module, integrating local self-attention and cross-temporal attention, conducts dynamic iteration on images differences, thereby targeting feature extraction of change areas. The Multi-scale Change-Enhanced Aggregator enables the model to adapt to various scales and complex shapes through local change enhancement and multi-scale aggregation strategies. To overcome the limitations of existing datasets regarding the scale diversity and morphological complexity of change areas, we construct the Mining Area Change Detection dataset. The dataset offers a diverse array of change areas that span multiple scales and exhibit complex shapes, providing a robust benchmark for change detection. Extensive experiments demonstrate that our model outperforms existing methods, especially for irregular and multi-scale changes. Codes and dataset are available at https://github.com/chh11/MCAT.
Huan Chen 0018, Tingfa Xu, Peifu Liu, Huiyan Bai, Jianan Li 0001
ACM Multimedia2
2024 Target-Guided Adversarial Point Cloud Transformer Towards Recognition Against Real-world Corruptions
abstract
Achieving robust 3D perception in the face of corrupted data presents an challenging hurdle within 3D vision research. Contemporary transformer-based point cloud recognition models, albeit advanced, tend to overfit to specific patterns, consequently undermining their robustness against corruption. In this work, we introduce the Target-Guided Adversarial Point Cloud Transformer, termed APCT, a novel architecture designed to augment global structure capture through an adversarial feature erasing mechanism predicated on patterns discerned at each step during training. Specifically, APCT integrates an Adversarial Significance Identifier and a Target-guided Promptor. The Adversarial Significance Identifier, is tasked with discerning token significance by integrating global contextual analysis, utilizing a structural salience index algorithm alongside an auxiliary supervisory mechanism. The Target-guided Promptor, is responsible for accentuating the propensity for token discard within the self-attention mechanism, utilizing the value derived above, consequently directing the model attention towards alternative segments in subsequent stages. By iteratively applying this strategy in multiple steps during training, the network progressively identifies and integrates an expanded array of object-associated patterns. Extensive experiments demonstrate that our method achieves state-of-the-art results on multiple corruption benchmarks.
Jie Wang 0097, Tingfa Xu, Lihe Ding, Jianan Li 0001
NeurIPS2
2024 Feature-Enhanced Convolutional Attention for Unstable Rock Detection in Aerial Images
abstract
This study introduces the Feature-Enhanced Convolutional Attention Integration (FEC-AI) module, an innovative convolutional neural network (CNN) algorithm specifically tailored for the meticulous detection of small, challenging objects in high-resolution remote sensing imagery, emphasizing unstable rock formation monitoring. FEC-AI significantly advances CNN-based feature extraction and representation through a combination of advanced techniques. These include the Cross-Layer Attention Module, which enriches feature maps with multi-scale contextual details; the Offset-aware Adjustment Module for precise spatial refinement of feature localization; and the Contextual Feature Aggregation process, which synergizes these refined features for enhanced detection efficacy. Concurrently, we introduce the RSUR-2D dataset, a comprehensive compilation of 1,557 rigorously annotated images depicting karst landscapes around Beijing, expressly designed to challenge and advance remote sensing algorithms in geological hazard analysis. Through extensive testing on the RSUR-2D and established COCO datasets, the FEC-AI module demonstrated outstanding performance in small object detection, achieving a mean Average Precision at mAP50of 66.0% and mAPsof 21.1% on the RSUR-2D dataset. The RSUR-2D dataset, a valuable resource for the research community, is publicly accessible at https://github.com/chenmu1204/czx.
Peiran Peng, Jianan Li 0001, Shuaihao Han, Tongtong Gao, Lang Hong, Tingfa Xu
IEEE Geosci. Remote. Sens. Lett.7
2024 Feature-Based Knowledge Distillation for Infrared Small Target Detection
abstract
Infrared small target detection is an extremely challenging task because of its low resolution, noise interference, and the weak thermal signal of small targets. Nevertheless, despite these difficulties, there is a growing interest in this field due to its significant application value in areas such as military, security, and unmanned aerial vehicles. In light of this, we propose a feature-based knowledge distillation method(IRKD) for infrared small target detection which can efficiently transfer detailed knowledge to students. The key idea behind IRKD is to assign varying importance to the features of teachers and students in different areas during the distillation process. Treating all features equally would negatively impact the distillation results. Therefore, We have developed a Unified Channel-Spatial Attention(UCSA) module that adaptively enhancing the crucial learning areas within the features. The experimental results show that compared to other knowledge distillation methods, our student detector achieved significant improvement in mean average precision(mAP). For example, IRKD improves ResNet50 based RetinaNet from 50.5% to 59.1% mAP, and improver ResNet50 based FCOS from 53.7% to 56.6% on Roboflow.
Jinglei Xue, Jianan Li 0001, Yuqi Han, Chenwei Deng, Tingfa Xu
IEEE Geosci. Remote. Sens. Lett.6
2024 Anti-UAV410: A Thermal Infrared Benchmark and Customized Scheme for Tracking Drones in the Wild
abstract
The perception of drones, also known as Unmanned Aerial Vehicles (UAVs), particularly in infrared videos, is crucial for effective anti-UAV tasks. However, existing datasets for UAV tracking have limitations in terms of target size and attribute distribution characteristics, which do not fully represent complex realistic scenes. To address this issue, we introduce a generalized infrared UAV tracking benchmark called Anti-UAV410. The benchmark comprises a total of 410 videos with over 438 K manually annotated bounding boxes. To tackle the challenges of UAV tracking in complex environments, we propose a novel method called Siamese drone tracker (SiamDT). SiamDT incorporates a dual-semantic feature extraction mechanism that explicitly models targets in dynamic background clutter, enabling effective tracking of small UAVs. The SiamDT method consists of three key steps: Dual-Semantic RPN Proposals (DS-RPN), Versatile R-CNN (VR-CNN), and Background Distractors Suppression. These steps are responsible for generating candidate proposals, refining prediction scores based on dual-semantic features, and enhancing the discriminative capacity of the trackers against dynamic background clutter, respectively. Extensive experiments conducted on the Anti-UAV410 dataset and three other large-scale benchmarks demonstrate the superior performance of the proposed SiamDT method compared to recent state-of-the-art trackers.
Bo Huang 0012, Jianan Li 0001, Gang Wang 0031, Jian Zhao 0006, Tingfa Xu
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 MsSVT++: Mixed-Scale Sparse Voxel Transformer With Center Voting for 3D Object Detection
abstract
Accurate 3D object detection in large-scale outdoor scenes, characterized by considerable variations in object scales, necessitates features rich in both long-range and fine-grained information. While recent detectors have utilized window-based transformers to model long-range dependencies, they tend to overlook fine-grained details. To bridge this gap, we propose MsSVT++, an innovative Mixed-scale Sparse Voxel Transformer that simultaneously captures both types of information through a divide-and-conquer approach. This approach involves explicitly dividing attention heads into multiple groups, each responsible for attending to information within a specific range. The outputs of these groups are subsequently merged to obtain final mixed-scale features. To mitigate the computational complexity associated with applying a window-based transformer in 3D voxel space, we introduce a novel Chessboard Sampling strategy and implement voxel sampling and gathering operations sparsely using a hash map. Moreover, an important challenge stems from the observation that non-empty voxels are primarily located on the surface of objects, which impedes the accurate estimation of bounding boxes. To overcome this challenge, we introduce a Center Voting module that integrates newly voted voxels enriched with mixed-scale contextual information towards the centers of the objects, thereby improving precise object localization. Extensive experiments demonstrate that our single-stage detector, built upon the foundation of MsSVT++, consistently delivers exceptional performance across diverse datasets.
Jianan Li 0001, Shaocong Dong, Lihe Ding, Tingfa Xu
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Spectral-Wise Implicit Neural Representation for Hyperspectral Image Reconstruction
abstract
Coded Aperture Snapshot Spectral Imaging (CASSI) reconstruction aims to recover the 3D spatial-spectral signal from 2D measurement. Existing methods for reconstructing Hyperspectral Image (HSI) typically involve learning mappings from a 2D compressed image to a predetermined set of discrete spectral bands. However, this approach overlooks the inherent continuity of the spectral information. In this study, we propose an innovative method called Spectral-wise Implicit Neural Representation (SINR) as a pioneering step toward addressing this limitation. SINR introduces a continuous spectral amplification process for HSI reconstruction, enabling spectral super-resolution with customizable magnification factors. To achieve this, we leverage the concept of implicit neural representation. Specifically, our approach introduces a spectral-wise attention mechanism that treats individual channels as distinct tokens, thereby capturing global spectral dependencies. Additionally, our approach incorporates two components, namely a Fourier coordinate encoder and a spectral scale factor module. The Fourier coordinate encoder enhances the SINR’s ability to emphasize high-frequency components, while the spectral scale factor module guides the SINR to adapt to the variable number of spectral channels. Notably, the SINR framework enhances the flexibility of CASSI reconstruction by accommodating an unlimited number of spectral bands in the desired output. Extensive experiments demonstrate that our SINR outperforms baseline methods. By enabling continuous reconstruction within the CASSI framework, we take the initial stride toward integrating implicit neural representation into the field.
Huan Chen 0018, Wangcai Zhao, Tingfa Xu, Guokai Shi, Shiyun Zhou, Peifu Liu, Jianan Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 BACTrack: Building Appearance Collection for Aerial Tracking
abstract
Siamese network-based trackers have shown remarkable success in aerial tracking. Most previous works, however, usually perform template matching only between the initial template and the search region and thus fail to deal with rapidly changing targets that often appear in aerial tracking. As a remedy, this work presents Building Appearance Collection Tracking (BACTrack). This simple yet effective tracking framework builds a dynamic collection of target templates online and performs efficient multi-template matching to achieve robust tracking. Specifically, BACTrack mainly comprises a Mixed-Temporal Transformer (MTT) and an appearance discriminator. The former is responsible for efficiently building relationships between the search region and multiple target templates in parallel through a mixed-temporal attention mechanism. At the same time, the appearance discriminator employs an online adaptive template-update strategy to ensure that the collected multiple templates remain reliable and diverse, allowing them to closely follow rapid changes in the target’s appearance and suppress background interference during tracking. Extensive experiments show that our BACTrack achieves top performance on four challenging aerial tracking benchmarks while maintaining an impressive speed of over 87 FPS on a single GPU. Speed tests on embedded platforms also validate our potential suitability for deployment on UAV platforms.
Xincong Liu, Tingfa Xu, Ying Wang 0064, Zhinong Yu, Xiaoying Yuan, Haolin Qin, Jianan Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Multi-Step Temporal Modeling for UAV Tracking
abstract
In the realm of unmanned aerial vehicle (UAV) tracking, Siamese-based approaches have gained traction due to their optimal balance between efficiency and precision. However, UAV scenarios often present challenges such as insufficient sampling resolution, fast motion and small objects with limited feature information. As a result, temporal context in UAV tracking tasks plays a pivotal role in target location, overshadowing the target’s precise features. In this paper, we introduce MT-Track, a streamlined and efficient multi-step temporal modeling framework designed to harness the temporal context from historical frames for enhanced UAV tracking. This temporal integration occurs in two steps: correlation map generation and correlation map refinement. Specifically, we unveil a unique temporal correlation module that dynamically assesses the interplay between the template and search region features. This module leverages temporal information to refresh the template feature, yielding a more precise correlation map. Subsequently, we propose a mutual transformer module to refine the correlation maps of historical and current frames by modeling the temporal knowledge in the tracking sequence. This method significantly trims computational demands compared to the raw transformer. The compact yet potent nature of our tracking framework ensures commendable tracking outcomes, particularly in extended tracking scenarios. Comprehensive tests across four renowned UAV benchmarks substantiate the superior efficacy of our approach, delivering real-time performance at 84.7 FPS on a single GPU. Real-world test on the NVIDIA AGX hardware platform achieves a speed exceeding 30 FPS, validating the practicality of our method.
Xiaoying Yuan, Tingfa Xu, Xincong Liu, Ying Wang 0064, Haolin Qin, Yuqiang Fang, Jianan Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Content-Driven Magnitude-Derivative Spectrum Complementary Learning for Hyperspectral Image Classification
abstract
Extracting discriminative information from complex spectral details in hyperspectral image (HSI) for HSI classification is pivotal. While current prevailing methods rely on spectral magnitude features, they could cause confusion in certain classes, resulting in misclassification and decreased accuracy. We find that the derivative spectrum proves more adept at capturing concealed information, thereby offering a distinct advantage in separating these confusion classes. Leveraging the complementarity between spectral magnitude and derivative features, we propose a content-driven spectrum complementary network (CSCN) based on magnitude-derivative dual encoder, employing these two features as combined inputs. To fully utilize their complementary information, we raise a content-adaptive pointwise fusion module (CPFM), enabling adaptive fusion of dual-encoder features in a pointwise selective manner, contingent upon feature representation. To preserve a rich source of complementary information while extracting more distinguishable features, we introduce a hybrid disparity-enhancing loss that enhances the differential expression of the features from the two branches and increases the interclass distance. As a result, our method achieves state-of-the-art results on the extensive WHU-OHS dataset and eight other benchmark datasets.
Huiyan Bai, Tingfa Xu, Huan Chen 0018, Peifu Liu, Jianan Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Edge Feature Enhancement for Fine-Grained Segmentation of Remote Sensing Images
abstract
Fine-grained segmentation of remote sensing mineral images plays a crucial role in the investigation and monitoring of mineral resource. In deep-learning methods, fine-grained segmentation and edge detection are closely related in both data construction and feature extraction. However, the open-pit mineral areas in remote sensing images are heavily affected by complex natural environmental interference, posing challenges for precise data annotation and dataset construction. In view of this, we introduce the fine-annotated remote sensing mineral image (Fine-RSMI) dataset, which includes a total of 10225 images with finely annotated edges, while also introducing challenges such as multiscale and edge irregularities. To tackle the challenge of fine-grained segmentation in irregular edges, we propose a hierarchical fusion edge feature enhancement framework. Our framework consists of an edge detail feature enhancement module (EDFEM) and an edge supervision module (ESM). EDFEM vertically cascades multiple feature fusion units to obtain high-order complementary information for refining edge features. ESM further supervises network reinforcement learning of mineral area edges using ground truth edge maps to improve edge segmentation performance. Both modules work in a plug-and-play manner, enabling effortless integration into existing segmentation networks. Our method achieves further performance improvement in many general remote sensing segmentation frameworks, reaching the best results of 74.12% mean intersection-over-union (mIoU) on Fine-RSMI dataset and 78.64% mean accuracy (mAcc) on WHDLD dataset. Fine-RSMI dataset and code will be available athttps://github.com/chenmu1204/czx.
Tingfa Xu, Yongzhuo Pan, Huan Chen 0018, Jianan Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Searching Region-Free and Template-Free Siamese Network for Tracking Drones in TIR Videos
abstract
With the growing threat of unmanned aerial vehicle (UAV) intrusions, the topic of anti-UAV tracking has received widespread attention from the community. Traditional Siamese trackers struggle with small UAV targets and are plagued by model degradation issues. To mitigate this, we propose a novel Searching Region-free and Template-free Siamese network (SiamSRT) to track UAV targets in thermal infrared (TIR) videos. The proposed tracker builds a two-stage Siamese architecture with the former providing detection of the first-frame groundtruth by using a cross-correlated region proposal network (C-C RPN) and the latter providing detection of previous-frame predictions via a similarity-learning region convolutional neural network (S-L RCNN). In both stage, global proposals are acquired by ROI alignment operation to break the limitation of searching region. Then, a spatial location consistency function is introduced to suppress background thermal distractors and a temporal memory bank is utilized to avoid template update degradation problem. Further, a single-category foreground detector (SCFD) is designed to independently predict the position of the UAV target. SCFD can re-initialize the tracker without the given target in the first frame, which can help to recover the tracking failures. Comprehensive experiments demonstrate that SiamSRT achieves the best performance compared to the most advanced algorithms in the anti-UAV tracking missions.
Bo Huang 0012, Zeyang Dou, Jianan Li 0001, Ying Wang 0064, Tingfa Xu
IEEE Trans. Geosci. Remote. Sens.7
2024 DMSSN: Distilled Mixed Spectral-Spatial Network for Hyperspectral Salient Object Detection
abstract
Hyperspectral salient object detection (HSOD) has exhibited remarkable promise across various applications, particularly in intricate scenarios where conventional RGB-based approaches fall short. Despite the considerable progress in HSOD method advancements, two critical challenges require immediate attention. Firstly, existing hyperspectral data dimension reduction techniques incur a loss of spectral information, which adversely affects detection accuracy. Secondly, previous methods insufficiently harness the inherent distinctive attributes of hyperspectral images (HSIs) during the feature extraction process. To address these challenges, we propose a novel approach termed the Distilled Mixed Spectral-Spatial Network (DMSSN), comprising a Distilled Spectral Encoding process and a Mixed Spectral-Spatial Transformer (MSST) feature extraction network. The encoding process utilizes knowledge distillation to construct a lightweight autoencoder for dimension reduction, striking a balance between robust encoding capabilities and low computational costs. The MSST extracts spectral-spatial features through multiple attention head groups, collaboratively enhancing its resistance to intricate scenarios. Moreover, we have created a large-scale HSOD dataset, HSOD-BIT, to tackle the issue of data scarcity in this field and meet the fundamental data requirements of deep network training. Extensive experiments demonstrate that our proposed DMSSN achieves state-of-the-art performance on multiple datasets. We will soon make the code and dataset publicly available on https://github.com/anonymous0519/HSOD-BIT.
Haolin Qin, Tingfa Xu, Peifu Liu, Jingxuan Xu, Jianan Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 A Lightweight Fusion Strategy With Enhanced Interlayer Feature Correlation for Small Object Detection
abstract
Detecting small objects in drone imagery is challenging due to low resolution and background blending, leading to limited feature information. Multiscale feature fusion can enhance detection by capturing information at different scales, but traditional strategies fall short. Simple concatenation or addition operations do not fully utilize multiscale fusion advantages, resulting in insufficient correlation between features. This inadequacy hinders the detection of small objects, especially in complex backgrounds and densely populated areas. To address this issue and efficiently utilize the limited computational resources, we propose a lightweight fusion strategy based on enhanced interlayer feature correlation (EFC) to replace the traditional feature fusion strategy in feature pyramid network (FPN). The semantic expressions of different layers in the feature pyramid are inconsistent. In EFC, the grouped feature focus unit (GFF) enhances the feature correlation of each layer by focusing on the contextual information of different features. The multilevel feature reconstruction module (MFR) effectively reconstructs and transforms the strength and weakness information of each layer in the pyramid to reduce redundant feature fusion and retain more information about small targets in deep networks. It is noteworthy that the proposed method is plug-and-play and can be widely applied to various base networks. Extensive experiments and comprehensive evaluations on VisDrone, unmanned aerial vehicle benchmark object detection and tracking (UAVDT), and microsoft common objects in context (COCO) demonstrate the effectiveness. Using generalized focal loss (GFL) as the baseline on the VisDrone dataset with a large number of small targets, the proposed method improves the detection mean average precision (mAP) by 1.7%, surpassing many lightweight state-of-the-art methods and significantly reducing the Params and GFLOPs at the neck end. The code will be available athttps://github.com/nuliweixiao/EFC.git.
Tingfa Xu, Yuqiang Fang, Jianan Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 PointGL: A Simple Global-Local Framework for Efficient Point Cloud Analysis
abstract
Efficient analysis of point clouds holds paramount significance in real-world 3D applications. Currently, prevailing point-based models adhere to the PointNet++ methodology, which involves embedding and abstracting point features within a sequence of spatially overlapping local point sets, resulting in noticeable computational redundancy. Drawing inspiration from the streamlined paradigm of pixel embedding followed by regional pooling in Convolutional Neural Networks (CNNs), we introduce a novel, uncomplicated yet potent architecture known as PointGL, crafted to facilitate efficient point cloud analysis. PointGL employs a hierarchical process of feature acquisition through two recursive steps. First, theGlobal Point Embeddingleverages straightforward residual Multilayer Perceptrons (MLPs) to effectuate feature embedding for each individual point. Second, the novelLocal Graph Poolingtechnique characterizes point-to-point relationships and abstracts regional representations through succinct local graphs. The harmonious fusion of one-time point embedding and parameter-free graph pooling contributes to PointGL's defining attributes of minimized model complexity and heightened efficiency. Our PointGL attains state-of-the-art accuracy on the ScanObjectNN dataset while exhibiting a runtime that is more than 5 times faster and utilizing only approximately 4% of the FLOPs and 30% of the parameters compared to the recent PointMLP model. The code for PointGL is available athttps://github.com/Roywangj/PointGL.
Jianan Li 0001, Jie Wang 0097, Tingfa Xu
IEEE Trans. Multim.3
2024 Spectrum-Driven Mixed-Frequency Network for Hyperspectral Salient Object Detection
abstract
Hyperspectral salient object detection (HSOD) aims to detect spectrally salient objects in hyperspectral images (HSIs). However, existing methods inadequately utilize spectral information by either converting HSIs into false-color images or converging neural networks with clustering. We propose a novel approach that fully leverages the spectral characteristics by extracting two distinct frequency components from the spectrum: low-frequency Spectral Saliency and high-frequency Spectral Edge. The Spectral Saliency approximates the region of salient objects, while the Spectral Edge captures edge information of salient objects. These two complementary components, crucial for HSOD, are derived by computing from the inter-layer spectral angular distance of the Gaussian pyramid and the intra-neighborhood spectral angular gradients, respectively. To effectively utilize this dual-frequency information, we introduce a novel lightweight Spectrum-driven Mixed-frequency Network (SMN). SMN incorporates two parameter-free plug-and-play operators, namely Spectral Saliency Generator and Spectral Edge Operator, to extract the Spectral Saliency and Spectral Edge components from the input HSI independently. Subsequently, the Mixed-frequency Attention module, comprised of two frequency-dependent heads, intelligently combines the embedded features of edge and saliency information, resulting in a mixed-frequency feature representation. Furthermore, a saliency-edge-aware decoder progressively scales up the mixed-frequency feature while preserving rich detail and saliency information for accurate salient object prediction. Extensive experiments conducted on the HS-SOD benchmark and our custom dataset HSOD-BIT demonstrate that our SMN outperforms state-of-the-art methods regarding HSOD performance. Code and dataset will be available athttps://github.com/laprf/SMN.
Peifu Liu, Tingfa Xu, Huan Chen 0018, Shiyun Zhou, Haolin Qin, Jianan Li 0001
IEEE Trans. Multim.2
2024 MetaSeg: Content-Aware Meta-Net for Omni-Supervised Semantic Segmentation
abstract
Noisy labels, inevitably existing in pseudo-segmentation labels generated from weak object-level annotations, severely hamper model optimization for semantic segmentation. Previous works often rely on massive handcrafted losses and carefully tuned hyperparameters to resist noise, suffering poor generalization capability and high model complexity. Inspired by recent advances in meta-learning, we argue that rather than struggling to tolerate noise hidden behind clean labels passively, a more feasible solution would be to find out the noisy regions actively, so as to simply ignore them during model optimization. With this in mind, this work presents a novel meta-learning-based semantic segmentation method, MetaSeg, that comprises a primary content-aware meta-net (CAM-Net) to serve as a noise indicator for an arbitrary segmentation model counterpart. Specifically, CAM-Net learns to generate pixel-wise weights to suppress noisy regions with incorrect pseudo-labels while highlighting clean ones by exploiting hybrid strengthened features from image content, providing straightforward and reliable guidance for optimizing the segmentation model. Moreover, to break the barrier of time-consuming training when applying meta-learning to common large segmentation models, we further present a new decoupled training strategy that optimizes different model layers in a divide-and-conquer manner. Extensive experiments on object, medical, remote sensing, and human segmentation show that our method achieves superior performance, approaching that of fully supervised settings, which paves a new promising way for omni-supervised semantic segmentation.
Shenwang Jiang, Jianan Li 0001, Ying Wang 0064, Jizhou Zhang, Bo Huang 0012, Tingfa Xu
IEEE Trans. Neural Networks Learn. Syst.7
2023 Rethinking Few-Shot Medical Segmentation: A Vector Quantization View
abstract
The existing few-shot medical segmentation networks share the same practice that the more prototypes, the better performance. This phenomenon can be theoretically interpreted in Vector Quantization (VQ) view: the more prototypes, the more clusters are separated from pixel-wise feature points distributed over the full space. However, as we further think about few-shot segmentation with this perspective, it is found that the clusterization of feature points and the adaptation to unseen tasks have not received enough attention. Motivated by the observation, we propose a learning VQ mechanism consisting of grid-format VQ (GFVQ), self-organized VQ (SOVQ) and residual oriented VQ (ROVQ). To be specific, GFVQ generates the prototype matrix by averaging square grids over the spatial extent, which uniformly quantizes the local details; SOVQ adaptively assigns the feature points to different local classes and creates a new representation space where the learnable local prototypes are updated with a global view; ROVQ introduces residual information to fine-tune the aforementioned learned local prototypes without retraining, which benefits the generalization performance for the irrelevance to the training task. We empirically show that our VQ framework yields the state-of-the-art performance over abdomen, cardiac and prostate MRI datasets and expect this work will provoke a rethink of the current few-shot medical segmentation model design. Our code will soon be publicly available.
Shiqi Huang 0001, Tingfa Xu, Feng Mu, Jianan Li 0001
CVPR2
2023 Sample-adaptive Augmentation for Point Cloud Recognition Against Real-world Corruptions
abstract
Robust 3D perception under corruption has become an essential task for the realm of 3D vision. While current data augmentation techniques usually perform random transformations on all point cloud objects in an offline way and ignore the structure of the samples, resulting in over-or-under enhancement. In this work, we propose an alternative to make sample-adaptive transformations based on the structure of the sample to cope with potential corruption via an auto-augmentation framework, named as Adapt-Point. Specially, we leverage a imitator, consisting of a Deformation Controller and a Mask Controller, respectively in charge of predicting deformation parameters and producing a per-point mask, based on the intrinsic structural information of the input point cloud, and then conduct corruption simulations on top. Then a discriminator is utilized to prevent the generation of excessive corruption that deviates from the original data distribution. In addition, a perception-guidance feedback mechanism is incorporated to guide the generation of samples with appropriate difficulty level. Furthermore, to address the paucity of real-world corrupted point cloud, we also introduce a new dataset ScanObjectNN-C, that exhibits greater similarity to actual data in real-world environments, especially when contrasted with preceding CAD datasets. Experiments show that our method achieves state-of-the-art results on multiple corruption benchmarks, including ModelNet-C, our ScanObjectNN-C, and ShapeNet-C.
Jie Wang 0097, Lihe Ding, Tingfa Xu, Shaocong Dong, Xinli Xu, Long Bai 0008, Jianan Li 0001
ICCV3
2023 Dynamic Loss for Robust Learning
abstract
Label noise and class imbalance are common challenges encountered in real-world datasets. Existing approaches for robust learning often focus on addressing either label noise or class imbalance individually, resulting in suboptimal performance when both biases are present. To bridge this gap, this work introduces a novel meta-learning-based dynamic loss that adapts the objective functions during the training process to effectively learn a classifier from long-tailed noisy data. Specifically, our dynamic loss consists of two components: a label corrector and a margin generator. The label corrector is responsible for correcting noisy labels, while the margin generator generates per-class classification margins by capturing the underlying data distribution and the learning state of the classifier. In addition, we employ a hierarchical sampling strategy that enriches a small amount of unbiased metadata with diverse and challenging samples. This enables the joint optimization of the two components in the dynamic loss through meta-learning, allowing the classifier to effectively adapt to clean and balanced test data. Extensive experiments conducted on multiple real-world and synthetic datasets with various types of data biases, including CIFAR-10/100, Animal-10N, ImageNet-LT, and Webvision, demonstrate that our method achieves state-of-the-art accuracy.
Shenwang Jiang, Jianan Li 0001, Jizhou Zhang, Ying Wang 0064, Tingfa Xu
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Expert-Guided Knowledge Distillation for Semi-Supervised Vessel Segmentation
abstract
In medical image analysis, blood vessel segmentation is of considerable clinical value for diagnosis and surgery. The predicaments of complex vascular structures obstruct the development of the field. Despite many algorithms have emerged to get off the tight corners, they rely excessively on careful annotations for tubular vessel extraction. A practical solution is to excavate the feature information distribution from unlabeled data. This work proposes a novel semi-supervised vessel segmentation framework, named EXP-Net, to navigate through finite annotations. Based on the training mechanism of the Mean Teacher model, we innovatively engage an expert network in EXP-Net to enhance knowledge distillation. The expert network comprises knowledge and connectivity enhancement modules, which are respectively in charge of modeling feature relationships from global and detailed perspectives. In particular, the knowledge enhancement module leverages the vision transformer to highlight the long-range dependencies among multi-level token components; the connectivity enhancement module maximizes the properties of topology and geometry by skeletonizing the vessel in a non-parametric manner. The key components are dedicated to the conditions of weak vessel connectivity and poor pixel contrast. Extensive evaluations show that our EXP-Net achieves state-of-the-art performance on subcutaneous vessel, retinal vessel, and coronary artery segmentations.
Tingfa Xu, Shiqi Huang 0001, Feng Mu, Jianan Li 0001
IEEE J. Biomed. Health Informatics2
2023 SCANet: A Unified Semi-Supervised Learning Framework for Vessel Segmentation
abstract
Automatic subcutaneous vessel imaging with near-infrared (NIR) optical apparatus can promote the accuracy of locating blood vessels, thus significantly contributing to clinical venipuncture research. Though deep learning models have achieved remarkable success in medical image segmentation, they still struggle in the subfield of subcutaneous vessel segmentation due to the scarcity and low-quality of annotated data. To relieve it, this work presents a novel semi-supervised learning framework, SCANet, that achieves accurate vessel segmentation through an alternate training strategy. The SCANet is composed of a multi-scale recurrent neural network that embeds coarse-to-fine features and two auxiliary branches, a consistency decoder and an adversarial learning branch, responsible for strengthening fine-grained details and eliminating differences between ground-truths and predictions, respectively. Equipped with a novel semi-supervised alternate training strategy, the three components work collaboratively, enabling SCANet to accurately segment vessel regions with only a handful of labeled data and abounding unlabeled data. Moreover, to mitigate the shortage of annotated data in this field, we provide a new subcutaneous vessel dataset, VESSEL-NIR. Extensive experiments on a wide variety of tasks, including the segmentation of subcutaneous vessels, retinal vessels, and skin lesions, well demonstrate the superiority and generality of our approach.
Tingfa Xu, Ziyang Bian, Shiqi Huang 0001, Feng Mu, Bo Huang 0012, Yuze Xiao, Jianan Li 0001
IEEE Trans. Medical Imaging2
2023 Learning Context Restrained Correlation Tracking Filters via Adversarial Negative Instance Generation
abstract
The tracking performance of discriminative correlation filters (DCFs) is often subject to unwanted boundary effects. Many attempts have already been made to address the above issue by enlarging searching regions over the last years. However, introducing excessive background information makes the discriminative filter prone to learn from the surrounding context rather than the target. In this article, we propose a novel context restrained correlation tracking filter (CRCTF) that can effectively suppress background interference via incorporating high-quality adversarial generative negative instances. Concretely, we first construct an adversarial context generation network to simulate the central target area with surrounding background information at the initial frame. Then, we suggest a coarse background estimation network to accelerate the background generation in subsequent frames. By introducing a suppression convolution term, we utilize generative background patches to reformulate the original ridge regression objective through circulant property of correlation and a cropping operator. Finally, our tracking filter is efficiently solved by the alternating direction method of multipliers (ADMM). CRCTF demonstrates the accuracy performance on par with several well-established and highly optimized baselines on multiple challenging tracking datasets, verifying the effectiveness of our proposed approach.
Bo Huang 0012, Tingfa Xu, Jianan Li 0001, Qingwang Qin
IEEE Trans. Neural Networks Learn. Syst.2
2022 Delving into Sample Loss Curve to Embrace Noisy and Imbalanced Data
abstract
Corrupted labels and class imbalance are commonly encountered in practically collected training data, which easily leads to over-fitting of deep neural networks (DNNs). Existing approaches alleviate these issues by adopting a sample re-weighting strategy, which is to re-weight sample by designing weighting function. However, it is only applicable for training data containing only either one type of data biases. In practice, however, biased samples with corrupted labels and of tailed classes commonly co-exist in training data. How to handle them simultaneously is a key but under-explored problem. In this paper, we find that these two types of biased samples, though have similar transient loss, have distinguishable trend and characteristics in loss curves, which could provide valuable priors for sample weight assignment. Motivated by this, we delve into the loss curves and propose a novel probe-and-allocate training strategy: In the probing stage, we train the network on the whole biased training data without intervention, and record the loss curve of each sample as an additional attribute; In the allocating stage, we feed the resulting attribute to a newly designed curve-perception network, named CurveNet, to learn to identify the bias type of each sample and assign proper weights through meta-learning adaptively. The training speed of meta learning also blocks its application. To solve it, we propose a method named skip layer meta optimization (SLMO) to accelerate training speed by skipping the bottom layers. Extensive synthetic and real experiments well validate the proposed method, which achieves state-of-the-art performance on multiple challenging benchmarks.
Shenwang Jiang, Jianan Li 0001, Ying Wang 0064, Bo Huang 0012, Tingfa Xu
AAAI6
2022 FH-Net: A Fast Hierarchical Network for Scene Flow Estimation on Real-World Point Clouds
Lihe Ding, Shaocong Dong, Tingfa Xu, Xinli Xu, Jie Wang 0097, Jianan Li 0001
ECCV (39)3
2022 MsSVT: Mixed-scale Sparse Voxel Transformer for 3D Object Detection on Point Clouds
abstract
3D object detection from the LiDAR point cloud is fundamental to autonomous driving. Large-scale outdoor scenes usually feature significant variance in instance scales, thus requiring features rich in long-range and fine-grained information to support accurate detection. Recent detectors leverage the power of window-based transformers to model long-range dependencies but tend to blur out fine-grained details. To mitigate this gap, we present a novel Mixed-scale Sparse Voxel Transformer, named MsSVT, which can well capture both types of information simultaneously by the divide-and-conquer philosophy. Specifically, MsSVT explicitly divides attention heads into multiple groups, each in charge of attending to information within a particular range. All groups' output is merged to obtain the final mixed-scale features. Moreover, we provide a novel chessboard sampling strategy to reduce the computational complexity of applying a window-based transformer in 3D voxel space. To improve efficiency, we also implement the voxel sampling and gathering operations sparsely with a hash map. Endowed by the powerful capability and high efficiency of modeling mixed-scale information, our single-stage detector built on top of MsSVT surprisingly outperforms state-of-the-art two-stage detectors on Waymo. Our project page: https://github.com/dscdyc/MsSVT.
Shaocong Dong, Lihe Ding, Tingfa Xu, Xinli Xu, Jie Wang 0097, Ziyang Bian, Ying Wang 0064, Jianan Li 0001
NeurIPS4
2022 ARTracker: Compute a More Accurate and Robust Correlation Filter for UAV Tracking
abstract
Unmanned aerial vehicle (UAV) tracking focus on tracking moving targets from flying platforms, where the target undergoes a lot of aspect ratio changes,i.e., viewpoint change, rotation. Discriminative correlation filter (DCF) based method shows a promising solution to UAV tracking due to its high computational efficiency. DCF based trackers apply a fixed-bandwidth Gaussian function label for model training and incremental update to adapt to the dramatic appearance changes in target during tracking. However, due to the poor regression ability of correlation filter, DCF trackers is unable to describe the target state when the aspect ratio changes, thus hurting the tracking performance. To alleviate this, we propose a novelty correlation filter, which constructs a Gaussian-like function label for correlation filter training. The label fully considers the aspect ratio distribution of the target, which facilitates the training of a more robust tracker. Furthermore, an accurate incremental update is proposed to mitigate model degradation by combining target samples with adaptive aspect ratios. Extensive experiments are conducted on three popular UAV benchmarks,i.e., VisDrone2018-test-dev, UAV20L and DTB70. Results well demonstrate the superiority of the proposed method over both DCF and deep based trackers. Code will be released soon.
Tingfa Xu, Bo Huang 0012, Ying Wang 0064, Jianan Li 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Pixel-Adaptive Field-of-View for Remote Sensing Image Segmentation
abstract
Mineral segmentation of satellite imagery is crucial to mining surveying and monitoring. Conventional deep segmentation networks extract features at every position with a fixed field-of-view. Nevertheless, the rich content in large mineral scenes, which causes dramatically different local characteristics across regions, may require features with spatially varying field-of-view to achieve accurate segmentation. In light of this, we propose a novel Pixel-Adaptive Field-of-View (PA-FoV) module to adjust the field-of-view of a given feature map in a pixel-wise manner. Specifically, it refines the features at each position by a weighted aggregation of the outputs from atrous convolutions with different dilation rates adaptively depending on the position-specific content. The module works in a plug-and-play manner and can be flexibly inserted into any arbitrary backbone network or segmentation head, to boost feature representation and in turn improve the result of segmentation. Moreover, in order to mitigate the scarcity of labeled data, we further establish a benchmark remote sensing mineral dataset, dubbed RSMI, to facilitate research in this field. Extensive experiments show a simple addition of our PA-FoV module provides solid improvements on top of strong baselines, achieving state-of-the-art performance.
Feng Mu, Jianan Li 0001, Shiqi Huang 0001, Yongzhuo Pan, Tingfa Xu
IEEE Geosci. Remote. Sens. Lett.6
2022 MLP-Based Efficient Stitching Method for UAV Images
abstract
Unmanned aerial vehicle (UAV) image stitching techniques based on position and attitude information have shown clear speed superiority over feature-based counterparts. However, how to improve stitching accuracy and robustness remains a great challenge since position and attitude parameters are sensitive to noise introduced by sensors and external environment. To mitigate this issue, this work presents a simple yet effective stitching algorithm for UAV images based on a coarse-to-fine strategy. Specifically, we first conduct coarse registration using the position and attitude information obtained from GPS, IMU, and altimeter. Then, we introduce a novel offline calibration phase that is designed to regress the obtained global transformation matrix to the optimal one computed from feature-based algorithms, by using multi-layer perceptron (MLP) neural networks for fast correction. Consequently, the proposed method well integrates the complementary strengths of both parameter and feature-based methods, achieving an ideal speed–accuracy tradeoff. Moreover, to facilitate research on this topic, we establish a new dataset, named UAV-AIRPAI, that comprises over 100 UAV image pairs with position and attitude annotations to the community, opening up a promising direction for UAV image stitching. Extensive experiments on the UAV-AIRPAI dataset show that our method achieves superior accuracy compared to priors while running at a real-time speed of 0.0124 s per image pair. Code and data will be available athttps://github.com/dededust/UAV-AIRPAI.
Moxuan Ren, Jianan Li 0001, Liqiang Song, Tingfa Xu
IEEE Geosci. Remote. Sens. Lett.5
2022 Four-dimensional compressed spectropolarimetric imaging
abstract
Polarized hyperspectral images contain rich information representing both surface texture and spectral signature of target scene. However, it is difficult to obtain the full-Stokes parameters of hyperspectral images directly. A compressive imaging system is proposed in this paper to recover the four-dimensional polarized hyperspectral images with full-Stokes parameters using quarter-wave plate (QWP) and liquid crystal tunable filter (LCTF). Changing the fast axis angle of QWP provides the degrees of freedom to modulate the polarization states. The LCTF serves as the broad-band spectral filter to modulate the spectral signatures. The output of LCTF is modulated by a coded aperture in spatial domain, and then the modulated polarized hyperspectral images are projected and multiplexed on a two-dimensional detector. Based on the compressive sensing theory, the full-Stokes polarized hyperspectral images can be reconstructed from several compressive measurements by solving the convex optimization problems with sparsity prior. The feasibility of proposed system is verified by simulations and experiments. Compared to the traditional compressive spectropolarimetric imaging methods, the proposed method is beneficial to reduce the compression rate, and thus shorten the data acquisition time. This work also paves a new way for the full-Stokes polarized hyperspectral imaging system with simple and compact structure.
Axin Fan, Tingfa Xu, Jianan Li 0001, Xi Wang 0035, Yuhan Zhang 0008, Chang Xu 0018
Signal Process.2
2022 Auto-Perceiving Correlation Filter for UAV Tracking
abstract
Discriminative correlation filter (DCF)-based methods have demonstrated superior performance in UAV tracking via fusing multiple types of features and updating models online. However, most DCF-based trackers simply cascade different features, failing to fully take advantage of their complementary strength. In addition, online update strategies are limited to using a single and fixed learning rate, which often leads to model degradation when suffering tracking challenges. In this paper, we present an Auto-Perceiving Correlation Filter (APCF) which explicitly models the target and context with a novel Target State and Background Perception (TSBP) feature. Concretely, we first propose a simple yet effective State Evaluation Metric (SEM) to estimate target states by analyzing the spatial distribution of responses. Based on SEM, we extract TSBP features by adaptively selecting effective features depending on the current target state. Accordingly, a new online model update strategy is also introduced to avoid model degradation. Moreover, we further introduce a perception regularization term to make the extracted feature emphasis more on the target rather than background. Extensive experiments on four widely-used UAV benchmarks have well demonstrated the superiority of the proposed method compared with both DCF and deep learning based trackers while running at a high speed of 76.7 FPS on a single CPU. In addition, APCF with deep features also performs favorably against state-of-the-art trackers.
Jianan Li 0001, Bo Huang 0012, Xiangmin Li, Jihui Wang, Tingfa Xu
IEEE Trans. Circuits Syst. Video Technol.7
2022 SiamATL: Online Update of Siamese Tracking Network via Attentional Transfer Learning
abstract
Visual object tracking with semantic deep features has recently attracted much attention in computer vision. Especially, Siamese trackers, which aim to learn a decision making-based similarity evaluation, are widely utilized in the tracking community. However, the online updating of the Siamese fashion is still a tricky issue due to the limitation, which is a tradeoff between model adaption and degradation. To address such an issue, in this article, we propose a novel attentional transfer learning-based Siamese network (SiamATL), which fully exploits the previous knowledge to inspire the current tracker learning in the decision-making module. First, we explicitly model the template and surroundings by using an attentional online update strategy to avoid template pollution. Then, we introduce an instance-transfer discriminative correlation filter (ITDCF) to enhance the distinguishing ability of the tracker. Finally, we suggest a mutual compensation mechanism that integrates cross-correlation matching and ITDCF detection into the decision-making subnetwork to achieve online tracking. Comprehensive experiments demonstrate that our approach outperforms state-of-the-art tracking algorithms on multiple large-scale tracking datasets.
Bo Huang 0012, Tingfa Xu, Ziyi Shen, Shenwang Jiang, Bingqing Zhao, Ziyang Bian
IEEE Trans. Cybern.2
2022 RTNet: Relation Transformer Network for Diabetic Retinopathy Multi-Lesion Segmentation
abstract
Automatic diabetic retinopathy (DR) lesions segmentation makes great sense of assisting ophthalmologists in diagnosis. Although many researches have been conducted on this task, most prior works paid too much attention to the designs of networks instead of considering the pathological association for lesions. Through investigating the pathogenic causes of DR lesions in advance, we found that certain lesions are closed to specific vessels and present relative patterns to each other. Motivated by the observation, we propose a relation transformer block (RTB) to incorporate attention mechanisms at two main levels: a self-attention transformer exploits global dependencies among lesion features, while a cross-attention transformer allows interactions between lesion and vessel features by integrating valuable vascular information to alleviate ambiguity in lesion detection caused by complex fundus structures. In addition, to capture the small lesion patterns first, we propose a global transformer block (GTB) which preserves detailed information in deep network. By integrating the above blocks of dual-branches, our network segments the four kinds of lesions simultaneously. Comprehensive experiments on IDRiD and DDR datasets well demonstrate the superiority of our approach, which achieves competitive performance compared to state-of-the-arts.
Shiqi Huang 0001, Jianan Li 0001, Yuze Xiao, Tingfa Xu
IEEE Trans. Medical Imaging5
2021 Pyramid Correlation based Deep Hough Voting for Visual Object Tracking
abstract
Most of the existing Siamese-based trackers treat tracking problem as a parallel task of classification and regression. However, some studies show that the sibling head structure could lead to suboptimal solutions during the network training. Through experiments we find that, without regression, the performance could be equally promising as long as we delicately design the network to suit the training objective. We introduce a novel voting-based classification-only tracking algorithm named Pyramid Correlation based Deep Hough Voting (short for PCDHV), to jointly locate the top-left and bottom-right corners of the target. Specifically we innovatively construct a Pyramid Correlation module to equip the embedded feature with fine-grained local structures and global spatial contexts; The elaborately designed Deep Hough Voting module further take over, integrating long-range dependencies of pixels to perceive corners; In addition, the prevalent discretization gap is simply yet effectively alleviated by increasing the spatial resolution of the feature maps while exploiting channel-space relationships. The algorithm is general, robust and simple. We demonstrate the effectiveness of the module through a series of ablation experiments. Without bells and whistles, our tracker achieves better or comparable performance to the SOTA algorithms on three challenging benchmarks (TrackingNet, GOT-10k and LaSOT) while running at a real-time speed of 80 FPS. Codes and models will be released.
Ying Wang 0064, Tingfa Xu, Shenwang Jiang, Jianan Li 0001
ACML2
2021 Adaptive Gaussian-Like Response Correlation Filter for UAV Tracking
Tingfa Xu, Jianan Li 0001, Ying Wang 0064, Xiangmin Li
ICIG (3)2
2021 Foreground-aware Siamese tracker with dynamic template in wireless sensor networks
Tingfa Xu, Bo Huang 0012, Jihui Wang, Xiangmin Li
Ad Hoc Networks2
2021 BSCF: Learning background suppressed correlation filter tracker for wireless multimedia sensor networks
Bo Huang 0012, Tingfa Xu, Ziyi Shen, Shenwang Jiang, Jianan Li 0001
Ad Hoc Networks2
2021 Salient object detection on hyperspectral images in wireless network using CNN and saliency optimization
Tingfa Xu, Chenguang Pan, Jianhua Hao, Xiangmin Li
Ad Hoc Networks2
2021 Multiplex Fourier ptychographic reconstruction with model-based neural network for Internet of Things
Jizhou Zhang, Tingfa Xu, Yiwen Chen 0002, Shushan Wang
Ad Hoc Networks2
2021 Deep Siamese Cross-Residual Learning for Robust Visual Tracking
abstract
The sixth-generation (6G) wireless technology contributes to the establishment of the Internet of Things (IoT). Recently, the IoT has become popular because of its smart architectures and various applications. Among these applications, intelligent urban surveillance systems for smart cities are becoming more and more important. Therefore, designing a robust visual tracking method has become an urgent task. Deep Siamese convolutional neural networks have been applied to visual tracking recently because of their advantageous abilities to learn a matching function between the template and the target candidate. Unlike traditional Siamese networks, which separately treat the two branches, we propose deep Siamese cross-residual learning to entangle the two branches from the beginning to the end of the Siamese network. This strategy can make the two branches exchange instance-specific information at different nodes of the network and learn a more compact representation of the target. In addition, we propose a combined loss function, which consists of two complementary tasks. One task is to learn a matching function directly and the other one is to learn a classification function. Moreover, our model does not need to load any pretrained weights and is trained with limited sequences from scratch. Plenty of experiments show that our tracker performs favorably against many state-of-the-art tracking methods.
Tingfa Xu, Jie Guo 0004, Bo Huang 0012, Chang Xu 0018, Jihui Wang, Xiangmin Li
IEEE Internet Things J.2
2021 LayoutGAN: Synthesizing Graphic Layouts With Vector-Wireframe Adversarial Networks
abstract
Layout is important for graphic design and scene generation. We propose a novel Generative Adversarial Network, called LayoutGAN, that synthesizes layouts by modeling geometric relations of different types of 2D elements. The generator of LayoutGAN takes as input a set of randomly-placed 2D graphic elements, represented by vectors and uses self-attention modules to refine their labels and geometric parameters jointly to produce a realistic layout. Accurate alignment is critical for good layouts. We, thus, propose a novel differentiable wireframe rendering layer that maps the generated layout to a wireframe image, upon which a CNN-based discriminator is used to optimize the layouts in image space. We validate the effectiveness of LayoutGAN in various experiments including MNIST digit generation, document layout generation, clipart abstract scene generation, tangram graphic design, mobile app layout design, and webpage layout optimization from hand-drawn sketches.
Jianan Li 0001, Jimei Yang, Aaron Hertzmann, Jianming Zhang 0001, Tingfa Xu
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 Attribute-Conditioned Layout GAN for Automatic Graphic Design
abstract
Modeling layout is an important first step for graphic design. Recently, methods for generating graphic layouts have progressed, particularly with Generative Adversarial Networks (GANs). However, the problem of specifying the locations and sizes of design elements usually involves constraints with respect to element attributes, such as area, aspect ratio and reading-order. Automating attribute conditional graphic layouts remains a complex and unsolved problem. In this article, we introduce Attribute-conditioned Layout GAN to incorporate the attributes of design elements for graphic layout generation by forcing both the generator and the discriminator to meet attribute conditions. Due to the complexity of graphic designs, we further propose an element dropout method to make the discriminator look at partial lists of elements and learn their local patterns. In addition, we introduce various loss designs following different design principles for layout optimization. We demonstrate that the proposed method can synthesize graphic layouts conditioned on different element attributes. It can also adjust well-designed layouts to new sizes while retaining elements' original reading-orders. The effectiveness of our method is validated through a user study.
Jianan Li 0001, Jimei Yang, Jianming Zhang 0001, Christina Wang, Tingfa Xu
IEEE Trans. Vis. Comput. Graph.6
2020 Exploiting Semantics for Face Image Deblurring
Ziyi Shen, Wei-Sheng Lai, Tingfa Xu, Jan Kautz, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.3
2020 Context constraint and pattern memory for long-term correlation tracking
Bo Huang 0012, Tingfa Xu, Bo Liu 0075, Bo Yuan 0001
Neurocomputing2
2020 Transfer learning-based discriminative correlation filter for visual tracking
abstract
Most Correlation Filter (CF)-based tracking methods can hardly handle occlusion or severe deformation, due to the lack of effective utilization of previous target information. To overcome this, we propose a novel Transfer Learning-based Discriminative Correlation Filter (TLDCF), which extracts knowledge from multiple previous tracking tasks and applies the knowledge for a new tracking task through Instance-Transfer Learning (ITL) and Probability-Transfer Learning (PTL). ITL applies knowledge of Gaussian Mixture Modelling (GMM) target representations and multi-channel filters learned in previous frames to directly train a new correlation filter. This improves the robustness of tracker for heavy occlusion and large appearance variations. Meanwhile, PTL encodes the spatio-temporal relationship predicted by Kalman Filter (KF) into a shared Gaussian prior to suppress huge location drift caused by similar targets. For optimization, we develop an efficient Alternating Direction Method of Multipliers (ADMM) based algorithm to calculate CFs on each independent channel in real time. Extensive experiments on OTB-2013 and OTB-2015 datasets well demonstrate the effectiveness of the proposed method. In particular, our method improves AUC score of the two datasets by 5.5% and 3.9% respectively compared to baseline, and achieves competitive performance against recent state-of-the-art deep trackers.
Bo Huang 0012, Tingfa Xu, Jianan Li 0001, Ziyi Shen, Yiwen Chen 0002
Pattern Recognit.2
2020 Robust Visual Tracking via Constrained Multi-Kernel Correlation Filters
abstract
Discriminative Correlation Filter (DCF) based trackers are quite efficient in tracking objects by exploiting the circulant structure. The kernel trick further improves the performance of such trackers. The unwanted boundary effects, however, are difficult to solve in the kernelized correlation models. In this paper, we propose a novel Constrained Multi-Kernel Correlation tracking Filter (CMKCF), which applies spatial constraints to address this drawback. We build the multi-kernel models for multi-channel features with three different attributes, and then employ a spatial cropping operator on the semi-kernel matrix to address the boundary effects. For the constrained optimization solution, we develop an Alternating Direction Method of Multipliers (ADMM) based algorithm to learn our multi-kernel filters efficiently in the frequency domain. In particular, we suggest an adaptive updating mechanism by exploiting the feedback from high-confidence tracking results to avoid corruption in the model. Extensive experimental results demonstrate that the proposed method performs favorably on OTB-2013, OTB-2015, VOT-2016 and VOT-2018 dataset against several state-of-the-art methods.
Bo Huang 0012, Tingfa Xu, Shenwang Jiang, Yiwen Chen 0002, Yu Bai 0009
IEEE Trans. Multim.2
2019 Human-Aware Motion Deblurring
abstract
This paper proposes a human-aware deblurring model that disentangles the motion blur between foreground (FG) humans and background (BG). The proposed model is based on a triple-branch encoder-decoder architecture. The first two branches are learned for sharpening FG humans and BG details, respectively; while the third one produces global, harmonious results by comprehensively fusing multi-scale deblurring information from the two domains. The proposed model is further endowed with a supervised, human-aware attention mechanism in an end-to-end fashion. It learns a soft mask that encodes FG human information and explicitly drives the FG/BG decoder-branches to focus on their specific domains. Above designs lead to a fully differentiable motion deblurring network, which can be trained end-to-end. To further benefit the research towards Human-aware Image Deblurring, we introduce a large-scale dataset, named HIDE, which consists of 8,422 blurry and sharp image pairs with 65,784 densely annotated FG human bounding boxes. HIDE is specifically built to span a broad range of scenes, human object sizes, motion patterns, and background complexities. Extensive experiments on public benchmarks and our dataset demonstrate that our model performs favorably against the state-of-the-art motion deblurring methods, especially in capturing semantic details.
Ziyi Shen, Wenguan Wang, Xiankai Lu, Jianbing Shen, Haibin Ling, Tingfa Xu, Ling Shao 0001
ICCV6
2019 LayoutGAN: Generating Graphic Layouts with Wireframe Discriminators
Jianan Li 0001, Jimei Yang, Aaron Hertzmann, Jianming Zhang 0001, Tingfa Xu
ICLR (Poster)5
2018 Deep Semantic Face Deblurring
abstract
In this paper, we present an effective and efficient face deblurring algorithm by exploiting semantic cues via deep convolutional neural networks (CNNs). As face images are highly structured and share several key semantic components (e.g., eyes and mouths), the semantic information of a face provides a strong prior for restoration. As such, we propose to incorporate global semantic priors as input and impose local structure losses to regularize the output within a multi-scale deep CNN. We train the network with perceptual and adversarial losses to generate photo-realistic results and develop an incremental training strategy to handle random blur kernels in the wild. Quantitative and qualitative evaluations demonstrate that the proposed face deblurring algorithm restores sharp images with more facial details and performs favorably against state-of-the-art methods in terms of restoration quality, face recognition and execution speed.
Ziyi Shen, Wei-Sheng Lai, Tingfa Xu, Jan Kautz, Ming-Hsuan Yang 0001
CVPR3
2018 Generating Reliable Online Adaptive Templates for Visual Tracking
abstract
Online adaption of visual tracking is a significant strategy to achieve good tracking performance since the appearance of the object target varies all along with the sequence. However, directly using the tracking results of previous frames to update the model will cause drifting, resulting in tracking failure. We propose a task-guided generative adversarial network (GAN), named TGGAN, to learn the general appearance distribution that a target may undergo through a sequence. Then the online adaption is simply to select templates from the images that are generated from the ground truth template in the first frame and a set of random vectors by the generator. This strategy helps the model alleviate drifting while still obtaining adaptivity. Tracking is treated as a template matching problem under a proposed Siamese matching network structure. Experiments show the effectiveness of the proposed online adaption strategy and the Siamese matching network.
Jie Guo 0004, Tingfa Xu, Shenwang Jiang, Ziyi Shen
ICIP2
2018 Non-uniform motion deblurring with Kernel grid regularization
Ziyi Shen, Tingfa Xu, Jinshan Pan, Jie Guo 0004
Signal Process. Image Commun.2
2018 Multistage Object Detection With Group Recursive Learning
abstract
Most existing detection pipelines treat object proposals independently and predict bounding box locations and classification scores over them separately. However, the important semantic and spatial layout correlations among proposals are often ignored, which are actually useful for more accurate object detection. In this paper, we propose a new EM-like group recursive learning approach to iteratively refine object proposals by incorporating such context of surrounding proposals and provide an optimal spatial configuration of object detections. In addition, we propose to incorporate the weakly supervised object segmentation cues and region-based object detection into a multistage architecture in order to fully exploit the learned segmentation features for better object detection in an end-toend way. The proposed architecture consists of three cascaded networks that, respectively, learn to perform weakly supervised object segmentation, object proposal generation, and recursive detection refinement. Combining the group recursive learning and the multistage architecture provides competitive mAPs of 78.7% and 74.9% on the PASCAL VOC2007 and VOC2012 datasets, respectively, which outperform many well-established baselines significantly.
Jianan Li 0001, Xiaodan Liang, Jianshu Li, Yunchao Wei, Tingfa Xu, Jiashi Feng, Shuicheng Yan
IEEE Trans. Multim.5
2018 Scale-Aware Fast R-CNN for Pedestrian Detection
abstract
In this paper, we consider the problem of pedestrian detection in natural scenes. Intuitively, instances of pedestrians with different spatial scales may exhibit dramatically different features. Thus, large variance in instance scales, which results in undesirable large intracategory variance in features, may severely hurt the performance of modern object instance detection methods. We argue that this issue can be substantially alleviated by the divide-and-conquer philosophy. Taking pedestrian detection as an example, we illustrate how we can leverage this philosophy to develop a Scale-Aware Fast R-CNN (SAF R-CNN) framework. The model introduces multiple built-in subnetworks which detect pedestrians with scales from disjoint ranges. Outputs from all of the subnetworks are then adaptively combined to generate the final detection results that are shown to be robust to large variance in instance scales, via a gate function defined over the sizes of object proposals. Extensive evaluations on several challenging pedestrian detection datasets well demonstrate the effectiveness of the proposed SAF R-CNN. Particularly, our method achieves state-of-the-art performance on Caltech, and obtains competitive results on INRIA, ETH, and KITTI.
Jianan Li 0001, Xiaodan Liang, Shengmei Shen, Tingfa Xu, Jiashi Feng, Shuicheng Yan
IEEE Trans. Multim.4
2017 Perceptual Generative Adversarial Networks for Small Object Detection
abstract
Detecting small objects is notoriously challenging due to their low resolution and noisy representation. Existing object detection pipelines usually detect small objects through learning representations of all the objects at multiple scales. However, the performance gain of such ad hoc architectures is usually limited to pay off the computational cost. In this work, we address the small object detection problem by developing a single architecture that internally lifts representations of small objects to super-resolved ones, achieving similar characteristics as large objects and thus more discriminative for detection. For this purpose, we propose a new Perceptual Generative Adversarial Network (Perceptual GAN) model that improves small object detection through narrowing representation difference of small objects from the large ones. Specifically, its generator learns to transfer perceived poor representations of the small objects to super-resolved ones that are similar enough to real large objects to fool a competing discriminator. Meanwhile its discriminator competes with the generator to identify the generated representation and imposes an additional perceptual requirement - generated representations of small objects must be beneficial for detection purpose - on the generator. Extensive evaluations on the challenging Tsinghua-Tencent 100K [45] and the Caltech [9] benchmark well demonstrate the superiority of Perceptual GAN in detecting small objects, including traffic signs and pedestrians, over well-established state-of-the-arts.
Jianan Li 0001, Xiaodan Liang, Yunchao Wei, Tingfa Xu, Jiashi Feng, Shuicheng Yan
CVPR4
2017 Deep Attribute-preserving Metric Learning for Natural Language Object Retrieval
abstract
Retrieving image content with a natural language expression is an emerging interdisciplinary problem at the intersection of multimedia, natural language processing and artificial intelligence. Existing methods tackle this challenging problem by learning features from the visual and linguistic domains independently while the critical semantic correlations bridging two domains have been under-explored in the feature learning process. In this paper, we propose to exploit sharable semantic attributes as "anchors" to ensure the learned features are well aligned across domains for better object retrieval. We define "attributes" as the common concepts that are informative for object retrieval and can be easily learned from both visual content and language expression. In particular, diverse and complex attributes (e.g., location, color, category, interaction between object and context) are modeled and incorporated to promote cross-domain alignment for feature learning from multiple perspectives. Based on the sharable attributes, we propose a deep Attribute-Preserving Metric learning (AP-Metric) framework that jointly generates unique query-sensitive region proposals and conducts novel cross-modal feature learning that explicitly pursues consistency over semantic attribute abstraction within both domains for deep metric learning. Benefiting from the cross-modal semantic correlations, our proposed framework can localize challenging visual objects to match complex query expressions within cluttered background accurately. The overall framework is end-to-end trainable. Extensive evaluations on popular datasets including ReferItGame, RefCOCO, and RefCOCO+ well demonstrate its superiority. Notably, it achieves state-of-the-art performance on the challenging ReferItGame dataset.
Jianan Li 0001, Yunchao Wei, Xiaodan Liang, Fang Zhao 0006, Jianshu Li, Tingfa Xu, Jiashi Feng
ACM Multimedia6
2017 Deep Ensemble Tracking
abstract
In this letter, we cast visual tracking as a template matching problem in a Siamese deep convolutional neural network architecture. In contrast to traditional or other deep feature-based tracking methods, the proposed model exploits multilevel convolutional features from a partial view. The model matches candidate patch and template patch from the feature dimension of convolutional features, leading to hundreds of thousands of base matchers. The base matchers from low-level convolutional features have small receptive fields which contain partial details of targets while the base matchers from high-level convolutional features have big receptive fields which capture semantic information of targets. The model achieves the final strong matcher as a weighted ensemble of all the base matchers. We design an effective weights propagation strategy to update the weights of base matchers. Moreover, we propose to use Cosine as the distance metric and a customized squared-loss function as cost function for robust. Experiments show that our tracker outperforms the state-of-the-art trackers in a wide range of tracking scenarios.
Jie Guo 0004, Tingfa Xu
IEEE Signal Process. Lett.2
2017 Visual Tracking Via Sparse Representation With Reliable Structure Constraint
abstract
In this letter, we present a novel visual tracking algorithm based on sparse representation. In contrast to just use the target templates and the trivial templates to sparsely represent the target, we propose to further constrain the model with a set of discriminative weight maps. These weight maps contain the reliable structures of the target object. They help the model penalize the trivial template coefficients depending on the reliable structures of the target object. Then, the target object can be well represented by a sparse set of target templates together with a sparse set of target weight maps. We propose a unified objective function to integrate these two sparse representation problems together. This optimization problem can be well solved by the proposed iteration manner and a customized accelerated proximal gradient method. Furthermore, a novel weight map constructing method is proposed based on consistent motion property and forward-backward errors. Plenty of qualitative and quantitative evaluations demonstrate that our method performs favorably against the state-of-the-art methods in a wide range of tracking scenarios.
Jie Guo 0004, Tingfa Xu, Ziyi Shen, Guokai Shi
IEEE Signal Process. Lett.2
2017 Real-Time Feature-Based Video Stabilization on FPGA
abstract
Digital video stabilization is an important video enhancement technology that aims to remove unwanted camera vibrations from video sequences. Trading off between stabilization performance and real-time hardware implementation feasibility, this paper presents a feature-based full-frame video stabilization method and a novel complete fully pipelined architectural design to implement it on field-programmable gate array (FPGA). In the proposed method, feature points are first extracted with the oriented features from accelerated segment test and rotated binary robust independent elementary features algorithm and matched between consecutive frames. Next, the matched point pairs are fitted to the affine transformation model using a random-sample consensus-based approach to estimate inter-frame motion robustly. Then, the estimated results are accumulated to compute the cumulative motion parameters between the current and reference frames, and the translational components are smoothed by a Kalman filter representing intentional camera movement. Finally, a mosaicked image is constructed based on cumulative motion parameters using an image mosaicking technique, and then a display window is created with the desired frame size according to the computed intentional camera movement to obtain a full motion-compensated frame. Using pipelining and parallel processing strategies, the whole process has been designed using a novel complete fully pipelined architecture and implemented on Altera's Cyclone III FPGA to build a real-time stabilization system. The experimental results have shown that the proposed system can deal with standard PAL video input including arbitrate translation and rotation and can produce full-frame stabilized output providing a better viewing experience at 22.37 ms/frame, thus achieving real-time processing performance.
Jianan Li 0001, Tingfa Xu
IEEE Trans. Circuits Syst. Video Technol.2
2017 Attentive Contexts for Object Detection
abstract
Modern deep neural network-based object detection methods typically classify candidate proposals using their interior features. However, global and local surrounding contexts that are believed to be valuable for object detection are not fully exploited by existing methods yet. In this work, we take a step towards understanding what is a robust practice to extract and utilize contextual information to facilitate object detection in practice. Specifically, we consider the following two questions: “how to identify useful global contextual information for detecting a certain object?” and “how to exploit local context surrounding a proposal for better inferring its contents?” We provide preliminary answers to these questions through developing a novel attention to context convolution neural network (AC-CNN)-based object detection model. AC-CNN effectively incorporates global and local contextual information into the region-based CNN (e.g., fast R-CNN and faster R-CNN) detection framework and provides better object detection performance. It consists of one attention-based global contextualized (AGC) subnetwork and one multi-scale local contextualized (MLC) subnetwork. To capture global context, the AGC subnetwork recurrently generates an attention map for an input image to highlight useful global contextual locations, through multiple stacked long short-term memory layers. For capturing surrounding local context, the MLC subnetwork exploits both the inside and outside contextual information of each specific proposal at multiple scales. The global and local context are then fused together for making the final decision for detection. Extensive experiments on PASCAL VOC 2007 and VOC 2012 well demonstrate the superiority of the proposed AC-CNN over well-established baselines.
Jianan Li 0001, Yunchao Wei, Xiaodan Liang, Jian Dong 0011, Tingfa Xu, Jiashi Feng, Shuicheng Yan
IEEE Trans. Multim.5
2006 Research on Algorithms of Gabor Wavelet Neural Network Based on Parallel Structure
Tingfa Xu, Zefeng Nie, Guoqiang Ni
PRIMA1