EDBT 2026 Demo / reviewers in the wild / expert
Chenglong Li 0002
dblp:83/7820-2
· DBLP profile ↗
165ranked-venue papers
24as first author
124since 2021 · last 2026
0000-0002-7233-2739ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 105 · 17 first-author · 74 since 2021Artificial intelligence and machine learning · 56 · 12 first-author · 42 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 19 since 2021Security and privacy · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-Teacher Interactive Knowledge Distillation Network for Text-to-Visible & Infrared Person RetrievalabstractText-to-visible & infrared person retrieval aims to retrieve the corresponding visible (RGB) and thermal infrared (TIR) images given the text descriptions. Existing methods perform semantic decoupling by aligning RGB and TIR features separately to different attributes, thereby facilitating the alignment between the fused multimodal representation and the text. However, insufficient TIR representation ability and cross-view representation capabilities of RGB and TIR modalities limit the retrieval accuracy and robustness. To address these issues, we propose a novel Dual-teacher Interactive Knowledge Distillation Network called DIKDNet, which performs the interactive knowledge distillation between two modality-specific teachers with rich cross-view representation capabilities to enhance TIR representations and the collaborative knowledge distillation from both teachers to the corresponding students to enhance the cross-modal cross-view representations, for robust text-to-visible & infrared person retrieval. Specifically, to enhance the representation ability of the TIR backbone network while preserving modality-specific characteristics, we design an Interactive Knowledge Distillation Module (IKDM), which introduces a boundary-constrained distillation strategy between RGB and TIR backbones, to transfer the semantic features of RGB backbone to TIR one. To enhance the cross-modal cross-view representation capability, we design a Collaborative Knowledge Distillation Module (CKDM) to transfer the cross-modal similarity relations and the cross-view multimodal representations from teacher networks to student ones. Experimental results demonstrate that our method consistently achieves significant performance gains on both the RGBT-PEDES and RGBNT201-PEDES datasets. The code will be released upon the acceptance. Chenglong Li 0002, Yifei Deng, Aihua Zheng |
AAAI | 1 |
| 2026 | Unaligned UAV RGBT Tracking: A Largescale Benchmark and a Novel ApproachabstractWith the rapid development of the low-altitude economy, multimodal visual tracking in UAV scenarios has attracted extensive attention. UAVs are typically equipped with independent visible (RGB) and thermal infrared (TIR) sensors, resulting in an inherent spatial misalignment between the two modalities. However, existing RGBT tracking methods generally rely on spatially aligned data inputs, making them unsuitable for unaligned RGBT tracking task in UAV scenarios. In this work, we introduce the new task called unaligned UAV RGBT tracking and construct the first large-scale unaligned RGB and TIR video dataset to promote the research and development of this field. The dataset contains 1,453 pairs of UAV-captured RGBT sequences with precise dual-modal bounding box annotations, and covers 42 object categories, 22 typical challenge attributes, and diverse spatial misalignment scales to better simulate real-world challenging scenarios. To address the limitations of existing methods that fail to handle the spatial misalignment issue in UAV scenarios, we propose the novel RGBT tracking approach. In particular, we design a mixture of shift estimation experts module to adaptively estimate the spatial shifts across two modalities at different scales, and a cross-modal alignment and fusion module to further compensate for nonlinear deformations and integrate multimodal information. Extensive experiments on the created dataset demonstrate that the proposed tracker significantly outperforms existing state-of-the-art tracking methods, validating its practicality and robustness in real-world unaligned UAV tracking scenarios. Yun Xiao 0003, Jiandong Jin, Wankang Zhang, Chenglong Li 0002 |
AAAI | 5 |
| 2026 | ProxyTTT: Proxy-driven Test-Time Training for Multi-modal Re-identificationabstractMulti-modal object re-identification (ReID) aims to retrieve specific targets by leveraging complementary cues from different sensing modalities. Despite recent progress, two key challenges remain: (1) the limited ability to jointly address both modality and viewpoint discrepancies, and (2) the difficulty of effectively leveraging reliable target-domain data to improve generalization. To address these challenges, we propose Proxy-driven Test-Time Training (ProxyTTT), a unified framework that enhances both multi-modal identity representation learning and model generalization. During training, we propose a Multi-Proxy Learning (MPL) mechanism to address the representation bias across different views and modalities. MPL disentangles fine-grained modality-specific and modality-common identity proxies as semantic anchors to align identity features across diverse perspectives and sensing modalities. This alignment strategy enables the model to learn robust and discriminative global identity representations under heterogeneous modality conditions. At test time, to reliably exploit target domain data, we propose Proxy-guided Entropy-based Selective Adaptation (PESA) for test-time training. Specifically, PESA leverages the semantic structure encoded by identity proxies to estimate prediction uncertainty via entropy, and selectively adapts the model using only high-confidence samples. This selective adaptation effectively mitigates the domain shift between training and deployment environments, improving the model’s generalization in real-world scenarios. Extensive experiments on four public multi-modal ReID benchmarks (RGBNT201, RGBNT100, MSVR310, and WMVeID863) demonstrate the effectiveness of ProxyTTT. Aihua Zheng, Zhaojun Liu, Xixi Wan, Chenglong Li 0002, Jin Tang 0001, Yan Yan 0002 |
AAAI | 4 |
| 2026 | Morphology-aware hierarchical mixture of experts for Chest X-ray anatomy segmentation
Lili Huang 0006, Yuanjun He, Chenglong Li 0002, Jin Tang 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Uncertainty-Aware RGBT Tracking
Zhaodong Ding, Chenglong Li 0002, Futian Wang, Jin Tang 0001 |
Int. J. Comput. Vis. | 2 |
| 2026 | ImageBind Guided Progressive Transformation Network for Alignment-free RGBT Video Object Detection
Zhengzheng Tu, Chuanwang Guo, Qishun Wang, Chenglong Li 0002, Jin Tang 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | Medical report generation via knowledge distillation and medical keywords
Lili Huang 0006, Chenglong Li 0002, Jin Tang 0001 |
Neurocomputing | 4 |
| 2026 | SequencePAR: Understanding pedestrian attributes via a sequence generation paradigm
Jiandong Jin, Xiao Wang 0014, Yin Lin, Chenglong Li 0002, Lili Huang 0006, Aihua Zheng, Jin Tang 0001 |
Pattern Recognit. | 4 |
| 2026 | RGBT tracking via supervised mutual guiding
Lei Liu 0049, Chenglong Li 0002, Jin Tang 0001, Changhe Li |
Pattern Recognit. | 2 |
| 2026 | Temporal multimodal knowledge distillation for modality-missing RGBT tracking
Rui Ruan, Yunlong Kang, Lei Liu 0049, Jingpeng Sun, Chenglong Li 0002, Jin Tang 0001 |
Pattern Recognit. | 5 |
| 2026 | Structure and progress aware diffusion for medical image segmentation
Siyuan Song, Guyue Hu 0001, Chenglong Li 0002, Dengdi Sun, Zhe Jin 0001, Jin Tang 0001 |
Pattern Recognit. | 3 |
| 2026 | Bidirectional intervention attention network for audio-visual matching
Jiaxiang Wang 0001, Aihua Zheng, Dequan Li, Chenglong Li 0002, Wenjuan Cheng, Ran He 0001 |
Pattern Recognit. | 4 |
| 2026 | Medical image segmentation via Attention-enhanced Mamba with learnable Symmetry scan and a benchmark
Chenglong Li 0002, Jin Tang 0001, Chuanfu Li |
Pattern Recognit. | 2 |
| 2026 | Fine-Grained and Granularity-Dynamic Framework for Referring Remote Sensing Image Segmentation
Duzhi Yuan, Guyue Hu 0001, Aihua Zheng, Chenglong Li 0002, Jin Tang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2026 | Revisiting Frequency Domain: Spatial-Frequency Joint Tuning for Referring Image SegmentationabstractReferring image segmentation aims at segmenting the target object referred by a natural language expression, which requires semantic-level object understanding and pixel-level contour segmentation. Existing methods are limited to the spatial domain, thus ignoring potential discriminability from the frequency domain and missing mutual boost between the spatial and frequency domains, also facing heavy high-frequency degeneration issues. In this paper, we revisit frequency domain and propose a novel lightweight spatial-frequency joint tuning (SFJT) plugin for referring image segmentation. We decompose 4 mainstream methodology paradigms into two general stages consisting of roughly semantic-level object understanding and precisely pixel-level contour segmentation, then respectively enhance them with parameter-efficient spatial-frequency tuning strategies. Specifically, we develop a spatial-frequency joint prompting technique (SFJ-Prompt) during early object understanding stage, which mines bidirectional spatial-frequency information to realize mutually spatial-frequency boosting, facilitating more comprehensive object understanding. Besides, we introduce a LoRA-based high-frequency auxiliary branch (HF-LoRA) during latter contour segmentation stage, which compensates for heavy high-frequency degeneration issues in spatial neural networks, facilitating more precise contour segmentation. Eventually, extensive experiments of 4 mainstream methodology paradigms for referring image segmentation on 4 large-scale datasets demonstrate the effectiveness and superiority of the proposed method. Guyue Hu 0001, Yuxing Tong, Dong Geng, Chenglong Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Vehicle-Centric Perception via Multimodal Structured Pre-TrainingabstractVehicle-centric perception plays a crucial role in many intelligent systems, including large-scale surveillance systems, intelligent transportation, and autonomous driving. Existing approaches typically employ general pre-trained weights to initialize backbone networks, followed by task-specific fine-tuning. However, these models lack effective learning of vehiclerelated knowledge during pre-training, resulting in poor capability for modeling general vehicle perception representations. To handle this problem, we propose VehicleMAE-V2, a novel vehicle-centric pre-trained large model. By exploring and exploiting vehicle-related multimodal structured priors to guide the masked token reconstruction process, our approach can significantly enhance the model’s capability to learn generalizable representations for vehicle-centric perception. Specifically, we design the Symmetry-guided Mask Module (SMM), Contour-guided Representation Module (CRM) and Semantics-guided Representation Module (SRM) to incorporate three kinds of structured priors into token reconstruction including symmetry, contour and semantics of vehicles respectively. SMM utilizes the vehicle symmetry constraints to avoid retaining symmetric patches and can thus select high-quality masked image patches and reduce information redundancy. CRM minimizes the prob23 ability distribution divergence between contour features and reconstructed features and can thus preserve holistic vehicle structure information during pixel-level reconstruction. SRM aligns image-text features through contrastive learning and cross-modal distillation to address the feature confusion caused by insufficient semantic understanding during masked reconstruction. To support the pre-training of VehicleMAE-V2, we construct Autobot4M, a large-scale dataset comprising approximately 4 million vehicle images and 12,693 text descriptions. Extensive experiments on five downstream tasks demonstrate the superior performance of VehicleMAE-V2. The source code, dataset, and pre-trained large models are available on https://github.com/Vehicle-AHU/VehicleMAE. Xiao Wang 0014, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Cross-Modal Person Retrieval With One-to-Many Relation ModelingabstractExisting text-image person retrieval methods are built upon fixed-point or distributional embeddings, but they typically perform single-point alignment across modalities, making it difficult to effectively capture the one-to-many cross-modal semantic associations. To address this, we propose a One-to-Many Relation modEling network (OMRE) that explicitly constructs one-to-many semantic matching structures across modalities, thereby modeling richer and more diverse semantic associations. Specifically, to achieve one-to-many matching modeling, we design a bidirectional one-to-many alignment module, which constructs cross-modal matching distributions by aggregating relations between the mean embedding and multiple sampled embeddings, and minimizes their discrepancy with the true distribution to capture complex semantic associations. To construct fine-grained one-to-many matching relationships, we propose a collaborative reconstruction-based similarity refinement module, which maximizes the semantic consistency between multiple reconstructed masked tokens and the original tokens, effectively achieving robust and precise one-to-many cross-modal fine-grained semantic alignment. Moreover, to enhance the discriminative capability of one-to-many semantic distributions, we introduce a Hard Negative Mining mechanism that focuses on semantically similar but mismatched samples, helping to refine distribution boundaries in the probabilistic space and suppress interference from hard negative samples. Extensive experiments on three public datasets demonstrate that our method not only achieves superior overall performance but also exhibits excellent generalization ability. The code will be released on https://github.com/Yifei-AHU/OMRE. Yifei Deng, Chenglong Li 0002, Guyue Hu 0001, Jin Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2026 | Adversarial Semantic and Label Perturbation Attack for Pedestrian Attribute RecognitionabstractPedestrian Attribute Recognition (PAR) is an indispensable task in human-centered research and has made great progress in recent years with the development of deep neural networks. However, the potential vulnerability and anti-interference ability have still not been fully explored. To bridge this gap, this paper proposes the first adversarial attack and defense framework for pedestrian attribute recognition. Specifically, we exploit both global- and patch-level attacks on the pedestrian images, based on the pre-trained CLIP-based PAR framework. It first divides the input pedestrian image into non-overlapping patches and embeds them into feature embeddings using a projection layer. Meanwhile, the attribute set is expanded into sentences using prompts and embedded into attribute features using a pre-trained CLIP text encoder. A multi-modal Transformer is adopted to fuse the obtained vision and text tokens, and a feed-forward network is utilized for attribute recognition. Based on the aforementioned PAR framework, we adopt the adversarial semantic and label-perturbation to generate the adversarial noise, termed ASL-PAR. We also design a semantic offset defense strategy to suppress the influence of adversarial attacks. Extensive experiments conducted on both digital domains (i.e., PETA, PA100K, MSP60K, RAPv2) and physical domains fully validated the effectiveness of our proposed adversarial attack and defense strategies for the pedestrian attribute recognition. The source code of this paper will be released on https://github.com/Event-AHU/OpenPAR. Weizhe Kong, Xiao Wang 0014, Ruichong Gao, Chenglong Li 0002, Yu Zhang 0091, Xing Yang 0004, Yaowei Wang 0001, Jin Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | Causality-Based Modality- and Platform-Invariant Representation Learning for Dynamic RGBT Tracking and a BenchmarkabstractEach sequence in existing RGBT tracking datasets is typically captured from a single platform equipped with both RGB (visible light) and TIR (thermal infrared) sensors. In real-world applications, tracking some objects requires cross-platform collaboration and these platforms might be equipped with different sensors. However, changes in modalities and platforms may cause significant variations in target appearance and abrupt position shifts, which existing RGBT trackers struggle to handle. To address these challenges, we define a new task, termed dynamic RGBT tracking, focusing on cross-platform and modality-variant scenarios. Considering the dynamic changes of modalities and platforms, we investigate dynamic RGBT tracking from a causal perspective, and assume that images consist of causal factors (target-relevant information) and non-causal factors (target-irrelevant information, i.e., modality/platform information), where only the former is conducive to stable tracking. Based on this assumption, we propose a novel causality-based modality&platform-invariant representation learning approach to capture robust invariant representations for dynamic RGBT tracking. In particular, to mitigate the challenges posed by modality variations, we design a causal consistency encoder that introduces an intervener to model feature uncertainty and simulate modal variations, compelling the model to focus on modality-invariant features to improve tracking robustness. To overcome the issue of abrupt view change and position shift, we design a platform-independent global searcher to re-localize the target whenever a platform switch occurs, which leverages an intervener to simulate the interference of platform changes on features, encouraging the searcher to learn platform-invariant representations for improved localization accuracy. In addition, to promote the research and development of dynamic RGBT tracking, we construct a dataset named DRGBT603, which consists of 603 sequences with a total of 1.49 M frame pairs. Extensive experiments on DRGBT603 dataset validate the effectiveness of the proposed method against other state-of-the-art methods. Our code and data are now available: https://github.com/dongdong2061/DRGBT. Zhaodong Ding, Chenglong Li 0002, Shengqing Miao, Jin Tang 0001 |
IEEE Trans. Image Process. | 2 |
| 2026 | Text-Visible/Infrared Person Retrieval: Attribute-Guided Feature Decoupling and Collaborative Alignment and a Unified BenchmarkabstractExisting research on text-to-image person retrieval primarily focuses on visible images, which are not suitable under low-light scenarios. Infrared imaging becomes necessary in many visual systems, and matching text with both visible and infrared images is required. However, visible and infrared images are heterogeneous with different visual characteristics, so matching text with them in a unified framework is very challenging. In this work, we design a new task called Text-Visible/Infrared person retrieval and contribute a novel approach and a unified benchmark to promote the research and development of this field. On one hand, we propose a novel Attribute-guided feature decoupling and Collaborative Alignment Network (ACANet) that pursues accurate alignment from the text modality to both visible and infrared modalities in a unified framework according to the texture and color attribute information of text descriptions. In particular, we decouple the color features of visible images supervised by the text labels and integrate them into the infrared features to eliminate the impact of the absence of color information in infrared images during cross-modal collaborative alignment. Moreover, we also decouple the texture information from visible images supervised by the text labels and perform the collaborative alignment of texture and infrared features with a fusion agent. In addition, we extend conventional masked language modeling to a cross-modal paradigm to help ACANet learn uniform fine-grained alignment in multiple image modalities. On the other hand, we contribute a unified high-quality MM01LLCM-Text dataset, which provides person images in both visible and infrared modalities paired with fine-grained text descriptions. Experimental results show that the proposed ACANet outperforms existing state-of-the-art methods on MM01LLCM-Text dataset. Chenglong Li 0002, Yifei Deng, Aihua Zheng, Jin Tang 0001 |
IEEE Trans. Image Process. | 1 |
| 2026 | Unveiling the Power of Multi-Modal Template Update in RGBT TrackingabstractTemplate update is essential for improving the adaptability of tracking algorithms to target appearance variations. While previous methods have leveraged the spatio-temporal complementarity of multi-modal templates for RGBT tracking, a comprehensive analysis of the template update mechanism remains underexplored. In this work, we propose a novel prototype-based framework that decomposes the multi-modal template update process from the perspective of prototype learning into four key components: multi-modal prototype, prototype integration, prototype evaluation, and prototype update algorithm. Our findings highlight that the multi-modal prototype is the most critical factor in enhancing tracking adaptability to appearance variations, leading to more robust target representations. While prototype integration is less crucial when the target representation is already robust, it still contributes to learning a more discriminative representation. Additionally, the accuracy of template updates is strongly influenced by prototype evaluation, which controls the accuracy of the update process. Finally, the prototype update algorithm, which determines when and how template updates occur, is key to maintaining tracking robustness. Building on these insights, we introduce the Multi-modal Prototype RGBT Tracker (MPTrack), which adapts dynamically to appearance variations through prototype learning. MPTrack combines a fixed template from the first frame with both modality-shared and modality-specific templates, forming a robust multi-modal prototype representation. It incorporates a prototype evaluation module that guides updates based on template reliability, and an adaptive update algorithm to manage templates effectively. Additionally, a prototype-guided cross-modal integration module enhances the discriminative power of multi-modal relation modeling. Experimental results on five challenging RGBT tracking benchmarks demonstrate that MPTrack consistently outperforms state-of-the-art methods, setting new performance records. The experimental data and source code will be made publicly available at: https://github.com/mmic-lcl/Datasets-and-benchmark-code. Lei Liu 0049, Chenglong Li 0002, Andong Lu, Yabin Zhu, Shoufei Han, Xinye Cai, Changhe Li |
IEEE Trans. Image Process. | 2 |
| 2026 | Pixel-Level RGBT Fusion Tracking via Heterogeneous Multi-Expert Distillation and Decoupled Representation LearningabstractPixel-level fusion is widely considered a lightweight yet limited strategy in RGB-Thermal (RGBT) tracking due to its shallow representational capacity. However, its actual limitations and potential remain largely unexplored. We systematically analyze fusion location, modality alignment, and tracking performance, revealing that despite lower modality gaps than feature-level fusion, pixel-level fusion lacks task-relevant discrimination, restricting its effectiveness. In this paper, we propose the Task-driven Pixel-level Fusion tracker (TPF), which preserves the efficiency of early fusion while enhancing discriminative capacity. Central to TPF is a lightweight pixel fusion adapter that ensures real-time image fusion with only 14.3KB extra parameters over the baseline at inference. To enhance its limited representational capacity, we propose a task-driven progressive learning framework consisting of two key stages. First, a heterogeneous multi-expert distillation scheme adaptively transfers image fusion knowledge from diverse models under tracking-guided evaluation, mitigating the generalization limitations of single-teacher distillation across varied tracking scenarios. Second, to overcome limited task discrimination caused by sparse, target-focused tracking supervision, we propose a decoupled representation learning strategy that offers dense, complementary guidance to improve target-background separation and fusion quality. A nearest-neighbor dynamic template update further enhances robustness to appearance changes. Extensive experiments on four RGBT tracking benchmarks show that TPF achieves competitive accuracy and speed, outperforming both feature-level and existing pixel-level fusion methods, offering new insights into efficient RGBT tracking. Andong Lu, Yuanzhi Guo, Kunpeng Wang 0005, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object DetectionabstractExisting multimodal UAV object detection methods often overlook the impact of semantic gaps between modalities, which makes it difficult to achieve accurate semantic and spatial alignments and ultimately limits detection performance. To address this problem, we propose a Large Language Model (LLM) guided Progressive feature Alignment Network called LPANet, which leverages the semantic features extracted from a large language model to guide the progressive semantic and spatial alignment between modalities for multimodal UAV object detection. To employ the powerful semantic representation of LLM, we generate the fine-grained text descriptions of each object category by ChatGPT and then extract the semantic features using the large language model MPNet, providing high-level semantic priors to guide multimodal alignment. Based on the semantic features, we guide the semantic and spatial alignments in a progressive manner as follows. First, we design the Semantic Alignment Module (SAM) to pull the semantic features and multimodal visual features of each object closer, alleviating the semantic differences of objects between modalities. Second, we design the Explicit Spatial Alignment Module (ESM) by integrating the semantic relations into the estimation of feature-level offsets, alleviating the coarse spatial misalignment between modalities. Finally, we design the Implicit Spatial alignment Module (ISM), which leverages the cross-modal correlations to aggregate key features from neighboring regions to achieve implicit spatial alignment. Comprehensive experiments on two public multimodal UAV object detection datasets demonstrate that our approach outperforms state-of-the-art multimodal UAV object detectors. The source code will be released on https://github.com/Vehicle-AHU/LPANet. Chenglong Li 0002, Xiao Wang 0014, Bin Luo 0001 |
IEEE Trans. Image Process. | 2 |
| 2026 | REMIND: Retrieval-Augmented Reconstruction With Dual Memories for Modality-Missing Object Re-IdentificationabstractTo address the modality-missing object Re-Identification (Re-ID) task, a common strategy is to compensate for absent information by exploiting available modalities. However, existing reconstruction-based approaches suffer from two major limitations: 1) they often overlook modality-specific cues inherent in the missing modality; 2) they typically adopt a single-path reconstruction strategy. These issues result in incomplete representations and constrain the capacity to model complex semantic mappings across heterogeneous modalities. To address these challenges, we propose REMIND, a novel framework for modality-missing object Re-Identification, namely REtrieval-AugMented ReconstructIoN With Dual Memories. Specifically, we design a Dual Memory Construction module that, guided by information-theoretic insights, extracts modality-specific and modality-common features through two complementary branches and stores them in dedicated memory banks. These memory banks serve as structured prior knowledge to guide the reconstruction process, ensuring that the features of missing modalities are preserved even under modality-missing conditions. In addition, we have developed a retrieval-augmented missing reconstruction module that enhances the expressiveness and robustness of the reconstruction through multi-path reconstruction and perturbation mechanisms. Adaptive fusion techniques are employed for integration, simultaneously improving the expressiveness and robustness of the reconstructed features. Through the synergy of information-theoretically motivated regularization and retrieval-enhanced reconstruction, REMIND achieves robust feature recovery and delivers highly discriminative representations for reliable modality-missing Re-ID. Extensive experiments on several multi-modal object Re-ID benchmarks demonstrate the effectiveness and superiority of REMIND under various missing modality scenarios. The code is publicly available at: https://github.com/skye-1201/REMIND. Zhendong Xu, Zi Wang 0013, Aihua Zheng, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Attribute-Guided Semantic Alignment With Pre-Trained Foundation Models for Vehicle DetectionabstractVehicle detection is a fundamental perception task in intelligent transportation systems and plays a crucial role in enabling reliable traffic perception and analysis. Existing vehicle detectors are typically obtained by training conventional object detection models (e.g., YOLO, RCNN, and DETR series) on vehicle images based on pre-trained backbone networks (e.g., ResNet and ViT). Although some studies introduce large-scale foundation models to improve detection performance, these models are not specifically designed for vehicle-centric scenarios and therefore tend to yield sub-optimal results in complex traffic environments. Moreover, most existing methods heavily rely on visual features and pay limited attention to the alignment between vehicle semantic information and visual representations. In this paper, we propose a novel vehicle detection paradigm, termed VFM-Det, which integrates a pre-trained vehicle foundation model (VehicleMAE) with a large language model (T5) to achieve semantically enhanced vehicle detection for intelligent transportation scenarios. Specifically, the proposed method follows a region proposal-based detection framework and employs VehicleMAE to enhance the features of each proposal. More importantly, we introduce a novel VAtt2Vec module to predict the vehicle semantic attributes corresponding to each proposal and transform them into feature vectors, which further enhance visual features through contrastive learning. Extensive experiments on three vehicle detection benchmark datasets thoroughly proved the effectiveness of our vehicle detector. Specifically, our model improves the baseline approach by +6.0%, +8.4% on the$AP_{0.5}$,$AP_{0.75}$metrics, respectively, on the Cityscapes dataset. The source code of this work will be released athttps://github.com/Event-AHU/VFM-Det Fanghua Hong, Xiao Wang 0014, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2026 | ICPL-ReID: Identity-Conditional Prompt Learning for Multi-Spectral Object Re-Identification
Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 2 |
| 2026 | DEEP: Decoupled Semantic Prompt Learning, Guiding and Embedding for Multi-Spectral Object Re-IdentificationabstractMulti-spectral object re-identification (ReID) captures diverse object semantics to robustly recognize identity in complex environments. However, without explicit semantic guidance (e.g., attributes, masks, and keypoints), existing modal fusion-based methods struggle to comprehensively capture person or vehicle semantics across spectra. Thanks to the large-scale vision-language pre-training, CLIP effectively aligns visual concepts across different image modalities to a unified semantic prompt. In this paper, we proposeDEEP, aDEcoupled sEmanticPrompt Learning, Guiding and Embedding framework for Multi-Spectral Object ReID. Specifically, to address the challenges posed by low-quality modality noise and spectral style discrepancies, we first propose a Decoupled Semantic Prompt (DSP) strategy, which explicitly decouples the semantic alignment into spectral-style learning with spectral-shared prompts and object content learning with instance-specific inversion token. Second, to lead the model focusing on semantically faithful regions, we propose a Semantic-Guided Spectral Fusion (SGSF) module that builds a semantic interaction bridge between spectra to explore complementary semantics across modalities. Finally, to further empower the spectral representation, we propose a Spectral Semantic Embedding (SSE) module constrained by semantic-aware structural consistency to refine the fine-grained identity semantics in each spectrum. Extensive experiments on five public benchmarks, RGBNT201, Market-MM, MSVR310, WMVEID863, and RGBNT100, demonstrate the proposed method outperforms the state-of-the-art methods. The source code is released at this link:https://github.com/lsh-ahu/DEEP-ReID. Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Cross-modulated Attention Transformer for RGBT TrackingabstractExisting Transformer-based RGBT trackers achieve remarkable performance benefits by leveraging self-attention to extract uni-modal features and cross-attention to enhance multi-modal feature interaction and search-template correlation. Nevertheless, the independent search-template correlation calculations are prone to be affected by low-quality data, which might result in contradictory and ambiguous correlation weights. It not only limits the intra-modal feature representation, but also harms the robustness of cross-attention for multi-modal feature interaction and search-template correlation computation. To address these issues, we propose a novel approach called Cross-modulated Attention Transformer (CAFormer), which innovatively integrates inter-modality interaction into the search-template correlation computation within typical attention mechanism, for RGBT tracking. In particular, we first independently generate correlation maps for each modality and feed them into the designed correlation modulated enhancement module, which can modify inaccurate correlation weights by seeking the consensus between modalities. Such kind of design unifies self-attention and cross-attention schemes, which not only alleviates inaccurate attention weight computation in self-attention but also eliminates redundant computation introduced by extra cross-attention scheme. In addition, we design a collaborative token elimination strategy to further improve tracking inference efficiency and accuracy. Experiments on five public RGBT tracking benchmarks show the outstanding performance of the proposed CAFormer against state-of-the-art methods. Yun Xiao 0003, Jiacong Zhao, Andong Lu, Chenglong Li 0002, Yin Lin, Cong Liu 0006 |
AAAI | 4 |
| 2025 | RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion MambaabstractExisting RGBT tracking methods often design various interaction models to perform cross-modal fusion of each layer, but can not execute the feature interactions among all layers, which plays a critical role in robust multimodal representation, due to large computational burden. To address this issue, this paper presents a novel All-layer multimodal Interaction Network, named AINet, which performs efficient and effective feature interactions of all modalities and layers in a progressive fusion Mamba, for robust RGBT tracking. Even though modality features in different layers are known to contain different cues, it is always challenging to build multimodal interactions in each layer due to struggling in balancing interaction capabilities and efficiency. Meanwhile, considering that the feature discrepancy between RGB and thermal modalities reflects their complementary information to some extent, we design a Difference-based Fusion Mamba (DFM) to achieve enhanced fusion of different modalities with linear complexity. When interacting with features from all layers, a huge number of token sequences (3840 tokens in this work) are involved and the computational burden is thus large. To handle this problem, we design an Order-dynamic Fusion Mamba (OFM) to execute efficient and effective feature interactions of all layers by dynamically adjusting the scan order of different layers in Mamba. Extensive experiments on four public RGBT tracking datasets show that AINet achieves leading performance against existing state-of-the-art methods. We will release the code upon acceptance of the paper. Andong Lu, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
AAAI | 3 |
| 2025 | Alignment-Free RGB-T Salient Object Detection: A Large-Scale Dataset and Progressive Correlation NetworkabstractAlignment-free RGB-Thermal (RGB-T) salient object detection (SOD) aims to achieve robust performance in complex scenes by directly leveraging the complementary information from unaligned visible-thermal image pairs, without requiring manual alignment. However, the labor-intensive process of collecting and annotating image pairs limits the scale of existing benchmarks, hindering the advancement of alignment-free RGB-T SOD. In this paper, we construct a large-scale and high-diversity unaligned RGB-T SOD dataset named UVT20K, comprising 20,000 image pairs, 407 scenes, and 1256 object categories. All samples are collected from real-world scenarios with various challenges, such as low illumination, image clutter, complex salient objects, and so on. To support the exploration for further research, each sample in UVT20K is annotated with a comprehensive set of ground truths, including saliency masks, scribbles, boundaries, and challenge attributes. In addition, we propose a Progressive Correlation Network (PCNet), which models inter- and intra-modal correlations on the basis of explicit alignment to achieve accurate predictions in unaligned image pairs. Extensive experiments conducted on two unaligned three weakly aligned three aligned datasets demonstrate the effectiveness of our method. Kunpeng Wang 0005, Keke Chen, Chenglong Li 0002, Zhengzheng Tu, Bin Luo 0001 |
AAAI | 3 |
| 2025 | Efficient RGBT Tracking via Early Fusion and Hierarchical Knowledge Distillation
Jinhu Wang, Mai Wen, Zhang Zhang 0001, Liang Wang 0001, Chenglong Li 0002 |
ICIG (2) | 5 |
| 2025 | Efficient RGBT Tracking via Heterogeneous Hierarchical Knowledge DistillationabstractThe increasing demand for real-time RGBT (RGB and Thermal) tracking in applications such as video surveillance, autonomous driving, and robotic navigation underscores the need for lightweight and efficient tracking frameworks. However, existing approaches require dual-stream architectures for repeated feature extraction, incurring high costs, while complex interaction strategies further reduce efficiency, limiting the real-time performance of RGBT trackers. To address this, we propose a novel Heterogeneous Hierarchical Knowledge Distillation framework (H2KD) to enable a single-stream RGBT tracker that maintains high efficiency while delivering performance comparable to existing dual-stream trackers. In particular, H2KD takes an existing dual-stream network tracker as the teacher and builds a simple single-stream network tracker as the student by concatenating the inputs. To inherit powerful representation of the heterogeneous teacher network, we expand the channel dimensions of single-stream networks to align with the fusion features of teacher network, and employ a hierarchical distillation strategy between their backbone networks. Moreover, H2KD also introduces architecture-independent prediction-level distillation between their prediction score maps to inherit the tracking capability of teachers more directly. Extensive experiments on three major RGBT tracking benchmarks and multiple dual-stream RGBT trackers demonstrate the effectiveness and generalization of the proposed method, which achieves competitive accuracy while achieving 120.2 FPS inference speed. Dengdi Sun, Chenglong Li 0002, Andong Lu |
ICME | 3 |
| 2025 | Template-based Uncertainty Multimodal Fusion Network for RGBT TrackingabstractRGBT tracking is to localize the predefined targets in video sequences by effectively leveraging the information from both visible light (RGB) and thermal infrared (TIR) modalities. However, the quality of different modalities changes dynamically in complex scenes, and effectively perceiving modal quality for multimodal fusion remains a significant challenge. To address this challenge, we propose to employ the reliability of initial template to explore the uncertainty across different modalities, and design a novel template-based uncertainty computation framework for robust multimodal fusion in RGBT tracking. In particular, we introduce an Uncertainty-aware Multimodal Fusion Module (UMFM), which constructs the uncertainty of each modality by leveraging the correlation between the template and search region in the Subjective Logic framework, aiming to achieve robust multimodal fusion. In addition, existing methods focus on dynamic template update while overlooking the potential role of a reliable initial template in the template updating process.To this end, we design a simple yet effective Contrastive Template Update Module (CTUM) to assess the reliability of the new template by comparing its quality with that of the initial template. Extensive experiments suggest that our method outperforms existing approaches on four RGBT tracking benchmarks. Zhaodong Ding, Chenglong Li 0002, Shengqing Miao, Jin Tang 0001 |
IJCAI | 2 |
| 2025 | Frequency-Domain Multi-modal Fusion for Language-Guided Medical Image Segmentation
Zetao Du, Yan Huang 0008, Chenglong Li 0002, Liang Wang 0001 |
MICCAI (9) | 5 |
| 2025 | Learning with Explicit Topological Priors for Chest X-Ray Rib Segmentation
Chenglong Li 0002, Jin Tang 0001, Chuanfu Li |
MICCAI (16) | 2 |
| 2025 | Learning Hierarchical Cross-modal Association with Intra-modal Context for Text-Image Person RetrievalabstractExisting Text-Image Person Retrieval (TIPR) methods have made substantial progress in modeling cross-modal associations via contrastive learning frameworks, but usually ignore the fine-grained differences in semantic relevance among different samples, which limits retrieval accuracy. To address this problem, we propose a novel Hierarchical Cross-modal Association framework HCA, which leverages the intra-modal fine-grained semantic relations distilled by single-modal pretrained models to constrain hierarchical cross-modal association between image and text modalities, for accurate TIPR. Specifically, to model hierarchical cross-modal semantic relationships, we propose a Hierarchical Relevance Matching (HRM) module. It partitions the matching strength of image-text pairs by jointly considering identity labels and cross-modal similarity, collaborating with unimodal similarity to construct a hierarchical relevance distribution that serves as a soft supervision signal. HRM not only helps the model better capture varying levels of semantic consistency between image-text pairs but also enhances the overall accuracy of cross-modal association learning. To enhance the ability to capture fine-grained cross-modal semantic relationships, we introduce an Image-guided Ambiguous text Token Modeling (IATM) module. It replaces original tokens with semantically ambiguous ones and leverages image guidance to detect and correct these tokens. This process further improves the fine-grained semantic alignment between images and texts. Experimental results demonstrate that HCA achieves new state-of-the-art performance across multiple datasets, thoroughly validating its effectiveness and advancement in cross-modal retrieval tasks. Yifei Deng, Chenglong Li 0002, Futian Wang, Jin Tang 0001 |
ACM Multimedia | 2 |
| 2025 | CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training FrameworkabstractEvent cameras have attracted increasing attention in recent years due to their advantages in high dynamic range, high temporal resolution, low power consumption, and low latency. Some researchers have begun exploring pre-training directly on event data. Nevertheless, these efforts often fail to establish strong connections with RGB frames, limiting their applicability in multi-modal fusion scenarios. To address these issues, we propose a novel CM3AE pre-training framework for the RGB-Event perception. This framework accepts multi-modalities/views of data as input, including RGB images, event images, and event voxels, providing robust support for both event-based and RGB-event fusion based downstream tasks. Specifically, we design a multi-modal fusion reconstruction module that reconstructs the original image from fused multi-modal features, explicitly enhancing the model's ability to aggregate cross-modal complementary information. Additionally, we employ a multi-modal contrastive learning strategy to align cross-modal feature representations in a shared latent space, which effectively enhances the model's capability for multi-modal understanding and capturing global dependencies. We construct a large-scale dataset containing 2,535,759 RGB-Event data pairs for the pre-training. Extensive experiments on five downstream tasks fully demonstrated the effectiveness of CM3AE. Source code and pre-trained models will be released on https://github.com/Event-AHU/CM3AE. Xiao Wang 0014, Chenglong Li 0002, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Qi Liu 0003 |
ACM Multimedia | 3 |
| 2025 | UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-IdentificationabstractMulti-modal object Re-IDentification (ReID) has gained considerable attention with the goal of retrieving specific targets across cameras using heterogeneous visual data sources. At present, multi-modal object ReID faces two core challenges: (1) learning robust features under fine-grained local noise caused by occlusion, frame loss, and other disruptions; and (2) effectively integrating heterogeneous modalities to enhance multi-modal representation. To address the above challenges, we propose a robust approach named Uncertainty-Guided Graph model for multi-modal object ReID (UGG-ReID). UGG-ReID is designed to mitigate noise interference and facilitate effective multi-modal fusion by estimating both local and sample-level aleatoric uncertainty and explicitly modeling their dependencies. Specifically, we first propose the Gaussian patch-graph representation model that leverages uncertainty to quantify fine-grained local cues and capture their structural relationships. This process boosts the expressiveness of modal-specific information, ensuring that the generated embeddings are both more informative and robust. Subsequently, we design an uncertainty-guided mixture of experts strategy that dynamically routes samples to experts exhibiting low uncertainty. This strategy effectively suppresses noise-induced instability, leading to enhanced robustness. Meanwhile, we design an uncertainty-guided routing to strengthen the multi-modal interaction, improving the performance. UGG-ReID is comprehensively evaluated on five representative multi-modal object ReID datasets, encompassing diverse spectral modalities. Experimental results show that the proposed method achieves excellent performance on all datasets and is significantly better than current methods in terms of noise immunity. Our code is available at https://github.com/wanxixi11/UGG-ReID. Xixi Wan, Aihua Zheng, Bo Jiang 0002, Beibei Wang 0006, Chenglong Li 0002, Jin Tang 0001 |
NeurIPS | 5 |
| 2025 | Erasure-based interaction network for red-green-blue and thermal object detection and a unified benchmark
Qishun Wang, Zhengzheng Tu, Chenglong Li 0002, Hongshun Wang, Kunpeng Wang 0005 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Keypoint-guided feature enhancement and alignment for cross-resolution vehicle re-identification
Aihua Zheng, Zi Wang 0013, Chenglong Li 0002, Xiaofei Sheng |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Modality-missing RGBT Tracking: Invertible Prompt Learning and High-quality Benchmarks
Andong Lu, Chenglong Li 0002, Jiacong Zhao, Jin Tang 0001, Bin Luo 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | CMCNet:Cross-directional morphology-aware convolution network for chest X-ray anatomy segmentation
Lili Huang 0006, Yuhan Feng, Chenglong Li 0002, Jin Tang 0001 |
Neurocomputing | 4 |
| 2025 | Visible-thermal multiple object tracking: Large-scale video dataset and progressive fusion approach
Yabin Zhu, Qianwu Wang, Chenglong Li 0002, Jin Tang 0001, Chengjie Gu, Zhixiang Huang |
Pattern Recognit. | 3 |
| 2025 | Pedestrian Attribute Recognition via CLIP-Based Prompt Vision-Language FusionabstractExisting pedestrian attribute recognition (PAR) algorithms adopt pre-trained CNN (e.g., ResNet) as their backbone network for visual feature learning, which might obtain sub-optimal results due to the insufficient employment of the relations between pedestrian images and attribute labels. In this paper, we formulate PAR as a vision-language fusion problem and fully exploit the relations between pedestrian images and attribute labels. Specifically, the attribute phrases are first expanded into sentences, and then the pre-trained vision-language model CLIP is adopted as our backbone for feature embedding of visual images and attribute descriptions. The contrastive learning objective connects the vision and language modalities well in the CLIP-based feature space, and the Transformer layers used in CLIP can capture the long-range relations between pixels. Then, a multi-modal Transformer is adopted to fuse the dual features effectively and feed-forward network is used to predict attributes. To optimize our network efficiently, we propose the region-aware prompt tuning technique to adjust very few parameters (i.e., only the prompt vectors and classification heads) and fix both the pre-trained VL model and multi-modal Transformer. Our proposed PAR algorithm only adjusts 0.75% learnable parameters compared with the fine-tuning strategy. It also achieves new state-of-the-art performance on both standard and zero-shot settings for PAR, including RAPv1, RAPv2, WIDER, PA100K, and PETA-ZS, RAP-ZS datasets. The source code and pre-trained models will be released onhttps://github.com/Event-AHU/OpenPAR. Xiao Wang 0014, Jiandong Jin, Chenglong Li 0002, Jin Tang 0001, Cheng Zhang 0010, Wei Wang 0115 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Unified-Modal Salient Object Detection via Adaptive Prompt LearningabstractExisting single-modal and multi-modal salient object detection (SOD) methods focus on designing specific architectures tailored for their respective tasks. However, developing completely different models for different tasks leads to labor and time consumption, as well as high computational and practical deployment costs. In this paper, we attempt to address both single-modal and multi-modal SOD in a unified framework called UniSOD, which fully exploits the overlapping prior knowledge between different tasks. Nevertheless, assigning appropriate strategies to modality variable inputs is challenging. To this end, UniSOD learns modality-aware prompts with task-specific hints through adaptive prompt learning, which are seamlessly plugged into the proposed pre-trained baseline SOD model to handle corresponding tasks, while only requiring few learnable parameters compared to training the entire model from scratch. In particular, each modality-aware prompt is solely generated from a homogeneous switchable prompt generation (SPG) block, which adaptively performs structural switching based on single-modal and multi-modal inputs without manual intervention, ensuring that the framework can effectively handle diverse input cases (e.g., RGB-only, RGB-D, RGB-T) with a unified approach. Through end-to-end joint training, UniSOD achieves ovrall competitive performance on 14 benchmark datasets, demonstrating its ability to efficiently unify single-modal and multi-modal SOD tasks. Code has been available athttps://github.com/Angknpng/UniSOD Kunpeng Wang 0005, Zhengzheng Tu, Chenglong Li 0002, Zhengyi Liu, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | UAV Video Vehicle Detection: Benchmark and BaselineabstractWith the increasing application of unmanned aerial vehicles (UAVs) in intelligent transportation systems, vehicle object detection in UAV videos has received increasing attention. Precise categorization and detection for vehicles in UAVs is important in many practical applications. However, existing object detection methods, tailored for natural images, often fall short of accurately identifying vehicle objects. Additionally, high-altitude UAV imaging mainly employs horizontal bounding box annotation, frequently leading to significant obstruction and overlapping. Hence, we propose a new task called UAV video vehicle detection (VVD) to achieve precise detection and categorization of vehicles in high-altitude UAV imaging environments. To facilitate the research and development of UAV VVD, we construct the first large-scale well-annotated benchmark UAV VVD dataset, which includes 70 UAV videos captured at a 500-m altitude, with 361489 vehicle instances annotated by the oriented bounding boxes and vehicle categories. Moreover, we introduce a novel category refinement network (CRNet) approach that extracts and refines vehicle object features from the bounding box of the detection results to classify vehicle categories. This approach effectively eliminates the interference of the background and other vehicle objects in candidate boxes. Notably, the vehicle object features are projected into subspace, enabling the category refinement module (CRM) to focus more on the distinctive characteristics of the vehicle object itself through normalization operations. We conduct extensive experiments on the proposed VVD dataset. Experimental results demonstrate the superiority and effectiveness of the proposed CRNet method. The relevant code and dataset are available athttps://github.com/mmic-lcl. Yun Xiao 0003, Jinfa Wang, Zhicheng Zhao 0002, Bo Jiang 0002, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Reflectance-Guided Progressive Feature Alignment Network for All-Day UAV Object DetectionabstractObject detection using visible-infrared images has become increasingly crucial for all-day applications of unmanned aerial vehicle (UAV). However, existing multi-modal detection methods face significant challenges in low-light conditions, where degraded visible image quality exacerbates weak alignment issues and compromises feature fusion effectiveness. Although recent approaches have attempted to address these issues through cross-attention mechanisms or feature alignment strategies, they often suffer from unstable performance and limited generalization capability in challenging nighttime scenarios. To address these limitations, we propose a novel Reflectance-Guided Progressive Feature Alignment Network (RGFNet) for robust UAV object detection. Our proposed method leverages the illumination-invariant characteristic of reflectance features decomposed from visible images via Retinex theory to guide cross-modal alignment and fusion. Specifically, we design a Reflectance-Guided Collaborative Alignment Module (RCAM) that utilizes reflectance guidance to perform bidirectional feature alignment between visible and infrared modalities, effectively reducing position misalignment under varying lighting conditions. Furthermore, we introduce a Light-Aware Selective Fusion Module (LSFM) that maps multi-modal features into a shared hidden state space through selective state space mechanism, enabling efficient feature interaction while maintaining linear computational complexity. Extensive experiments on two challenging UAV detection benchmarks, DroneVehicle and DVTOD, demonstrate the superiority of our method. RGFNet achieves state-of-the-art performance with 81.4% mAP on DroneVehicle and 88.5% mAP on DVTOD. The code is available at https://github.com/uavdet/RGFNet. Zhicheng Zhao 0002, Wei Zhang 0393, Yun Xiao 0003, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Federated Client-Tailored Adapter for Medical Image SegmentationabstractMedical image segmentation in X-ray images is beneficial for computer-aided diagnosis and lesion localization. Existing methods mainly fall into a centralized learning paradigm, which is inapplicable in the practical medical scenario that only has access to distributed data islands. Federated Learning has the potential to offer a distributed solution but struggles with heavy training instability due to client-wise domain heterogeneity (including distribution diversity and class imbalance). In this paper, we propose a novel Federated Client-tailored Adapter (FCA) framework for medical image segmentation, which achieves stable and client-tailored adaptive segmentation without sharing sensitive local data. Specifically, the federated adapter stirs universal knowledge in off-the-shelf medical foundation models to stabilize the federated training process. In addition, we develop two client-tailored federated updating strategies that adaptively decompose the adapter into common and individual components, then globally and independently update the parameter groups associated with common client-invariant and individual client-specific units, respectively. They further stabilize the heterogeneous federated learning process and realize optimal client-tailored instead of sub-optimal global-compromised segmentation models. Extensive experiments on three large-scale datasets demonstrate the effectiveness and superiority of the proposed FCA framework for federated medical segmentation. Guyue Hu 0001, Siyuan Song, Yukun Kang, Zhu Yin, Gangming Zhao, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Nighttime Person Re-Identification via Collaborative Enhancement Network With Multi-Domain LearningabstractPrevalent nighttime person re-identification (ReID) methods typically combine image relighting and ReID networks in a sequential manner. However, their performance (recognition accuracy) is limited by the quality of relighting images and insufficient collaboration between image relighting and ReID tasks. To handle these problems, we propose a novel Collaborative Enhancement Network called CENet, which performs the multilevel feature interactions in a parallel framework, for nighttime person ReID. In particular, the designed parallel structure of CENet can not only avoid the impact of the quality of relighting images on ReID performance, but also allow us to mine the collaborative relations between image relighting and person ReID tasks. To this end, we integrate the multilevel feature interactions in CENet, where we first share the Transformer encoder to build the low-level feature interaction, and then perform the feature distillation that transfers the high-level features from image relighting to ReID, thereby alleviating the severe image degradation issue caused by the nighttime scenario while avoiding the impact of relighting images. In addition, the sizes of existing real-world nighttime person ReID datasets are limited, and large-scale synthetic ones exhibit substantial domain gaps with real-world data. To leverage both small-scale real-world and large-scale synthetic training data, we develop a multi-domain learning algorithm, which alternately utilizes both kinds of data to reduce the inter-domain difference in training procedure. Extensive experiments on two real nighttime datasets,Night600andRGBNT201rgb, and a synthetic nighttime ReID dataset are conducted to validate the effectiveness of CENet. We release the code and synthetic dataset at: https://github.com/Alexadlu/CENet. Andong Lu, Chenglong Li 0002, Tianrui Zha, Xiaofeng Wang 0009, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Prototype-Based Diversity and Integrity Learning for All-Day Multi-Modal Person Re-IdentificationabstractRecent multi-modal person re-identification methods have improved model performance by leveraging complementary information from multiple spectra. However, existing methods cannot ensure feature stability under varying illumination and rely on inflexible paired data, remaining inadequate against real-world cross-time retrieval and modality-missing challenges. To solve these, we first propose diversity representation that augments illumination-sensitive images to simulate diverse lighting conditions via illumination augmentation and enriches instance features using modality-specific prototypes via multiple interaction modules. Secondly, we propose integrity reconstruction that leverages prototypes and available instance features to recover information, the reconstruction module effectively utilizes identity and modality cues to address unpredictable missing problems. In addition, we build a more comprehensive dataset (AllDay843) to alleviate the inadequate dataset diversity, which comprises 91,371 images of 843 identities captured by multi-modal cameras across various periods throughout the day, while incorporating numerous real-world challenges. By integrating diversity representation and integrity reconstruction, the proposed Prototype-Based Diversity and Integrity learning network (PDINet) establishes excellence on the AllDay843 dataset, surpassing existing state-of-the-art approaches. The data and codes are available in https://github.com/ziwang1121/PDINet. Zi Wang 0013, Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Adaptive Interaction and Correction Attention Network for Audio-Visual MatchingabstractAudio-visual matching techniques aim to recognize and match information across different identities by learning a similarity metric across modalities. However, modal differences arise from insufficient cross-modal correlations and noise interference, which substantially hinder the performance of traditional deep metric learning methods in audio-visual matching tasks. To address the modal differences issue, we propose a novel Adaptive Interactive and Correction Attention Network (AICANet). This network efficiently captures deep information connections, generating modality-consistent feature embeddings within a unified metric framework. The core of AICANet is its two-pronged approach to reducing modal differences. First, we propose the Adaptive Interactive Attention (AIA) module, which flexibly establishes associations among cross-modal local features using dynamically generated pseudo-labels. Second, we propose the Adaptive Correction Attention (ACA) mechanism, which employs an adaptive threshold to de-interference effectively and accurately adjust the representation of local feature associations. Notably, the ACA mechanism is suitable for both intra-modal and inter-modal refined attention correction. Additionally, we design a relative distance stretching metric loss (LRDSM), which reinforces the similarity invariance of feature embeddings in a uniform space and enhances matching accuracy. Extensive tests on the VoxCeleb and VoxCeleb2 datasets demonstrate that AICANet outperforms leading existing algorithms across several evaluation metrics, validating its superior performance. The codes can be found at https://github.com/w1018979952/AICANet. Jiaxiang Wang 0001, Aihua Zheng, Lei Liu 0049, Chenglong Li 0002, Ran He 0001, Jin Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Quality-Aware Spatio-Temporal Transformer Network for RGBT TrackingabstractTransformer-based RGBT tracking has attracted much attention due to the strong modeling capacity of self attention and cross attention mechanisms. These attention mechanisms utilize the correlations among tokens to construct powerful feature representations, but are easily affected by low-quality tokens. To address this issue, we propose a novel Quality-aware Spatio-temporal Transformer Network (QSTNet), which calculates the quality weights of tokens in search regions based on the correlation with multimodal template tokens to suppress the negative effects of low-quality tokens in spatio-temporal feature representations, for robust RGBT tracking. In particular, we argue that the correlation between search tokens of one modality and multimodal template tokens could reflect the quality of these search tokens, and thus design the Quality-aware Token Weighting Module (QTWM) based on the correlation matrix of search and template tokens to suppress the negative effects of low-quality tokens. Specifically, we calculate the difference matrix derived from the attention matrices of the search tokens from both modalities and the multimodal template tokens, and then assign the quality weight for each search token based on the difference matrix, which reflects the relative correlation of search tokens from different modalities to multimodal template tokens. In addition, we propose the Prompt-based Spatio-temporal Encoder Module (PSEM) to utilize spatio-temporal multimodal information while alleviating the impact of low-quality spatio-temporal features. Extensive experiments on four RGBT benchmark datasets demonstrate that the proposed QSTNet exhibits superior performance compared to other state-of-the-art tracking methods. Our code and supplementary video are now available: https://zhaodongah.github.io/QSTNet. Zhaodong Ding, Chenglong Li 0002, Futian Wang |
IEEE Trans. Image Process. | 2 |
| 2025 | AFTER: Attention-Based Fusion Router for RGBT TrackingabstractMulti-modal feature fusion as a core investigative component of RGBT tracking emerges numerous fusion studies in recent years. However, existing RGBT tracking methods widely adopt fixed fusion structures to integrate multi-modal feature, which are hard to handle various challenges in dynamic scenarios. To address this problem, this work presents a novel Attention-based Fusion router called AFTER, which optimizes the fusion structure to adapt to the dynamic challenging scenarios, for robust RGBT tracking. In particular, we design a fusion structure space based on the hierarchical attention network, each attention-based fusion unit corresponding to a fusion operation and a combination of these attention units corresponding to a fusion structure. Through optimizing the combination of attention-based fusion units, we can dynamically select the fusion structure to adapt to various challenging scenarios. Unlike complex search of different structures in neural architecture search algorithms, we develop a dynamic routing algorithm, which equips each attention-based fusion unit with a router, to predict the combination weights for efficient optimization of the fusion structure. Extensive experiments on five mainstream RGBT tracking datasets demonstrate the superior performance of the proposed AFTER against state-of-the-art RGBT trackers. We release the code in https://github.com/Alexadlu/AFter. Andong Lu, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Dynamic Strip Convolution and Adaptive Morphology Perception Plugin for Medical Anatomy SegmentationabstractMedical anatomy segmentation is essential for computer-aided diagnosis and lesion localization in medical images. For example, segmenting individual ribs benefits localizing the lung lesions and providing vital medical measurements (such as rib spacing) for generating medical reports. Existing methods segment shape-different anatomies (such as striped ribs, bulky lungs, and angular scapula) with the same network architecture, the morphology heterogeneity is heavily overlooked. Although some shape-aware operators like deformable convolution and dynamic snake convolution have been introduced to cater to specific object morphology, they still struggle with orientation-varying strip structures, such as 24 ribs and 2 clavicles. In this paper, we propose a novel convolution plugin (DSC-AMP) for medical anatomy segmentation, which is comprised of a dynamic strip convolution (DSC) operator and an adaptive morphology perception (AMP) strategy. Specifically, the dynamic strip convolution customizes gradually varying directions and offsets for each local region, achieving dynamic striped receptive fields. Additionally, the adaptive morphology perception strategy incorporates insights from various shape-aware convolutional kernels, enabling the model to discern and integrate crucial representations corresponding to heterogeneous anatomies. Extensive experiments on two large-scale datasets demonstrate the effectiveness and superiority of the proposed approach for tackling heterogeneous medical anatomy segmentation. Guyue Hu 0001, Yukun Kang, Gangming Zhao, Zhe Jin 0001, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Knowledge-Guided Cross-Modal Alignment and Progressive Fusion for Chest X-Ray Report GenerationabstractThe task of chest X-ray report generation, which aims to simulate the diagnosis process of doctors, has received widespread attention. Compared with the image caption task, chest X-ray report generation is more challenging since it needs to generate a longer and more accurate description of each diagnostic part in chest X-ray images. Most of existing works focus on how to extract better visual features or more accurate text expression based on existing reports. However, they ignore the interactions between visual and text modalities and are thus obviously not in line with human thinking. A small part of works explore the interactions of visual and text modalities, but data-driven learning of cross-modal information mapping can not break the semantic gap between different modalities. In this work, we propose a novel approach called Knowledge-guided Cross-modal Alignment and Progressive fusion (KCAP), which takes the knowledge words from a created medical knowledge dictionary as the bridge to guide the cross-modal feature alignment and fusion, for accurate chest X-ray report generation. In particular, we create the medical knowledge dictionary by extracting medical phrases from the training set and then selecting some phrases with substantive meanings as knowledge words based on their frequency of occurrence. Based on the knowledge words from the medical knowledge dictionary, the visual and text modalities are interacted by a mapping layer for the enhancement of the features of two modalities, and then the alignment fusion module is introduced to mitigate the semantic gap between visual and text modalities. To retain the important details of the original information, we design a progressive fusion scheme to integrate the advantages of both salient fused and original features to generate better medical reports. The experimental results on IU-Xray and MIMIC datasets demonstrate the effectiveness of the proposed KCAP. Lili Huang 0006, Pengcheng Jia, Chenglong Li 0002, Jin Tang 0001, Chuanfu Li |
IEEE Trans. Multim. | 4 |
| 2025 | CRSOT: Cross-Resolution Object Tracking Using Unaligned Frame and Event CamerasabstractExisting datasets for RGB-DVS tracking are collected with DVS346 camera and their resolution ($346 \times 260$) is low for practical applications. Actually, only visible cameras are deployed in many practical systems, and the newly designed neuromorphic cameras may have different resolutions. The latest neuromorphic sensors can output high-definition event streams, but it is very difficult to achieve strict alignment between events and frames on both spatial and temporal views. Therefore, how to achieve accurate tracking with unaligned neuromorphic and visible sensors is a valuable but unresearched problem. In this work, we formally propose the task of object tracking using unaligned neuromorphic and visible cameras. We build the first unaligned frame-event dataset CRSOT collected with a specially built data acquisition system, which contains 1,030 high-definition RGB-Event video pairs, 304,974 video frames. In addition, we propose a novel unaligned object tracking framework that can realize robust tracking even using the loosely aligned RGB-Event data. This proposed method utilizes uncertainty perception techniques, which can effectively reduce the negative impact of noise (especially noise in event data) on tracking performance. Specifically, we extract the template and search regions of RGB and Event data and feed them into a unified ViT backbone for feature embedding. Next, we propose uncertainty perception modules to encode the RGB and Event features, respectively, then, we propose a modality uncertainty fusion module to aggregate the two modalities. These three branches are jointly optimized in the training phase. Extensive experiments demonstrate that our tracker can collaborate the dual modalities for high-performance tracking even without strictly temporal and spatial alignment. Yabin Zhu, Xiao Wang 0014, Chenglong Li 0002, Bo Jiang 0002, Lin Zhu 0012, Zhixiang Huang, Yonghong Tian 0001, Jin Tang 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Cross-Modal Object Tracking via Modality-Aware Fusion Network and a Large-Scale DatasetabstractVisual object tracking often faces challenges such as invalid targets and decreased performance in low-light conditions when relying solely on RGB image sequences. While incorporating additional modalities like depth and infrared data has proven effective, existing multimodal imaging platforms are complex and lack real-world applicability. In contrast, near-infrared (NIR) imaging, commonly used in surveillance cameras, can switch between RGB and NIR based on light intensity. However, tracking objects across these heterogeneous modalities poses significant challenges, particularly due to the absence of modality switch signals during tracking. To address these challenges, we propose an adaptive cross-modal object tracking algorithm called modality-aware fusion network (MAFNet). MAFNet efficiently integrates information from both RGB and NIR modalities using an adaptive weighting mechanism, effectively bridging the appearance gap and enabling a modality-aware target representation. It consists of two key components: an adaptive weighting module and a modality-specific representation module. The adaptive weighting module predicts fusion weights to dynamically adjust the contribution of each modality, while the modality-specific representation module captures discriminative features specific to RGB and NIR modalities. MAFNet offers great flexibility as it can effortlessly integrate into diverse tracking frameworks. With its simplicity, effectiveness, and efficiency, MAFNet outperforms state-of-the-art methods in cross-modal object tracking. To validate the effectiveness of our algorithm and overcome the scarcity of data in this field, we introduce CMOTB, a comprehensive and extensive benchmark dataset for cross-modal object tracking. CMOTB consists of 61 categories and 1000 video sequences, comprising a total of over 799K frames. We believe that our proposed method and dataset offer a strong foundation for advancing cross-modal object-tracking research. The dataset, toolkit, experimental data, and source code will be publicly available at: https://github.com/mmic-lcl/ Datasets-and-benchmark-code. Lei Liu 0049, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Duality-Gated Mutual Condition Network for RGBT TrackingabstractLow-quality modalities contain not only a lot of noisy information but also some discriminative features in RGB-Thermal (RGBT) tracking. However, the potentials of low-quality modalities are not well explored in existing RGBT tracking algorithms. In this work, we propose a novel duality-gated mutual condition network to fully exploit the discriminative information of all modalities while suppressing the effects of data noise. In specific, we design a mutual condition module, which takes the discriminative information of a modality as the condition to guide feature learning of target appearance in another modality. Such a module can effectively enhance target representations of all modalities even in the presence of low-quality modalities. To improve the quality of conditions and further reduce data noise, we propose a duality-gated mechanism and integrate it into the mutual condition module. To deal with the tracking failure caused by sudden camera motion, which often occurs in RGBT tracking, we design a resampling strategy based on optical flow. It does not increase much computational cost since we perform optical flow calculation only when the model prediction is unreliable and then execute resampling when the sudden camera motion is detected. Extensive experiments on four RGBT tracking benchmark datasets show that our method performs favorably against the state-of-the-art tracking algorithms. Andong Lu, Cun Qian, Chenglong Li 0002, Jin Tang 0001, Liang Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Structural Information Guided Multimodal Pre-training for Vehicle-Centric PerceptionabstractUnderstanding vehicles in images is important for various applications such as intelligent transportation and self-driving system. Existing vehicle-centric works typically pre-train models on large-scale classification datasets and then fine-tune them for specific downstream tasks. However, they neglect the specific characteristics of vehicle perception in different tasks and might thus lead to sub-optimal performance. To address this issue, we propose a novel vehicle-centric pre-training framework called VehicleMAE, which incorporates the structural information including the spatial structure from vehicle profile information and the semantic structure from informative high-level natural language descriptions for effective masked vehicle appearance reconstruction. To be specific, we explicitly extract the sketch lines of vehicles as a form of the spatial structure to guide vehicle reconstruction. The more comprehensive knowledge distilled from the CLIP big model based on the similarity between the paired/unpaired vehicle image-text sample is further taken into consideration to help achieve a better understanding of vehicles. A large-scale dataset is built to pre-train our model, termed Autobot1M, which contains about 1M vehicle images and 12693 text information. Extensive experiments on four vehicle-based downstream tasks fully validated the effectiveness of our VehicleMAE. The source code and pre-trained models will be released at https://github.com/Event-AHU/VehicleMAE. Xiao Wang 0014, Chenglong Li 0002, Zhicheng Zhao 0001, Zhe Chen 0013, Yukai Shi, Jin Tang 0001 |
AAAI | 3 |
| 2024 | DI-MVS: Learning Efficient Multi-View Stereo With Depth-Aware IterationsabstractLearning-based Multi-View Stereo (MVS) methods aim to reconstruct 3D scenes from a set of 2D calibrated images. However, existing learning-based MVS methods often overlook depth maps that include the geometric shapes of the scene when constructing the cost volume. This can result in suboptimal reconstructions, particularly in low-texture or repetitive-texture regions where valuable geometric information is absent. To address this issue, we develop DI-MVS, a coarse-to-fine framework that effectively incorporates context-guided depth geometry into the cost volume using a depth-aware iterator. First, we employ the proposed depth-aware cost completion module to update the cost volume, followed by 2D ConvGRUs to iteratively optimize depth maps efficiently. Second, we propose a hybrid loss strategy that combines two loss functions’ strengths to improve depth estimation’s robustness. Extensive experiments demonstrate that DI-MVS outperforms state-of-the-art methods on the DTU dataset and the Tanks & Temples benchmark. The source code is available at: https://github.com/JianfeiJ/DI-MVS. Jianfei Jiang 0005, Mingwei Cao, Chenglong Li 0002 |
ICASSP | 4 |
| 2024 | Parallel Augmentation and Dual Enhancement for Occluded Person Re-IdentificationabstractOccluded person re-identification (Re-ID), the task of searching for the same person’s images in occluded environments, has attracted lots of attention in the past decades. Recent approaches concentrate on improving performance on occluded data by data/feature augmentation or using extra models to predict occlusions. However, they ignore the imbalance problem in this task and can not fully utilize the information from the training data. To alleviate these two issues, we propose a simple yet effective method with Parallel Augmentation and Dual Enhancement (PADE), which is robust on both occluded and non-occluded data and does not require any auxiliary clues. First, we design a parallel augmentation mechanism (PAM) to generate more suitable occluded data to mitigate the negative effects of unbalanced data. Second, we propose the global and local dual enhancement strategy (DES) to promote the context information and details. Experimental results on three widely used occluded datasets and two non-occluded datasets validate the effectiveness of our method. The code is available at PADE (GitHub). Zi Wang 0013, Huaibo Huang, Aihua Zheng, Chenglong Li 0002, Ran He 0001 |
ICASSP | 4 |
| 2024 | Breaking Modality Gap in RGBT Tracking: Coupled Knowledge DistillationabstractModality gap between RGB and thermal infrared (TIR) images is a crucial issue but often overlooked in existing RGBT tracking methods. It can be observed that modality gap mainly lies in the image style difference. In this work, we propose a novel Coupled Knowledge Distillation framework called CKD, which pursues common styles of different modalities to break modality gap, for high performance RGBT tracking. In particular, we introduce two student networks and employ the style distillation loss to make their style features consistent as much as possible. Through alleviating the style difference of two student networks, we can break modality gap of different modalities well. However, the distillation of style features might harm to the content representations of two modalities in student networks. To handle this issue, we take original RGB and TIR networks as the teachers, and distill their content knowledge into two student networks respectively by the style-content orthogonal feature decoupling scheme. We couple the above two distillation processes in an online optimization framework to form new feature representations of RGB and thermal modalities without modality gap. In addition, we design a masked modeling strategy and a multi-modal candidate token elimination strategy into CKD to improve tracking robustness and efficiency respectively. Extensive experiments on five standard RGBT tracking datasets validate the effectiveness of the proposed method against state-of-the-art methods while achieving the fastest tracking speed of 96.4 FPS. Andong Lu, Jiacong Zhao, Chenglong Li 0002, Yun Xiao 0003, Bin Luo 0001 |
ACM Multimedia | 3 |
| 2024 | Semantics Guided Disentangled GAN for Chest X-Ray Image Rib Segmentation
Lili Huang 0006, Dexin Ma, Chenglong Li 0002, Haifeng Zhao 0001, Jin Tang 0001, Chuanfu Li |
PRCV (14) | 4 |
| 2024 | Disentangled generation network for enlarged license plate recognition and a unified dataset
Chenglong Li 0002, Xiaobin Yang, Guohao Wang, Aihua Zheng, Jin Tang 0001 |
Comput. Vis. Image Underst. | 1 |
| 2024 | Lane detection via disentangled representation network with slope consistency loss
Zhaodong Ding, Yifei Deng, Chenglong Li 0002, Rui Ruan, Jin Tang 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Class Hierarchy-Guided Generalized Few-Shot Ship Detection in Remote Sensing ImagesabstractFine-grained ship detection in remote sensing images (RSIs) depends heavily on numerous training data with expensive manual annotations. Learning novel ship categories from very few labeled samples and without forgetting the learned knowledge of seen categories is important to real-world applications. In this letter, we formulate fine-grained ship detection in RSIs as a problem of generalized few-shot object detection (G-FSOD). Existing methods often neglect the structured information in ship taxonomy, and thus result in mutually exclusive representations between base and novel classes and hinder the transfer of the learned knowledge to the novel concepts under the few-shot settings. To handle this problem, we propose to incorporate the inherent hierarchical taxonomy in ship classes into the generalized few-shot ship detection to leverage the shared knowledge among base and novel classes. In particular, a ship detector is trained based on the coarsest class labels and a multitask classification network is built to distinguish various ships at both coarse and fine-grained levels on base classes, which leads to a generalized ship representation between base classes to novel classes. To build the classifier of novel classes, a prototype bank is constructed with the few-shot samples of novel classes, without the wreck of the feature extractor so as to maintain the performance on base classes. Extensive experiments on two large-scale ship detection datasets demonstrate the effectiveness of our method against state-of-the-art methods. Shuangqing Zhang, Zhang Zhang 0001, Da Li 0003, Chenglong Li 0002, Liang Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | UAV-Ground Visual Tracking: A Unified Dataset and Collaborative Learning ApproachabstractVisual tracking from the ground view and the UAV view has received increasing attention due to its wide range of practical applications. These two tasks have strong complementary benefits in the description of the target object, such as detailed appearance in the ground view and global motion information in the UAV view, and their combination has the potential to allow the tracking system to be more robust. However, no work has studied this problem in-depth, and it is challenging to accurately combine the ground view information and the UAV view information. To fill the gap and address the challenge, we propose a new computer vision task called UAV-Ground visual tracking. Considering the lack of relevant data and methods, we first propose a unified video dataset called UGVT, which includes 210 pairs of UAV and ground high-resolution video sequences with a total of more than 204K frames, which can be used as a comprehensive evaluation platform for relevant tracking methods. Secondly, based on the newly constructed dataset, we propose a co-learning method called MvCL to fuse the information of ground and UAV views. It first associates the same tracking target in the two views based on cross-attention operation and then fuses the complementary information of the two views. In particular, as a plug-and-play module based on Transformer structure, this method can be flexibly embedded into different tracking frameworks. Extensive experiments are conducted on the newly created dataset. The results demonstrate the effectiveness of the proposed method in improving the robustness of the tracking system compared with 10 state-of-the-art tracking methods and also indicate the prospect and significance of potential UAV-Ground visual tracking research. The dataset is available at:https://github.com/mmic-lcl/Datasets-and-benchmark-code/. Dengdi Sun, Leilei Cheng, Chenglong Li 0002, Yun Xiao 0003, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Transformer RGBT Tracking With Spatio-Temporal Multimodal TokensabstractMany RGBT tracking researches primarily focus on modal fusion design, while overlooking the effective handling of target appearance changes. While some approaches have introduced historical frames or fuse and replace initial templates to incorporate temporal information, they have the risk of disrupting the original target appearance and accumulating errors over time. To alleviate these limitations, we propose a novel Transformer RGBT tracking approach, which mixes spatio-temporal multimodal tokens from the static multimodal templates and multimodal search regions in Transformer to handle target appearance changes, for robust RGBT tracking. We introduce independent dynamic template tokens to interact with the search region, embedding temporal information to address appearance changes, while also retaining the involvement of the initial static template tokens in the joint feature extraction process to ensure the preservation of the original reliable target appearance information that prevent deviations from the target appearance caused by traditional temporal updates. We also use attention mechanisms to enhance the target features of multimodal template tokens by incorporating supplementary modal cues, and make the multimodal search region tokens interact with multimodal dynamic template tokens via attention mechanisms, which facilitates the conveyance of multimodal-enhanced target change information. Our module is inserted into the transformer backbone network and inherits joint feature extraction, search-template matching, and cross-modal interaction. Extensive experiments on three RGBT benchmark datasets show that the proposed approach maintains competitive performance compared to other state-of-the-art tracking algorithms while running at 39.1 FPS. The project-related materials are available at:https://github.com/yinghaidada/STMT. Dengdi Sun, Yajie Pan, Andong Lu, Chenglong Li 0002, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Learning Adaptive Fusion Bank for Multi-Modal Salient Object DetectionabstractMulti-modal salient object detection (MSOD) aims to boost saliency detection performance by integrating visible sources with depth or thermal infrared ones. Existing methods generally design different fusion schemes to handle certain issues or challenges. Although these fusion schemes are effective at addressing specific issues or challenges, they may struggle to handle multiple complex challenges simultaneously. To solve this problem, we propose a novel adaptive fusion bank that makes full use of the complementary benefits from a set of basic fusion schemes to handle different challenges simultaneously for robust MSOD. We focus on handling five major challenges in MSOD, namely center bias, scale variation, image clutter, low illumination, and thermal crossover or depth ambiguity. The fusion bank proposed consists of five representative fusion schemes, which are specifically designed based on the characteristics of each challenge, respectively. The bank is scalable, and more fusion schemes could be incorporated into the bank for more challenges. To adaptively select the appropriate fusion scheme for multi-modal input, we introduce an adaptive ensemble module that forms the adaptive fusion bank, which is embedded into hierarchical layers for sufficient fusion of different source data. Moreover, we design an indirect interactive guidance module to accurately detect salient hollow objects via the skip integration of high-level semantic information and low-level spatial details. Extensive experiments on three RGBT datasets and seven RGBD datasets demonstrate that the proposed method achieves the outstanding performance compared to the state-of-the-art methods. Kunpeng Wang 0005, Zhengzheng Tu, Chenglong Li 0002, Cheng Zhang 0010, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Public-Private Attributes-Based Variational Adversarial Network for Audio-Visual Cross-Modal MatchingabstractExisting audio-visual cross-modal matching methods focus on mitigating cross-modal heterogeneity but ignore the impact of intra-class discrepancy of the same identity in different scenarios, which might greatly limit the matching performance. To simultaneously handle both problems of intra-class discrepancy and cross-modal heterogeneity, we propose a novel public-private attributes-based variational adversarial network (P2VANet), which captures the consistency within and between classes, for audio-visual cross-modal matching. In particular,P2VANet first uses a variational auto-encoder, which captures the inherent global information in diverse scenarios from the hidden variable through reconstruction, to reduce the intra-class discrepancy. Then it integrates a public attributes guidance module to capture the consistency of audio and visual by supervision of the common high-level semantic information to mitigate cross-modal heterogeneity. In addition,P2VANet designs private attributes embedding module to enhance the discriminative features inherent in each class to decrease inter-class similarity. Extensive experiments on audio-visual cross-modal matching demonstrate the effectiveness of the proposed approach compared with the state-of-the-art methods. Aihua Zheng, Jiaxiang Wang 0001, Chao Tang 0002, Chenglong Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | RGBT Tracking via Progressive Fusion Transformer With Dynamically Guided LearningabstractExisting Transformer-based RGB-Thermal (RGBT) tracking methods either use cross-attention to fuse the two modalities, or use self-attention and cross-attention to model both modality-specific and modality-sharing information. However, the significant appearance gap between modalities limits the feature representation ability of certain modalities during the fusion process. To address this problem, we propose a novel Progressive Fusion Transformer called ProFormer, which progressively integrates single-modality information into the multimodal representation for robust RGBT tracking. In particular, ProFormer first uses a self-attention module to collaboratively extract the multimodal representation. Then, ProFormer introduces two cross-attention modules to interact it with the features of the dual modalities for enhancing modality-specific information in the multimodal representation. In addition, we propose a dynamically guided learning algorithm that adaptively employs the well-performing branches to guide the learning of other branches, to improve the representation ability of each branch. Extensive experiments demonstrate that our proposed ProFormer achieves a new state-of-the-art performance on RGBT210, RGBT234, LasHeR, and VTUAV datasets. Yabin Zhu, Chenglong Li 0002, Xiao Wang 0014, Jin Tang 0001, Zhixiang Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Dense Tiny Object Detection: A Scene Context Guided Approach and a Unified BenchmarkabstractWith the continuous advancement of remote sensing observation technology, wide-area observation and high-resolution imaging make remote sensing images contain a large number of dense tiny objects. The detection of dense tiny objects is a very challenging task since these objects are with very low resolution and might stick together. Existing work lacks further exploration of the contextual scene information and inherent characteristics of dense tiny objects, which are crucial for performance improvement of dense tiny object detection. In this work, we propose a novel Scene Contextualized Detection Network (SCDNet) by decoupling scene contextual information through a dedicated scene classification sub-network, thereby enabling an enhanced exploration of the relationship between tiny objects and their surrounding environments. In particular, we design a lightweight scene context guided fusion module in SCDNet to incorporate scene context information around dense tiny objects more effectively. Moreover, we further develop the scene context guided foreground enhancement module to suppress the background information while enhancing the foreground information based on the scene information. In addition, this research field still lacks a large-scale benchmark dataset with dense tiny objects, which is crucial for the training and comprehensive evaluation of detection methods. To this end, we construct a large-scale dataset for dense tiny object detection. It contains 11,600 images with 1,019,800 instances, the average absolute size of objects is smaller than 13 pixels, and each image contains 88 objects on average. Extensive experiments are conducted on the proposed dataset, and the results demonstrate the superiority and effectiveness of SCDNet compared to existing methods. The dataset and evaluation code are available at https://github.com/mmic-lcl. Zhicheng Zhao 0002, Chenglong Li 0002, Yun Xiao 0003, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Modality Conversion Meets Superresolution: A Collaborative Framework for High- Resolution Thermal UAV Image GenerationabstractDue to the limitations and costs of thermal sensors, unmanned aerial vehicle (UAV) platforms often equip with high-resolution (HR) visible imaging and low-resolution (LR) thermal imaging cameras for all-day monitoring capability. Existing works generate the high-resolution thermal UAV images by either super-resolution (SR) from high-resolution visible and low-resolution thermal images or modality conversion (MC) from high-resolution visible images. However, the modality gap between visible and thermal sources might degrade the generation quality. We observe that the MC task is beneficial in addressing the cross-modal gap in the SR task, while the SR task can provide the condition of thermal information to boost the MC task. Moreover, these two tasks have the same output and can thus be carried out simultaneously without any additional annotation. Based on this observation, we propose a collaborative enhancement network (CENet), which performs thermal UAV image SR and visible image MC in a joint manner, for high-resolution thermal UAV image generation. In particular, we design a mutual guidance module to interact the features from SR and MC tasks in an alternating bidirectional manner. Considering that low-level vision tasks are position-sensitive, to further enhance the feature alignment between the two tasks, we design a bidirectional alignment fusion module to maintain feature consistency of the MC and SR branches. The proposed collaborative framework not only achieves joint and unified training of the two tasks, but also generates two types of complementary high-resolution images. Extensive experiments on public datasets demonstrate that the proposed CENet outperforms current state-of-the-art super-resolution (SR) methods in generating high-resolution thermal UAV images, as quantified by peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM). Zhicheng Zhao 0002, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Long-Term Motion-Assisted Remote Sensing Object TrackingabstractRemote sensing object tracking has gained significant attention due to its wide range of applications including surveillance and motion analysis. However, it faces various challenges such as low resolution, low contrast, blurring, and occlusion, which impede its development at a significantly slower pace compared to object tracking methods for general scenes. The challenges of low resolution, low contrast, and blurring result in weak target features, while the occlusion challenge poses a problem for target search range and tracker discrimination in subsequent frames. To address these issues, we propose a novel long-term motion-assisted framework, which can effectively mine long-term motion information and use an evaluation scheme for robust remote sensing object tracking. Specifically, we design a long-term motion feature mining module (LMFM), which efficiently calculates the long-term motion information by integrating previous motion features in a temporal-iterative manner to alleviate the problem of weak features caused by low resolution, low contrast, and blurring. Moreover, we design an evaluation scheme that combines the motion trajectory model, target classification scores, and predicted target positions to handle the issue of massive occlusion or target loss. Extensive experiments on the SatSOT, SV248S, and VISO datasets show that our approach outperforms state-of-the-art (SOTA) trackers. The source code, trained models, and raw results are released athttps://github.com/zhaoxingle/LMANet. Yabin Zhu, Xingle Zhao, Chenglong Li 0002, Jin Tang 0001, Zhixiang Huang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | RGBT Tracking via Challenge-Based Appearance Disentanglement and InteractionabstractRGB and thermal source data suffer from both shared and specific challenges, and how to explore and exploit them plays a critical role in representing the target appearance in RGBT tracking. In this paper, we propose a novel approach, which performs target appearance representation disentanglement and interaction via both modality-shared and modality-specific challenge attributes, for robust RGBT tracking. In particular, we disentangle the target appearance representations via five challenge-based branches with different structures according to their properties, including three parameter-shared branches to model modality-shared challenges and two parameter-independent branches to model modality-specific challenges. Considering the complementary advantages between modality-specific cues, we propose a guidance interaction module to transfer discriminative features from one modality to another one to enhance the discriminative ability of weak modality. Moreover, we design an aggregation interaction module to combine all challenge-based target representations, which could form more discriminative target representations and fit the challenge-agnostic tracking process. These challenge-based branches are able to model the target appearance under certain challenges so that the target representations can be learned by a few parameters even in the situation of insufficient training data. In addition, to relieve labor costs and avoid label ambiguity, we design a generation strategy to generate training data with different challenge attributes. Comprehensive experiments demonstrate the superiority of the proposed tracker against the state-of-the-art methods on four benchmark datasets. Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003, Rui Ruan, Minghao Fan |
IEEE Trans. Image Process. | 2 |
| 2024 | Text-to-Image Vehicle Re-Identification: Multi-Scale Multi-View Cross-Modal Alignment Network and a Unified BenchmarkabstractVehicle Re-IDentification (Re-ID) aims to retrieve the most similar images with a given query vehicle image from a set of images captured by non-overlapping cameras, and plays a crucial role in intelligent transportation systems and has made impressive advancements in recent years. In real-world scenarios, we can often acquire the text descriptions of target vehicle through witness accounts, and then manually search the image queries for vehicle Re-ID, which is time-consuming and labor-intensive. To solve this problem, this paper introduces a new fine-grained cross-modal retrieval task called text-to-image vehicle re-identification, which seeks to retrieve target vehicle images based on the given text descriptions. To bridge the significant gap between language and visual modalities, we propose a novel Multi-scale multi-view Cross-modal Alignment Network (MCANet). In particular, we incorporate view masks and multi-scale features to align image and text features in a progressive way. In addition, we design the Masked Bidirectional InfoNCE (MB-InfoNCE) loss to enhance the training stability and make the best use of negative samples. To provide an evaluation platform for text-to-image vehicle re-identification, we create a Text-to-Image Vehicle Re-Identification dataset (T2I VeRi), which contains 2465 image-text pairs from 776 vehicles with an average sentence length of 26.8 words. Extensive experiments conducted on T2I VeRi demonstrate MCANet outperforms the current state-of-art (SOTA) method by 2.2% in rank-1 accuracy. Leqi Ding, Lei Liu 0049, Yan Huang 0008, Chenglong Li 0002, Cheng Zhang 0010, Wei Wang 0115, Liang Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Collaborative License Plate Recognition via Association Enhancement Network With Auxiliary Learning and a Unified BenchmarkabstractSince the standard license plate of large vehicle is easily affected by occlusion and stain, the traffic management department introduces the enlarged license plate at the rear of the large vehicle to assist license plate recognition. However, current researches regards standard license plate recognition and enlarged license plate recognition as independent tasks, and do not take advantage of the complementary benefits from the two types of license plates. In this work, we propose a new computer vision task called collaborative license plate recognition, aiming to leverage the complementary advantages of standard and enlarged license plates for achieving more accurate license plate recognition. To achieve this goal, we propose an Association Enhancement Network (AENet), which achieves robust collaborative licence plate recognition by capturing the correlations between characters within a single licence plate and enhancing the associations between two license plates. In particular, we design an association enhancement branch, which supervises the fusion of two licence plate information using the complete licence plate number to mine the association between them. To enhance the representation ability of each type of licence plates, we design an auxiliary learning branch in the training stage, which supervises the learning of individual license plates in the association enhancement between two license plates. In addition, we contribute a comprehensive benchmark dataset called CLPR, which consists of a total of 19,782 standard and enlarged licence plates from 24 provinces in China and covers most of the challenges in real scenarios, for collaborative license plate recognition. Extensive experiments on the proposed CLPR dataset demonstrate the effectiveness of the proposed AENet against several state-of-the-art methods. Yifei Deng, Guohao Wang, Chenglong Li 0002, Wei Wang 0115, Cheng Zhang 0010, Jin Tang 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Illumination Distillation Framework for Nighttime Person Re-Identification and a New BenchmarkabstractNighttime person Re-ID (person re-identification in the nighttime) is a very important and challenging task for visual surveillance but it has not been thoroughly investigated. Under the low illumination condition, the performance of person Re-ID methods usually sharply deteriorates. To address the low illumination challenge in nighttime person Re-ID, this paper proposes an Illumination Distillation Framework (IDF), which utilizes illumination enhancement and illumination distillation schemes to promote the learning of Re-ID models. Specifically, IDF consists of a master branch, an illumination enhancement branch, and an illumination distillation module. The master branch is used to extract the features from a nighttime image. The illumination enhancement branch first estimates an enhanced image from the nighttime image using a nonlinear curve mapping method and then extracts the enhanced features. However, nighttime and enhanced features usually contain data noise due to unstable lighting conditions and enhancement failures. To fully exploit the complementary benefits of nighttime and enhanced features while suppressing data noise, we propose an illumination distillation module. In particular, the illumination distillation module fuses the features from two branches through a bottleneck fusion model and then uses the fused features to guide the learning of both branches in a distillation manner. In addition, we build a real-world nighttime person Re-ID dataset, namedNight600, which contains 600 identities captured from different viewpoints and nighttime illumination conditions under complex outdoor environments. Experimental results demonstrate that our IDF can achieve state-of-the-art performance on two nighttime person Re-ID datasets (i.e.,Night600andKnight). We will release our code and dataset athttps://github.com/Alexadlu/IDF. Andong Lu, Zhang Zhang 0001, Yan Huang 0023, Yifan Zhang 0004, Chenglong Li 0002, Jin Tang 0001, Liang Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Alignment-Free RGBT Salient Object Detection: Semantics-Guided Asymmetric Correlation Network and a Unified BenchmarkabstractRGB and Thermal (RGBT) Salient Object Detection (SOD) aims to achieve high-quality saliency prediction by exploiting the complementary information of visible and thermal image pairs, which are initially captured in an unaligned manner. However, existing methods are tailored for manually aligned image pairs, which are labor-intensive, and directly applying these methods to original unaligned image pairs could significantly degrade their performance. In this paper, we make the first attempt to address RGBT SOD for initially captured RGB and thermal image pairs without manual alignment. Specifically, we propose a Semantics-guided Asymmetric Correlation Network (SACNet) that consists of two novel components: 1) an asymmetric correlation module utilizing semantics-guided attention to model cross-modal correlations specific to unaligned salient regions; 2) an associated feature sampling module to sample relevant thermal features according to the corresponding RGB features for multi-modal feature integration. In addition, we construct a unified benchmark dataset called UVT2000, containing 2000 RGB and thermal image pairs directly captured from various real-world scenes without any alignment, to facilitate research on alignment-free RGBT SOD. Extensive experiments on both aligned and unaligned datasets demonstrate the effectiveness and superior performance of our method. Kunpeng Wang 0005, Danying Lin, Chenglong Li 0002, Zhengzheng Tu, Bin Luo 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Group Multi-View Transformer for 3D Shape Analysis With Spatial EncodingabstractIn recent years, the results of view-based 3D shape recognition methods have saturated, and models with excellent performance cannot be deployed on memory-limited devices due to their huge size of parameters. To address this problem, we introduce a compression method based on knowledge distillation for this field, which largely reduces the number of parameters while preserving model performance as much as possible. Specifically, to enhance the capabilities of smaller models, we design a high-performing large model called Group Multi-view Vision Transformer (GMViT). In GMViT, the view-level ViT first establishes relationships between view-level features. Additionally, to capture deeper features, we employ the grouping module to enhance view-level features into group-level features. Finally, the group-level ViT aggregates group-level features into complete, well-formed 3D shape descriptors. Notably, in both ViTs, we introduce spatial encoding of camera coordinates as innovative position embeddings. Furthermore, we propose two compressed versions based on GMViT, namely GMViT-simple and GMViT-mini. To enhance the training effectiveness of the small models, we introduce a knowledge distillation method throughout the GMViT process, where the key outputs of each GMViT component serve as distillation targets. Extensive experiments demonstrate the efficacy of the proposed method. The large model GMViT achieves excellent 3D classification and retrieval results on the benchmark datasets ModelNet, ShapeNetCore55, and MCB. The smaller models, GMViT-simple and GMViT-mini, reduce the parameter size by 8 and 17.6 times, respectively, and improve shape recognition speed by 1.5 times on average, while preserving at least 90% of the recognition performance. Lixiang Xu, Qingzhe Cui, Richang Hong, Enhong Chen, Xin Yuan 0008, Chenglong Li 0002, Yuan Yan Tang |
IEEE Trans. Multim. | 7 |
| 2024 | Tiny Object Tracking: A Large-Scale Dataset and a BaselineabstractTiny objects, frequently appearing in practical applications, have weak appearance and features, and receive increasing interests in many vision tasks, such as object detection and segmentation. To promote the research and development of tiny object tracking, we create a large-scale video dataset, which contains 434 sequences with a total of more than 217K frames. Each frame is carefully annotated with a high-quality bounding box. In data creation, we take 12 challenge attributes into account to cover a broad range of viewpoints and scene complexities, and annotate these attributes for facilitating the attribute-based performance analysis. To provide a strong baseline in tiny object tracking, we propose a novel multilevel knowledge distillation network (MKDNet), which pursues three-level knowledge distillations in a unified framework to effectively enhance the feature representation, discrimination, and localization abilities in tracking tiny objects. Extensive experiments are performed on the proposed dataset, and the results prove the superiority and effectiveness of MKDNet compared with state-of-the-art methods. The dataset, the algorithm code, and the evaluation code are available at https://github.com/mmic-lcl/Datasets-and-benchmark-code. Yabin Zhu, Chenglong Li 0002, Xiao Wang 0014, Jin Tang 0001, Bin Luo 0001, Zhixiang Huang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Quality-Aware RGBT Tracking via Supervised Reliability Learning and Weighted Residual GuidanceabstractRGB and thermal infrared (TIR) data have different visual properties, which make their fusion essential for effective object tracking in diverse environments and scenes. Existing RGBT tracking methods commonly use attention mechanisms to generate reliability weights for multi-modal feature fusion. However, without explicit supervision, these weights may be unreliably estimated, especially in complex scenarios. To address this problem, we propose a novel Quality-Aware RGBT Tracker (QAT) for robust RGBT tracking. QAT learns reliable weights for each modality in a supervised manner and performs weighted residual guidance to extract and leverage useful features from both modalities. We address the issue of the lack of labels for reliability learning by designing an efficient three-branch network that generates reliable pseudo labels, and a simple binary classification scheme that estimates high-accuracy reliability weights, mitigating the effect of noisy pseudo labels. To propagate useful features between modalities while reducing the influence of noisy modal features on the migrated information, we design a weighted residual guidance module based on the estimated weights and residual connections. We evaluate our proposed QAT on five benchmark datasets, including GTOT, RGBT210, RGBT234, LasHeR, and VTUAV, and demonstrate its excellent performance compared to state-of-the-art methods. Experimental results show that QAT outperforms existing RGBT tracking methods in various challenging scenarios, demonstrating its efficacy in improving the reliability and accuracy of RGBT tracking. Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003, Jin Tang 0001 |
ACM Multimedia | 2 |
| 2023 | Siamese transformer RGBT tracking
Futian Wang, Lei Liu 0049, Chenglong Li 0002, Jing Tang 0001 |
Appl. Intell. | 4 |
| 2023 | Multimodal salient object detection via adversarial learning with collaborative generator
Zhengzheng Tu, Wenfang Yang, Kunpeng Wang 0005, Amir Hussain 0001, Bin Luo 0001, Chenglong Li 0002 |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | SiamON: Siamese Occlusion-Aware Network for Visual TrackingabstractOcclusion has been proven to be one of the most challenging factors faced by most visual trackers. There are mainly two difficulties, the first one is that the number of occlusion samples are very limited even though collecting a large-scale training data set, and another one is how to correctly learn the features of the target when comes to occlusion situations. In this paper, we tried to solve these two problems together in our proposed model. To this end, we propose a novel Siamese Occlusion-aware Network (SiamON) for high-performance visual tracking. In particular, we predefine some soft-masks to solve the problem of fewer occlusion samples, which perceive patterns of occlusion contents at different locations and take these masks as the conditions to guide occlusion-aware feature learning. Meanwhile, we propose a target-aware attention mechanism allows the model to pay more attention to the target and further weaken the impact of occlusion. Extensive experiments on several popular benchmarks show that our tracking method exceeds many state-of-the-art trackers especially in the presence of occlusion and meets the requirements of real-time. Chao Fan 0001, Hongyuan Yu, Yan Huang 0008, Caifeng Shan, Liang Wang 0001, Chenglong Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Diag-IoU Loss for Object DetectionabstractExisting IoU-based loss functions have achieved promising performance for bounding box regression in object detection. However, they cannot fully reflect the relation between the predicted and target boxes in the case of box inclusions, and might thus deteriorate detection accuracy and efficiency. In this paper, we design a novel similarity measurement based on the box diagonal called Diag-IoU to well represent the divergence between the predicted and target boxes even in the case of box inclusions, and thus achieve superior localization accuracy and fast convergence. In particular, we equivalently represent a rectangular box with its box diagonal, which contains exclusive and informative geometrical factors, and define the Diag-IoU based on the similarities of a set of sampled point pairs from the predicted and target box diagonals. Based on the Diag-IoU, we design a general Diag-IoU loss, which can provide holistic information in measuring two boxes and thus differentiate the two boxes in the case of box inclusions. To validate the effectiveness of the proposed method, we apply the Diag-IoU loss to several representative object detectors, including YOLO v5s, Faster R-CNN, and FCOS. Extensive experiments on the synthetic data and two challenging object detection benchmark datasets, i.e., MS COCO and PASCAL VOC, demonstrate the superior performance of the proposed Diag-IoU loss compared to previous IoU-based losses as well as other metrics. Shuangqing Zhang, Chenglong Li 0002, Lei Liu 0049, Zhang Zhang 0001, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Category-Oriented Localization Distillation for SAR Object Detection and a Unified BenchmarkabstractDespite much research progress in synthetic aperture radar (SAR) object detection, the performance of SAR object detection has encountered a bottleneck limited by the imaging mechanism of SAR. In this work, we investigate how to perform robust SAR object detection by distilling the category knowledge from optical images in the training stage. To this end, we propose a novel knowledge distillation method called Category-oriented Localization Distillation (CoLD), which employs the optical object detection network as the teacher to guide the SAR object detection network. To introduce the category prior knowledge of the teacher network in the localization knowledge transferring, a category-oriented partition module is designed in CoLD to decouple candidate bounding boxes into target and non-target ones according to the category information in optical images. Through box decoupling, the accuracy and efficiency of SAR object detection can be significantly improved. Moreover, an IoU-based weighting module is introduced in CoLD to guide the student network focusing more on high-quality candidate boxes by adaptively changing the weight of each candidate bounding box based on the corresponding IoU score in the teacher network. In addition, a unified benchmark dataset is created for the evaluation of optical information guided SAR object detection, which consists of 14,665 optical and SAR image pairs in the training set and 3,666 SAR images in the testing set. Extensive experiments on the dataset demonstrate the effectiveness of our CoLD against state-of-the-art methods. The dataset is available at: https://github.com/mmic-lcl/Datasets-and-benchmark-code. Rui Ruan, Zhicheng Zhao 0002, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Thermal UAV Image Super-Resolution Guided by Multiple Visible CuesabstractUnmanned aerial vehicle (UAV) thermal-imaging has received much attention, but the insufficient image resolution caused by thermal imaging systems is still a crucial problem that limits the understanding of thermal UAV images. However, high-resolution visible images are relatively easy to access, and it is thus valuable for exploring useful information from visible image to assist thermal UAV image super-resolution (SR). In this article, we propose a novel multiconditioned guidance network (MGNet) to effectively mine the information of visible images for thermal UAV image SR. High-resolution visible UAV images usually contain salient appearance, semantic, and edge information, which plays a critical role in boosting the performance of thermal UAV image SR. Therefore, we design an effective multicue guidance module (MGM) to leverage the appearance, edge, and semantic cues from visible images to guide thermal UAV image SR. In addition, we build the first benchmark dataset for the task of thermal UAV image SR guided by visible images. It is collected by a multimodal UAV platform and composes of 1025 pairs of manually aligned visible and thermal images. Extensive experiments on the built dataset show that our MGNet can effectively leverage useful information from visible images to improve the performance of thermal UAV image SR and perform well against several state-of-the-art methods. The dataset is available at:https://github.com/mmic-lcl/Datasets-and-benchmark-code. Zhicheng Zhao 0002, Chenglong Li 0002, Yun Xiao 0003, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Multi-Query Vehicle Re-Identification: Viewpoint-Conditioned Network, Unified Dataset and New MetricabstractExisting vehicle re-identification methods mainly rely on the single query, which has limited information for vehicle representation and thus significantly hinders the performance of vehicle Re-ID in complicated surveillance networks. In this paper, we propose a more realistic and easily accessible task, called multi-query vehicle Re-ID, which leverages multiple queries to overcome viewpoint limitation of single one. Based on this task, we make three major contributions. First, we design a novel viewpoint-conditioned network (VCNet), which adaptively combines the complementary information from different vehicle viewpoints, for multi-query vehicle Re-ID. Moreover, to deal with the problem of missing vehicle viewpoints, we propose a cross-view feature recovery module which recovers the features of the missing viewpoints by learnt the correlation between the features of available and missing viewpoints. Second, we create a unified benchmark dataset, taken by 6142 cameras from a real-life transportation surveillance system, with comprehensive viewpoints and large number of crossed scenes of each vehicle for multi-query vehicle Re-ID evaluation. Finally, we design a new evaluation metric, called mean cross-scene precision (mCSP), which measures the ability of cross-scene recognition by suppressing the positive samples with similar viewpoints from the same camera. Comprehensive experiments validate the superiority of the proposed method against other methods, as well as the effectiveness of the designed metric in the evaluation of multi-query vehicle Re-ID. The codes and dataset are available at: https://github.com/zhangchaobin001/VCNet. Aihua Zheng, Chaobin Zhang, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | RGBT Salient Object Detection: A Large-Scale Dataset and BenchmarkabstractSalient object detection in complex scenes and environments is a challenging research topic. Most works focus on RGB-based salient object detection, which limits its performance of real-life applications when confronted with adverse conditions such as dark environments and complex backgrounds. Taking advantage of RGB and thermal infrared(RGBT) images becomes a new research direction for detecting salient objects in complex scenes, since the thermal infrared spectrum provides the complementary information and has been used in many computer vision tasks. However, current research for RGBT salient object detection is limited by the lack of a large-scale dataset and comprehensive benchmark. This work contributes such a RGBT image dataset named VT5000, including 5000 spatially aligned RGBT image pairs with ground truth annotations. VT5000 has 11 challenges collected in different scenes and environments for exploring the robustness of algorithms. With this dataset, we propose a powerful baseline approach, which extracts multilevel features of each modality and aggregates these features of all modalities with the attention mechanism for accurate RGBT salient object detection. To further solve the problem of blur boundaries of salient objects, we also use an edge loss to refine the boundaries. Extensive experiments show that the proposed baseline approach outperforms the state-of-the-art methods on VT5000 dataset and other two public datasets. In addition, we carry out a comprehensive analysis of different algorithms of RGBT salient object detection on VT5000 dataset, and then make several valuable conclusions and provide some potential research directions for RGBT salient object detection. Our new VT5000 dataset is made publicly available at https://github.com/lz118/RGBT-Salient-Object-Detection. Zhengzheng Tu, Chenglong Li 0002, Jieming Xu |
IEEE Trans. Multim. | 4 |
| 2023 | Looking and Hearing Into Details: Dual-Enhanced Siamese Adversarial Network for Audio-Visual MatchingabstractAudio-visual cross-modal matching aims to explore the intrinsic correspondence between face images and audio clips. Existing methods usually focus on the salient features of identities between visual images and voice clips, while neglecting their subtle differences, which are crucial to distinguishing cross-modal samples. To deal with this problem, we propose a novel Dual-enhanced Siamese Adversarial Network (DSANet), which pursues the adversarial dual enhancement to highlight both salient and subtle features for robust audio-visual cross-modal matching. First, we designed a dual enhancement mechanism to enhance potential subtle features by randomly selecting a region feature for salient feature suppression, while enhancing salient features in the corresponding region to ensure the global discriminative ability. Second, to establish the correlation of subtle features in the process of eliminating cross-modal heterogeneity, we design a siamese adversarial structure to perform modal heterogeneity elimination for both enhanced salient and subtle features in a parallel manner. Moreover, we propose an adaptive masked cross-entropy loss to force the network to focus on the feature differences among hard classes. Experiments on public benchmark datasets validate the effectiveness of the proposed algorithm. Jiaxiang Wang 0001, Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Cross-Modal Object Tracking: Modality-Aware Representations and a Unified BenchmarkabstractIn many visual systems, visual tracking often bases on RGB image sequences, in which some targets are invalid in low-light conditions, and tracking performance is thus affected significantly. Introducing other modalities such as depth and infrared data is an effective way to handle imaging limitations of individual sources, but multi-modal imaging platforms usually require elaborate designs and cannot be applied in many real-world applications at present. Near-infrared (NIR) imaging becomes an essential part of many surveillance cameras, whose imaging is switchable between RGB and NIR based on the light intensity. These two modalities are heterogeneous with very different visual properties and thus bring big challenges for visual tracking. However, existing works have not studied this challenging problem. In this work, we address the cross-modal object tracking problem and contribute a new video dataset, including 654 cross-modal image sequences with over 481K frames in total, and the average video length is more than 735 frames. To promote the research and development of cross-modal object tracking, we propose a new algorithm, which learns the modality-aware target representation to mitigate the appearance gap between RGB and NIR modalities in the tracking process. It is plug-and-play and could thus be flexibly embedded into different tracking frameworks. Extensive experiments on the dataset are conducted, and we demonstrate the effectiveness of the proposed algorithm in two representative tracking frameworks against 19 state-of-the-art tracking methods. Dataset, code, model and results are available at https://github.com/mmic-lcl/source-code. Chenglong Li 0002, Tianhao Zhu, Lei Liu 0049, Xiaonan Si, Zilin Fan, Sulan Zhai |
AAAI | 1 |
| 2022 | Interact, Embed, and EnlargE: Boosting Modality-Specific Representations for Multi-Modal Person Re-identificationabstractMulti-modal person Re-ID introduces more complementary information to assist the traditional Re-ID task. Existing multi-modal methods ignore the importance of modality-specific information in the feature fusion stage. To this end, we propose a novel method to boost modality-specific representations for multi-modal person Re-ID: Interact, Embed, and EnlargE (IEEE). First, we propose a cross-modal interacting module to exchange useful information between different modalities in the feature extraction phase. Second, we propose a relation-based embedding module to enhance the richness of feature descriptors by embedding the global feature into the fine-grained local information. Finally, we propose multi-modal margin loss to force the network to learn modality-specific information for each modality by enlarging the intra-class discrepancy. Superior performance on multi-modal Re-ID dataset RGBNT201 and three constructed Re-ID datasets validate the effectiveness of the proposed method compared with the state-of-the-art approaches. Zi Wang 0013, Chenglong Li 0002, Aihua Zheng, Ran He 0001, Jin Tang 0001 |
AAAI | 2 |
| 2022 | Attribute-Based Progressive Fusion Network for RGBT TrackingabstractRGBT tracking usually suffers from various challenge factors, such as fast motion, scale variation, illumination variation, thermal crossover and occlusion, to name a few. Existing works often study fusion models to solve all challenges simultaneously, and it requires fusion models complex enough and training data large enough, which are usually difficult to be constructed in real-world scenarios. In this work, we disentangle the fusion process via the challenge attributes, and thus propose a novel Attribute-based Progressive Fusion Network (APFNet) to increase the fusion capacity with a small number of parameters while reducing the dependence on large-scale training data. In particular, we design five attribute-specific fusion branches to integrate RGB and thermal features under the challenges of thermal crossover, illumination variation, scale variation, occlusion and fast motion respectively. By disentangling the fusion process, we can use a small number of parameters for each branch to achieve robust fusion of different modalities and train each branch using the small training subset with the corresponding attribute annotation. Then, to adaptive fuse features of all branches, we design an aggregation fusion module based on SKNet. Finally, we also design an enhancement fusion transformer to strengthen the aggregated feature and modality-specific features. Experimental results on benchmark datasets demonstrate the effectiveness of our APFNet against other state-of-the-art methods. Yun Xiao 0003, Chenglong Li 0002, Lei Liu 0049, Jin Tang 0001 |
AAAI | 3 |
| 2022 | Fusion Tree Network for RGBT TrackingabstractRGBT tracking is often affected by complex scenes (i.e., occlusions, scale changes, noisy background, etc). Existing works usually adopt a single-strategy RGBT tracking fusion scheme to handle modality fusion in all scenarios. However, due to the limitation of fusion model capacity, it is difficult to fully integrate the discriminative features between different modalities. To tackle this problem, we propose a Fusion Tree Network (FTNet), which provides a multi-strategy fusion model with high capacity to efficiently fuse different modalities. Specifically, we combine three kinds of attention modules (i.e., channel attention, spatial attention, and location attention) in a tree structure to achieve multi-path hybrid attention in the deeper convolutional stages of the object tracking network. Extensive experiments are performed on three RGBT tracking datasets, and the results show that our method achieves superior performance among state-of-the-art RGBT tracking models. Zhiyuan Cheng 0012, Andong Lu, Zhang Zhang 0001, Chenglong Li 0002, Liang Wang 0001 |
AVSS | 4 |
| 2022 | The First Challenge on Moving Object Detection and Tracking in Satellite Videos: Methods and ResultsabstractIn this paper, we briefly summarize the first challenge on moving object detection and tracking in satellite videos (SatVideoDT). This challenge has three tracks related to satellite video analysis, including moving object detection (Track 1), single object tracking (Track 2), and multiple-object tracking (Track 3). 123, 89, and 70 participants successfully registered, while 37, 42, and 29 teams submitted their final results on the test datasets for Tracks 1-3, respectively. The top-performing methods and their results in each track are described with details. This challenge establishes a new benchmark for satellite video analysis. Yulan Guo, Qingyong Hu, Feng Zhang 0046, Ye Zhang 0037, Hanyun Wang, Chenguang Dai, Weilong Guo, Xiyu Qi, Kelong Tu, Shudan Zhu, Lai Chen, Bin Lin 0013, Chaocan Xue, Jinlei Zheng, Limei Qin, Ying Li 0017, Manqi Zhao, Lu Ruan 0003, Mingpeng Cui, Guanchen Ding, Guangwei Jiang, Zhenzhong Chen 0001, Kaiyang Cao, Lingyu Kong, Shaodong Chen, Zhicheng Zhao 0001, Qin Shen, Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003 |
ICPR | 35 |
| 2022 | Dynamic Collaboration Convolution for Robust RGBT TrackingabstractLearning powerful representation of individual modality is critical for RGBT tracking. Recent works mainly focus on utilizing multiple convolutions to model feature representations of each modality. However, they usually leverage static convolutions to extract features, which are hard to handle complex input data. To deal with this problem, we propose a dynamic collaboration convolution, named DC-Conv, including a set of static convolutions and a weight-router module, for robust RGBT tracking. In specific, we set four static convolutions to each modality in every layer to model each modality, and design a weight-router module to fuse these static convolutions using learned dynamic weights. Such a dynamic weighting scheme makes the convolutions can be adapted to the variations of input data, and thus greatly improves the tracking performance. In addition, we propose an effective progressive learning algorithm to maximize the role of each convolution to make it capture discriminative representations. We evaluate our method on two public RGBT tracking benchmarks, and the results demonstrate the effectiveness of our tracker against state-of-the-art methods. Andong Lu, Chenglong Li 0002, Yan Huang 0023, Liang Wang 0001 |
ICPR | 3 |
| 2022 | Progressive Attribute Embedding for Accurate Cross-modality Person Re-IDabstractAttributes are important information to bridge the appearance gap across modalities, but have not been well explored in cross-modality person ReID. This paper proposes a progressive attribute embedding module (PAE) to effectively fuse the fine-grained semantic attribute information and the global structural visual information. Through a novel cascade way, we use attribute information to learn the relationship between the person images in different modalities, which significantly relieves the modality heterogeneity. Meanwhile, by embedding attribute information to guide more discriminative image feature generation, it simultaneously reduces the inter-class similarity and the intra-class discrepancy. In addition, we propose an attribute-based auxiliary learning strategy (AAL) to supervise the network to learn modality-invariant and identity-specific local features by joint attribute and identity classification losses. The PAE and AAL are jointly optimized in an end-to-end framework, namely, progressive attribute embedding network (PAENet). One can plug PAE and AAL into current mainstream models, as we implement them in five cross-modality person ReID frameworks to further boost the performance. Extensive experiments on public datasets demonstrate the effectiveness of the proposed method against the state-of-the-art cross-modality person ReID methods. Aihua Zheng, Chenglong Li 0002, Bin Luo 0001, Ruoran Jia |
ACM Multimedia | 4 |
| 2022 | Efficient License Plate Recognition via Parallel Position-Aware Attention
Wenzhong Wang, Chenglong Li 0002, Jin Tang 0001 |
PRCV (3) | 3 |
| 2022 | EllipseIoU: A General Metric for Aerial Object Detection
Xinbo Yang, Chenglong Li 0002, Rui Ruan, Lei Liu 0049, Bin Luo 0001 |
PRCV (3) | 2 |
| 2022 | RGBT tracking via reliable feature configuration
Zhengzheng Tu, Wenli Pan, Yunsheng Duan, Jin Tang 0001, Chenglong Li 0002 |
Sci. China Inf. Sci. | 5 |
| 2022 | Joint Token and Feature Alignment Framework for Text-Based Person SearchabstractText-based person search is a challenging crossmodal retrieval task. Existing works reduce the inter-modality and intra-class gaps by aligning local features extracted from image and text modalities, which easily lead to mismatching problems due to the lack of annotation information. Besides, it is sub-optimal to reduce two gaps simultaneously in the same feature space. This work proposes a novel joint token and feature alignment framework to reduce the inter-modality and intraclass gaps progressively. Specifically, we first build a dual-path feature learning network to extract features and conduct feature alignment to reduce the inter-modality gap. Second, we design a text generation module to generate token sequences using visual features, and then token alignment is performed to reduce the intra-class gap. Last, a fusion interaction module is introduced to further eliminate the modality heterogeneity using the strategy of multi-stage feature fusion. Extensive experiments on the CUHKPEDES dataset demonstrate the effectiveness of our model, which significantly outperforms previous state-of-the-art methods. Shangze Li, Andong Lu, Yan Huang 0008, Chenglong Li 0002, Liang Wang 0001 |
IEEE Signal Process. Lett. | 4 |
| 2022 | RGBT Tracking by Trident Fusion NetworkabstractIn recent years, RGBT tracking has become a hot topic in the field of visual tracking, and made great progress. In this paper, we propose a novel Trident Fusion Network (TFNet) to achieve effective fusion of different modalities for robust RGBT tracking. In specific, to deploy the complementarity of features of all convolutional layers, we propose a recursive strategy to densely aggregate these features that yield robust representations of target objects in two modalities. Moreover, we design a trident architecture to integrate the fused features and both modality-specific features for robust target representations. There are three main advantages. First, retaining the classification layer of each modality is beneficial to enhance feature learning of single modality, and compared with aggregate branches, single-modality branches pay more attention to the mining of modal specific information. Second, when some modality is noisy or invalid, the modality-specific branches would capture more discriminative features for RGBT tracking. Finally, the integration of aggregation branches and single-modality branches is beneficial to the complementary learning of different modalities. In addition, we also introduce a feature pruning module in each branch to prune the redundant features and avoid network overfitting. Experimental results on four RGBT tracking benchmark datasets suggest that our tracker achieves superior performance against the state-of-the-art RGBT tracking methods. Yabin Zhu, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | ORSI Salient Object Detection via Multiscale Joint Region and Boundary ModelabstractSalient object detection (SOD) in optical remote sense images (ORSIs) is a valuable and challenging task. The factors in ORSI, such as background clutter, lighting shadows, imaging blur, and low resolution, significantly degrade the completeness and accuracy of salient objects. To handle this problem, we propose a novel model to learn robust multiscale region features of salient objects by simultaneously optimizing their boundaries. First, we extract multiscale region features of salient objects through a hierarchical attention module. Second, we generate the boundary features by combining the local cues and the global information generated by pyramid pooling. Finally, we embed the boundary features into region features at multiple scales. In particular, we design a joint learning scheme based on a bidirectional feature transformation to optimize boundary and region features simultaneously for accurate ORSI SOD. To provide a comprehensive evaluation platform, we construct a new dataset called ORSI-4199 for ORSI SOD. It contains 4199 finely annotated image pairs with diverse scenes, in which nine attributes (i.e., challenge types) are annotated to facilitate analyzing the strengths and weaknesses of SOD models from different perspectives. Extensive experiments on the public dataset ORSSD, EORRSD, and the newly created dataset ORSI-4199 show that the proposed approach achieves promising results against state-of-the-art methods.https://github.com/wchao1213/ORSI-SOD. Zhengzheng Tu, Chenglong Li 0002, Minghao Fan, Haifeng Zhao 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Category-Wise Fusion and Enhancement Learning for Multimodal Remote Sensing Image Semantic SegmentationabstractThis paper presents a simple yet effective method called Category-wise Fusion and Enhancement learning (CaFE), which leverages the category priors to achieve effective feature fusion and imbalance learning, for multi-modal remote sensing image semantic segmentation. In particular, we disentangle the feature fusion process via the categories to achieve the category-wise fusion based on the fact that the feature fusion in the same category regions tends to have similar characteristics. The disentangled fusion would also increase the fusion capacity with a small number of parameters while reducing the dependence on large-scale training data. For the sample imbalance problem, we design a simple yet effective category-wise enhancement learning scheme. In particular, we assign the weight for each category region based on the proportion of samples in this region over the whole image. By this way, the learning algorithm would focus more on the regions with smaller proportion. Note that both category-wise feature fusion and imbalance learning are only performed in the training stage, and the segmentation efficiency is thus not affected. Experimental results on two benchmark datasets demonstrate the effectiveness of our CaFE against other state-of-the-art methods. Aihua Zheng, Jinbo He, Chenglong Li 0002, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Entropy Guided Adversarial Domain Adaptation for Aerial Image Semantic SegmentationabstractRecent advances on aerial image semantic segmentation mainly employ the domain adaption to transfer knowledge from the source domain to the target domain. Despite the remarkable achievement, most methods focus on the global marginal distribution alignment to reduce the domain shift between source and target domains, leading to a wrong mapping of the well-aligned features. In this article, we propose an effective unsupervised domain adaptation approach, which relies on a novel entropy guided adversarial learning algorithm, for aerial image semantic segmentation. In specific, we perform local feature alignment between domains by learning a self-adaptive weight from the target prediction probability map to measure the interdomain discrepancy. To exploit the meaningful structure information among semantic regions, we propose to utilize the graph convolutions for long-range semantic reasoning. Comprehensive experimental results on the benchmark dataset of aerial image semantic segmentation and natural scenes demonstrate the superior performance of the proposed method compared to the state-of-the-art methods. Aihua Zheng, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Attribute and State Guided Structural Embedding Network for Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID) is a crucial task in smart city and intelligent transportation, aiming to match vehicle images across non-overlapping surveillance camera scenarios. However, the images of different vehicles may have small visual discrepancies when they have the same/similar attributes, e.g., the same/similar color, type, and manufacturer. Meanwhile, the images from a vehicle may have large visual discrepancies with different states, e.g., different camera views, vehicle viewpoints, and capture time. In this paper, we propose an attribute and state guided structural embedding network (ASSEN) to achieve discriminative feature learning by attribute-based enhancement and state-based weakening for vehicle Re-ID. First, we propose an attribute-based enhancement and expanding module to enhance the discrimination of vehicle features through identity-related attribute information, and we design an attribute-based expanding loss to increase the feature gap between different vehicles. Second, we design a state-based weakening and shrinking module, which not only weakens the state information that interferes with identification but also reduces the intra-class feature gap by a state-based shrinking loss. Third, we propose a global structural embedding module that exploits the attribute information and state information to explore hierarchical relationships between vehicle features, then we use these relationships for feature embedding to learn more robust vehicle features. Extensive experiments on benchmark datasets VeRi-776, VehicleID, and VERI-Wild demonstrate the superior performance and generalization of the proposed method against state-of-the-art vehicle Re-ID methods. The code is available at https://github.com/ttaalle/fast_assen. Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | LasHeR: A Large-Scale High-Diversity Benchmark for RGBT TrackingabstractRGBT tracking receives a surge of interest in the computer vision community, but this research field lacks a large-scale and high-diversity benchmark dataset, which is essential for both the training of deep RGBT trackers and the comprehensive evaluation of RGBT tracking methods. To this end, we present a La rge- s cale H igh-diversity [Formula: see text]nchmark for short-term R GBT tracking (LasHeR) in this work. LasHeR consists of 1224 visible and thermal infrared video pairs with more than 730K frame pairs in total. Each frame pair is spatially aligned and manually annotated with a bounding box, making the dataset well and densely annotated. LasHeR is highly diverse capturing from a broad range of object categories, camera viewpoints, scene complexities and environmental factors across seasons, weathers, day and night. We conduct a comprehensive performance evaluation of 12 RGBT tracking algorithms on the LasHeR dataset and present detailed analysis. In addition, we release the unaligned version of LasHeR to attract the research interest for alignment-free RGBT tracking, which is a more practical task in real-world applications. The datasets and evaluation protocols are available at: https://github.com/mmic-lcl/Datasets-and-benchmark-code. Chenglong Li 0002, Wanlin Xue, Yaqing Jia, Zhichen Qu, Bin Luo 0001, Jin Tang 0001, Dengdi Sun |
IEEE Trans. Image Process. | 1 |
| 2022 | Weakly Alignment-Free RGBT Salient Object Detection With Deep Correlation NetworkabstractRGBT Salient Object Detection (SOD) focuses on common salient regions of a pair of visible and thermal infrared images. Existing methods perform on the well-aligned RGBT image pairs, but the captured image pairs are always unaligned and aligning them requires much labor cost. To handle this problem, we propose a novel deep correlation network (DCNet), which explores the correlations across RGB and thermal modalities, for weakly alignment-free RGBT SOD. In particular, DCNet includes a modality alignment module based on the spatial affine transformation, the feature-wise affine transformation and the dynamic convolution to model the strong correlation of two modalities. Moreover, we propose a novel bi-directional decoder model, which combines the coarse-to-fine and fine-to-coarse processes for better feature enhancement. In particular, we design a modality correlation ConvLSTM by adding the first two components of modality alignment module and a global context reinforcement module into ConvLSTM, which is used to decode hierarchical features in both top-down and button-up manners. Extensive experiments on three public benchmark datasets show the remarkable performance of our method against state-of-the-art methods. Zhengzheng Tu, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | M5L: Multi-Modal Multi-Margin Metric Learning for RGBT TrackingabstractClassifying hard samples in the course of RGBT tracking is a quite challenging problem. Existing methods only focus on enlarging the boundary between positive and negative samples, but ignore the relations of multilevel hard samples, which are crucial for the robustness of hard sample classification. To handle this problem, we propose a novel Multi-Modal Multi-Margin Metric Learning framework named M5L for RGBT tracking. In particular, we divided all samples into four parts including normal positive, normal negative, hard positive and hard negative ones, and aim to leverage their relations to improve the robustness of feature embeddings, e.g., normal positive samples are closer to the ground truth than hard positive ones. To this end, we design a multi-modal multi-margin structural loss to preserve the relations of multilevel hard samples in the training stage. In addition, we introduce an attention-based fusion module to achieve quality-aware integration of different source data. Extensive experiments on large-scale datasets testify that our framework clearly improves the tracking performance and performs favorably the state-of-the-art RGBT trackers. Zhengzheng Tu, Chun Lin, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | MsKAT: Multi-Scale Knowledge-Aware Transformer for Vehicle Re-IdentificationabstractExisting vehicle re-identification (Re-ID) methods usually suffer from intra-instance discrepancy and inter-instance similarity. The key to solving this problem lies in filtering out identity-irrelevant interference and collecting identity-relevant vehicle details. In this paper, we aim to design a robust vehicle Re-ID framework that trains a model guided by knowledge vectors yet is able to disentangle the identity-relevant features and identity-irrelevant features. Toward this end, we propose a novel Multi-scale Knowledge-Aware Transformer (MsKAT) to build a knowledge-guided multi-scale feature alignment framework. First, we construct a Knowledge-Aware Transformer (KAT) to interact with semantic knowledge and visual feature. KAT mainly includes State elimination Transformer (SeT) to eliminate state (camera, viewpoint) interference and Attribute aggregation Transformer (AaT) to gather attribute (color, type) information. Second, to learn the knowledge-guided sample differences, we propose to encourage the separation of identity-relevant features and identity-irrelevant features by a Knowledge-Guided Alignment loss ($\mathcal {L}_{KGA}$). Specifically,$\mathcal {L}_{KGA}$suppresses the difference between knowledge-guided positive pairs and the similarity between knowledge-guided negative pairs. Third, with the multi-scale settings of KAT and$\mathcal {L}_{KGA}$, our model can capture knowledge-guided visual consistency features at different scales. Extensive evidence demonstrates our approach achieves new state-of-the-art on three widely-used vehicle re-identification benchmarks. Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Viewpoint-Aware Progressive Clustering for Unsupervised Vehicle Re-IdentificationabstractVehicle re-identification (Re-ID) is an active task due to its importance in large-scale intelligent monitoring in smart cities. Despite the rapid progress in recent years, most existing methods handle vehicle Re-ID task in a supervised manner, which is both time and labor-consuming and limits their application to real-life scenarios. Recently, unsupervised person Re-ID methods achieve impressive performance by exploring domain adaption or clustering-based techniques. However, one cannot directly generalize these methods to vehicle Re-ID since vehicle images present huge appearance variations in different viewpoints. To handle this problem, we propose a novel viewpoint-aware clustering algorithm for unsupervised vehicle Re-ID. In particular, we first divide the entire feature space into different subspaces according to the predicted viewpoints and then perform a progressive clustering to mine the accurate relationship among samples. Comprehensive experiments against the state-of-the-art methods on two multi-viewpoint benchmark datasets VeRi-776 and VeRi-Wild validate the promising performance of the proposed method in both with and without domain adaption scenarios while handling unsupervised vehicle Re-ID. Aihua Zheng, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Multimodal Cross-Layer Bilinear Pooling for RGBT TrackingabstractHierarchical deep features can provide multilevel abstractions of target objects, which play an important role in target localization and classification. Determining how to effectively aggregate abstract information from different levels in RGB and thermal modalities is the key to exploiting their complementary advantages for robust RGBT tracking. However, existing RGBT tracking algorithms either focus on the semantic information of the last layer or aggregate hierarchical deep features from each modal using simple operations (e.g., summation and concatenation), which limit the capability of the multimodal tracker. To address these issues, in this paper, we propose a novel multimodal cross-layer bilinear pooling network for RGBT tracking. In our network, firstly, to boost the performance of the tracker, we use a channel attention mechanism to implement the adaptive calibration of feature channels for all convolutional layer features before realizing hierarchical feature fusion. Then, a bilinear pooling operation is performed on any two layers through the cross product, which is a second-order computation that effectively aggregates the deep semantic and shallow texture information of the target. Finally, a quality-aware fusion module is designed to aggregate the bilinear pooling features of different layer interactions between different modalities in an adaptive manner. The results of a large number of experiments on two public benchmark datasets demonstrate the effectiveness of our tracker compared with other state-of-the-art tracking methods. Yiming Mei, Jinpei Liu, Chenglong Li 0002 |
IEEE Trans. Multim. | 4 |
| 2022 | RGBT Tracking via Noise-Robust Cross-Modal RankingabstractExisting RGBT tracking methods usually localize a target object with a bounding box, in which the trackers are often affected by the inclusion of background clutter. To address this issue, this article presents a novel algorithm, called noise-robust cross-modal ranking, to suppress background effects in target bounding boxes for RGBT tracking. In particular, we handle the noise interference in cross-modal fusion and seed labels from the following two aspects. First, the soft cross-modality consistency is proposed to allow the sparse inconsistency in fusing different modalities, aiming to take both collaboration and heterogeneity of different modalities into account for more effective fusion. Second, the optimal seed learning is designed to handle label noises of ranking seeds caused by some problems, such as irregular object shape and occlusion. In addition, to deploy the complementarity and maintain the structural information of different features within each modality, we perform an individual ranking for each feature and employ a cross-feature consistency to pursue their collaboration. A unified optimization framework with an efficient convergence speed is developed to solve the proposed model. Extensive experiments demonstrate the effectiveness and efficiency of the proposed approach comparing with state-of-the-art tracking methods on GTOT and RGBT234 benchmark data sets. Chenglong Li 0002, Zhiqiang Xiang, Jin Tang 0001, Bin Luo 0001, Futian Wang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Robust Multi-Modality Person Re-identificationabstractTo avoid the illumination limitation in visible person re-identification (Re-ID) and the heterogeneous issue in cross-modality Re-ID, we propose to utilize complementary advantages of multiple modalities including visible (RGB), near infrared (NI) and thermal infrared (TI) ones for robust person Re-ID. A novel progressive fusion network is designed to learn effective multi-modal features from single to multiple modalities and from local to global views. Our method works well in diversely challenging scenarios even in the presence of missing modalities. Moreover, we contribute a comprehensive benchmark dataset, RGBNT201, including 201 identities captured from various challenging conditions, to facilitate the research of RGB-NI-TI multi-modality person Re-ID. Comprehensive experiments on RGBNT201 dataset comparing to the state-of-the-art methods demonstrate the contribution of multi-modality person Re-ID and the effectiveness of the proposed approach, which launch a new benchmark and a new baseline for multi-modality person Re-ID. Aihua Zheng, Zi Wang 0013, Zi-Han Chen, Chenglong Li 0002, Jin Tang 0001 |
AAAI | 4 |
| 2021 | Cross-Modal Collaborative Representation Learning and a Large-Scale RGBT Benchmark for Crowd CountingabstractCrowd counting is a fundamental yet challenging task, which desires rich information to generate pixel-wise crowd density maps. However, most previous methods only used the limited information of RGB images and cannot well discover potential pedestrians in unconstrained scenarios. In this work, we find that incorporating optical and thermal information can greatly help to recognize pedestrians. To promote future researches in this field, we introduce a large-scale RGBT Crowd Counting (RGBT-CC) benchmark, which contains 2,030 pairs of RGB-thermal images with 138,389 annotated people. Furthermore, to facilitate the multimodal crowd counting, we propose a cross-modal collaborative representation learning framework, which consists of multiple modality-specific branches, a modality-shared branch, and an Information Aggregation-Distribution Module (IADM) to capture the complementary information of different modalities fully. Specifically, our IADM incorporates two collaborative information transfers to dynamically enhance the modality-shared and modality-specific representations with a dual information propagation mechanism. Extensive experiments conducted on the RGBT-CC benchmark demonstrate the effectiveness of our framework for RGBT crowd counting. Moreover, the proposed approach is universal for multimodal crowd counting and is also capable to achieve superior performance on the ShanghaiTechRGBD [22] dataset. Finally, our source code and benchmark have been released at http://lingboliu.com/RGBT_Crowd_Counting.html. Lingbo Liu, Hefeng Wu, Guanbin Li, Chenglong Li 0002, Liang Lin 0004 |
CVPR | 5 |
| 2021 | Progressive Fusion Network for Safety Protection Detection
Futian Wang, Lugang Wang, Jin Tang 0001, Chenglong Li 0002 |
ICIG (1) | 4 |
| 2021 | Joint Learning Appearance and Motion Models for Visual Tracking
Wenmei Xu, Hongyuan Yu, Wei Wang 0115, Chenglong Li 0002, Liang Wang 0001 |
PRCV (1) | 4 |
| 2021 | Learning spatio-temporal correlation filter for visual tracking
Youmin Yan, Xixian Guo, Jin Tang 0001, Chenglong Li 0002, Xin Wang 0013 |
Neurocomputing | 4 |
| 2021 | RGBT tracking via cross-modality message passing
Xiao Wang 0014, Chenglong Li 0002, Jinmin Hu, Jin Tang 0001 |
Neurocomputing | 3 |
| 2021 | Edge-Guided Non-Local Fully Convolutional Network for Salient Object DetectionabstractFully Convolutional Neural Network (FCN) has been widely applied to salient object detection recently by virtue of high-level semantic feature extraction, but existing FCN-based methods still suffer from continuous striding and pooling operations leading to loss of spatial structure and blurred edges. To maintain the clear edge structure of salient objects, we propose a novel Edge-guided Non-local FCN (ENFNet) to perform edge-guided feature learning for accurate salient object detection. In a specific, we extract hierarchical global and local information in FCN to incorporate non-local features for effective feature representations. To preserve good boundaries of salient objects, we propose a guidance block to embed edge prior knowledge into hierarchical feature maps. The guidance block not only performs feature-wise manipulation but also spatial-wise transformation for effective edge embeddings. Our model is trained on the MSRA-B dataset and tested on five popular benchmark datasets. Comparing with the state-of-the-art methods, the proposed method performance well on five datasets. Zhengzheng Tu, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | RGBT Tracking via Multi-Adapter Network with Hierarchical Divergence LossabstractRGBT tracking has attracted increasing attention since RGB and thermal infrared data have strong complementary advantages, which could make trackers all-day and all-weather work. Existing works usually focus on extracting modality-shared or modality-specific information, but the potentials of these two cues are not well explored and exploited in RGBT tracking. In this paper, we propose a novel multi-adapter network to jointly perform modality-shared, modality-specific and instance-aware target representation learning for RGBT tracking. To this end, we design three kinds of adapters within an end-to-end deep learning framework. In specific, we use the modified VGG-M as the generality adapter to extract the modality-shared target representations. To extract the modality-specific features while reducing the computational complexity, we design a modality adapter, which adds a small block to the generality adapter in each layer and each modality in a parallel manner. Such a design could learn multilevel modality-specific representations with a modest number of parameters as the vast majority of parameters are shared with the generality adapter. We also design instance adapter to capture the appearance properties and temporal variations of a certain target. Moreover, to enhance the shared and specific features, we employ the loss of multiple kernel maximum mean discrepancy to measure the distribution divergence of different modal features and integrate it into each layer for more robust representation learning. Extensive experiments on two RGBT tracking benchmark datasets demonstrate the outstanding performance of the proposed tracker against the state-of-the-art methods. Andong Lu, Chenglong Li 0002, Yuqing Yan, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Multi-Interactive Dual-Decoder for RGB-Thermal Salient Object DetectionabstractRGB-thermal salient object detection (SOD) aims to segment the common prominent regions of visible image and corresponding thermal infrared image that we call it RGBT SOD. Existing methods don't fully explore and exploit the potentials of complementarity of different modalities and multi-type cues of image contents, which play a vital role in achieving accurate results. In this paper, we propose a multi-interactive dual-decoder to mine and model the multi-type interactions for accurate RGBT SOD. In specific, we first encode two modalities into multi-level multi-modal feature representations. Then, we design a novel dual-decoder to conduct the interactions of multi-level features, two modalities and global contexts. With these interactions, our method works well in diversely challenging scenarios even in the presence of invalid modality. Finally, we carry out extensive experiments on public RGBT and RGBD SOD datasets, and the results show that the proposed method achieves the outstanding performance against state-of-the-art algorithms. The source code has been released at: https://github.com/lz118/Multi-interactive-Dual-decoder. Zhengzheng Tu, Chenglong Li 0002, Yang Lang, Jin Tang 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Segmenting Objects in Day and Night: Edge-Conditioned CNN for Thermal Image Semantic SegmentationabstractDespite much research progress in image semantic segmentation, it remains challenging under adverse environmental conditions caused by imaging limitations of the visible spectrum, while thermal infrared cameras have several advantages over cameras for the visible spectrum, such as operating in total darkness, insensitive to illumination variations, robust to shadow effects, and strong ability to penetrate haze and smog. These advantages of thermal infrared cameras make the segmentation of semantic objects in day and night. In this article, we propose a novel network architecture, called edge-conditioned convolutional neural network (EC-CNN), for thermal image semantic segmentation. Particularly, we elaborately design a gated featurewise transform layer in EC-CNN to adaptively incorporate edge prior knowledge. The whole EC-CNN is end-to-end trained and can generate high-quality segmentation results with edge guidance. Meanwhile, we also introduce a new benchmark data set named "Segmenting Objects in Day And night" (SODA) for comprehensive evaluations in thermal image semantic segmentation. SODA contains over 7168 manually annotated and synthetically generated thermal images with 20 semantic region labels and from a broad range of viewpoints and scene complexities. Extensive experiments on SODA demonstrate the effectiveness of the proposed EC-CNN against state-of-the-art methods. Chenglong Li 0002, Yan Yan 0002, Bin Luo 0001, Jin Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Multi-Spectral Vehicle Re-Identification: A ChallengeabstractVehicle re-identification (Re-ID) is a crucial task in smart city and intelligent transportation, aiming to match vehicle images across non-overlapping surveillance camera views. Currently, most works focus on RGB-based vehicle Re-ID, which limits its capability of real-life applications in adverse environments such as dark environments and bad weathers. IR (Infrared) spectrum imaging offers complementary information to relieve the illumination issue in computer vision tasks. Furthermore, vehicle Re-ID suffers a big challenge of the diverse appearance with different views, such as trucks. In this work, we address the RGB and IR vehicle Re-ID problem and contribute a multi-spectral vehicle Re-ID benchmark named RGBN300, including RGB and NIR (Near Infrared) vehicle images of 300 identities from 8 camera views, giving in total 50125 RGB images and 50125 NIR images respectively. In addition, we have acquired additional TIR (Thermal Infrared) data for 100 vehicles from RGBN300 to form another dataset for three-spectral vehicle Re-ID. Furthermore, we propose a Heterogeneity-collaboration Aware Multi-stream convolutional Network (HAMNet) towards automatically fusing different spectrum features in an end-to-end learning framework. Comprehensive experiments on prevalent networks show that our HAMNet can effectively integrate multi-spectral data for robust vehicle Re-ID in day and night. Our work provides a benchmark dataset for RGB-NIR and RGB-NIR-TIR multi-spectral vehicle Re-ID and a baseline network for both research and industrial communities. The dataset and baseline codes are available at: https://github.com/ttaalle/multi-modal-vehicle-Re-ID. Chenglong Li 0002, Xianpeng Zhu, Aihua Zheng, Bin Luo 0001 |
AAAI | 2 |
| 2020 | Challenge-Aware RGBT Tracking
Chenglong Li 0002, Lei Liu 0049, Andong Lu, Jin Tang 0001 |
ECCV (22) | 1 |
| 2020 | LSOTB-TIR: A Large-Scale High-Diversity Thermal Infrared Object Tracking BenchmarkabstractIn this paper, we present a Large-Scale and high-diversity general Thermal InfraRed (TIR) Object Tracking Benchmark, called LSOTB-TIR, which consists of an evaluation dataset and a training dataset with a total of 1,400 TIR sequences and more than 600K frames. We annotate the bounding box of objects in every frame of all sequences and generate over 730K bounding boxes in total. To the best of our knowledge, LSOTB-TIR is the largest and most diverse TIR object tracking benchmark to date. To evaluate a tracker on different attributes, we define 4 scenario attributes and 12 challenge attributes in the evaluation dataset. By releasing LSOTB-TIR, we encourage the community to develop deep learning based TIR trackers and evaluate them fairly and comprehensively. We evaluate and analyze more than 30 trackers on LSOTB-TIR to provide a series of baselines, and the results show that deep trackers achieve promising performance. Furthermore, we re-train several representative deep trackers on LSOTB-TIR, and their results demonstrate that the proposed training dataset significantly improves the performance of deep TIR trackers. Codes and dataset are available at https://github.com/QiaoLiuHit/LSOTB-TIR. Qiao Liu 0001, Xin Li 0034, Zhenyu He 0001, Chenglong Li 0002, Zikun Zhou, Di Yuan 0002, Jing Li 0071, Kai Yang 0018, Nana Fan, Feng Zheng 0001 |
ACM Multimedia | 4 |
| 2020 | Synthesizing Large-Scale Datasets for License Plate Detection and Recognition in the Wild
Chaochen Wang, Wenzhong Wang, Chenglong Li 0002, Jin Tang 0001 |
PRCV (3) | 3 |
| 2020 | Multi-modal foreground detection via inter- and intra-modality-consistent low-rank separation
Aihua Zheng, Naipeng Ye, Chenglong Li 0002, Xiao Wang 0014, Jin Tang 0001 |
Neurocomputing | 3 |
| 2020 | RGBT Salient Object Detection: Benchmark and A Novel Cooperative Ranking ApproachabstractDespite significant progress, image saliency detection still remains a challenging task in complex scenes and environments. Integrating multiple different but complementary cues, like RGB and Thermal infrared (RGBT), may be an effective way for boosting saliency detection performance. This work contributes a RGBT image dataset, which includes 821 spatially aligned RGBT image pairs and their ground truth annotations for saliency detection purpose. Moreover, 11 challenges are annotated on these image pairs for performing the challenge-sensitive analysis and 3 kinds of baseline methods are implemented to provide a comprehensive comparison platform. With this benchmark, we propose a novel approach based on a cooperative ranking algorithm for RGBT saliency detection. In particular, we introduce a weight for each modality to describe the reliability and a ℓ1-based cross-modal consistency in a unified ranking model, and design an efficient solver to iteratively optimize several subproblems with closed-form solutions. Extensive experiments against baseline methods demonstrate the effectiveness of the proposed approach on both our introduced dataset and a public dataset. Jin Tang 0001, Dongzhe Fan, Xiaoxiao Wang 0003, Zhengzheng Tu, Chenglong Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | RGB-T Image Saliency Detection via Collaborative Graph LearningabstractImage saliency detection is an active research topic in the community of computer vision and multimedia. Fusing complementary RGB and thermal infrared data has been proven to be effective for image saliency detection. In this paper, we propose an effective approach for RGB-T image saliency detection. Our approach relies on a novel collaborative graph learning algorithm. In particular, we take superpixels as graph nodes, and collaboratively use hierarchical deep features to jointly learn graph affinity and node saliency in a unified optimization framework. Moreover, we contribute a more challenging dataset for the purpose of RGB-T image saliency detection, which contains 1000 spatially aligned RGB-T image pairs and their ground truth annotations. Extensive experiments on the public dataset and the newly created dataset suggest that the proposed approach performs favorably against the state-of-the-art RGB-T saliency detection methods. Zhengzheng Tu, Chenglong Li 0002, Xiaoxiao Wang 0003, Jin Tang 0001 |
IEEE Trans. Multim. | 3 |
| 2020 | A Subspace Learning Approach to Multishot Person ReidentificationabstractThis paper addresses the challenging problem of multishot person reidentification (Re-ID) in real world uncontrolled surveillance systems. A key issue is how to effectively represent and process the multiple data with various appearance information due to the variations of pose, occlusions, and viewpoints. To this end, this paper develops a novel subspace learning approach, which pursues regularized low-rank and sparse representation for multishot person Re-ID. For the images of a person crossing a certain camera, we assume that the appearances of those subset images with similar viewpoints against a camera draw from the same low-rank subspace, and all the images of a person under a camera lie on a union of low-rank subspaces. Based on this assumption, we propose to learn a nonnegative low-rank and sparse graph to represent the person images. Moreover, the recurring pattern prior is integrated into our model to refine the affinities among images. Extensive experiments on four public benchmark datasets yield impressive performance by improving 22.9% on imagery library for intelligent detection systems video re identification (iLIDS-VID), 42.4% on person RE-ID (PRID) dataset 2011, 39.7% and 30.6% on speech, audio, image, and video technology-SoftBio camera 3/8 and camera 5/8, respectively, and 1.6% on motion analysis and re identification set compared to the state-of-the-art methods. Aihua Zheng, Xuehan Zhang, Bo Jiang 0002, Bin Luo 0001, Chenglong Li 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2019 | Visual Tracking Via Siamese Network With Global SimilarityabstractVisual tracking is a very important and challenging problem in the field of computer vision. In recent years, Siamese networks have been widely used for visual tracking due to their fast tracking speed, but many trackers based on Siamese network train their networks by utilizing either pairwise loss or triplet loss, which easily leads to over-fitting. In addition, it is difficult to distinguish some hard samples in the training samples. In this paper, we propose a novel global similarity loss to train the network. Specifically, we utilize two Gaussian distributions to simulate and optimize the distribution of positive and negative samples in the train set and add constraint on the hard samples. In experiments, without any other modification, we apply the proposed method to the Siamese network. And the results on several popular tracking benchmarks show our method achieves superior tracking performance than the baseline. Chao Fan 0001, Chenglong Li 0002, Jin Tang 0001 |
ICIP | 3 |
| 2019 | Learning Target-Oriented Dual Attention for Robust RGB-T TrackingabstractRGB-Thermal object tracking attempts to locate target object using complementary visual and thermal infrared data. Existing RGB-T trackers fuse different modalities by robust feature representation learning or adaptive modal weighting. However, how to integrate dual attention mechanism for visual tracking is still a subject that has not been studied yet. In this paper, we propose two visual attention mechanisms for robust RGB-T object tracking. Specifically, the local attention is implemented by exploiting the common visual attention of RGB and thermal data to train deep classifiers. We also introduce the global attention, which is a multimodal target-driven attention estimation network. It can provide global proposals for the classifier together with local proposals extracted from previous tracking result. Extensive experiments on two RGB-T benchmark datasets validated the effectiveness of our proposed algorithm. Yabin Zhu, Xiao Wang 0014, Chenglong Li 0002, Jin Tang 0001 |
ICIP | 4 |
| 2019 | Dense Feature Aggregation and Pruning for RGBT TrackingabstractHow to perform effective information fusion of different modalities is a core factor in boosting the performance of RGBT tracking. This paper presents a novel deep fusion algorithm based on the representations from an end-to-end trained convolutional neural network. To deploy the complementarity of features of all layers, we propose a recursive strategy to densely aggregate these features that yield robust representations of target objects in each modality. In different modalities, we propose to prune the densely aggregated features of all modalities in a collaborative way. In a specific, we employ the operations of global average pooling and weighted random selection to perform channel scoring and selection, which could remove redundant and noisy features to achieve more robust feature representation. Experimental results on two RGBT tracking benchmark datasets suggest that our tracker achieves clear state-of-the-art against other RGB and RGBT tracking methods. Yabin Zhu, Chenglong Li 0002, Bin Luo 0001, Jin Tang 0001, Xiao Wang 0014 |
ACM Multimedia | 2 |
| 2019 | Robust visual tracking via Laplacian Regularized Random Walk Ranking
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Chenglong Li 0002 |
Neurocomputing | 5 |
| 2019 | Background subtraction with multi-scale structured low-rank and sparse factorization
Aihua Zheng, Tian Zou, Yumiao Zhao, Bo Jiang 0002, Jin Tang 0001, Chenglong Li 0002 |
Neurocomputing | 6 |
| 2019 | Visual Tracking via Dynamic Graph LearningabstractExisting visual tracking methods usually localize a target object with a bounding box, in which the performance of the foreground object trackers or detectors is often affected by the inclusion of background clutter. To handle this problem, we learn a patch-based graph representation for visual tracking. The tracked object is modeled by with a graph by taking a set of non-overlapping image patches as nodes, in which the weight of each node indicates how likely it belongs to the foreground and edges are weighted for indicating the appearance compatibility of two neighboring nodes. This graph is dynamically learned and applied in object tracking and model updating. During the tracking process, the proposed algorithm performs three main steps in each frame. First, the graph is initialized by assigning binary weights of some image patches to indicate the object and background patches according to the predicted bounding box. Second, the graph is optimized to refine the patch weights by using a novel alternating direction method of multipliers. Third, the object feature representation is updated by imposing the weights of patches on the extracted image features. The object location is predicted by maximizing the classification score in the structured support vector machine. Extensive experiments show that the proposed tracking algorithm performs well against the state-of-the-art methods on large-scale benchmark datasets. Chenglong Li 0002, Liang Lin 0004, Wangmeng Zuo, Jin Tang 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | RGB-T object tracking: Benchmark and baseline
Chenglong Li 0002, Xinyan Liang, Yijuan Lu, Jin Tang 0001 |
Pattern Recognit. | 1 |
| 2019 | Quality-aware dual-modal saliency detection via deep reinforcement learning
Xiao Wang 0014, Chenglong Li 0002, Bin Luo 0001, Jin Tang 0001 |
Signal Process. Image Commun. | 4 |
| 2019 | Learning Local-Global Multi-Graph Descriptors for RGB-T Object TrackingabstractRGB-thermal (RGB-T) object tracking, which has attracted much recent attention, uses thermal infrared information to assist object tracking with visible light information. However, it still faces many challenging problems, especially the background inclusion in the target bounding box which easily results in model drifting. To handle this problem, we propose a novel and general approach to learn a local-global multi-graph descriptor to suppress background effects for RGB-T tracking. Our approach relies on a novel graph learning algorithm. First, the object is represented with multiple graphs, with a set of multi-modal image patches as nodes, for the robustness to prevent deformation and partial occlusion. Second, we dynamically learn a joint graph over time with both local and global considerations using spatial smoothness and low-rank representation. In particular, we design a single unified alternating direction method of multipliers-based optimization framework to learn graph structure, edge weights, and node weights simultaneously. Third, we combine multi-graph information with corresponding graph node weights to form a robust object descriptor, and tracking is finally carried out by adopting the structured support vector machine. Extensive experiments conducted on the tracking benchmark data sets demonstrate the effectiveness of the proposed approach against the state-of-the-art RGB-T trackers. Chenglong Li 0002, Chengli Zhu, Justin Jian Zhang, Bin Luo 0001, Xiaohao Wu, Jin Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | SINT++: Robust Visual Tracking via Adversarial Positive Instance GenerationabstractExisting visual trackers are easily disturbed by occlusion, blur and large deformation. We think the performance of existing visual trackers may be limited due to the following issues: i) Adopting the dense sampling strategy to generate positive examples will make them less diverse; ii) The training data with different challenging factors are limited, even through collecting large training dataset. Collecting even larger training dataset is the most intuitive paradigm, but it may still can not cover all situations and the positive samples are still monotonous. In this paper, we propose to generate hard positive samples via adversarial learning for visual tracking. Specifically speaking, we assume the target objects all lie on a manifold, hence, we introduce the positive samples generation network (PSGN) to sampling massive diverse training data through traversing over the constructed target object manifold. The generated diverse target object images can enrich the training dataset and enhance the robustness of visual trackers. To make the tracker more robust to occlusion, we adopt the hard positive transformation network (HPTN) which can generate hard samples for tracking algorithm to recognize. We train this network with deep reinforcement learning to automatically occlude the target object with a negative patch. Based on the generated hard positive samples, we train a Siamese network for visual tracking and our experiments validate the effectiveness of the introduced algorithm. The project page of this paper can be found from the website1. Xiao Wang 0014, Chenglong Li 0002, Bin Luo 0001, Jin Tang 0001 |
CVPR | 2 |
| 2018 | Cross-Modal Ranking with Soft Consistency and Noisy Labels for Robust RGB-T Tracking
Chenglong Li 0002, Chengli Zhu, Yan Huang 0008, Jin Tang 0001, Liang Wang 0001 |
ECCV (13) | 1 |
| 2018 | Exploring Scene Geometry for Scale Adaptive Object Tracking in Surveillance VideosabstractObject tracking is a key technology in video surveillance. Reliable tracker must be adaptive to the constantly changing object sizes. Most of the state-of-the-art methods estimate the object scales using their appearances. Those methods are vulnerable to occlusion, object deformation, illumination change and background clutter. In this paper, we propose to use the geometric context of the surveillance site as a strong clue for scale adaptation. With three reasonable assumptions on the video cameras and the surveillance sites, we deduce a simple geometric model for object scales. The parameters of this model are learned without any human intervention. Then we integrate this model into baseline trackers for robust scale adaptive object tracking. Experimental results on challenging surveillance videos indicate that our approach favorably improves the performance of single-scale baselines, and performs better or comparative to the state-of-the-art multi-scale trackers while significantly improve the speed. Ran Zhong, Wenzhong Wang, Chenglong Li 0002, Aihua Zheng, Jin Tang 0001 |
ICIP | 3 |
| 2018 | Learning Soft-Consistent Correlation Filters for RGB-T Object Tracking
Chenglong Li 0002, Jin Tang 0001 |
PRCV (4) | 2 |
| 2018 | Non-negative Dual Graph Regularized Sparse Ranking for Multi-shot Person Re-identification
Aihua Zheng, Bo Jiang 0002, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
PRCV (1) | 4 |
| 2018 | Fusing two-stream convolutional neural networks for RGB-T object tracking
Chenglong Li 0002, Xiaohao Wu, Xiaochun Cao, Jin Tang 0001 |
Neurocomputing | 1 |
| 2018 | Moving object detection via robust background modeling with recurring patterns voting
Chenglong Li 0002, Zhimin Bao, Xiao Wang 0014, Jin Tang 0001 |
Multim. Tools Appl. | 1 |
| 2018 | Two-stage modality-graphs regularized manifold ranking for RGB-T tracking
Chenglong Li 0002, Chengli Zhu, Shaofei Zheng, Bin Luo 0001, Jing Tang 0001 |
Signal Process. Image Commun. | 1 |
| 2018 | Fast Grayscale-Thermal Foreground Detection With Collaborative Low-Rank DecompositionabstractThis paper investigates how to perform efficient and robust foreground detection in challenging scenarios by leveraging multiple source data. We propose a novel approach, called collaborative low-rank decomposition (CLoD), for grayscale-thermal foreground detection. Given two data matrices by accumulating sequential frames from the grayscale and the thermal videos, CLoD detects the foreground objects as sparse noises against the backgrounds with collaborative low rank structure, and also incorporates modality weights to achieve adaptive fusion of different source data. For the optimization, CLoD seeks a sub-optimal solution by making the background matrix rank explicitly determined. In particular, the background matrix with the fixed rank can be decomposed into two sub-matrices of low rank, and then, we iteratively optimize them and the modality weights with closed-form solutions. For improving the efficiency, we design a block-based accelerated algorithm to speed up CLoD while employing the edge-preserving algorithm to keep the accuracy. Extensive experiments on the recently public benchmark grayscale-thermal foreground detection suggest that our approach achieves comparable performance in terms of both accuracy and efficiency against other state-of-the-art methods. Bin Luo 0001, Chenglong Li 0002, Guizhao Wang, Jin Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Learning Patch-Based Dynamic Graph for Visual TrackingabstractExisting visual tracking methods usually localize the object with a bounding box, in which the foreground object trackers/detectors are often disturbed by the introduced background information. To handle this problem, we aim to learn a more robust object representation for visual tracking. In particular, the tracked object is represented with a graph structure (i.e., a set of non-overlapping image patches), in which the weight of each node (patch) indicates how likely it belongs to the foreground and edges are also weighed for indicating the appearance compatibility of two neighboring nodes. This graph is dynamically learnt (i.e., the nodes and edges received weights) and applied in object tracking and model updating. We constrain the graph learning from two aspects: i) the global low-rank structure over all nodes and ii) the local sparseness of node neighbors. During the tracking process, our method performs the following steps at each frame. First, the graph is initialized by assigning either 1 or 0 to the weights of some image patches according to the predicted bounding box. Second, the graph is optimized through designing a new ALM (Augmented Lagrange Multiplier) based algorithm. Third, the object feature representation is updated by imposing the weights of patches on the extracted image features. The object location is finally predicted by adopting the Struck tracker. Extensive experiments show that our approach outperforms the state-of-the-art tracking methods on two standard benchmarks, i.e., OTB100 and NUS-PRO. Chenglong Li 0002, Liang Lin 0004, Wangmeng Zuo, Jin Tang 0001 |
AAAI | 1 |
| 2017 | Selecting attentive frames from visually coherent video chunks for surveillance video summarizationabstractThis paper investigates how to extract key-frames from surveillance video while maximizing their diversity and representational ability. We solve this problem by two steps, i.e., video partition and frame selection. The first step is to partition a surveillance video into visually coherent video chunks, which have high intra-chunk similarity and interchunk dissimilarity. In particular, we propose an object-based frame metric to measure the relevance of two frames, and apply the Normalized Cut algorithm to achieve video partition. The second step is to select the attentive frames from the partitioned video chunks. We propose an attention score based on the content completeness and the visual satisfaction for each frame, and select most attentive frame with highest attention score in each chunk. Extensive experiments on both public and our newly created datasets suggest that our approach significantly outperforms other video summarization methods. Wenzhong Wang, Qiaoqiao Zhang, Bin Luo 0001, Jin Tang 0001, Rui Ruan, Chenglong Li 0002 |
ICIP | 6 |
| 2017 | ReGLe: Spatially Regularized Graph Learning for Visual TrackingabstractWeighted patch representation of the target object has been proven to be effective for suppressing the background effects in visual tracking. In this paper, we propose a novel approach, called spatially Regularized Graph Learning (ReGLe), to automatically explore the intrinsic relationship among patches both with global and local cues for robust object representation. In particular, the target object bounding box is partitioned into a set of non-overlapping image patches, which are taken as graph nodes, and each of them is associated with a weight to represent how likely it belongs to the target object. To improve the accuracy of node weight computation, we dynamically learn the edge weights (i.e., the appearance compatibility of two nodes) according to both global and local relationship among patches. First, we pursue the low-rank representation for capturing the global low-dimensional subspace structure of patches. Second, we encode the local information into the low-rank representation by exploiting the fact that neighboring nodes usually have similar appearance. Finally, we utilize the representations to learn their affinities (i.e., graph edge weights). The node and edge weights are jointly optimized by a designed ADMM (Alternating Direction Method of Multipliers) algorithm, the object feature representation is updated by imposing the weights of patches on the extracted image features. The object location is finally predicted by maximizing the classification score in the structured SVM. Extensive experiments demonstrate the effectiveness of the proposed approach on the tracking benchmark datasets: OTB100 and Temple-Color. Chenglong Li 0002, Xiaohao Wu, Zhimin Bao, Jin Tang 0001 |
ACM Multimedia | 1 |
| 2017 | Weighted Sparse Representation Regularized Graph Learning for RGB-T Object TrackingabstractIn this paper, we propose a novel graph model, called weighted sparse representation regularized graph, to learn a robust object representation using multispectral (RGB and thermal) data for visual tracking. In particular, the tracked object is represented with a graph with image patches as nodes. This graph is dynamically learned from two aspects. First, the graph affinity (i.e., graph structure and edge weights) that indicates the appearance compatibility of two neighboring nodes is optimized based on the weighted sparse representation, in which the modality weight is introduced to leverage RGB and thermal information adaptively. Second, each node weight that indicates how likely it belongs to the foreground is propagated from others along with graph affinity. The optimized patch weights are then imposed on the extracted RGB and thermal features, and the target object is finally located by adopting the structured SVM algorithm. Moreover, we also contribute a comprehensive dataset for RGB-T tracking purpose. Comparing with existing ones, the new dataset has the following advantages: 1) Its size is sufficiently large for large-scale performance evaluation (total frame number: 210K, maximum frames per video pair: 8K). 2) The alignment between RGB-T video pairs is highly accurate, which does not need pre- and post-processing. 3) The occlusion levels are annotated for analyzing the occlusion-sensitive performance of different methods. Extensive experiments on both public and newly created datasets demonstrate the effectiveness of the proposed tracker against several state-of-the-art tracking methods. Chenglong Li 0002, Yijuan Lu, Chengli Zhu, Jin Tang 0001 |
ACM Multimedia | 1 |
| 2017 | Local-to-global background modeling for moving object detection from non-static cameras
Aihua Zheng, Lei Zhang 0074, Wei Zhang 0012, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
Multim. Tools Appl. | 4 |
| 2017 | Weighted Low-Rank Decomposition for Robust Grayscale-Thermal Foreground DetectionabstractThis paper investigates how to fuse grayscale and thermal video data for detecting foreground objects in challenging scenarios. To this end, we propose an intuitive yet effective method called weighted low-rank decomposition (WELD), which adaptively pursues the cross-modality low-rank representation. Specifically, we form two data matrices by accumulating sequential frames from the grayscale and the thermal videos, respectively. Within these two observing matrices, WELD detects moving foreground pixels as sparse outliers against the low-rank structure background and incorporates the weight variables to make the models of two modalities complementary to each other. The smoothness constraints of object motion are also introduced in WELD to further improve the robustness to noises. For optimization, we propose an iterative algorithm to efficiently solve the low-rank models with three subproblems. Moreover, we utilize an edge-preserving filtering-based method to substantially speed up WELD while preserving its accuracy. To provide a comprehensive evaluation benchmark of grayscale-thermal foreground detection, we create a new data set including 25 aligned grayscale-thermal video pairs with high diversity. Our extensive experiments on both the newly created data set and the public data set OSU3 suggest that WELD achieves superior performance and comparable efficiency against other state-of-the-art approaches. Chenglong Li 0002, Xiao Wang 0014, Lei Zhang 0074, Jin Tang 0001, Hejun Wu, Liang Lin 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Grayscale-Thermal Object Tracking via Multitask Laplacian Sparse RepresentationabstractThis paper studies the problem of object tracking in challenging scenarios by leveraging multimodal visual data. We propose a grayscale-thermal object tracking method in Bayesian filtering framework based on multitask Laplacian sparse representation. Given one bounding box, we extract a set of overlapping local patches within it, and pursue the multitask joint sparse representation for grayscale and thermal modalities. Then, the representation coefficients of the two modalities are concatenated into a vector to represent the feature of the bounding box. Moreover, the similarity between each patch pair is deployed to refine their representation coefficients in the sparse representation, which can be formulated as the Laplacian sparse representation. We also incorporate the modal reliability into the Laplacian sparse representation to achieve an adaptive fusion of different source data. Experiments on two grayscale-thermal datasets suggest that the proposed approach outperforms both grayscale and grayscale-thermal tracking approaches. Chenglong Li 0002, Xiao Wang 0014, Lei Zhang 0074, Jin Tang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2016 | Real-Time Grayscale-Thermal Tracking via Laplacian Sparse Representation
Chenglong Li 0002, Shiyi Hu, Sihan Gao, Jin Tang 0001 |
MMM (2) | 1 |
| 2016 | Detection-Free Multiobject Tracking by Reconfigurable Inference With Bundle RepresentationsabstractThis paper presents a conceptually simple but effective approach to track multiobject in videos without requiring elaborate supervision (i.e., training object detectors or templates offline). Our framework performs a bi-layer inference of spatio-temporal grouping to exploit rich appearance and motion information in the observed sequence. First, we generate a robust middle-level video representation based on clustered point tracks, namely video bundles. Each bundle encapsulates a chunk of point tracks satisfying both spatial proximity and temporal coherency. Taking the video bundles as vertices, we build a spatio-temporal graph that incorporates both competitive and compatible relations among vertices. The multiobject tracking can be then phrased as a graph partition problem under the Bayesian framework, and we solve it by developing a reconfigurable belief propagation (BP) algorithm. This algorithm improves the traditional BP method by allowing a converged solution to be reconfigured during optimization, so that the inference can be reactivated once it gets stuck in local minima and thus conduct more reliable results. In the experiments, we demonstrate the superior performances of our approach on the challenging benchmarks compared with other state-of-the-art methods. Liang Lin 0004, Yongyi Lu, Chenglong Li 0002, Wangmeng Zuo |
IEEE Trans. Cybern. | 3 |
| 2016 | Inference With Collaborative Model for Interactive Tumor Segmentation in Medical Image SequencesabstractSegmenting organisms or tumors from medical data (e.g., computed tomography volumetric images, ultrasound, or magnetic resonance imaging images/image sequences) is one of the fundamental tasks in medical image analysis and diagnosis, and has received long-term attentions. This paper studies a novel computational framework of interactive segmentation for extracting liver tumors from image sequences, and it is suitable for different types of medical data. The main contributions are twofold. First, we propose a collaborative model to jointly formulate the tumor segmentation from two aspects: 1) region partition and 2) boundary presence. The two terms are complementary but simultaneously competing: the former extracts the tumor based on its appearance/texture information, while the latter searches for the palpable tumor boundary. Moreover, in order to adapt the data variations, we allow the model to be discriminatively trained based on both the seed pixels traced by the Lucas-Kanade algorithm and the scribbles placed by the user. Second, we present an effective inference algorithm that iterates to: 1) solve tumor segmentation using the augmented Lagrangian method and 2) propagate the segmentation across the image sequence by searching for distinctive matches between images. We keep the collaborative model updated during the inference in order to well capture the tumor variations over time. We have verified our system for segmenting liver tumors from a number of clinical data, and have achieved very promising results. The software developed with this paper can be found at http://vision.sysu.edu.cn/projects/med-interactive-seg/. Liang Lin 0004, Wei Yang 0019, Chenglong Li 0002, Jin Tang 0001, Xiaochun Cao |
IEEE Trans. Cybern. | 3 |
| 2016 | Learning Collaborative Sparse Representation for Grayscale-Thermal TrackingabstractIntegrating multiple different yet complementary feature representations has been proved to be an effective way for boosting tracking performance. This paper investigates how to perform robust object tracking in challenging scenarios by adaptively incorporating information from grayscale and thermal videos, and proposes a novel collaborative algorithm for online tracking. In particular, an adaptive fusion scheme is proposed based on collaborative sparse representation in Bayesian filtering framework. We jointly optimize sparse codes and the reliable weights of different modalities in an online way. In addition, this paper contributes a comprehensive video benchmark, which includes 50 grayscale-thermal sequences and their ground truth annotations for tracking purpose. The videos are with high diversity and the annotations were finished by one single person to guarantee consistency. Extensive experiments against other state-of-the-art trackers with both grayscale and grayscale-thermal inputs demonstrate the effectiveness of the proposed tracking approach. Through analyzing quantitative results, we also provide basic insights and potential future research directions in grayscale-thermal tracking. Chenglong Li 0002, Shiyi Hu, Xiaobai Liu, Jin Tang 0001, Liang Lin 0004 |
IEEE Trans. Image Process. | 1 |
| 2016 | An Approach to Streaming Video Segmentation With Sub-Optimal Low-Rank DecompositionabstractThis paper investigates how to perform robust and efficient video segmentation while suppressing the effects of data noises and/or corruptions, and an effective approach is introduced to this end. First, a general algorithm, called sub-optimal low-rank decomposition (SOLD), is proposed to pursue the low-rank representation for video segmentation. Given the data matrix formed by supervoxel features of an observed video sequence, SOLD seeks a sub-optimal solution by making the matrix rank explicitly determined. In particular, the representation coefficient matrix with the fixed rank can be decomposed into two sub-matrices of low rank, and then we iteratively optimize them with closed-form solutions. Moreover, we incorporate a discriminative replication prior into SOLD based on the observation that small-size video patterns tend to recur frequently within the same object. Second, based on SOLD, we present an efficient inference algorithm to perform streaming video segmentation in both unsupervised and interactive scenarios. More specifically, the constrained normalized-cut algorithm is adopted by incorporating the low-rank representation with other low level cues and temporal consistent constraints for spatio-temporal segmentation. Extensive experiments on two public challenging data sets VSB100 and SegTrack suggest that our approach outperforms other video segmentation approaches in both accuracy and efficiency. Chenglong Li 0002, Liang Lin 0004, Wangmeng Zuo, Wenzhong Wang, Jin Tang 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | SOLD: Sub-optimal low-rank decomposition for efficient video segmentationabstractThis paper investigates how to perform robust and efficient unsupervised video segmentation while suppressing the effects of data noises and/or corruptions. We propose a general algorithm, called Sub-Optimal Low-rank Decomposition (SOLD), which pursues the low-rank representation for video segmentation. Given the supervoxels affinity matrix of an observed video sequence, SOLD seeks a sub-optimal solution by making the matrix rank explicitly determined. In particular, the affinity matrix with the rank fixed can be decomposed into two sub-matrices of low rank, and then we iteratively optimize them with closed-form solutions. Moreover, we incorporate a discriminative replication prior into our framework based on the obervation that small-size video patterns tend to recur frequently within the same object. The video can be segmented into several spatio-temporal regions by applying the Normalized-Cut (NCut) algorithm with the solved low-rank representation. To process the streaming videos, we apply our algorithm sequentially over a batch of frames over time, in which we also develop several temporal consistent constraints improving the robustness. Extensive experiments on the public benchmarks demonstrate superior performance of our framework over other state-of-the-art approaches. Chenglong Li 0002, Liang Lin 0004, Wangmeng Zuo, Shuicheng Yan, Jin Tang 0001 |
CVPR | 1 |
| 2015 | Person Re-identification with Density-Distance Unsupervised Salience Learning
Baoliang Zhou, Aihua Zheng, Bo Jiang 0002, Chenglong Li 0002, Jin Tang 0001 |
ICIG (3) | 4 |
| 2015 | PISA: Pixelwise Image Saliency by Aggregating Complementary Appearance Contrast Measures With Edge-Preserving CoherenceabstractDriven by recent vision and graphics applications such as image segmentation and object recognition, computing pixel-accurate saliency values to uniformly highlight foreground objects becomes increasingly important. In this paper, we propose a unified framework called pixelwise image saliency aggregating (PISA) various bottom-up cues and priors. It generates spatially coherent yet detail-preserving, pixel-accurate, and fine-grained saliency, and overcomes the limitations of previous methods, which use homogeneous superpixel based and color only treatment. PISA aggregates multiple saliency cues in a global context, such as complementary color and structure contrast measures, with their spatial priors in the image domain. The saliency confidence is further jointly modeled with a neighborhood consistence constraint into an energy minimization formulation, in which each pixel will be evaluated with multiple hypothetical saliency levels. Instead of using global discrete optimization methods, we employ the cost-volume filtering technique to solve our formulation, assigning the saliency levels smoothly while preserving the edge-aware structure details. In addition, a faster version of PISA is developed using a gradient-driven image subsampling strategy to greatly improve the runtime efficiency while keeping comparable detection accuracy. Extensive experiments on a number of public data sets suggest that PISA convincingly outperforms other state-of-the-art approaches. In addition, with this work, we also create a new data set containing 800 commodity images for evaluating saliency detection. Keze Wang, Liang Lin 0004, Jiangbo Lu, Chenglong Li 0002, Keyang Shi |
IEEE Trans. Image Process. | 4 |