Jin Tang 0001

dblp:56/4951-1 · DBLP profile ↗
← Back
264ranked-venue papers
4as first author
176since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 127 · 1 first-author · 73 since 2021Artificial intelligence and machine learning · 106 · 3 first-author · 66 since 2021Applied, interdisciplinary, general and emerging computing · 45 · 44 since 2021Security and privacy · 11 · 10 since 2021Software engineering, systems software and programming languages · 3Computer networks · 2Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 ProxyTTT: Proxy-driven Test-Time Training for Multi-modal Re-identification
abstract
Multi-modal object re-identification (ReID) aims to retrieve specific targets by leveraging complementary cues from different sensing modalities. Despite recent progress, two key challenges remain: (1) the limited ability to jointly address both modality and viewpoint discrepancies, and (2) the difficulty of effectively leveraging reliable target-domain data to improve generalization. To address these challenges, we propose Proxy-driven Test-Time Training (ProxyTTT), a unified framework that enhances both multi-modal identity representation learning and model generalization. During training, we propose a Multi-Proxy Learning (MPL) mechanism to address the representation bias across different views and modalities. MPL disentangles fine-grained modality-specific and modality-common identity proxies as semantic anchors to align identity features across diverse perspectives and sensing modalities. This alignment strategy enables the model to learn robust and discriminative global identity representations under heterogeneous modality conditions. At test time, to reliably exploit target domain data, we propose Proxy-guided Entropy-based Selective Adaptation (PESA) for test-time training. Specifically, PESA leverages the semantic structure encoded by identity proxies to estimate prediction uncertainty via entropy, and selectively adapts the model using only high-confidence samples. This selective adaptation effectively mitigates the domain shift between training and deployment environments, improving the model’s generalization in real-world scenarios. Extensive experiments on four public multi-modal ReID benchmarks (RGBNT201, RGBNT100, MSVR310, and WMVeID863) demonstrate the effectiveness of ProxyTTT.
Aihua Zheng, Zhaojun Liu, Xixi Wan, Chenglong Li 0002, Jin Tang 0001, Yan Yan 0002
AAAI5
2026 Progressive Multi-modal Knowledge Distillation for Multi-spectral Object Re-identification
abstract
In the field of multi-spectral object re-identification (ReID), multi-modal knowledge and modal-specific knowledge exhibit complementary advantages when handling hard samples, but existing methods rarely integrate this collaborative information. Knowledge distillation is a direct approach for transferring information, however, heterogeneity in model architectures and variations in sample hardness can undermine the stability and controllability of knowledge transfer. To alleviate these limitations, we propose the novel Progressive Multi-modal Knowledge Distillation (PMKD) framework that enables multi-stage knowledge transfer guided by hard sample awareness. In the multi-modal knowledge transfer stage, the source model (pre-trained on multi-modal data) disseminates its learned multi-modal collaborative knowledge to multiple independently modal-specific target models, guiding their adaptation to hard samples within training batches. In the modal-specific knowledge retention stage, the independent models enriched with multi-modal knowledge guide the training phase. The architectural consistency between source-target models ensures more lossless knowledge transfer, effectively mitigating the risk of capability drift, and preserving inherent competence. Moreover, the entire progressive multi-modal knowledge distillation is regulated by the proposed hardness-aware distillation loss, which automatically adapts distillation intensity through hard sample mining, thereby ensuring stable transfer of hard sample handling capabilities. Extensive experiments on benchmark multi-spectral ReID datasets validate the effectiveness and superior performance of the proposed method.
Aihua Zheng, Zi Wang 0013, Jin Tang 0001
AAAI4
2026 Semantic-Driven Visual Progressive Refinement for Aerial-Ground Person ReID: A Challenging Large-Scale Benchmark
abstract
Aerial-Ground Person Re-IDentification (AGPReID) aims to extract identity-discriminative representations from heterogeneous perspectives across different platforms in complex real-world environments. However, existing methods primarily focus on visual appearance modeling and make insufficient use of semantic attribute priors, which limits their ability to bridge the aerial-ground view gap. To address this limitation, we propose a Semantic-driven Visual Progressive Refinement framework for AGPReID (SVPR-ReID), which effectively leverages textual attribute priors to guide the extraction of fine-grained visual cues. Specifically, we design a View-Decoupled Feature Extractor that incorporates view-aware textual prompts to decouple view-invariant identity features. Then, to alleviate inter-class ambiguity, we propose an Attribute-Scattered Mixture-of-Experts module that integrates attribute semantics into the visual space, thereby improving discrimination among visually similar pedestrians. Finally, we design a Context-Vision Progressive Refinement module for progressive refinement of attribute and view-invariant features, obtaining robust cross-view identity representations. In particular, we contribute a comprehensive benchmark for AGPReID, named CP2108, which contains 142,817 images of 2,108 identities annotated with 22 attributes. Notably, it includes 191 identities captured across different times, enabling both short- and long-term ReID evaluation, addressing the limitation of existing datasets that focus only on short-term scenarios. Extensive experimental results validate the effectiveness of our SVPR-ReID on four AGPReID datasets.
Aihua Zheng, Xixi Wan, Zi Wang 0013, Jin Tang 0001, Bin Luo 0001
AAAI6
2026 Morphology-aware hierarchical mixture of experts for Chest X-ray anatomy segmentation
Lili Huang 0006, Yuanjun He, Chenglong Li 0002, Jin Tang 0001
Eng. Appl. Artif. Intell.5
2026 Uncertainty-Aware RGBT Tracking
Zhaodong Ding, Chenglong Li 0002, Futian Wang, Jin Tang 0001
Int. J. Comput. Vis.4
2026 ImageBind Guided Progressive Transformation Network for Alignment-free RGBT Video Object Detection
Zhengzheng Tu, Chuanwang Guo, Qishun Wang, Chenglong Li 0002, Jin Tang 0001
Int. J. Comput. Vis.5
2026 Medical report generation via knowledge distillation and medical keywords
Lili Huang 0006, Chenglong Li 0002, Jin Tang 0001
Neurocomputing5
2026 ESTR-CoT: Towards explainable and accurate event stream based scene text recognition with chain-of-thought reasoning
Xiao Wang 0014, Jingtao Jiang, Qiang Chen 0007, Lan Chen 0003, Lin Zhu 0012, Yaowei Wang 0001, Yonghong Tian 0001, Jin Tang 0001
Neurocomputing8
2026 Hierarchical long and short-term preference modeling with denoising Mamba for sequential recommendation
Wei Jiang 0049, Yongquan Fan, Jin Tang 0001, Xianyong Li, Yajun Du
Inf. Process. Manag.3
2026 Multi-level alignment network for unsupervised domain adaptive multi-modality object re-identification
Yusong Sheng, Yuhe Ding, Aihua Zheng, Zi Wang 0013, Jin Tang 0001
Knowl. Based Syst.6
2026 YOFOR : You only focus on object regions for tiny object detection in aerial images
Heng Hu, Hao-Zhe Wang, Sibao Chen 0001, Jin Tang 0001
Neural Networks4
2026 Reliable and Compact Graph Fine-Tuning via Graph Sparse Prompting
abstract
Recently, graph prompt learning has garnered increasing attention in adapting pre-trained GNN models for downstream graph learning tasks. However, existing works generally conduct prompting over all graph elements (e.g., nodes, edges, node attributes, etc.), which is suboptimal and obviously redundant. To address this issue, we propose exploiting sparse representation theory for graph prompting and present Graph Sparse Prompting (GSP). GSP aims to adaptively and sparsely select the optimal elements (e.g., certain node attributes) to achieve compact prompting for downstream tasks. Specifically, we propose two kinds of GSP models, termed Graph Sparse Feature Prompting (GSFP) and Graph Sparse multi-Feature Prompting (GSmFP). Both GSFP and GSmFP provide a general scheme for tuning any specific pre-trained GNNs that can select some desired attributes for prompting by employing sparsity-guided prompt learning. A simple yet effective algorithm has been designed for solving GSFP and GSmFP models. Experiments on 16 widely-used benchmark datasets validate the effectiveness and advantages of the proposed GSFPs.
Bo Jiang 0002, Beibei Wang 0006, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Revisiting Deformable Convolution on Graphs: Large-Range Modeling and Robustness
abstract
Graph Convolution Networks (GCNs) have achieved remarkable success in representation of structured graph data. As we know that traditional GCNs are generally defined on the fixed first-order neighborhood receptive field which makes them be incapable to capture the long-range dependencies between distant nodes and also vulnerable to graph attacks and noises. To address these limitations, we revisit deformable convolution on graphs and propose a novel deformable graph convolution, termed Neighborhood-Deformable Graph Convolution (NDGC). The core of NDGC is to explicitly achieve the deformable convolution on graphs by introducing virtual neighbors which encode large-range information via the offsetting and interpolation function. That is, the introduced virtual neighbors can provide a larger receptive field with deformable receptive shape for graph convolution definition. Also, NDGC conducts message aggregation on the deformable virtual neighbors which thus performs more robustly w.r.t. graph attacks and noises. In particular, NDGC provides a general neighborhood deformable scheme, seamlessly integrating with many graph convolution definitions to derive their deformable variants. Experimental results validate the effectiveness and advantages of the proposed NDGC networks on several graph learning tasks.
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 SequencePAR: Understanding pedestrian attributes via a sequence generation paradigm
Jiandong Jin, Xiao Wang 0014, Yin Lin, Chenglong Li 0002, Lili Huang 0006, Aihua Zheng, Jin Tang 0001
Pattern Recognit.7
2026 RGBT tracking via supervised mutual guiding
Lei Liu 0049, Chenglong Li 0002, Jin Tang 0001, Changhe Li
Pattern Recognit.3
2026 Temporal multimodal knowledge distillation for modality-missing RGBT tracking
Rui Ruan, Yunlong Kang, Lei Liu 0049, Jingpeng Sun, Chenglong Li 0002, Jin Tang 0001
Pattern Recognit.6
2026 Semantic change detection of roads and bridges: A fine-grained dataset and multimodal frequency-driven detector
Qing-Ling Shu, Sibao Chen 0001, Xiao Wang 0014, Zhi-Hui You, Wei Lu 0032, Jin Tang 0001, Bin Luo 0001
Pattern Recognit.6
2026 Structure and progress aware diffusion for medical image segmentation
Siyuan Song, Guyue Hu 0001, Chenglong Li 0002, Dengdi Sun, Zhe Jin 0001, Jin Tang 0001
Pattern Recognit.6
2026 CLNS: Camera-aware label noise suppression for unsupervised visible-infrared person re-identification
Sicheng Zhao, Wei Lu 0032, Sibao Chen 0001, Chris Ding, Futian Wang, Jin Tang 0001, Bin Luo 0001
Pattern Recognit.7
2026 Medical image segmentation via Attention-enhanced Mamba with learnable Symmetry scan and a benchmark
Chenglong Li 0002, Jin Tang 0001, Chuanfu Li
Pattern Recognit.3
2026 Fine-Grained and Granularity-Dynamic Framework for Referring Remote Sensing Image Segmentation
Duzhi Yuan, Guyue Hu 0001, Aihua Zheng, Chenglong Li 0002, Jin Tang 0001
IEEE Signal Process. Lett.6
2026 BHGraphAdapter: Parameter-Efficient VLMs Tuning Meets Hyper-Graph Learning
abstract
Adapter-based fine-tuning methods for Visual-Language Models (VLMs) have shown promising performance for feature adaptation in limited data scenarios. However, existing adapters generallyeitheremploy parameterized transformation for multi-modality feature refiningorexploit pairwise relationships between classes (i.e., GraphAdapter) for text enhancement, which ignore the inherent high-order correlations among data samples in the adaptation process. In this paper, for the first time, we propose to exploit the high-order relationships of visual samples within each mini-batch for fine-tuning VLMs and develop a novel Batch HyperGraph Adapter (BHGraphAdapter) to fine-tune VLMs. The core idea of BHGraphAdapter is to conduct feature adapter learning by capturing the inherent high-order semantic information of different samples within each mini-batch, which thus can fully exploit the complex context information in adaptation. Specifically, we first construct a Batch HyperGraph (BHGraph) to model the high-order correlation of samples within each mini-batch. Then, we introduce a message propagation module on BHGraph to update the node embeddings by aggregating information from their high-order neighbors, thereby capturing semantic relationships to enrich feature representation. Finally, we incorporate the proposed BHGraph learning into the pre-trained CLIP framework to achieve the feature adaptation for the downstream tasks. Extensive experiments on 11 benchmark datasets show that our proposed BHGraphAdapter outperforms the SOTA adapter tuning methods. The source code and data will be released at https://github.com/LiuMeilin7195/BHGraphAdapter.
Xixi Wang 0005, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Vehicle-Centric Perception via Multimodal Structured Pre-Training
abstract
Vehicle-centric perception plays a crucial role in many intelligent systems, including large-scale surveillance systems, intelligent transportation, and autonomous driving. Existing approaches typically employ general pre-trained weights to initialize backbone networks, followed by task-specific fine-tuning. However, these models lack effective learning of vehiclerelated knowledge during pre-training, resulting in poor capability for modeling general vehicle perception representations. To handle this problem, we propose VehicleMAE-V2, a novel vehicle-centric pre-trained large model. By exploring and exploiting vehicle-related multimodal structured priors to guide the masked token reconstruction process, our approach can significantly enhance the model’s capability to learn generalizable representations for vehicle-centric perception. Specifically, we design the Symmetry-guided Mask Module (SMM), Contour-guided Representation Module (CRM) and Semantics-guided Representation Module (SRM) to incorporate three kinds of structured priors into token reconstruction including symmetry, contour and semantics of vehicles respectively. SMM utilizes the vehicle symmetry constraints to avoid retaining symmetric patches and can thus select high-quality masked image patches and reduce information redundancy. CRM minimizes the prob23 ability distribution divergence between contour features and reconstructed features and can thus preserve holistic vehicle structure information during pixel-level reconstruction. SRM aligns image-text features through contrastive learning and cross-modal distillation to address the feature confusion caused by insufficient semantic understanding during masked reconstruction. To support the pre-training of VehicleMAE-V2, we construct Autobot4M, a large-scale dataset comprising approximately 4 million vehicle images and 12,693 text descriptions. Extensive experiments on five downstream tasks demonstrate the superior performance of VehicleMAE-V2. The source code, dataset, and pre-trained large models are available on https://github.com/Vehicle-AHU/VehicleMAE.
Xiao Wang 0014, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Cross-Modal Person Retrieval With One-to-Many Relation Modeling
abstract
Existing text-image person retrieval methods are built upon fixed-point or distributional embeddings, but they typically perform single-point alignment across modalities, making it difficult to effectively capture the one-to-many cross-modal semantic associations. To address this, we propose a One-to-Many Relation modEling network (OMRE) that explicitly constructs one-to-many semantic matching structures across modalities, thereby modeling richer and more diverse semantic associations. Specifically, to achieve one-to-many matching modeling, we design a bidirectional one-to-many alignment module, which constructs cross-modal matching distributions by aggregating relations between the mean embedding and multiple sampled embeddings, and minimizes their discrepancy with the true distribution to capture complex semantic associations. To construct fine-grained one-to-many matching relationships, we propose a collaborative reconstruction-based similarity refinement module, which maximizes the semantic consistency between multiple reconstructed masked tokens and the original tokens, effectively achieving robust and precise one-to-many cross-modal fine-grained semantic alignment. Moreover, to enhance the discriminative capability of one-to-many semantic distributions, we introduce a Hard Negative Mining mechanism that focuses on semantically similar but mismatched samples, helping to refine distribution boundaries in the probabilistic space and suppress interference from hard negative samples. Extensive experiments on three public datasets demonstrate that our method not only achieves superior overall performance but also exhibits excellent generalization ability. The code will be released on https://github.com/Yifei-AHU/OMRE.
Yifei Deng, Chenglong Li 0002, Guyue Hu 0001, Jin Tang 0001
IEEE Trans. Inf. Forensics Secur.5
2026 Adversarial Semantic and Label Perturbation Attack for Pedestrian Attribute Recognition
abstract
Pedestrian Attribute Recognition (PAR) is an indispensable task in human-centered research and has made great progress in recent years with the development of deep neural networks. However, the potential vulnerability and anti-interference ability have still not been fully explored. To bridge this gap, this paper proposes the first adversarial attack and defense framework for pedestrian attribute recognition. Specifically, we exploit both global- and patch-level attacks on the pedestrian images, based on the pre-trained CLIP-based PAR framework. It first divides the input pedestrian image into non-overlapping patches and embeds them into feature embeddings using a projection layer. Meanwhile, the attribute set is expanded into sentences using prompts and embedded into attribute features using a pre-trained CLIP text encoder. A multi-modal Transformer is adopted to fuse the obtained vision and text tokens, and a feed-forward network is utilized for attribute recognition. Based on the aforementioned PAR framework, we adopt the adversarial semantic and label-perturbation to generate the adversarial noise, termed ASL-PAR. We also design a semantic offset defense strategy to suppress the influence of adversarial attacks. Extensive experiments conducted on both digital domains (i.e., PETA, PA100K, MSP60K, RAPv2) and physical domains fully validated the effectiveness of our proposed adversarial attack and defense strategies for the pedestrian attribute recognition. The source code of this paper will be released on https://github.com/Event-AHU/OpenPAR.
Weizhe Kong, Xiao Wang 0014, Ruichong Gao, Chenglong Li 0002, Yu Zhang 0091, Xing Yang 0004, Yaowei Wang 0001, Jin Tang 0001
IEEE Trans. Inf. Forensics Secur.8
2026 Reliable Multi-Modal Object Re-Identification via Modality-Aware Graph Reasoning
abstract
Multi-modal data provides abundant and diverse object information, crucial for effective modal interactions in Re-Identification (ReID) task. However, existing approaches often overlook the quality variations in local features and fail to fully leverage the complementary information across modalities, particularly in cases where features are of low quality. In this paper, we propose to address this issue by leveraging a novel graph reasoning model, termed the Modality-aware Graph Reasoning Network (MGRNet). Specifically, we first construct modality-aware graphs to enhance the extraction of fine-grained local details by effectively capturing and modeling the relationships between patches. Subsequently, the selective graph nodes swap operation is employed to alleviate the adverse effects of low-quality local features by considering both local and global information, enhancing the representation of discriminative information. Finally, the swapped modality-aware graphs are fed into the local-aware graph reasoning module, which propagates multi-modal information to yield a reliable feature representation. Another advantage of the proposed graph reasoning approach is its ability to reconstruct missing modal information by exploiting inherent structural relationships, thereby minimizing disparities between different modalities. Experimental results on four benchmarks (RGBNT201, Market1501-MM, RGBNT100, MSVR310) indicate that the proposed method achieves state-of-the-art performance in multi-modal object ReID. The code for our method will be available upon acceptance.
Xixi Wan, Aihua Zheng, Zi Wang 0013, Bo Jiang 0002, Jin Tang 0001, Jixin Ma 0001
IEEE Trans. Inf. Forensics Secur.5
2026 Causality-Based Modality- and Platform-Invariant Representation Learning for Dynamic RGBT Tracking and a Benchmark
abstract
Each sequence in existing RGBT tracking datasets is typically captured from a single platform equipped with both RGB (visible light) and TIR (thermal infrared) sensors. In real-world applications, tracking some objects requires cross-platform collaboration and these platforms might be equipped with different sensors. However, changes in modalities and platforms may cause significant variations in target appearance and abrupt position shifts, which existing RGBT trackers struggle to handle. To address these challenges, we define a new task, termed dynamic RGBT tracking, focusing on cross-platform and modality-variant scenarios. Considering the dynamic changes of modalities and platforms, we investigate dynamic RGBT tracking from a causal perspective, and assume that images consist of causal factors (target-relevant information) and non-causal factors (target-irrelevant information, i.e., modality/platform information), where only the former is conducive to stable tracking. Based on this assumption, we propose a novel causality-based modality&platform-invariant representation learning approach to capture robust invariant representations for dynamic RGBT tracking. In particular, to mitigate the challenges posed by modality variations, we design a causal consistency encoder that introduces an intervener to model feature uncertainty and simulate modal variations, compelling the model to focus on modality-invariant features to improve tracking robustness. To overcome the issue of abrupt view change and position shift, we design a platform-independent global searcher to re-localize the target whenever a platform switch occurs, which leverages an intervener to simulate the interference of platform changes on features, encouraging the searcher to learn platform-invariant representations for improved localization accuracy. In addition, to promote the research and development of dynamic RGBT tracking, we construct a dataset named DRGBT603, which consists of 603 sequences with a total of 1.49 M frame pairs. Extensive experiments on DRGBT603 dataset validate the effectiveness of the proposed method against other state-of-the-art methods. Our code and data are now available: https://github.com/dongdong2061/DRGBT.
Zhaodong Ding, Chenglong Li 0002, Shengqing Miao, Jin Tang 0001
IEEE Trans. Image Process.4
2026 Text-Visible/Infrared Person Retrieval: Attribute-Guided Feature Decoupling and Collaborative Alignment and a Unified Benchmark
abstract
Existing research on text-to-image person retrieval primarily focuses on visible images, which are not suitable under low-light scenarios. Infrared imaging becomes necessary in many visual systems, and matching text with both visible and infrared images is required. However, visible and infrared images are heterogeneous with different visual characteristics, so matching text with them in a unified framework is very challenging. In this work, we design a new task called Text-Visible/Infrared person retrieval and contribute a novel approach and a unified benchmark to promote the research and development of this field. On one hand, we propose a novel Attribute-guided feature decoupling and Collaborative Alignment Network (ACANet) that pursues accurate alignment from the text modality to both visible and infrared modalities in a unified framework according to the texture and color attribute information of text descriptions. In particular, we decouple the color features of visible images supervised by the text labels and integrate them into the infrared features to eliminate the impact of the absence of color information in infrared images during cross-modal collaborative alignment. Moreover, we also decouple the texture information from visible images supervised by the text labels and perform the collaborative alignment of texture and infrared features with a fusion agent. In addition, we extend conventional masked language modeling to a cross-modal paradigm to help ACANet learn uniform fine-grained alignment in multiple image modalities. On the other hand, we contribute a unified high-quality MM01LLCM-Text dataset, which provides person images in both visible and infrared modalities paired with fine-grained text descriptions. Experimental results show that the proposed ACANet outperforms existing state-of-the-art methods on MM01LLCM-Text dataset.
Chenglong Li 0002, Yifei Deng, Aihua Zheng, Jin Tang 0001
IEEE Trans. Image Process.5
2026 Pixel-Level RGBT Fusion Tracking via Heterogeneous Multi-Expert Distillation and Decoupled Representation Learning
abstract
Pixel-level fusion is widely considered a lightweight yet limited strategy in RGB-Thermal (RGBT) tracking due to its shallow representational capacity. However, its actual limitations and potential remain largely unexplored. We systematically analyze fusion location, modality alignment, and tracking performance, revealing that despite lower modality gaps than feature-level fusion, pixel-level fusion lacks task-relevant discrimination, restricting its effectiveness. In this paper, we propose the Task-driven Pixel-level Fusion tracker (TPF), which preserves the efficiency of early fusion while enhancing discriminative capacity. Central to TPF is a lightweight pixel fusion adapter that ensures real-time image fusion with only 14.3KB extra parameters over the baseline at inference. To enhance its limited representational capacity, we propose a task-driven progressive learning framework consisting of two key stages. First, a heterogeneous multi-expert distillation scheme adaptively transfers image fusion knowledge from diverse models under tracking-guided evaluation, mitigating the generalization limitations of single-teacher distillation across varied tracking scenarios. Second, to overcome limited task discrimination caused by sparse, target-focused tracking supervision, we propose a decoupled representation learning strategy that offers dense, complementary guidance to improve target-background separation and fusion quality. A nearest-neighbor dynamic template update further enhances robustness to appearance changes. Extensive experiments on four RGBT tracking benchmarks show that TPF achieves competitive accuracy and speed, outperforming both feature-level and existing pixel-level fusion methods, offering new insights into efficient RGBT tracking.
Andong Lu, Yuanzhi Guo, Kunpeng Wang 0005, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Image Process.5
2026 REMIND: Retrieval-Augmented Reconstruction With Dual Memories for Modality-Missing Object Re-Identification
abstract
To address the modality-missing object Re-Identification (Re-ID) task, a common strategy is to compensate for absent information by exploiting available modalities. However, existing reconstruction-based approaches suffer from two major limitations: 1) they often overlook modality-specific cues inherent in the missing modality; 2) they typically adopt a single-path reconstruction strategy. These issues result in incomplete representations and constrain the capacity to model complex semantic mappings across heterogeneous modalities. To address these challenges, we propose REMIND, a novel framework for modality-missing object Re-Identification, namely REtrieval-AugMented ReconstructIoN With Dual Memories. Specifically, we design a Dual Memory Construction module that, guided by information-theoretic insights, extracts modality-specific and modality-common features through two complementary branches and stores them in dedicated memory banks. These memory banks serve as structured prior knowledge to guide the reconstruction process, ensuring that the features of missing modalities are preserved even under modality-missing conditions. In addition, we have developed a retrieval-augmented missing reconstruction module that enhances the expressiveness and robustness of the reconstruction through multi-path reconstruction and perturbation mechanisms. Adaptive fusion techniques are employed for integration, simultaneously improving the expressiveness and robustness of the reconstructed features. Through the synergy of information-theoretically motivated regularization and retrieval-enhanced reconstruction, REMIND achieves robust feature recovery and delivers highly discriminative representations for reliable modality-missing Re-ID. Extensive experiments on several multi-modal object Re-ID benchmarks demonstrate the effectiveness and superiority of REMIND under various missing modality scenarios. The code is publicly available at: https://github.com/skye-1201/REMIND.
Zhendong Xu, Zi Wang 0013, Aihua Zheng, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Image Process.5
2026 Attribute-Guided Semantic Alignment With Pre-Trained Foundation Models for Vehicle Detection
abstract
Vehicle detection is a fundamental perception task in intelligent transportation systems and plays a crucial role in enabling reliable traffic perception and analysis. Existing vehicle detectors are typically obtained by training conventional object detection models (e.g., YOLO, RCNN, and DETR series) on vehicle images based on pre-trained backbone networks (e.g., ResNet and ViT). Although some studies introduce large-scale foundation models to improve detection performance, these models are not specifically designed for vehicle-centric scenarios and therefore tend to yield sub-optimal results in complex traffic environments. Moreover, most existing methods heavily rely on visual features and pay limited attention to the alignment between vehicle semantic information and visual representations. In this paper, we propose a novel vehicle detection paradigm, termed VFM-Det, which integrates a pre-trained vehicle foundation model (VehicleMAE) with a large language model (T5) to achieve semantically enhanced vehicle detection for intelligent transportation scenarios. Specifically, the proposed method follows a region proposal-based detection framework and employs VehicleMAE to enhance the features of each proposal. More importantly, we introduce a novel VAtt2Vec module to predict the vehicle semantic attributes corresponding to each proposal and transform them into feature vectors, which further enhance visual features through contrastive learning. Extensive experiments on three vehicle detection benchmark datasets thoroughly proved the effectiveness of our vehicle detector. Specifically, our model improves the baseline approach by +6.0%, +8.4% on the$AP_{0.5}$,$AP_{0.75}$metrics, respectively, on the Cityscapes dataset. The source code of this work will be released athttps://github.com/Event-AHU/VFM-Det
Fanghua Hong, Xiao Wang 0014, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Intell. Transp. Syst.5
2026 Activating Associative Disease-Aware Vision Token Memory for LLM-Based X-Ray Report Generation
abstract
X-ray image based medical report generation achieves significant progress in recent years with the help of large language models, however, these models have not fully exploited the effective information in visual image regions, resulting in reports that are linguistically sound but insufficient in describing key diseases. In this paper, we propose a novel associative memory-enhanced X-ray report generation model that effectively mimics the process of professional doctors writing medical reports. It considers both the mining of global and local visual information and associates historical report information to better complete the writing of the current report. Specifically, given an X-ray image, we first utilize a classification model along with its activation maps to accomplish the mining of visual regions highly associated with diseases and the learning of disease query tokens. Then, we employ a visual Hopfield network to establish memory associations for disease-related tokens, and a report Hopfield network to retrieve report memory information. This process facilitates the generation of high-quality reports based on a large language model and achieves state-of-the-art performance on multiple benchmark datasets, including the IU X-ray, MIMIC-CXR, and Chexpert Plus. The source code and pre-trained models of this work have been released on https://github.com/Event-AHU/Medical_Image_Analysis.
Xiao Wang 0014, Fuling Wang, Bo Jiang 0002, Chuanfu Li, Yaowei Wang 0001, Yonghong Tian 0001, Jin Tang 0001
IEEE Trans. Medical Imaging8
2026 ICPL-ReID: Identity-Conditional Prompt Learning for Multi-Spectral Object Re-Identification
Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.4
2026 DEEP: Decoupled Semantic Prompt Learning, Guiding and Embedding for Multi-Spectral Object Re-Identification
abstract
Multi-spectral object re-identification (ReID) captures diverse object semantics to robustly recognize identity in complex environments. However, without explicit semantic guidance (e.g., attributes, masks, and keypoints), existing modal fusion-based methods struggle to comprehensively capture person or vehicle semantics across spectra. Thanks to the large-scale vision-language pre-training, CLIP effectively aligns visual concepts across different image modalities to a unified semantic prompt. In this paper, we proposeDEEP, aDEcoupled sEmanticPrompt Learning, Guiding and Embedding framework for Multi-Spectral Object ReID. Specifically, to address the challenges posed by low-quality modality noise and spectral style discrepancies, we first propose a Decoupled Semantic Prompt (DSP) strategy, which explicitly decouples the semantic alignment into spectral-style learning with spectral-shared prompts and object content learning with instance-specific inversion token. Second, to lead the model focusing on semantically faithful regions, we propose a Semantic-Guided Spectral Fusion (SGSF) module that builds a semantic interaction bridge between spectra to explore complementary semantics across modalities. Finally, to further empower the spectral representation, we propose a Spectral Semantic Embedding (SSE) module constrained by semantic-aware structural consistency to refine the fine-grained identity semantics in each spectrum. Extensive experiments on five public benchmarks, RGBNT201, Market-MM, MSVR310, WMVEID863, and RGBNT100, demonstrate the proposed method outperforms the state-of-the-art methods. The source code is released at this link:https://github.com/lsh-ahu/DEEP-ReID.
Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.4
2026 RCNet: Reliable Co-Training Network for Weakly Supervised Change Detection
abstract
Fully supervised change detection (CD) methods in remote sensing (RS) perform well but depend on costly and time-consuming pixel-level annotations, which are impractical to obtain at scale. Therefore, it is essential to develop annotation-efficient alternatives that can narrow the performance gap with fully supervised methods. To this end, we propose a novel weakly supervised CD framework, named RCNet, which employs dual networks to implement reliable co-training using image-level annotations. Our framework is grounded in multi-view learning of co-training and the localization ability of class activation mapping (CAM). In our approach, two sub-nets with the same architecture perform image-level change classification and pixel-level segmentation from different views. Although CAM roughly localizes changes, ambiguity and noise in its pseudo labels may cause confirmation bias, limiting performance. Our approach mitigates this bias by introducing a feature discrepancy loss to enable cross-supervision between two sub-nets. Meanwhile, CAM tends to highlight a single object, but RS images commonly contain many dense and small changed objects with complexity, resulting in decreased reliability of pseudo labels. Therefore, we present an IoU-based reliable pseudo label screening (RPLS) strategy, which minimizes the likelihood of changed areas being misidentified as unchanged, enhancing the reliability of changed information obtained. Besides, to further improve boundary fineness and internal integrity of changed areas, we incorporate an additional strong perturbation branch for each sub-net and develop a consistency regularization loss. Extensive experiments on three challenging RS image CD datasets demonstrate that our RCNet achieves competitive performance with image-level labels. The source code is available athttps://github.com/Youzhihui/RCNet.
Zhi-Hui You, Sibao Chen 0001, Chris Ding, Lili Huang 0006, Jia-Xin Wang, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.6
2025 RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba
abstract
Existing RGBT tracking methods often design various interaction models to perform cross-modal fusion of each layer, but can not execute the feature interactions among all layers, which plays a critical role in robust multimodal representation, due to large computational burden. To address this issue, this paper presents a novel All-layer multimodal Interaction Network, named AINet, which performs efficient and effective feature interactions of all modalities and layers in a progressive fusion Mamba, for robust RGBT tracking. Even though modality features in different layers are known to contain different cues, it is always challenging to build multimodal interactions in each layer due to struggling in balancing interaction capabilities and efficiency. Meanwhile, considering that the feature discrepancy between RGB and thermal modalities reflects their complementary information to some extent, we design a Difference-based Fusion Mamba (DFM) to achieve enhanced fusion of different modalities with linear complexity. When interacting with features from all layers, a huge number of token sequences (3840 tokens in this work) are involved and the computational burden is thus large. To handle this problem, we design an Order-dynamic Fusion Mamba (OFM) to execute efficient and effective feature interactions of all layers by dynamically adjusting the scan order of different layers in Mamba. Extensive experiments on four public RGBT tracking datasets show that AINet achieves leading performance against existing state-of-the-art methods. We will release the code upon acceptance of the paper.
Andong Lu, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
AAAI4
2025 CXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus Dataset
abstract
X-ray image-based medical report generation (MRG) is a pivotal area in artificial intelligence that can significantly reduce diagnostic burdens and patient wait times. Despite significant progress, we believe that the task has reached a bottleneck due to the limited benchmark datasets and the existing large models’ insufficient capability enhancements in this specialized domain. Specifically, the recently released CheXpert Plus dataset lacks comparative evaluation algorithms and their results, providing only the dataset itself. This situation makes the training, evaluation, and comparison of subsequent algorithms challenging. Thus, we conduct a comprehensive benchmarking of existing mainstream X-ray report generation models and large language models (LLMs), on the CheXpert Plus dataset. We believe that the proposed benchmark can provide a solid comparative basis for subsequent algorithms and serve as a guide for researchers to quickly grasp the state-of-the-art models in this field. More importantly, we propose a large model for the X-ray image report generation using a multi-stage pre-training strategy, including self-supervised autoregressive generation and Xray-report contrastive learning, and supervised fine-tuning. Extensive experimental results indicate that the autoregressive pre-training based on Mamba effectively encodes X-ray images, and the image-text contrastive pre-training further aligns the feature spaces, achieving better experimental results. Source code can be found on https://github.com/Event-AHU/Medical_Image_Analysis.
Xiao Wang 0014, Fuling Wang, Yuehang Li, Qingchuan Ma, Shiao Wang, Bo Jiang 0002, Jin Tang 0001
CVPR7
2025 CLIP Guided Multimodal Prototype Learning for One-Shot Semantic Segmentation
abstract
One-shot semantic segmentation is to segment the object regions of unseen categories with only one annotated example as the supervision. Existing methods often adopt the multimodal pre-trained model CLIP to generate some key priors for one-shot semantic segmentation, but still face two main challenges. First, the prototype merely based on the visual features still struggles with similar inter-class object. Second, the prior-guided prototypes are hard to represent various intra-class objects due to the insufficient query information. To address these challenges, this paper presents a CLIP guided multimodal prototype learning approach called CLIP-MP, which enhances the generalization of prototype and captures more reliable query features, for accurate one-shot semantic segmentation. Specifically, a Multimodal Prototype Learning Module (MPLM) is proposed to combine the textual prototype with the support visual prototype for enhancing semantic representation of the multimodal prototype. Furthermore, a Prototype Refinement Module (PRM) is designed to refine the learned multimodal prototype by incorporating query-specific information. Extensive experiments on the PASCAL-5iand COCO-20idatasets demonstrate the effectiveness of the proposed CLIP-MP in enhancing segmentation accuracy and generalization of one-shot setting.
Yulei Jian, Lingma Sun, Jin Tang 0001
ICME4
2025 Template-based Uncertainty Multimodal Fusion Network for RGBT Tracking
abstract
RGBT tracking is to localize the predefined targets in video sequences by effectively leveraging the information from both visible light (RGB) and thermal infrared (TIR) modalities. However, the quality of different modalities changes dynamically in complex scenes, and effectively perceiving modal quality for multimodal fusion remains a significant challenge. To address this challenge, we propose to employ the reliability of initial template to explore the uncertainty across different modalities, and design a novel template-based uncertainty computation framework for robust multimodal fusion in RGBT tracking. In particular, we introduce an Uncertainty-aware Multimodal Fusion Module (UMFM), which constructs the uncertainty of each modality by leveraging the correlation between the template and search region in the Subjective Logic framework, aiming to achieve robust multimodal fusion. In addition, existing methods focus on dynamic template update while overlooking the potential role of a reliable initial template in the template updating process.To this end, we design a simple yet effective Contrastive Template Update Module (CTUM) to assess the reliability of the new template by comparing its quality with that of the initial template. Extensive experiments suggest that our method outperforms existing approaches on four RGBT tracking benchmarks.
Zhaodong Ding, Chenglong Li 0002, Shengqing Miao, Jin Tang 0001
IJCAI4
2025 Learning with Explicit Topological Priors for Chest X-Ray Rib Segmentation
Chenglong Li 0002, Jin Tang 0001, Chuanfu Li
MICCAI (16)3
2025 Learning Hierarchical Cross-modal Association with Intra-modal Context for Text-Image Person Retrieval
abstract
Existing Text-Image Person Retrieval (TIPR) methods have made substantial progress in modeling cross-modal associations via contrastive learning frameworks, but usually ignore the fine-grained differences in semantic relevance among different samples, which limits retrieval accuracy. To address this problem, we propose a novel Hierarchical Cross-modal Association framework HCA, which leverages the intra-modal fine-grained semantic relations distilled by single-modal pretrained models to constrain hierarchical cross-modal association between image and text modalities, for accurate TIPR. Specifically, to model hierarchical cross-modal semantic relationships, we propose a Hierarchical Relevance Matching (HRM) module. It partitions the matching strength of image-text pairs by jointly considering identity labels and cross-modal similarity, collaborating with unimodal similarity to construct a hierarchical relevance distribution that serves as a soft supervision signal. HRM not only helps the model better capture varying levels of semantic consistency between image-text pairs but also enhances the overall accuracy of cross-modal association learning. To enhance the ability to capture fine-grained cross-modal semantic relationships, we introduce an Image-guided Ambiguous text Token Modeling (IATM) module. It replaces original tokens with semantically ambiguous ones and leverages image guidance to detect and correct these tokens. This process further improves the fine-grained semantic alignment between images and texts. Experimental results demonstrate that HCA achieves new state-of-the-art performance across multiple datasets, thoroughly validating its effectiveness and advancement in cross-modal retrieval tasks.
Yifei Deng, Chenglong Li 0002, Futian Wang, Jin Tang 0001
ACM Multimedia4
2025 CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework
abstract
Event cameras have attracted increasing attention in recent years due to their advantages in high dynamic range, high temporal resolution, low power consumption, and low latency. Some researchers have begun exploring pre-training directly on event data. Nevertheless, these efforts often fail to establish strong connections with RGB frames, limiting their applicability in multi-modal fusion scenarios. To address these issues, we propose a novel CM3AE pre-training framework for the RGB-Event perception. This framework accepts multi-modalities/views of data as input, including RGB images, event images, and event voxels, providing robust support for both event-based and RGB-event fusion based downstream tasks. Specifically, we design a multi-modal fusion reconstruction module that reconstructs the original image from fused multi-modal features, explicitly enhancing the model's ability to aggregate cross-modal complementary information. Additionally, we employ a multi-modal contrastive learning strategy to align cross-modal feature representations in a shared latent space, which effectively enhances the model's capability for multi-modal understanding and capturing global dependencies. We construct a large-scale dataset containing 2,535,759 RGB-Event data pairs for the pre-training. Extensive experiments on five downstream tasks fully demonstrated the effectiveness of CM3AE. Source code and pre-trained models will be released on https://github.com/Event-AHU/CM3AE.
Xiao Wang 0014, Chenglong Li 0002, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Qi Liu 0003
ACM Multimedia5
2025 UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-Identification
abstract
Multi-modal object Re-IDentification (ReID) has gained considerable attention with the goal of retrieving specific targets across cameras using heterogeneous visual data sources. At present, multi-modal object ReID faces two core challenges: (1) learning robust features under fine-grained local noise caused by occlusion, frame loss, and other disruptions; and (2) effectively integrating heterogeneous modalities to enhance multi-modal representation. To address the above challenges, we propose a robust approach named Uncertainty-Guided Graph model for multi-modal object ReID (UGG-ReID). UGG-ReID is designed to mitigate noise interference and facilitate effective multi-modal fusion by estimating both local and sample-level aleatoric uncertainty and explicitly modeling their dependencies. Specifically, we first propose the Gaussian patch-graph representation model that leverages uncertainty to quantify fine-grained local cues and capture their structural relationships. This process boosts the expressiveness of modal-specific information, ensuring that the generated embeddings are both more informative and robust. Subsequently, we design an uncertainty-guided mixture of experts strategy that dynamically routes samples to experts exhibiting low uncertainty. This strategy effectively suppresses noise-induced instability, leading to enhanced robustness. Meanwhile, we design an uncertainty-guided routing to strengthen the multi-modal interaction, improving the performance. UGG-ReID is comprehensively evaluated on five representative multi-modal object ReID datasets, encompassing diverse spectral modalities. Experimental results show that the proposed method achieves excellent performance on all datasets and is significantly better than current methods in terms of noise immunity. Our code is available at https://github.com/wanxixi11/UGG-ReID.
Xixi Wan, Aihua Zheng, Bo Jiang 0002, Beibei Wang 0006, Chenglong Li 0002, Jin Tang 0001
NeurIPS6
2025 AttentionGRN: a functional and directed graph transformer for gene regulatory network reconstruction from scRNA-seq data
abstract
Single-cell RNA sequencing (scRNA-seq) enables the reconstruction of cell type-specific gene regulatory networks (GRNs), offering detailed insights into gene regulation at high resolution. While graph neural networks have become widely used for GRN inference, their message-passing mechanisms are often limited by issues such as over-smoothing and over-squashing, which hinder the preservation of essential network structure. To address these challenges, we propose a novel graph transformer-based model, AttentionGRN, which leverages soft encoding to enhance model expressiveness and improve the accuracy of GRN inference from scRNA-seq data. Furthermore, the GRN-oriented message aggregation strategies are designed to capture both the directed network structure information and functional information inherent in GRNs. Specifically, we design directed structure encoding to facilitate the learning of directed network topologies and employ functional gene sampling to capture key functional modules and global network structure. Our extensive experiments, conducted on 88 datasets across two distinct tasks, demonstrate that AttentionGRN consistently outperforms existing methods. Furthermore, AttentionGRN has been successfully applied to reconstruct cell type-specific GRNs for human mature hepatocytes, revealing novel hub genes and previously unidentified transcription factor-target gene regulatory associations.
Yansen Su, Jin Tang 0001, Huaiwan Jin, Yun Ding, Pi-Jing Wei, Chun-Hou Zheng 0001
Briefings Bioinform.3
2025 Visual and text prompt learning for multi-modal brain disease diagnosis
Yumiao Zhao, Bo Jiang 0002, Yuhe Ding, Xixi Wan, Jin Tang 0001
Sci. China Inf. Sci.5
2025 Hybrid graph-based radiology report generation
Dengdi Sun, Chaofan Mu, Xuyang Fan, Jin Tang 0001, Zegeng Li, Zhuanlian Ding
Expert Syst. Appl.4
2025 Modality-missing RGBT Tracking: Invertible Prompt Learning and High-quality Benchmarks
Andong Lu, Chenglong Li 0002, Jiacong Zhao, Jin Tang 0001, Bin Luo 0001
Int. J. Comput. Vis.4
2025 Lightweight oriented object detection with Dynamic Smooth Feature Fusion Network
Wei Lu 0032, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001
Neurocomputing4
2025 CMCNet:Cross-directional morphology-aware convolution network for chest X-ray anatomy segmentation
Lili Huang 0006, Yuhan Feng, Chenglong Li 0002, Jin Tang 0001
Neurocomputing5
2025 Multi-view enhanced truck re-identification
Xue-Yan Wang, Run-Sen Xia, Sibao Chen 0001, Jin Tang 0001
Knowl. Based Syst.5
2025 Graph Spiking Attention Network: Sparsity, Efficiency and Robustness
abstract
Existing Graph Attention Networks (GATs) generally adopt the self-attention mechanism to learn graph edge attention, which usually return dense attention coefficients over all neighbors and thus are prone to be sensitive to graph edge noises. To overcome this problem, sparse GATs are desirable and have garnered increasing interest in recent years. However, existing sparse GATs usually suffer from high training complexity and are also not straightforward for inductive learning tasks. To address these issues, we propose to learn sparse GATs by exploiting spiking neuron (SN) mechanism, termed Graph Spiking Attention (GSAT). Specifically, it is known that spiking neuron can perform inexpensive information processing by transmitting the input data into discrete spike trains and return sparse outputs. Inspired by it, this work attempts to exploit spiking neuron to learn sparse attention coefficients, resulting in edge-sparsified graph for GNNs. Therefore, GSAT can perform message passing on the selective neighbors naturally, which makes GSAT perform compactly and robustly w.r.t graph noises. Moreover, GSAT can be used straightforwardly for inductive learning tasks. Extensive experiments on both transductive and inductive tasks demonstrate the effectiveness, robustness and efficiency of GSAT.
Beibei Wang 0006, Bo Jiang 0002, Jin Tang 0001, Lu Bai 0001, Bin Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Unifying Graph Contrastive Learning via Graph Message Augmentation
abstract
Graph contrastive learning is usually performed by first conducting Graph Data Augmentation (GDA) and then employing a contrastive learning pipeline to train GNNs. As we know that GDA is an important issue for graph contrastive learning. Various GDAs have been developed recently which mainly involve dropping or perturbing edges, nodes, node attributes and edge attributes. However, to our knowledge, it still lacks a universal and effective augmentor that is suitable for different types of graph data. To address this issue, in this paper, we first introduce the graph message representation of graph data. Based on it, we then propose a novel Graph Message Augmentation (GMA), a universal scheme for reformulating many existing GDAs. The proposed unified GMA not only gives a new perspective to understand many existing GDAs but also provides a universal and more effective graph data augmentation for graph self-supervised learning tasks. Moreover, GMA introduces an easy way to implement the mixup augmentor which is natural for images but usually challengeable for graphs. Based on the proposed GMA, we then propose a unified graph contrastive learning, termed Graph Message Contrastive Learning (GMCL), that employs attribution-guided universal GMA for graph contrastive learning. Experiments on many graph learning tasks demonstrate the effectiveness and benefits of the proposed GMA and GMCL approaches.
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 FSENet: Feature suppression and enhancement network for tiny object detection
Heng Hu, Sibao Chen 0001, Zhi-Hui You, Jin Tang 0001
Pattern Recognit.4
2025 Visible-thermal multiple object tracking: Large-scale video dataset and progressive fusion approach
Yabin Zhu, Qianwu Wang, Chenglong Li 0002, Jin Tang 0001, Chengjie Gu, Zhixiang Huang
Pattern Recognit.4
2025 Pedestrian Attribute Recognition via CLIP-Based Prompt Vision-Language Fusion
abstract
Existing pedestrian attribute recognition (PAR) algorithms adopt pre-trained CNN (e.g., ResNet) as their backbone network for visual feature learning, which might obtain sub-optimal results due to the insufficient employment of the relations between pedestrian images and attribute labels. In this paper, we formulate PAR as a vision-language fusion problem and fully exploit the relations between pedestrian images and attribute labels. Specifically, the attribute phrases are first expanded into sentences, and then the pre-trained vision-language model CLIP is adopted as our backbone for feature embedding of visual images and attribute descriptions. The contrastive learning objective connects the vision and language modalities well in the CLIP-based feature space, and the Transformer layers used in CLIP can capture the long-range relations between pixels. Then, a multi-modal Transformer is adopted to fuse the dual features effectively and feed-forward network is used to predict attributes. To optimize our network efficiently, we propose the region-aware prompt tuning technique to adjust very few parameters (i.e., only the prompt vectors and classification heads) and fix both the pre-trained VL model and multi-modal Transformer. Our proposed PAR algorithm only adjusts 0.75% learnable parameters compared with the fine-tuning strategy. It also achieves new state-of-the-art performance on both standard and zero-shot settings for PAR, including RAPv1, RAPv2, WIDER, PA100K, and PETA-ZS, RAP-ZS datasets. The source code and pre-trained models will be released onhttps://github.com/Event-AHU/OpenPAR.
Xiao Wang 0014, Jiandong Jin, Chenglong Li 0002, Jin Tang 0001, Cheng Zhang 0010, Wei Wang 0115
IEEE Trans. Circuits Syst. Video Technol.4
2025 Camera-Proxy Enhanced Identity-Recalibration Learning for Unsupervised Visible-Infrared Person Re-Identification
abstract
Visible-Infrared person Re-Identification (VI-ReID) involves querying images of the same person across visible and infrared modalities. To minimize annotation costs, Unsupervised Visible-Infrared person Re-Identification (UVI-ReID) using pseudo-label contrastive learning has emerged. Traditional UVI-ReID approaches often neglected camera domain information and relied on inadequate update strategies during training, only using cosine distance for testing, which led to incorrect mapping of cross-modal relationships. To address these issues, we propose Camera-proxy Enhanced Identity-recalibration Learning (CEIL). It consists of two main stages: first, it employs intra-modal contrastive learning in conjunction with the camera-proxy, updates the memory bank using our innovative Difficulty-aware Cluster-based Memory Updating (DCMU) strategy, and applies Camera Domain-driven Local correlation (CDL) Loss to enhance the learning process. Then utilizes cross-modal contrastive learning, featuring our Proxy-enhanced Cross-modal Mapping (PCM) module, to recalibrate the identity relationships between different modalities. Graph network-based Camera constraint adjustment Re-ranking (GCR) method is adopted during test, utilizing camera domain information to recalibrate the correspondence between identities. Extensive experiments have demonstrated that CEIL achieving state-of-the-art performance on the SYSU-MM01, RegDB, and LLCM datasets and the GCR, as a general unsupervised re-ranking method, can further enhance performance of model on these datasets. The code will be released athttps://github.com/maybeextra/CEIL.
Run-Sen Xia, Xue-Yan Wang, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 CFENet: Contextual Feature Enhancement Network for Tiny Object Detection in Aerial Images
abstract
With the development of deep learning techniques and object detectors, the performance of object detection has been rapidly improved. However, since tiny objects contain only a small number of pixels and lack appearance information, this creates difficulties for detector recognition. Although existing research has improved detection performance by fusing different feature layers to enhance feature information of objects, this also leads to the problem of mixed feature information, especially for tiny objects where features are easily covered, which exacerbates the difficulty of recognition. To solve the above problems, we propose a contextual feature enhancement network (CFENet), which is an efficient framework built on anchor-based object detectors. In CFENet, to effectively utilize contextual information around an object to enhance the detection of tiny objects, we use poolFormer to build a backbone to extract object features. To alleviate the feature blending problem caused by feature fusion, we propose a feature suppression module (FSM) that effectively suppresses background information and redundant features to enhance tiny object features. In addition, we utilize the improved Gaussian Wasserstein distance loss to modify the loss function to obtain high-quality bounding boxes, and we further manipulate the shallow feature layer of the output and then add a detection head to enhance the detection of tiny objects. We have conducted extensive experiments on the public datasets AI-TOD, VisDrone, and DOTA to demonstrate the effectiveness of our approach.
Heng Hu, Sibao Chen 0001, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Multiscale Adaptive Decoder and Diversity Selection Network for Road Extraction in Remote Sensing Image
abstract
Road extraction has been a common and challenging task in the field of remote sensing images. Due to factors such as the high resolution of remote sensing images and the subtle visibility of road features, existing methods often miss certain areas during detection and extraction. These methods struggle to capture contextual information effectively and tend to exhibit false positives and false negatives when handling objects of varying sizes. This article proposes a network based on a multi-scale adaptive decoder and diverse selection (MADSNet) to address the issue of inadequate contextual information capture. By leveraging feature diverse selection, the method minimizes errors in distinguishing between road features and background interference. Specifically, the multi-scale feature flexible extraction (MFFE) decoder utilizes the relevance inquiry attention (RIA) module and scope flexible fusion (SFF) module to enhance the ability to capture contextual information with relatively low computational demands. The optimal choice graph attention (OCGA) module aggregates neighboring nodes with similar features in a graph structure, improving focus on the single class of roads. Furthermore, a multi-level feature selection (MFS) module is proposed to activate the features relevant to the current stage while suppressing features from other stages and interfering with noise. Quantitative and qualitative experimental results on three public datasets demonstrate that the proposed MADSNet outperforms currently popular methods in terms of performance. The code will be available at https://github.com/Talent02/MADSNet.
Zhen-Tao Hua, Sibao Chen 0001, Wei Lu 0032, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Multidimensional Remote Sensing Change Detection Based on Siamese Dual-Branch Networks
abstract
Deep learning models, particularly convolutional neural networks (CNNs), have demonstrated outstanding feature learning capabilities, leading to remarkable performance in remote sensing change detection (RSCD) tasks. However, their most critical drawback lies in the lack of effective modeling of global information. This deficiency affects the model’s understanding of the overall context and structure of the entire image, making it difficult to distinguish between background and target areas, thereby leading to the erroneous identification of change regions. Second, features extracted by traditional backbone networks contain a significant amount of noise, resulting in blurred boundaries of changed objects. The challenge of effectively fusing detailed and semantic information to accurately differentiate pseudo changes remains significant. Furthermore, how to fully exploit multiscale information is another issue worth considering. We propose a full-scale multidimensional interaction network called SDSN, which enhances feature representation by leveraging both detail and semantic branches. Initially, bi-temporal images are processed by the encoder to extract coarse multiscale features. The semantic branch guides shallow-scale features, while the detail branch focuses on deep-scale features. Multikernel receptive module (MRM) aggregates global information. The detail branch utilizes a diversity variance module (DVM) and differential operations to generate refined change maps with noise reduction and background suppression. A multidimensional cross-perception module (MCM) guides the fusion of these change maps, establishing multidimensional dependencies to enrich feature representation. Compared with previous methods, SDSN demonstrates greater performance under complex environmental conditions, particularly noteworthy for its fewer parameters (4.03 M) and lower computational costs (7.94 G). The code is publicly available athttps://github.com/dpt000121/dpt.
Li-Rong Shen, Sibao Chen 0001, Lili Huang 0006, Zhi-Hui You, Chris Ding, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 UAV Video Vehicle Detection: Benchmark and Baseline
abstract
With the increasing application of unmanned aerial vehicles (UAVs) in intelligent transportation systems, vehicle object detection in UAV videos has received increasing attention. Precise categorization and detection for vehicles in UAVs is important in many practical applications. However, existing object detection methods, tailored for natural images, often fall short of accurately identifying vehicle objects. Additionally, high-altitude UAV imaging mainly employs horizontal bounding box annotation, frequently leading to significant obstruction and overlapping. Hence, we propose a new task called UAV video vehicle detection (VVD) to achieve precise detection and categorization of vehicles in high-altitude UAV imaging environments. To facilitate the research and development of UAV VVD, we construct the first large-scale well-annotated benchmark UAV VVD dataset, which includes 70 UAV videos captured at a 500-m altitude, with 361489 vehicle instances annotated by the oriented bounding boxes and vehicle categories. Moreover, we introduce a novel category refinement network (CRNet) approach that extracts and refines vehicle object features from the bounding box of the detection results to classify vehicle categories. This approach effectively eliminates the interference of the background and other vehicle objects in candidate boxes. Notably, the vehicle object features are projected into subspace, enabling the category refinement module (CRM) to focus more on the distinctive characteristics of the vehicle object itself through normalization operations. We conduct extensive experiments on the proposed VVD dataset. Experimental results demonstrate the superiority and effectiveness of the proposed CRNet method. The relevant code and dataset are available athttps://github.com/mmic-lcl.
Yun Xiao 0003, Jinfa Wang, Zhicheng Zhao 0002, Bo Jiang 0002, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 Multimodal Remote Sensing Image Registration via Modality Perception and Self-Supervised Position Estimation
abstract
Multi-modal remote sensing images registration ensures that images from different sensors or modalities are spatial and informational consistent for effective comparison and analysis. However, due to the non-linear modality gaps that exist between images, making it difficult to focus only on the spatial position differences of the images and ignore the modality gaps. In this paper, to address this issue, we propose a new framework for Multi-Modal remote sensing image Registration, named MMRNet. The proposed framework comprises the following main aspects. First, a novel self-supervised Positional Misalignment Estimator (PME) is designed for multi-modal image registration. PME is able to efficiently overcome the modality gaps and learn the positional differences between multi-modal images more reliably, optimizing the registration loss by minimizing the positional differences directly. Then, a new paradigm of modality translation, termed Modality Perception Module (MPM), is introduced to effectively learn modality gaps and perform modality translation in the case of positional misalignment. Finally, we further design the modality perception guidance loss to supervise the modality translation task, which can encourage the fidelity of the generated pseudo-modality images. Our registration network integrates both rigid registration model and non-rigid registration model. Experimental results demonstrate that the proposed registration framework can obtain obviously superior performance in both rigid and non-rigid image registration tasks on optical-SAR data, optical-map data and optical-infrared data. The code and relevant dataset will be made publicly available at https://github.com/Ahuer-Lei/MMRNet.
Yun Xiao 0003, Bo Jiang 0002, Yuan Chen 0012, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 S2FCNet: Semantic and Spatial Feature Compensation Network for Tiny-Object Detection
abstract
Tiny object detection (TOD) in remote sensing images remains an extremely challenging task, primarily due to the severely limited feature availability and susceptibility to interference from complex background. Recently, the multi-scale feature based methods have demonstrated effectiveness in tiny object detection. However, they often neglect that the low-level features struggle to activate the discriminative local semantics and lack global semantic information due to the limited local receptive fields. To address these issues, this paper proposes a Semantic and Spatial Feature Compensation Network (S2FCNet) for tiny object detection. To mitigate the gradual degradation of semantic information from high-level to low-level features in multi-scale representations, we propose a Local Semantic Reactivation Module (LSRM), which reactivates low-level local semantic features through top-down guidance from high-level semantic features. To enhance spatial perception capabilities, we develop a Foreground Spatial Sense Module (FSSM) that captures precise spatial location information, effectively suppresses the background noise and enhances the foreground features. Meanwhile, we introduce a Spatial Guidance Mechanism (SGM) to compensate for the loss of spatial awareness caused by downsampling operations. Additionally, we synergistically combine semantic and spatial features through a Multi-level Fusion Mechanism (MLFM), enabling more accurate detection of tiny objects. Extensive experiments on three challenging datasets demonstrate the effectiveness and superiority of the S2FCNet in comparison with the state-of-the-art methods. The code will be released at https://github.com/DetectionTiny/S2FCNet.
Yuhui Zhang 0005, Zhicheng Zhao 0001, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Reflectance-Guided Progressive Feature Alignment Network for All-Day UAV Object Detection
abstract
Object detection using visible-infrared images has become increasingly crucial for all-day applications of unmanned aerial vehicle (UAV). However, existing multi-modal detection methods face significant challenges in low-light conditions, where degraded visible image quality exacerbates weak alignment issues and compromises feature fusion effectiveness. Although recent approaches have attempted to address these issues through cross-attention mechanisms or feature alignment strategies, they often suffer from unstable performance and limited generalization capability in challenging nighttime scenarios. To address these limitations, we propose a novel Reflectance-Guided Progressive Feature Alignment Network (RGFNet) for robust UAV object detection. Our proposed method leverages the illumination-invariant characteristic of reflectance features decomposed from visible images via Retinex theory to guide cross-modal alignment and fusion. Specifically, we design a Reflectance-Guided Collaborative Alignment Module (RCAM) that utilizes reflectance guidance to perform bidirectional feature alignment between visible and infrared modalities, effectively reducing position misalignment under varying lighting conditions. Furthermore, we introduce a Light-Aware Selective Fusion Module (LSFM) that maps multi-modal features into a shared hidden state space through selective state space mechanism, enabling efficient feature interaction while maintaining linear computational complexity. Extensive experiments on two challenging UAV detection benchmarks, DroneVehicle and DVTOD, demonstrate the superiority of our method. RGFNet achieves state-of-the-art performance with 81.4% mAP on DroneVehicle and 88.5% mAP on DVTOD. The code is available at https://github.com/uavdet/RGFNet.
Zhicheng Zhao 0002, Wei Zhang 0393, Yun Xiao 0003, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Real-World Remote Sensing Image Dehazing: Benchmark and Baseline
abstract
Remote Sensing Image Dehazing (RSID) poses significant challenges in real-world scenarios due to the complex atmospheric conditions and severe color distortions that degrade image quality. The scarcity of real-world remote sensing hazy image pairs has compelled existing methods to rely primarily on synthetic datasets. However, these methods struggle with real-world applications due to the inherent domain gap between synthetic and real data. To address this, we introduce Real-World Remote Sensing Hazy Image Dataset (RRSHID), the first large-scale dataset featuring real-world hazy and hazy-free image pairs across diverse atmospheric conditions. Based on this, we propose MCAF-Net, a novel framework tailored for real-world RSID. Its effectiveness arises from three innovative components: Multi-branch Feature Integration Block Aggregator (MFIBA), which enables robust feature extraction through cascaded integration blocks and parallel multi-branch processing; Color-Calibrated Self-Supervised Attention Module (CSAM), which mitigates complex color distortions via self-supervised learning and attention-guided refinement; and Multi-Scale Feature Adaptive Fusion Module (MFAFM), which integrates features effectively while preserving local details and global context. Extensive experiments validate that MCAF-Net demonstrates state-of-the-art performance in real-world RSID, while maintaining competitive performance on synthetic datasets. The introduction of RRSHID and MCAF-Net sets new benchmarks for real-world RSID research, advancing practical solutions for this complex task. The code and dataset are publicly available at here.
Zeng-Hui Zhu, Wei Lu 0032, Sibao Chen 0001, Chris Ding, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Federated Client-Tailored Adapter for Medical Image Segmentation
abstract
Medical image segmentation in X-ray images is beneficial for computer-aided diagnosis and lesion localization. Existing methods mainly fall into a centralized learning paradigm, which is inapplicable in the practical medical scenario that only has access to distributed data islands. Federated Learning has the potential to offer a distributed solution but struggles with heavy training instability due to client-wise domain heterogeneity (including distribution diversity and class imbalance). In this paper, we propose a novel Federated Client-tailored Adapter (FCA) framework for medical image segmentation, which achieves stable and client-tailored adaptive segmentation without sharing sensitive local data. Specifically, the federated adapter stirs universal knowledge in off-the-shelf medical foundation models to stabilize the federated training process. In addition, we develop two client-tailored federated updating strategies that adaptively decompose the adapter into common and individual components, then globally and independently update the parameter groups associated with common client-invariant and individual client-specific units, respectively. They further stabilize the heterogeneous federated learning process and realize optimal client-tailored instead of sub-optimal global-compromised segmentation models. Extensive experiments on three large-scale datasets demonstrate the effectiveness and superiority of the proposed FCA framework for federated medical segmentation.
Guyue Hu 0001, Siyuan Song, Yukun Kang, Zhu Yin, Gangming Zhao, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Inf. Forensics Secur.7
2025 Nighttime Person Re-Identification via Collaborative Enhancement Network With Multi-Domain Learning
abstract
Prevalent nighttime person re-identification (ReID) methods typically combine image relighting and ReID networks in a sequential manner. However, their performance (recognition accuracy) is limited by the quality of relighting images and insufficient collaboration between image relighting and ReID tasks. To handle these problems, we propose a novel Collaborative Enhancement Network called CENet, which performs the multilevel feature interactions in a parallel framework, for nighttime person ReID. In particular, the designed parallel structure of CENet can not only avoid the impact of the quality of relighting images on ReID performance, but also allow us to mine the collaborative relations between image relighting and person ReID tasks. To this end, we integrate the multilevel feature interactions in CENet, where we first share the Transformer encoder to build the low-level feature interaction, and then perform the feature distillation that transfers the high-level features from image relighting to ReID, thereby alleviating the severe image degradation issue caused by the nighttime scenario while avoiding the impact of relighting images. In addition, the sizes of existing real-world nighttime person ReID datasets are limited, and large-scale synthetic ones exhibit substantial domain gaps with real-world data. To leverage both small-scale real-world and large-scale synthetic training data, we develop a multi-domain learning algorithm, which alternately utilizes both kinds of data to reduce the inter-domain difference in training procedure. Extensive experiments on two real nighttime datasets,Night600andRGBNT201rgb, and a synthetic nighttime ReID dataset are conducted to validate the effectiveness of CENet. We release the code and synthetic dataset at: https://github.com/Alexadlu/CENet.
Andong Lu, Chenglong Li 0002, Tianrui Zha, Xiaofeng Wang 0009, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Inf. Forensics Secur.5
2025 Prototype-Based Diversity and Integrity Learning for All-Day Multi-Modal Person Re-Identification
abstract
Recent multi-modal person re-identification methods have improved model performance by leveraging complementary information from multiple spectra. However, existing methods cannot ensure feature stability under varying illumination and rely on inflexible paired data, remaining inadequate against real-world cross-time retrieval and modality-missing challenges. To solve these, we first propose diversity representation that augments illumination-sensitive images to simulate diverse lighting conditions via illumination augmentation and enriches instance features using modality-specific prototypes via multiple interaction modules. Secondly, we propose integrity reconstruction that leverages prototypes and available instance features to recover information, the reconstruction module effectively utilizes identity and modality cues to address unpredictable missing problems. In addition, we build a more comprehensive dataset (AllDay843) to alleviate the inadequate dataset diversity, which comprises 91,371 images of 843 identities captured by multi-modal cameras across various periods throughout the day, while incorporating numerous real-world challenges. By integrating diversity representation and integrity reconstruction, the proposed Prototype-Based Diversity and Integrity learning network (PDINet) establishes excellence on the AllDay843 dataset, surpassing existing state-of-the-art approaches. The data and codes are available in https://github.com/ziwang1121/PDINet.
Zi Wang 0013, Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Inf. Forensics Secur.5
2025 Adaptive Interaction and Correction Attention Network for Audio-Visual Matching
abstract
Audio-visual matching techniques aim to recognize and match information across different identities by learning a similarity metric across modalities. However, modal differences arise from insufficient cross-modal correlations and noise interference, which substantially hinder the performance of traditional deep metric learning methods in audio-visual matching tasks. To address the modal differences issue, we propose a novel Adaptive Interactive and Correction Attention Network (AICANet). This network efficiently captures deep information connections, generating modality-consistent feature embeddings within a unified metric framework. The core of AICANet is its two-pronged approach to reducing modal differences. First, we propose the Adaptive Interactive Attention (AIA) module, which flexibly establishes associations among cross-modal local features using dynamically generated pseudo-labels. Second, we propose the Adaptive Correction Attention (ACA) mechanism, which employs an adaptive threshold to de-interference effectively and accurately adjust the representation of local feature associations. Notably, the ACA mechanism is suitable for both intra-modal and inter-modal refined attention correction. Additionally, we design a relative distance stretching metric loss (LRDSM), which reinforces the similarity invariance of feature embeddings in a uniform space and enhances matching accuracy. Extensive tests on the VoxCeleb and VoxCeleb2 datasets demonstrate that AICANet outperforms leading existing algorithms across several evaluation metrics, validating its superior performance. The codes can be found at https://github.com/w1018979952/AICANet.
Jiaxiang Wang 0001, Aihua Zheng, Lei Liu 0049, Chenglong Li 0002, Ran He 0001, Jin Tang 0001
IEEE Trans. Inf. Forensics Secur.6
2025 AFTER: Attention-Based Fusion Router for RGBT Tracking
abstract
Multi-modal feature fusion as a core investigative component of RGBT tracking emerges numerous fusion studies in recent years. However, existing RGBT tracking methods widely adopt fixed fusion structures to integrate multi-modal feature, which are hard to handle various challenges in dynamic scenarios. To address this problem, this work presents a novel Attention-based Fusion router called AFTER, which optimizes the fusion structure to adapt to the dynamic challenging scenarios, for robust RGBT tracking. In particular, we design a fusion structure space based on the hierarchical attention network, each attention-based fusion unit corresponding to a fusion operation and a combination of these attention units corresponding to a fusion structure. Through optimizing the combination of attention-based fusion units, we can dynamically select the fusion structure to adapt to various challenging scenarios. Unlike complex search of different structures in neural architecture search algorithms, we develop a dynamic routing algorithm, which equips each attention-based fusion unit with a router, to predict the combination weights for efficient optimization of the fusion structure. Extensive experiments on five mainstream RGBT tracking datasets demonstrate the superior performance of the proposed AFTER against state-of-the-art RGBT trackers. We release the code in https://github.com/Alexadlu/AFter.
Andong Lu, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Image Process.4
2025 Dynamic Strip Convolution and Adaptive Morphology Perception Plugin for Medical Anatomy Segmentation
abstract
Medical anatomy segmentation is essential for computer-aided diagnosis and lesion localization in medical images. For example, segmenting individual ribs benefits localizing the lung lesions and providing vital medical measurements (such as rib spacing) for generating medical reports. Existing methods segment shape-different anatomies (such as striped ribs, bulky lungs, and angular scapula) with the same network architecture, the morphology heterogeneity is heavily overlooked. Although some shape-aware operators like deformable convolution and dynamic snake convolution have been introduced to cater to specific object morphology, they still struggle with orientation-varying strip structures, such as 24 ribs and 2 clavicles. In this paper, we propose a novel convolution plugin (DSC-AMP) for medical anatomy segmentation, which is comprised of a dynamic strip convolution (DSC) operator and an adaptive morphology perception (AMP) strategy. Specifically, the dynamic strip convolution customizes gradually varying directions and offsets for each local region, achieving dynamic striped receptive fields. Additionally, the adaptive morphology perception strategy incorporates insights from various shape-aware convolutional kernels, enabling the model to discern and integrate crucial representations corresponding to heterogeneous anatomies. Extensive experiments on two large-scale datasets demonstrate the effectiveness and superiority of the proposed approach for tackling heterogeneous medical anatomy segmentation.
Guyue Hu 0001, Yukun Kang, Gangming Zhao, Zhe Jin 0001, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Medical Imaging6
2025 Retain, Blend, and Exchange: A Quality-Aware Spatial-Stereo Fusion Approach for Event Stream Recognition
abstract
Current event stream-based pattern recognition models typically present the event stream as the point cloud, voxel, image, and the like, and formulate multiple deep neural networks to acquire their features. Although considerable results can be achieved in simple cases, however, the performance of the model might be restricted by monotonous modality expressions, sub-optimal fusion, and readout mechanisms. In this article, we put forward a novel dual-stream framework for event stream-based pattern recognition through differentiated fusion, which is called EFV++. It models two common event representations simultaneously, i.e., event images and event voxels. The spatial and three-dimensional stereo information can be separately learned by making use of Transformer and Graph Neural Network (GNN). We believe the features of each representation still contain both efficient and redundant features and a sub-optimal solution may be obtained if we directly fuse them without differentiation. Thus, we divide each feature into three levels and retain high-quality features, blend medium-quality features, and exchange low-quality features. The enhanced dual features will be provided to the fusion Transformer together with bottleneck features. In addition, we introduce a novel hybrid interaction readout mechanism to enhance the diversity of features as final representations. Comprehensive experiments validate that the framework we have proposed attains cutting-edge performance on a variety of extensively utilized event stream-based classification datasets. Particularly, we have realized a freshly pioneering performance on the Bullying10 k dataset, precisely 90.51%, and this outpaces the runner-up by$+2.21\%$.
Lan Chen 0003, Xiao Wang 0014, Pengpeng Shao, Wei Zhang 0161, Yaowei Wang 0001, Yonghong Tian 0001, Jin Tang 0001
IEEE Trans. Multim.8
2025 Knowledge-Guided Cross-Modal Alignment and Progressive Fusion for Chest X-Ray Report Generation
abstract
The task of chest X-ray report generation, which aims to simulate the diagnosis process of doctors, has received widespread attention. Compared with the image caption task, chest X-ray report generation is more challenging since it needs to generate a longer and more accurate description of each diagnostic part in chest X-ray images. Most of existing works focus on how to extract better visual features or more accurate text expression based on existing reports. However, they ignore the interactions between visual and text modalities and are thus obviously not in line with human thinking. A small part of works explore the interactions of visual and text modalities, but data-driven learning of cross-modal information mapping can not break the semantic gap between different modalities. In this work, we propose a novel approach called Knowledge-guided Cross-modal Alignment and Progressive fusion (KCAP), which takes the knowledge words from a created medical knowledge dictionary as the bridge to guide the cross-modal feature alignment and fusion, for accurate chest X-ray report generation. In particular, we create the medical knowledge dictionary by extracting medical phrases from the training set and then selecting some phrases with substantive meanings as knowledge words based on their frequency of occurrence. Based on the knowledge words from the medical knowledge dictionary, the visual and text modalities are interacted by a mapping layer for the enhancement of the features of two modalities, and then the alignment fusion module is introduced to mitigate the semantic gap between visual and text modalities. To retain the important details of the original information, we design a progressive fusion scheme to integrate the advantages of both salient fused and original features to generate better medical reports. The experimental results on IU-Xray and MIMIC datasets demonstrate the effectiveness of the proposed KCAP.
Lili Huang 0006, Pengcheng Jia, Chenglong Li 0002, Jin Tang 0001, Chuanfu Li
IEEE Trans. Multim.5
2025 CRSOT: Cross-Resolution Object Tracking Using Unaligned Frame and Event Cameras
abstract
Existing datasets for RGB-DVS tracking are collected with DVS346 camera and their resolution ($346 \times 260$) is low for practical applications. Actually, only visible cameras are deployed in many practical systems, and the newly designed neuromorphic cameras may have different resolutions. The latest neuromorphic sensors can output high-definition event streams, but it is very difficult to achieve strict alignment between events and frames on both spatial and temporal views. Therefore, how to achieve accurate tracking with unaligned neuromorphic and visible sensors is a valuable but unresearched problem. In this work, we formally propose the task of object tracking using unaligned neuromorphic and visible cameras. We build the first unaligned frame-event dataset CRSOT collected with a specially built data acquisition system, which contains 1,030 high-definition RGB-Event video pairs, 304,974 video frames. In addition, we propose a novel unaligned object tracking framework that can realize robust tracking even using the loosely aligned RGB-Event data. This proposed method utilizes uncertainty perception techniques, which can effectively reduce the negative impact of noise (especially noise in event data) on tracking performance. Specifically, we extract the template and search regions of RGB and Event data and feed them into a unified ViT backbone for feature embedding. Next, we propose uncertainty perception modules to encode the RGB and Event features, respectively, then, we propose a modality uncertainty fusion module to aggregate the two modalities. These three branches are jointly optimized in the training phase. Extensive experiments demonstrate that our tracker can collaborate the dual modalities for high-performance tracking even without strictly temporal and spatial alignment.
Yabin Zhu, Xiao Wang 0014, Chenglong Li 0002, Bo Jiang 0002, Lin Zhu 0012, Zhixiang Huang, Yonghong Tian 0001, Jin Tang 0001
IEEE Trans. Multim.8
2025 Cross-Modal Object Tracking via Modality-Aware Fusion Network and a Large-Scale Dataset
abstract
Visual object tracking often faces challenges such as invalid targets and decreased performance in low-light conditions when relying solely on RGB image sequences. While incorporating additional modalities like depth and infrared data has proven effective, existing multimodal imaging platforms are complex and lack real-world applicability. In contrast, near-infrared (NIR) imaging, commonly used in surveillance cameras, can switch between RGB and NIR based on light intensity. However, tracking objects across these heterogeneous modalities poses significant challenges, particularly due to the absence of modality switch signals during tracking. To address these challenges, we propose an adaptive cross-modal object tracking algorithm called modality-aware fusion network (MAFNet). MAFNet efficiently integrates information from both RGB and NIR modalities using an adaptive weighting mechanism, effectively bridging the appearance gap and enabling a modality-aware target representation. It consists of two key components: an adaptive weighting module and a modality-specific representation module. The adaptive weighting module predicts fusion weights to dynamically adjust the contribution of each modality, while the modality-specific representation module captures discriminative features specific to RGB and NIR modalities. MAFNet offers great flexibility as it can effortlessly integrate into diverse tracking frameworks. With its simplicity, effectiveness, and efficiency, MAFNet outperforms state-of-the-art methods in cross-modal object tracking. To validate the effectiveness of our algorithm and overcome the scarcity of data in this field, we introduce CMOTB, a comprehensive and extensive benchmark dataset for cross-modal object tracking. CMOTB consists of 61 categories and 1000 video sequences, comprising a total of over 799K frames. We believe that our proposed method and dataset offer a strong foundation for advancing cross-modal object-tracking research. The dataset, toolkit, experimental data, and source code will be publicly available at: https://github.com/mmic-lcl/ Datasets-and-benchmark-code.
Lei Liu 0049, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 Duality-Gated Mutual Condition Network for RGBT Tracking
abstract
Low-quality modalities contain not only a lot of noisy information but also some discriminative features in RGB-Thermal (RGBT) tracking. However, the potentials of low-quality modalities are not well explored in existing RGBT tracking algorithms. In this work, we propose a novel duality-gated mutual condition network to fully exploit the discriminative information of all modalities while suppressing the effects of data noise. In specific, we design a mutual condition module, which takes the discriminative information of a modality as the condition to guide feature learning of target appearance in another modality. Such a module can effectively enhance target representations of all modalities even in the presence of low-quality modalities. To improve the quality of conditions and further reduce data noise, we propose a duality-gated mechanism and integrate it into the mutual condition module. To deal with the tracking failure caused by sudden camera motion, which often occurs in RGBT tracking, we design a resampling strategy based on optical flow. It does not increase much computational cost since we perform optical flow calculation only when the model prediction is unreliable and then execute resampling when the sudden camera motion is detected. Extensive experiments on four RGBT tracking benchmark datasets show that our method performs favorably against the state-of-the-art tracking algorithms.
Andong Lu, Cun Qian, Chenglong Li 0002, Jin Tang 0001, Liang Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Structural Information Guided Multimodal Pre-training for Vehicle-Centric Perception
abstract
Understanding vehicles in images is important for various applications such as intelligent transportation and self-driving system. Existing vehicle-centric works typically pre-train models on large-scale classification datasets and then fine-tune them for specific downstream tasks. However, they neglect the specific characteristics of vehicle perception in different tasks and might thus lead to sub-optimal performance. To address this issue, we propose a novel vehicle-centric pre-training framework called VehicleMAE, which incorporates the structural information including the spatial structure from vehicle profile information and the semantic structure from informative high-level natural language descriptions for effective masked vehicle appearance reconstruction. To be specific, we explicitly extract the sketch lines of vehicles as a form of the spatial structure to guide vehicle reconstruction. The more comprehensive knowledge distilled from the CLIP big model based on the similarity between the paired/unpaired vehicle image-text sample is further taken into consideration to help achieve a better understanding of vehicles. A large-scale dataset is built to pre-train our model, termed Autobot1M, which contains about 1M vehicle images and 12693 text information. Extensive experiments on four vehicle-based downstream tasks fully validated the effectiveness of our VehicleMAE. The source code and pre-trained models will be released at https://github.com/Event-AHU/VehicleMAE.
Xiao Wang 0014, Chenglong Li 0002, Zhicheng Zhao 0001, Zhe Chen 0013, Yukai Shi, Jin Tang 0001
AAAI7
2024 Cross-Covariate Gait Recognition: A Benchmark
abstract
Gait datasets are essential for gait research. However, this paper observes that present benchmarks, whether conventional constrained or emerging real-world datasets, fall short regarding covariate diversity. To bridge this gap, we undertake an arduous 20-month effort to collect a cross-covariate gait recognition (CCGR) dataset. The CCGR dataset has 970 subjects and about 1.6 million sequences; almost every subject has 33 views and 53 different covariates. Compared to existing datasets, CCGR has both population and individual-level diversity. In addition, the views and covariates are well labeled, enabling the analysis of the effects of different factors. CCGR provides multiple types of gait data, including RGB, parsing, silhouette, and pose, offering researchers a comprehensive resource for exploration. In order to delve deeper into addressing cross-covariate gait recognition, we propose parsing-based gait recognition (ParsingGait) by utilizing the newly proposed parsing data. We have conducted extensive experiments. Our main results show: 1) Cross-covariate emerges as a pivotal challenge for practical applications of gait recognition. 2) ParsingGait demonstrates remarkable potential for further advancement. 3) Alarmingly, existing SOTA methods achieve less than 43% accuracy on the CCGR, highlighting the urgency of exploring cross-covariate gait recognition. Link: https://github.com/ShinanZou/CCGR.
Shinan Zou, Chao Fan 0001, Jianbo Xiong, Chuanfu Shen, Shiqi Yu 0001, Jin Tang 0001
AAAI6
2024 Event Stream-Based Visual Object Tracking: A High-Resolution Benchmark Dataset and A Novel Baseline
abstract
Tracking with bio-inspired event cameras has garnered increasing interest in recent years. Existing works either utilize aligned RGB and event data for accurate tracking or directly learn an event-based tracker. The former incurs higher inference costs while the latter may be susceptible to the impact of noisy events or sparse spatial resolution. In this paper, we propose a novel hierarchical knowledge distillation framework that can fully utilize multimodal / multi-view information during training to facilitate knowledge transfer, enabling us to achieve high-speed and low-latency visual tracking during testing by using only event signals. Specifically, a teacher Transformer-based multimodal tracking framework is first trained by feeding the RGB frame and event stream simultaneously. Then, we design a new hierarchical knowledge distillation strategy which includes pairwise similarity, feature representation, and response maps-based knowledge distillation to guide the learning of the student Transformer network. In particular, since existing event-based tracking datasets are all low-resolution (346 × 260), we propose the first large-scale high-resolution (1280 × 720) dataset named EventVOT. It contains 1141 videos and covers a wide range of categories such as pedestrians, vehicles, UAVs, ping pong, etc. Ex-tensive experiments on both low-resolution (FE240hz, Vi-sEvent, COESOT), and our newly proposed high-resolution EventVOT dataset fully validated the effectiveness of our proposed method.
Xiao Wang 0014, Shiao Wang, Chuanming Tang, Lin Zhu 0012, Bo Jiang 0002, Yonghong Tian 0001, Jin Tang 0001
CVPR7
2024 MGDR: Multi-modal Graph Disentangled Representation for Brain Disease Prediction
Bo Jiang 0002, Xixi Wan, Yuan Chen 0012, Zhengzheng Tu, Yumiao Zhao, Jin Tang 0001
MICCAI (2)7
2024 Knowledge-Driven Subspace Fusion and Gradient Coordination for Multi-modal Learning
Xiaofei Wang 0004, Fangliangzi Meng, Jin Tang 0001, Chao Li 0031
MICCAI (4)4
2024 DFGait: Decomposition Fusion Representation Learning for Multimodal Gait Recognition
Jianbo Xiong, Shinan Zou, Jin Tang 0001
MMM (3)3
2024 Semantics Guided Disentangled GAN for Chest X-Ray Image Rib Segmentation
Lili Huang 0006, Dexin Ma, Chenglong Li 0002, Haifeng Zhao 0001, Jin Tang 0001, Chuanfu Li
PRCV (14)6
2024 CNN-Transformer with Stepped Distillation for Fine-Grained Visual Classification
Lili Huang 0006, Jin Tang 0001
PRCV (9)5
2024 Disentangled generation network for enlarged license plate recognition and a unified dataset
Chenglong Li 0002, Xiaobin Yang, Guohao Wang, Aihua Zheng, Jin Tang 0001
Comput. Vis. Image Underst.6
2024 Lane detection via disentangled representation network with slope consistency loss
Zhaodong Ding, Yifei Deng, Chenglong Li 0002, Rui Ruan, Jin Tang 0001
Eng. Appl. Artif. Intell.5
2024 MutualFormer: Multi-modal Representation Learning via Cross-Diffusion Attention
Xixi Wang 0005, Xiao Wang 0014, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
Int. J. Comput. Vis.4
2024 DEGANet: Road Extraction Using Dual-Branch Encoder With Gated Attention Mechanism
abstract
Automatic identification and extraction of roads from high-resolution remote sensing images (RSIs) are important in remote sensing and computer vision. Advancements in remote sensing technology have increased the information in images, making road extraction more challenging. Conventional convolutional methods have limitations, such as loss of spatial details and inadequate fusion of multiscale features. To address these challenges, the letter introduces a novel encoder-decoder architecture called dual-branch encoder with gated attention mechanism network (DEGANet), for extracting road networks in remote sensing image (RSI). First, we propose a multigated informative self-attention (MGSA) module that combines information from dual-branch encoders. By integrating the ResNet and the dynamic snake convolution (DSC) block, which conforms to road shapes, the module emphasizes slender structures similar to roads, thus enhancing the extraction of road features and focusing on capturing more road details. Second, we also introduce the cascade receptive field enhancement (CRFE) module, which optimizes both accuracy and computational complexity. This module combines various receptive field enhancement modules to improve capture long-range dependencies and spatial information perception. Comprehensive experiments conducted on various public remote sensing road datasets demonstrate that our network attains greater segmentation accuracy (intersection over union (IoU) and$F1$score) and connectivity [average path length similarity (APLS)], validating the effectiveness of our proposed method.
Sibao Chen 0001, Lili Huang 0006, Chris Ding, Jin Tang 0001, Bin Luo 0001
IEEE Geosci. Remote. Sens. Lett.5
2024 Few-Shot Object Detection in Remote Sensing Images With Multiscale Spatial Selective Attention
abstract
Few-shot object detection (FSOD) leverages limited labeled data and substantial unlabeled data for detection. However, these approaches mainly target natural images and ignore the spatial relationships and contextual information between objects in remote sensing images (RSIs). To overcome these challenges, this letter introduces a novel method for detecting few-shot objects in RSI. First, we propose a new attention, called multiscale spatial selective attention (MSSSA). This attention spatially selects feature maps from convolution kernels of different scales through spatial selection, focusing the network on the most relevant region of spatial context. Then, our proposed pixel-level feature extractor module (PLFEM) was used in the first stage of FSOD, providing pixel-level object position information to reduce false and missed detection. To evaluate the proposed method, we carry out comprehensive experiments on the DIOR dataset. The results show that the novel class mAP of our method reaches 38.2% in ten shots, an increase of 3.0% compared with the baseline, significantly improving the accuracy of FSOD in RSI.
Yingnan Yu, Sibao Chen 0001, Lili Huang 0006, Jin Tang 0001, Bin Luo 0001
IEEE Geosci. Remote. Sens. Lett.4
2024 RGBT Tracking based on modality feature enhancement
Sulan Zhai, Lei Liu 0049, Jin Tang 0001
Multim. Tools Appl.4
2024 Learning Graph Attentions via Replicator Dynamics
abstract
Graph Attention (GA) which aims to learn the attention coefficients for graph edges has achieved impressive performance in GNNs on many graph learning tasks. However, existing GAs are usually learned based on edges' (or connected nodes') features which fail to fully capture the rich structural information of edges. Some recent research attempts to incorporate the structural information into GA learning but how to fully exploit them in GA learning is still a challenging problem. To address this challenge, in this work, we propose to leverage a new Replicator Dynamics model for graph attention learning, termed Graph Replicator Attention (GRA). The core of GRA is our derivation of replicator dynamics based sparse attention diffusion which can explicitly learn context-aware and sparse preserved graph attentions via a simple self-supervised way. Moreover, GRA can be theoretically explained from an energy minimization model. This provides a more theoretical justification for the proposed GRA method. Experiments on several graph learning tasks demonstrate the effectiveness and advantages of the proposed GRA method on ten benchmark datasets.
Bo Jiang 0002, Sheng Ge, Beibei Wang 0006, Xiao Wang 0014, Jin Tang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 RGBT Tracking via Progressive Fusion Transformer With Dynamically Guided Learning
abstract
Existing Transformer-based RGB-Thermal (RGBT) tracking methods either use cross-attention to fuse the two modalities, or use self-attention and cross-attention to model both modality-specific and modality-sharing information. However, the significant appearance gap between modalities limits the feature representation ability of certain modalities during the fusion process. To address this problem, we propose a novel Progressive Fusion Transformer called ProFormer, which progressively integrates single-modality information into the multimodal representation for robust RGBT tracking. In particular, ProFormer first uses a self-attention module to collaboratively extract the multimodal representation. Then, ProFormer introduces two cross-attention modules to interact it with the features of the dual modalities for enhancing modality-specific information in the multimodal representation. In addition, we propose a dynamically guided learning algorithm that adaptively employs the well-performing branches to guide the learning of other branches, to improve the representation ability of each branch. Extensive experiments demonstrate that our proposed ProFormer achieves a new state-of-the-art performance on RGBT210, RGBT234, LasHeR, and VTUAV datasets.
Yabin Zhu, Chenglong Li 0002, Xiao Wang 0014, Jin Tang 0001, Zhixiang Huang
IEEE Trans. Circuits Syst. Video Technol.4
2024 An Oriented Object Detector for Hazy Remote Sensing Images
abstract
Currently, a lot of work is focused on aerial object detection and has achieved good results. Though these methods have achieved promising results on the conventional datasets, it is still challenging to locate objects from the low-quality images captured in adverse weather conditions. Currently, there are limited approaches that combine aerial object detection with hazy conditions, and there are few publicly available datasets for real hazy weather based on aerial images. For this purpose, we propose a dataset HRSI, hazy remote sensing images in the real world, which is mainly divided into three categories: airport, large vehicle, and ship. All images in HRSI are from real hazy conditions. In addition, we propose an object detection model DFENet, a dehazing feature enhancement model for hazy remote sensing images, which is suitable for hazy weather. DFENet consists of a two-branch and a dehazing module. The two-branch structure helps to fully learn hazy and dehazing features. In order to avoid the impact of noise caused by the dehezing module, we also designed a haze-predict module (HPM) to predict the information containing haze in the image. We introduce the cross-fuse module (CFM) to utilize the information of haze to guide the feature fusion of two branches. By utilizing the information of haze, DFENet can dynamically adjust the feature weight in the two-branch to avoid the impact of noise generated by the dehazing module. Compared with traditional object detection methods, DFENet not only has good performance in hazy conditions but also improves performance in clear conditions. We tested DFENet on DOTA, HRSI, and Foggy-DOTA to demonstrate that DFENet performs better under hazy conditions.
Sibao Chen 0001, Jia-Xin Wang, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 DecoupleNet: A Lightweight Backbone Network With Efficient Feature Decoupling for Remote Sensing Visual Tasks
abstract
In the realm of computer vision (CV), balancing speed and accuracy remains a significant challenge. Recent efforts have focused on developing lightweight networks that optimize computational efficiency and feature extraction. However, in remote sensing (RS) imagery, where small and multiscale object detection is critical, these networks often fall short in performance. To address these challenges, DecoupleNet is proposed, an innovative lightweight backbone network specifically designed for RS visual tasks in resource-constrained environments. DecoupleNet incorporates two key modules: the feature integration downsampling (FID) module and the multibranch feature decoupling (MBFD) module. The FID module preserves small object features during downsampling, while the MBFD module enhances small and multiscale object feature representation through a novel decoupling approach. Comprehensive evaluations on three RS visual tasks demonstrate DecoupleNet’s superior balance of accuracy and computational efficiency compared to existing lightweight networks. On the NWPU-RESISC45 classification dataset, DecoupleNet achieves a top-1 accuracy of 95.30%, surpassing FasterNet by 2%, with fewer parameters and lower computational overhead. In object detection tasks using the DOTA 1.0 test set, DecoupleNet records an accuracy of 78.04%, outperforming ARC-R50 by 0.69%. For semantic segmentation on the LoveDA test set, DecoupleNet achieves 53.1% accuracy, surpassing UnetFormer by 0.70%. These findings open new avenues for advancing RS image analysis on resource-constrained devices, addressing a pivotal gap in the field. The code and pretrained models are publicly available athttps://github.com/lwCVer/DecoupleNet.
Wei Lu 0032, Sibao Chen 0001, Qing-Ling Shu, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Prior Guidance and Principal Attention Network for Remote Sensing Image Change Detection
abstract
In the field of remote sensing (RS) image change detection (CD), the conventional encoder-decoder architecture networks often encounter three significant challenges. First, noise in the features extracted from traditional backbone networks leads to blurred boundaries of change objects. Second, upsampling techniques employed in the decoder, such as interpolation or deconvolution, are limited by their finite receptive fields, making it challenging to accurately distinguish pseudo-changes. Furthermore, how to merge encoder and decoder features with possible semantic gaps for the fine-grained details is a topic worth considering. To address these challenges, we introduce a prior guidance (PG) module that effectively aggregates prior high-level features as a semantic guidance map to guide encoder features for the enhancement of boundary detection. In addition, we design a principal attention (PA) module, which aggregates global information from principal regions through sparse operations and adaptively allocates this information to the upsampled and encoder features. This not only addresses the deficiency of global information in the upsampled features but also reduces the semantic gap between the encoder and decoder by establishing channel dependencies. PA does not divert attention to irrelevant regions, demonstrating excellent performance and computational efficiency. By integrating these two modules into our method, a novel PG and PA network (PGPANet) is elaborately designed. A wide range of experiments confirms the validity of our method, showcasing outstanding detection accuracy on three publicly available CD datasets: LEVIR-CD, SYSU-CD, and WHU-CD. The demo code of this work is publicly available athttps://github.com/DaGuangDaGuang/PGPANet.
Qing-Ling Shu, Sibao Chen 0001, Zhi-Hui You, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Attention-Aware Sobel Graph Convolutional Network for Remote Sensing Image Change Detection
abstract
In the study of remote sensing images, the problem of change detection (CD) is crucial. Convolutional neural networks (CNNs) are well-liked feature extraction structures that are frequently used in CD. On the other hand, graph convolutional networks (GCNs) are effective in building contextual structure information. Compared with CNN, GCN can make full use of the graph structure information to capture the changing features between different areas in the graph by learning the connections and interactions between nodes. In contrast, traditional pixel-based CNNs may have difficulty modeling semantic relationships and temporal variations among features and are susceptible to noise interference. So in this article, we extract optimization information using a GCN structure. Due to the particularity of remote sensing images, edge information is often ignored, which is useful in the field of CD. In this article, we propose an attention-aware Sobel GCN (ASGCN) for remote sensing image CD. First, we use a Siamese CNN to extract primary multilevel features. Then, a dual-branch attention module (DAM) including coordinate attention and multiscale local attention module (MLAM) is proposed to focus on informative pixels, we use Sobel operator to construct graph, and the graph convolutional module can expand receptive field and extract edge information. Attention fusion module (AFM) is adopted at decoder to perform effective feature fusion. Extensive comparative experiments on three CD datasets, LEVIR-CD, WHU-CD, and DSIFN-CD, verify the effectiveness of the proposed ASGCN.
Lei Wang 0095, Zhi-Hui You, Wei Lu 0032, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 ADRNet: Affine and Deformable Registration Networks for Multimodal Remote Sensing Images
abstract
Multi-modal remote sensing images registration ensures the consistency of the spatial positions for different images. It can provide the accurate geographic information and supports the fusion of multi-source data for geospatial analyses and applications. Rigid registration method shows high performance in dealing with large-scale deformation, but it is difficult to achieve high-precision image registration. In contrast, non-rigid registration method is suitable for processing local differences, but cannot effectively deal with large-scale deformation differences. Therefore, the combination of rigid and non-rigid registration methods becomes a necessary strategy to address such issues. In this paper, we propose a novel ADRNet method for multi-modal remote sensing images registration. The proposed ADRNet method contains three main modules: affine registration module, deformable registration module, and spatial transformer module that integrates the affine and deformable transformation parameters to obtain the final aligned images. Meanwhile, we design a new feature enhancement module and an attention module with dilated convolutions which have different dilation rates, which are used to alleviate the limitations imposed by receptive fields in the convolution operation. Moreover, we propose a specific symmetric loss function to optimize the whole network from the perspective of inverse consistency. To assess the efficiency and performance of the network, we extend the experimental data, ranging from cross-modal images in a conventional viewpoint to cross-modal images in a remote sensing viewpoint. The experimental results show that our method exhibits excellent performance for the images with different viewpoints and deformation scales. The relevant code will be released at: https://github.com/Ahuer-Lei/ADRNet.
Yun Xiao 0003, Yuan Chen 0012, Bo Jiang 0002, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Prototype Discriminative Learning for Semi-Supervised Change Detection in Remote Sensing Images
abstract
With the continuous progress of deep learning in remote sensing (RS) visual tasks, considerable advancements have been achieved in RS image change detection (CD). However, prevailing CD methods heavily rely on extensive sets of fully pixelwise hand-annotated training data, a time-consuming and costly process, and they fail to fully harness the potential benefits of deep feature representations within the deep feature domain. To tackle the mentioned issues, we propose a novel semi-supervised CD method called PDLCD, which strategically leverages useful information from massive unlabeled data to complement labeled data with just a few samples. Specifically, changed objects and unchanged backgrounds of bitemporal RS images are various and complex, our approach advocates dividing each category into multiple subclasses in the deep feature domain. In this scheme, the high-level feature of each subclass follows a Gaussian distribution. Then, the prototype discriminative learning (PDL) is introduced to explicitly encourage deep features of samples closer to the nearest prototype within their respective category, and away from all prototypes of other categories. We design feature discriminative loss (FDL) to implement PDL for constructing more pronounced intraclass compactness and interclass variability. Finally, we compute the supervised loss based on a limited set of labeled data, incorporate the unsupervised loss leveraging a substantial volume of unlabeled data, and include FDL within the deep feature domain to collectively optimize the model. Extensive experiments carried out on three challenging RS image CD datasets illustrate that our proposed semi-supervised CD method obtains better CD performance than previous counterparts. The source code is available at:https://github.com/Youzhihui/PDLCD.
Zhi-Hui You, Sibao Chen 0001, Jia-Xin Wang, Chris Ding, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Dense Tiny Object Detection: A Scene Context Guided Approach and a Unified Benchmark
abstract
With the continuous advancement of remote sensing observation technology, wide-area observation and high-resolution imaging make remote sensing images contain a large number of dense tiny objects. The detection of dense tiny objects is a very challenging task since these objects are with very low resolution and might stick together. Existing work lacks further exploration of the contextual scene information and inherent characteristics of dense tiny objects, which are crucial for performance improvement of dense tiny object detection. In this work, we propose a novel Scene Contextualized Detection Network (SCDNet) by decoupling scene contextual information through a dedicated scene classification sub-network, thereby enabling an enhanced exploration of the relationship between tiny objects and their surrounding environments. In particular, we design a lightweight scene context guided fusion module in SCDNet to incorporate scene context information around dense tiny objects more effectively. Moreover, we further develop the scene context guided foreground enhancement module to suppress the background information while enhancing the foreground information based on the scene information. In addition, this research field still lacks a large-scale benchmark dataset with dense tiny objects, which is crucial for the training and comprehensive evaluation of detection methods. To this end, we construct a large-scale dataset for dense tiny object detection. It contains 11,600 images with 1,019,800 instances, the average absolute size of objects is smaller than 13 pixels, and each image contains 88 objects on average. Extensive experiments are conducted on the proposed dataset, and the results demonstrate the superiority and effectiveness of SCDNet compared to existing methods. The dataset and evaluation code are available at https://github.com/mmic-lcl.
Zhicheng Zhao 0002, Chenglong Li 0002, Yun Xiao 0003, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Modality Conversion Meets Superresolution: A Collaborative Framework for High- Resolution Thermal UAV Image Generation
abstract
Due to the limitations and costs of thermal sensors, unmanned aerial vehicle (UAV) platforms often equip with high-resolution (HR) visible imaging and low-resolution (LR) thermal imaging cameras for all-day monitoring capability. Existing works generate the high-resolution thermal UAV images by either super-resolution (SR) from high-resolution visible and low-resolution thermal images or modality conversion (MC) from high-resolution visible images. However, the modality gap between visible and thermal sources might degrade the generation quality. We observe that the MC task is beneficial in addressing the cross-modal gap in the SR task, while the SR task can provide the condition of thermal information to boost the MC task. Moreover, these two tasks have the same output and can thus be carried out simultaneously without any additional annotation. Based on this observation, we propose a collaborative enhancement network (CENet), which performs thermal UAV image SR and visible image MC in a joint manner, for high-resolution thermal UAV image generation. In particular, we design a mutual guidance module to interact the features from SR and MC tasks in an alternating bidirectional manner. Considering that low-level vision tasks are position-sensitive, to further enhance the feature alignment between the two tasks, we design a bidirectional alignment fusion module to maintain feature consistency of the MC and SR branches. The proposed collaborative framework not only achieves joint and unified training of the two tasks, but also generates two types of complementary high-resolution images. Extensive experiments on public datasets demonstrate that the proposed CENet outperforms current state-of-the-art super-resolution (SR) methods in generating high-resolution thermal UAV images, as quantified by peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM).
Zhicheng Zhao 0002, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Long-Term Motion-Assisted Remote Sensing Object Tracking
abstract
Remote sensing object tracking has gained significant attention due to its wide range of applications including surveillance and motion analysis. However, it faces various challenges such as low resolution, low contrast, blurring, and occlusion, which impede its development at a significantly slower pace compared to object tracking methods for general scenes. The challenges of low resolution, low contrast, and blurring result in weak target features, while the occlusion challenge poses a problem for target search range and tracker discrimination in subsequent frames. To address these issues, we propose a novel long-term motion-assisted framework, which can effectively mine long-term motion information and use an evaluation scheme for robust remote sensing object tracking. Specifically, we design a long-term motion feature mining module (LMFM), which efficiently calculates the long-term motion information by integrating previous motion features in a temporal-iterative manner to alleviate the problem of weak features caused by low resolution, low contrast, and blurring. Moreover, we design an evaluation scheme that combines the motion trajectory model, target classification scores, and predicted target positions to handle the issue of massive occlusion or target loss. Extensive experiments on the SatSOT, SV248S, and VISO datasets show that our approach outperforms state-of-the-art (SOTA) trackers. The source code, trained models, and raw results are released athttps://github.com/zhaoxingle/LMANet.
Yabin Zhu, Xingle Zhao, Chenglong Li 0002, Jin Tang 0001, Zhixiang Huang
IEEE Trans. Geosci. Remote. Sens.4
2024 Attribute-Guided Cross-Modal Interaction and Enhancement for Audio-Visual Matching
abstract
Audio-visual matching is an essential task that measures the correlation between audio clips and visual images. However, current methods rely solely on the joint embedding of global features from audio clips and face image pairs to learn semantic correlations. This approach overlooks the importance of high-confidence correlations and discrepancies of local subtle features, which are crucial for cross-modal matching. To address this issue, we propose a novel Attribute-guided Cross-modal Interaction and Enhancement Network (ACIENet), which employs multiple attributes to explore the associations of different key local subtle features. The ACIENet contains two novel modules: the Attribute-guided Interaction (AGI) module and the Attribute-guided Enhancement (AGE) module. The AGI module employs global feature alignment similarity to guide cross-modal local feature interactions, which enhances cross-modal association features for the same identity and expands cross-modal distinctive features for different identities. Additionally, the interactive features and original features are fused to ensure intra-class discriminability and inter-class correspondence. The AGE module captures subtle attribute-related features by using an attribute-driven network, thereby enhancing discrimination at the attribute level. Specifically, it strengthens the combined attribute-related features of gender and nationality. To prevent interference between multiple attribute features, we design a multi-attribute learning network as a parallel framework. Experiments conducted on a public benchmark dataset demonstrate the efficacy of the ACIENet method in different scenarios. Code and models are available at https://github.com/w1018979952/ACIENet.
Jiaxiang Wang 0001, Aihua Zheng, Yan Yan 0002, Ran He 0001, Jin Tang 0001
IEEE Trans. Inf. Forensics Secur.5
2024 Collaborative License Plate Recognition via Association Enhancement Network With Auxiliary Learning and a Unified Benchmark
abstract
Since the standard license plate of large vehicle is easily affected by occlusion and stain, the traffic management department introduces the enlarged license plate at the rear of the large vehicle to assist license plate recognition. However, current researches regards standard license plate recognition and enlarged license plate recognition as independent tasks, and do not take advantage of the complementary benefits from the two types of license plates. In this work, we propose a new computer vision task called collaborative license plate recognition, aiming to leverage the complementary advantages of standard and enlarged license plates for achieving more accurate license plate recognition. To achieve this goal, we propose an Association Enhancement Network (AENet), which achieves robust collaborative licence plate recognition by capturing the correlations between characters within a single licence plate and enhancing the associations between two license plates. In particular, we design an association enhancement branch, which supervises the fusion of two licence plate information using the complete licence plate number to mine the association between them. To enhance the representation ability of each type of licence plates, we design an auxiliary learning branch in the training stage, which supervises the learning of individual license plates in the association enhancement between two license plates. In addition, we contribute a comprehensive benchmark dataset called CLPR, which consists of a total of 19,782 standard and enlarged licence plates from 24 provinces in China and covers most of the challenges in real scenarios, for collaborative license plate recognition. Extensive experiments on the proposed CLPR dataset demonstrate the effectiveness of the proposed AENet against several state-of-the-art methods.
Yifei Deng, Guohao Wang, Chenglong Li 0002, Wei Wang 0115, Cheng Zhang 0010, Jin Tang 0001
IEEE Trans. Multim.6
2024 AMatFormer: Efficient Feature Matching via Anchor Matching Transformer
abstract
Learning based feature matching methods have been commonly studied in recent years. The core issue for learning feature matching is to how to learn (1) discriminative representations for feature points (or regions) within each intra-image and (2) consensus representations for feature points across inter-images. Recently, self- and cross-attention models have been exploited to address this issue. However, in many scenes, features are coming with large-scale, redundant and outliers contaminated. Previous self-/cross-attention models generally conduct message passing on all primal features which thus lead to redundant learning and high computational cost. To mitigate limitations, inspired by recent seed matching methods, in this article, we propose a novel efficient Anchor Matching Transformer (AMatFormer) for the feature matching problem. AMatFormer has two main aspects: First, it mainly conducts self-/cross-attention on some anchor features and leverages these anchor features as message bottleneck to learn the representations for all primal features. Thus, it can be implemented efficiently and compactly. Second, AMatFormer adopts a shared FFN module to further embed the features of two images into the common domain and thus learn the consensus feature representations for the matching problem. Experiments on several benchmarks demonstrate the effectiveness and efficiency of the proposed AMatFormer matching approach.
Bo Jiang 0002, Shuxian Luo, Xiao Wang 0014, Chuanfu Li, Jin Tang 0001
IEEE Trans. Multim.5
2024 Illumination Distillation Framework for Nighttime Person Re-Identification and a New Benchmark
abstract
Nighttime person Re-ID (person re-identification in the nighttime) is a very important and challenging task for visual surveillance but it has not been thoroughly investigated. Under the low illumination condition, the performance of person Re-ID methods usually sharply deteriorates. To address the low illumination challenge in nighttime person Re-ID, this paper proposes an Illumination Distillation Framework (IDF), which utilizes illumination enhancement and illumination distillation schemes to promote the learning of Re-ID models. Specifically, IDF consists of a master branch, an illumination enhancement branch, and an illumination distillation module. The master branch is used to extract the features from a nighttime image. The illumination enhancement branch first estimates an enhanced image from the nighttime image using a nonlinear curve mapping method and then extracts the enhanced features. However, nighttime and enhanced features usually contain data noise due to unstable lighting conditions and enhancement failures. To fully exploit the complementary benefits of nighttime and enhanced features while suppressing data noise, we propose an illumination distillation module. In particular, the illumination distillation module fuses the features from two branches through a bottleneck fusion model and then uses the fused features to guide the learning of both branches in a distillation manner. In addition, we build a real-world nighttime person Re-ID dataset, namedNight600, which contains 600 identities captured from different viewpoints and nighttime illumination conditions under complex outdoor environments. Experimental results demonstrate that our IDF can achieve state-of-the-art performance on two nighttime person Re-ID datasets (i.e.,Night600andKnight). We will release our code and dataset athttps://github.com/Alexadlu/IDF.
Andong Lu, Zhang Zhang 0001, Yan Huang 0023, Yifan Zhang 0004, Chenglong Li 0002, Jin Tang 0001, Liang Wang 0001
IEEE Trans. Multim.6
2024 GDCNet: Graph Enrichment Learning via Graph Dropping Convolutional Networks
abstract
Graph convolutional networks (GCNs) have been widely studied to address graph data representation and learning. In contrast to traditional convolutional neural networks (CNNs) that employ many various (spatial) convolution filters to obtain rich feature descriptors to encode complex patterns of image data, GCNs, however, are defined on the input observed graph G(X,A) and usually adopt the single fixed spatial convolution filter for graph data feature extraction. This limits the capacity of the existing GCNs to encode the complex patterns of graph data. To overcome this issue, inspired by depthwise separable convolution and DropEdge operation, we first propose to generate various graph convolution filters by randomly dropping out some edges from the input graph A . Then, we propose a novel graph-dropping convolution layer (GDCLayer) to produce rich feature descriptors for graph data. Using GDCLayer, we finally design a new end-to-end network architecture, that is, a graph-dropping convolutional network (GDCNet), for graph data learning. Experiments on several datasets demonstrate the effectiveness of the proposed GDCNet.
Bo Jiang 0002, Beibei Wang 0006, Haiyun Xu, Jin Tang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Tiny Object Tracking: A Large-Scale Dataset and a Baseline
abstract
Tiny objects, frequently appearing in practical applications, have weak appearance and features, and receive increasing interests in many vision tasks, such as object detection and segmentation. To promote the research and development of tiny object tracking, we create a large-scale video dataset, which contains 434 sequences with a total of more than 217K frames. Each frame is carefully annotated with a high-quality bounding box. In data creation, we take 12 challenge attributes into account to cover a broad range of viewpoints and scene complexities, and annotate these attributes for facilitating the attribute-based performance analysis. To provide a strong baseline in tiny object tracking, we propose a novel multilevel knowledge distillation network (MKDNet), which pursues three-level knowledge distillations in a unified framework to effectively enhance the feature representation, discrimination, and localization abilities in tracking tiny objects. Extensive experiments are performed on the proposed dataset, and the results prove the superiority and effectiveness of MKDNet compared with state-of-the-art methods. The dataset, the algorithm code, and the evaluation code are available at https://github.com/mmic-lcl/Datasets-and-benchmark-code.
Yabin Zhu, Chenglong Li 0002, Xiao Wang 0014, Jin Tang 0001, Bin Luo 0001, Zhixiang Huang
IEEE Trans. Neural Networks Learn. Syst.5
2024 MCDGait: multimodal co-learning distillation network with spatial-temporal graph reasoning for gait recognition in the wild
Jianbo Xiong, Shinan Zou, Jin Tang 0001, Tardi Tjahjadi
Vis. Comput.3
2023 A Multi-Stage Adaptive Feature Fusion Neural Network for Multimodal Gait Recognition
abstract
Gait recognition is a biometric technology that has received extensive attention. Most existing gait recognition algorithms are unimodal, and a few multimodal gait recognition algorithms perform multimodal fusion only once. None of these algorithms may fully exploit the complementary advantages of the multiple modalities. In this paper, by considering the temporal and spatial characteristics of gait data, we propose a multi-stage feature fusion strategy (MSFFS), which performs multimodal fusions at different stages in the feature extraction process. Also, we propose an adaptive feature fusion module (AFFM) that considers the semantic association between silhouettes and skeletons. The fusion process fuses different silhouette areas with their more related skeleton joints. Since visual appearance changes and time passage co-occur in a gait period, we propose a multiscale spatial-temporal feature extractor (MSSTFE) to learn the spatial-temporal linkage features thoroughly. Specifically, MSSTFE extracts and aggregates spatial-temporal linkages information at different spatial scales. Combining the strategy and modules mentioned above, we propose a multi-stage adaptive feature fusion (MSAFF) neural network, which shows state-of-the-art performance in many experiments on three datasets. Besides, MSAFF is equipped with feature dimensional pooling (FD Pooling), which can significantly reduce the dimension of the gait representations without hindering the accuracy.
Shinan Zou, Jianbo Xiong, Chao Fan 0001, Shiqi Yu 0001, Jin Tang 0001
IJCB5
2023 Quality-Aware RGBT Tracking via Supervised Reliability Learning and Weighted Residual Guidance
abstract
RGB and thermal infrared (TIR) data have different visual properties, which make their fusion essential for effective object tracking in diverse environments and scenes. Existing RGBT tracking methods commonly use attention mechanisms to generate reliability weights for multi-modal feature fusion. However, without explicit supervision, these weights may be unreliably estimated, especially in complex scenarios. To address this problem, we propose a novel Quality-Aware RGBT Tracker (QAT) for robust RGBT tracking. QAT learns reliable weights for each modality in a supervised manner and performs weighted residual guidance to extract and leverage useful features from both modalities. We address the issue of the lack of labels for reliability learning by designing an efficient three-branch network that generates reliable pseudo labels, and a simple binary classification scheme that estimates high-accuracy reliability weights, mitigating the effect of noisy pseudo labels. To propagate useful features between modalities while reducing the influence of noisy modal features on the migrated information, we design a weighted residual guidance module based on the estimated weights and residual connections. We evaluate our proposed QAT on five benchmark datasets, including GTOT, RGBT210, RGBT234, LasHeR, and VTUAV, and demonstrate its excellent performance compared to state-of-the-art methods. Experimental results show that QAT outperforms existing RGBT tracking methods in various challenging scenarios, demonstrating its efficacy in improving the reliability and accuracy of RGBT tracking.
Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003, Jin Tang 0001
ACM Multimedia4
2023 Graph context-attention network via low and high order aggregation
Haiyun Xu, Shaojie Zhang 0002, Bo Jiang 0002, Jin Tang 0001
Neurocomputing4
2023 Road Extraction by Multiscale Deformable Transformer From Remote Sensing Images
abstract
Rapid progress has been made in the research of high-resolution remote sensing road extraction tasks in the past years, but due to the diversity of road types and the complexity of road context, extracting the perfect road network is still fraught with difficulties and challenges. Many Convolutional Neural Networks (CNNs) based on encoder-decoder structures have demonstrated their effectiveness. Transformer’s self-attention mechanism shows more powerful performance than CNNs in modeling global feature dependencies. In this paper, we propose a Multi-scale Deformable Transformer Network (MDTNet) based on encoder-decoder structure to extract road networks from remote sensing images. The core of MDTNet is our proposed Multi-scale Deformable Self-Attention (MDSA) mechanism. MDSA can capture more comprehensive features than conventional self-attention. In addition, roads are not present in certain blocks of areas like other objects, but are interwoven throughout the image in such a long, linear fashion that information about certain road segments may be overlooked. To minimize residual errors in road segmentations, our MDSA incorporates a deformable design on feature maps, which effectively enhances the salience of road features relative to their surroundings. Extensive experiments on several public remote sensing road datasets show that our MDTNet achieves higher segmentation [F1 score and Intersection over Union (IoU)] and connectivity [Average Path Length Similarity (APLS)] accuracy, which verifies the effectiveness of our approach.
Pengcheng Hu 0001, Sibao Chen 0001, Lili Huang 0006, Guizhou Wang, Jin Tang 0001, Bin Luo 0001
IEEE Geosci. Remote. Sens. Lett.5
2023 Multi-granularity cross attention network for person re-identification
Chengmei Han, Bo Jiang 0002, Jin Tang 0001
Multim. Tools Appl.3
2023 Graph Neural Network Meets Sparse Representation: Graph Sparse Neural Networks via Exclusive Group Lasso
abstract
Existing GNNs usually conduct the layer-wise message propagation via the 'full' aggregation of all neighborhood information which are usually sensitive to the structural noises existed in the graphs, such as incorrect or undesired redundant edge connections. To overcome this issue, we propose to exploit Sparse Representation (SR) theory into GNNs and propose Graph Sparse Neural Networks (GSNNs) which conduct sparse aggregation to select reliable neighbors for message aggregation. GSNNs problem contains discrete/sparse constraint which is difficult to be optimized. Thus, we then develop a tight continuous relaxation model Exclusive Group Lasso GNNs (EGLassoGNNs) for GSNNs. An effective algorithm is derived to optimize the proposed EGLassoGNNs model. Experimental results on several benchmark datasets demonstrate the better performance and robustness of the proposed EGLassoGNNs model.
Bo Jiang 0002, Beibei Wang 0006, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Generalizing Aggregation Functions in GNNs: Building High Capacity and Robust GNNs via Nonlinear Aggregation
abstract
The main aspect powering GNNs is the multi-layer network architecture to learn the nonlinear representation for graph learning task. The core operation in GNNs is the message propagation in which each node updates its information by aggregating the information from its neighbors. Existing GNNs usually adopt either linear neighborhood aggregation (e.g. mean, sum) or max aggregator in their message propagation. 1) For linear aggregators, the whole nonlinearity and network's capacity of GNNs are generally limited because deeper GNNs usually suffer from the over-smoothing issue due to their inherent information propagation mechanism. Also, linear aggregators are usually vulnerable to the spatial perturbations. 2) For max aggregator, it usually fails to be aware of the detailed information of node representations within neighborhood. To overcome these issues, we re-think the message propagation mechanism in GNNs and develop the new general nonlinear aggregators for neighborhood information aggregation in GNNs. One main aspect of our nonlinear aggregators is that they all provide the optimally balanced aggregator between max and mean/sum aggregators. Thus, they can inherit both i) high nonlinearity that enhances network's capacity, robustness and ii) detail-sensitivity that is aware of the detailed information of node representations in GNNs' message propagation. Promising experiments show the effectiveness, high capacity and robustness of the proposed methods.
Beibei Wang 0006, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 SCRDet++: Detecting Small, Cluttered and Rotated Objects via Instance-Level Feature Denoising and Rotation Loss Smoothing
abstract
Small and cluttered objects are common in real-world which are challenging for detection. The difficulty is further pronounced when the objects are rotated, as traditional detectors often routinely locate the objects in horizontal bounding box such that the region of interest is contaminated with background or nearby interleaved objects. In this paper, we first innovatively introduce the idea of denoising to object detection. Instance-level denoising on the feature map is performed to enhance the detection to small and cluttered objects. To handle the rotation variation, we also add a novel IoU constant factor to the smooth L1 loss to address the long standing boundary problem, which to our analysis, is mainly caused by the periodicity of angular (PoA) and exchangeability of edges (EoE). By combing these two features, our proposed detector is termed as SCRDet++. Extensive experiments are performed on large aerial images public datasets DOTA, DIOR, UCAS-AOD as well as natural image dataset COCO, scene text dataset ICDAR2015, small traffic light dataset BSTLD and our released S$^{2}$TLD by this paper. The results show the effectiveness of our approach. The released dataset S$^{2}$TLD is made public available, which contains 5,786 images with 14,130 traffic light instances across five categories.
Xue Yang 0005, Junchi Yan, Wenlong Liao, Xiaokang Yang 0001, Jin Tang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Detecting Rotated Objects as Gaussian Distributions and its 3-D Generalization
abstract
Existing detection methods commonly use a parameterized bounding box (BBox) to model and detect (horizontal) objects and an additional rotation angle parameter is used for rotated objects. We argue that such a mechanism has fundamental limitations in building an effective regression loss for rotation detection, especially for high-precision detection with high IoU (e.g., 0.75). Instead, we propose to model the rotated objects as Gaussian distributions. A direct advantage is that our new regression loss regarding the distance between two Gaussians e.g., Kullback-Leibler Divergence (KLD), can well align the actual detection performance metric, which is not well addressed in existing methods. Moreover, the two bottlenecks i.e., boundary discontinuity and square-like problem also disappear. We also propose an efficient Gaussian metric-based label assignment strategy to further boost the performance. Interestingly, by analyzing the BBox parameters' gradients under our Gaussian-based KLD loss, we show that these parameters are dynamically updated with interpretable physical meaning, which help explain the effectiveness of our approach, especially for high-precision detection. We extend our approach from 2-D to 3-D with a tailored algorithm design to handle the heading estimation, and experimental results on twelve public datasets (2-D/3-D, aerial/text/face images) with various base detectors show its superiority.
Xue Yang 0005, Gefan Zhang, Xiaojiang Yang, Yue Zhou 0005, Wentao Wang 0009, Jin Tang 0001, Junchi Yan
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 CNNGRN: A Convolutional Neural Network-Based Method for Gene Regulatory Network Inference From Bulk Time-Series Expression Data
abstract
Gene regulatory networks (GRNs) participate in many biological processes, and reconstructing them plays an important role in systems biology. Although many advanced methods have been proposed for GRN reconstruction, their predictive performance is far from the ideal standard, so it is urgent to design a more effective method to reconstruct GRN. Moreover, most methods only consider the gene expression data, ignoring the network structure information contained in GRN. In this study, we propose a supervised model named CNNGRN, which infers GRN from bulk time-series expression data via convolutional neural network (CNN) model, with a more informative feature. Bulk time series gene expression data imply the intricate regulatory associations between genes, and the network structure feature of ground-truth GRN contains rich neighbor information. Hence, CNNGRN integrates the above two features as model inputs. In addition, CNN is adopted to extract intricate features of genes and infer the potential associations between regulators and target genes. Moreover, feature importance visualization experiments are implemented to seek the key features. Experimental results show that CNNGRN achieved competitive performance on benchmark datasets compared to the state-of-the-art computational methods. Finally, hub genes identified based on CNNGRN have been confirmed to be involved in biological processes through literature.
Jin Tang 0001, Junfeng Xia, Chun-Hou Zheng 0001, Pi-Jing Wei
IEEE ACM Trans. Comput. Biol. Bioinform.2
2023 Category-Oriented Localization Distillation for SAR Object Detection and a Unified Benchmark
abstract
Despite much research progress in synthetic aperture radar (SAR) object detection, the performance of SAR object detection has encountered a bottleneck limited by the imaging mechanism of SAR. In this work, we investigate how to perform robust SAR object detection by distilling the category knowledge from optical images in the training stage. To this end, we propose a novel knowledge distillation method called Category-oriented Localization Distillation (CoLD), which employs the optical object detection network as the teacher to guide the SAR object detection network. To introduce the category prior knowledge of the teacher network in the localization knowledge transferring, a category-oriented partition module is designed in CoLD to decouple candidate bounding boxes into target and non-target ones according to the category information in optical images. Through box decoupling, the accuracy and efficiency of SAR object detection can be significantly improved. Moreover, an IoU-based weighting module is introduced in CoLD to guide the student network focusing more on high-quality candidate boxes by adaptively changing the weight of each candidate bounding box based on the corresponding IoU score in the teacher network. In addition, a unified benchmark dataset is created for the evaluation of optical information guided SAR object detection, which consists of 14,665 optical and SAR image pairs in the training set and 3,666 SAR images in the testing set. Extensive experiments on the dataset demonstrate the effectiveness of our CoLD against state-of-the-art methods. The dataset is available at: https://github.com/mmic-lcl/Datasets-and-benchmark-code.
Rui Ruan, Zhicheng Zhao 0002, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 A Robust Feature Downsampling Module for Remote-Sensing Visual Tasks
abstract
Remote sensing (RS) images present unique challenges for computer vision due to lower resolution, smaller objects, and fewer features. Mainstream backbone networks show promising results for traditional visual tasks. However, they use convolution to reduce feature map dimensionality, which can result in information loss for small objects in RS images and decreased performance. To address this problem, we propose a new and universal downsampling module named Robust Feature Downsampling (RFD). RFD fuses multiple feature maps extracted by different downsampling techniques, creating a more robust feature map with a complementary set of features. Leveraging this, we overcome the limitations of conventional convolutional downsampling, resulting in more accurate and robust analysis of RS images. We develop two versions of RFD module, Shallow RFD (SRFD) and Deep RFD (DRFD), tailored to adapt to different stages of feature capture and improve feature robustness. We replace the downsampling layers of existing mainstream backbones with RFD module and conduct comparative experiments on several public RS image datasets. The results show significant improvements compared to baseline approaches in RS image classification, object detection, and semantic segmentation. Specifically, our RFD module achieved an average performance gain of 1.5% on NWPU-RESISC45 classification dataset without utilizing any additional pretraining data, resulting in state-of-the-art performance on this dataset. Moreover, in detection and segmentation tasks on DOTA and iSAID datasets, our RFD module outperforms the baseline approaches by 2-7% when utilizing pretraining data from NWPU-RESISC45. These results highlight the value of RFD module in enhancing the performance of RS visual tasks.
Wei Lu 0032, Sibao Chen 0001, Jin Tang 0001, Chris Ding, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 AS3ITransUNet: Spatial-Spectral Interactive Transformer U-Net With Alternating Sampling for Hyperspectral Image Super-Resolution
abstract
Single hyperspectral image (HSI) super-resolution (SR) is an important topic in remote sensing field. However, existing HSI SR methods mainly use the feed-forward upsampling technique and convolutional neural network (CNN) to learn the feature representation, failing to learn the complex mapping relationship between low-resolution (LR) and high-resolution (HR) and long-range joint spectral and spatial features. To address this issue, in this paper, we propose the Spatial-Spectral Interactive Transformer U-Net with Alternating Sampling (AS3ITransUNet) for the HSI SR task. In this method, to mitigate the computational burden resulting from the high spectral dimension of HSI, a group reconstruction strategy is adopted. To effectively explore the hierarchical features of HSI, the U-Net with alternating upsampling and downsampling is designed that allocates the task of learning the complex mapping relationship to each stage of U-Net. To fully extract the spatial-spectral features of HSI, we propose the spatial-spectral interactive transformer (SSIT) block and integrate it into the encoder and decoder of U-Net. The SSIT block contains a cross-branch bidirectional interaction module, which further captures the complementary information between spatial and spectral dimensions. Moreover, the multi-stage complementary information learning (MFEL) is proposed to capture the complementary information in the adjacent HSI groups for recovering the absent details in the current HSI group. The experiments on the three benchmark datasets demonstrate that the proposed AS3ITransUNet can effectively improve the spatial resolution and preserve the spectral information at different scales. Models and code are available at https://github.com/liushiji666/AS3-ITransUNet.
Shiji Liu, Bo Jiang 0002, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Crossed Siamese Vision Graph Neural Network for Remote-Sensing Image Change Detection
abstract
The development of deep learning in remote sensing (RS) visual tasks has led to remarkable progress in RS image change detection (CD). However, RS bi-temporal images cover complex and confusing scenes due to natural environmental factors, which presents challenges for CD task. How to effectively exploit long-range dependencies and sensitively discriminate real-changes with various scales from pseudo-changes are urgent problems. It is especially obvious for the changes of building structures man-made. This paper presents a CD approach named CSViG, which utilizes Siamese Vision Graph neural network (SViG) with crossed feature fusion. SViG acts as a feature extractor to capture richer short- and long-range dependencies. Crossed feature fusion consists of a horizontal feature fusion module (HFFM) and a vertical feature fusion module (VFFM). HFFM designs cross-concatenation (CC) way to reveal real-changes from pseudo-change in the same horizontal stage, after which global and local features are extracted by using attention mechanism and multi-scale depth-wise separable convolution. VFFM further fuses complementary content from vertical multiple stages to effectively represent change regions of different sizes (tiny or huge) by using attention mechanism. Extensive comparative experiments conducted on three available building change detection datasets demonstrate that the proposed method achieves better CD performance than previous counterparts.
Zhi-Hui You, Jia-Xin Wang, Sibao Chen 0001, Chris Ding, Guizhou Wang, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Thermal UAV Image Super-Resolution Guided by Multiple Visible Cues
abstract
Unmanned aerial vehicle (UAV) thermal-imaging has received much attention, but the insufficient image resolution caused by thermal imaging systems is still a crucial problem that limits the understanding of thermal UAV images. However, high-resolution visible images are relatively easy to access, and it is thus valuable for exploring useful information from visible image to assist thermal UAV image super-resolution (SR). In this article, we propose a novel multiconditioned guidance network (MGNet) to effectively mine the information of visible images for thermal UAV image SR. High-resolution visible UAV images usually contain salient appearance, semantic, and edge information, which plays a critical role in boosting the performance of thermal UAV image SR. Therefore, we design an effective multicue guidance module (MGM) to leverage the appearance, edge, and semantic cues from visible images to guide thermal UAV image SR. In addition, we build the first benchmark dataset for the task of thermal UAV image SR guided by visible images. It is collected by a multimodal UAV platform and composes of 1025 pairs of manually aligned visible and thermal images. Extensive experiments on the built dataset show that our MGNet can effectively leverage useful information from visible images to improve the performance of thermal UAV image SR and perform well against several state-of-the-art methods. The dataset is available at:https://github.com/mmic-lcl/Datasets-and-benchmark-code.
Zhicheng Zhao 0002, Chenglong Li 0002, Yun Xiao 0003, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Multi-Query Vehicle Re-Identification: Viewpoint-Conditioned Network, Unified Dataset and New Metric
abstract
Existing vehicle re-identification methods mainly rely on the single query, which has limited information for vehicle representation and thus significantly hinders the performance of vehicle Re-ID in complicated surveillance networks. In this paper, we propose a more realistic and easily accessible task, called multi-query vehicle Re-ID, which leverages multiple queries to overcome viewpoint limitation of single one. Based on this task, we make three major contributions. First, we design a novel viewpoint-conditioned network (VCNet), which adaptively combines the complementary information from different vehicle viewpoints, for multi-query vehicle Re-ID. Moreover, to deal with the problem of missing vehicle viewpoints, we propose a cross-view feature recovery module which recovers the features of the missing viewpoints by learnt the correlation between the features of available and missing viewpoints. Second, we create a unified benchmark dataset, taken by 6142 cameras from a real-life transportation surveillance system, with comprehensive viewpoints and large number of crossed scenes of each vehicle for multi-query vehicle Re-ID evaluation. Finally, we design a new evaluation metric, called mean cross-scene precision (mCSP), which measures the ability of cross-scene recognition by suppressing the positive samples with similar viewpoints from the same camera. Comprehensive experiments validate the superiority of the proposed method against other methods, as well as the effectiveness of the designed metric in the evaluation of multi-query vehicle Re-ID. The codes and dataset are available at: https://github.com/zhangchaobin001/VCNet.
Aihua Zheng, Chaobin Zhang, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Image Process.4
2023 Looking and Hearing Into Details: Dual-Enhanced Siamese Adversarial Network for Audio-Visual Matching
abstract
Audio-visual cross-modal matching aims to explore the intrinsic correspondence between face images and audio clips. Existing methods usually focus on the salient features of identities between visual images and voice clips, while neglecting their subtle differences, which are crucial to distinguishing cross-modal samples. To deal with this problem, we propose a novel Dual-enhanced Siamese Adversarial Network (DSANet), which pursues the adversarial dual enhancement to highlight both salient and subtle features for robust audio-visual cross-modal matching. First, we designed a dual enhancement mechanism to enhance potential subtle features by randomly selecting a region feature for salient feature suppression, while enhancing salient features in the corresponding region to ensure the global discriminative ability. Second, to establish the correlation of subtle features in the process of eliminating cross-modal heterogeneity, we design a siamese adversarial structure to perform modal heterogeneity elimination for both enhanced salient and subtle features in a parallel manner. Moreover, we propose an adaptive masked cross-entropy loss to force the network to focus on the feature differences among hard classes. Experiments on public benchmark datasets validate the effectiveness of the proposed algorithm.
Jiaxiang Wang 0001, Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.4
2023 GPENs: Graph Data Learning With Graph Propagation-Embedding Networks
abstract
Compact representation of graph data is a fundamental problem in pattern recognition and machine learning area. Recently, graph neural networks (GNNs) have been widely studied for graph-structured data representation and learning tasks, such as graph semi-supervised learning, clustering, and low-dimensional embedding. In this article, we present graph propagation-embedding networks (GPENs), a new model for graph-structured data representation and learning problem. GPENs are mainly motivated by 1) revisiting of traditional graph propagation techniques for graph node context-aware feature representation and 2) recent studies on deeply graph embedding and neural network architecture. GPENs integrate both feature propagation on graph and low-dimensional embedding simultaneously into a unified network using a novel propagation-embedding architecture. GPENs have two main advantages. First, GPENs can be well-motivated and explained from feature propagation and deeply learning architecture. Second, the equilibrium representation of the propagation-embedding operation in GPENs has both exact and approximate formulations, both of which have simple closed-form solutions. This guarantees the compactivity and efficiency of GPENs. Third, GPENs can be naturally extended to multiple GPENs (M-GPENs) to address the data with multiple graph structures. Experiments on various semi-supervised learning tasks on several benchmark datasets demonstrate the effectiveness and benefits of the proposed GPENs and M-GPENs.
Bo Jiang 0002, Leiling Wang, Jian Cheng 0001, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Seamless Texture Optimization for RGB-D Reconstruction
abstract
Restoring high-fidelity textures for 3D reconstructed models are an increasing demand in AR/VR, cultural heritage protection, entertainment, and other relevant fields. Due to geometric errors and camera pose drifting, existing texture mapping algorithms are either plagued by blurring and ghosting or suffer from undesirable visual seams. In this paper, we propose a novel tri-directional similarity texture synthesis method to eliminate the texture inconsistency in RGB-D 3D reconstruction and generate visually realistic texture mapping results. In addition to RGB color information, we incorporate a novel color image texture detail layer serving as an additional context to improve the effectiveness and robustness of the proposed method. First, we select an optimal texture image for each triangle face of the reconstructed model to avoid texture blurring and ghosting. During the selection procedure, the texture details are weighted to avoid generating texture chart partitions across high-frequency areas. Then, we optimize the camera pose of each texture image to align with the reconstructed 3D shape. Next, we propose a tri-directional similarity function to resynthesize the image context within the boundary stripe of texture charts, which can significantly diminish the occurrence of texture seams. Finally, we introduce a global color harmonization method to address the color inconsistency between texture images captured from different viewpoints. The experimental results demonstrate that the proposed method outperforms state-of-the-art texture mapping methods and effectively overcomes texture tearing, blurring, and ghosting artifacts.
Yanping Fu, Qingan Yan, Huajian Zhou, Jin Tang 0001, Chunxia Xiao
IEEE Trans. Vis. Comput. Graph.5
2022 Interact, Embed, and EnlargE: Boosting Modality-Specific Representations for Multi-Modal Person Re-identification
abstract
Multi-modal person Re-ID introduces more complementary information to assist the traditional Re-ID task. Existing multi-modal methods ignore the importance of modality-specific information in the feature fusion stage. To this end, we propose a novel method to boost modality-specific representations for multi-modal person Re-ID: Interact, Embed, and EnlargE (IEEE). First, we propose a cross-modal interacting module to exchange useful information between different modalities in the feature extraction phase. Second, we propose a relation-based embedding module to enhance the richness of feature descriptors by embedding the global feature into the fine-grained local information. Finally, we propose multi-modal margin loss to force the network to learn modality-specific information for each modality by enlarging the intra-class discrepancy. Superior performance on multi-modal Re-ID dataset RGBNT201 and three constructed Re-ID datasets validate the effectiveness of the proposed method compared with the state-of-the-art approaches.
Zi Wang 0013, Chenglong Li 0002, Aihua Zheng, Ran He 0001, Jin Tang 0001
AAAI5
2022 Attribute-Based Progressive Fusion Network for RGBT Tracking
abstract
RGBT tracking usually suffers from various challenge factors, such as fast motion, scale variation, illumination variation, thermal crossover and occlusion, to name a few. Existing works often study fusion models to solve all challenges simultaneously, and it requires fusion models complex enough and training data large enough, which are usually difficult to be constructed in real-world scenarios. In this work, we disentangle the fusion process via the challenge attributes, and thus propose a novel Attribute-based Progressive Fusion Network (APFNet) to increase the fusion capacity with a small number of parameters while reducing the dependence on large-scale training data. In particular, we design five attribute-specific fusion branches to integrate RGB and thermal features under the challenges of thermal crossover, illumination variation, scale variation, occlusion and fast motion respectively. By disentangling the fusion process, we can use a small number of parameters for each branch to achieve robust fusion of different modalities and train each branch using the small training subset with the corresponding attribute annotation. Then, to adaptive fuse features of all branches, we design an aggregation fusion module based on SKNet. Finally, we also design an enhancement fusion transformer to strengthen the aggregated feature and modality-specific features. Experimental results on benchmark datasets demonstrate the effectiveness of our APFNet against other state-of-the-art methods.
Yun Xiao 0003, Chenglong Li 0002, Lei Liu 0049, Jin Tang 0001
AAAI5
2022 Semi-supervised Learning via Multiple Layer Graph Regularized Perception
abstract
Recently, Graph Neural Networks (GNNs) have made remarkable achievements in semi-supervised classification tasks. Nevertheless, GNNs usually rely on a specific graph convolution which has high computational complexity. To overcome this issue, recent works attempt to implicitly use adjacency matrix to guide message propagation in multi-layer perception (MLP) via neighboring contrastive loss. However, existing works accomplish implicit message passing only, without considering multi-order graph topology information. In this paper, we propose a novel method called Multiple Layer Graph Regularized Perception (MLGP). The main advantage of MLGP is to incorporate multi-order neighboring information into MLP. Further, inspired by gated mechanism, we design a linear gating to capture important features of nodes. More discriminant features can be obtained to alleviate over-smoothing. MLGP is more effective and more robust than existing works when dealing with large-scale graph data and noisy adjacency information. The comparative experiment results show that our model achieves better performance and strong robustness.
Haiyun Xu, Lili Huang 0006, Bo Jiang 0002, Jin Tang 0001, Shaojie Zhang 0002
ICPR4
2022 Efficient License Plate Recognition via Parallel Position-Aware Attention
Wenzhong Wang, Chenglong Li 0002, Jin Tang 0001
PRCV (3)4
2022 Depth-Aware Shadow Removal
abstract
Abstract Shadow removal from a single image is an ill‐posed problem because shadow generation is affected by the complex interactions of geometry, albedo, and illumination. Most recent deep learning‐based methods try to directly estimate the mapping between the non‐shadow and shadow image pairs to predict the shadow‐free image. However, they are not very effective for shadow images with complex shadows or messy backgrounds. In this paper, we propose a novel end‐to‐end depth‐aware shadow removal method without using depth images, which estimates depth information from RGB images and leverages the depth feature as guidance to enhance shadow removal and refinement. The proposed framework consists of three components, including depth prediction, shadow removal, and boundary refinement. First, the depth prediction module is used to predict the corresponding depth map of the input shadow image. Then, we propose a new generative adversarial network (GAN) method integrated with depth information to remove shadows in the RGB image. Finally, we propose an effective boundary refinement framework to alleviate the artifact around boundaries after shadow removal by depth cues. We conduct experiments on several public datasets and real‐world shadow images. The experimental results demonstrate the efficiency of the proposed method and superior performance against state‐of‐the‐art methods.
Yanping Fu, Zhenyu Gai, Haifeng Zhao 0001, Shaojie Zhang 0002, Ying Shan, Yang Wu 0001, Jin Tang 0001
Comput. Graph. Forum7
2022 RGBT tracking via reliable feature configuration
Zhengzheng Tu, Wenli Pan, Yunsheng Duan, Jin Tang 0001, Chenglong Li 0002
Sci. China Inf. Sci.4
2022 PISA: Pixel skipping-based attentional black-box adversarial attack
Jie Wang 0050, Zhao-Xia Yin, Jing Jiang 0021, Jin Tang 0001, Bin Luo 0001
Comput. Secur.4
2022 RGBT tracking based on cooperative low-rank graph model
Longfeng Shen, Xiaoxiao Wang 0003, Lei Liu 0049, Bin Hou, Yulei Jian, Jin Tang 0001, Bin Luo 0001
Neurocomputing6
2022 SILP-autoencoder for face de-occlusion
Dengdi Sun, Wandong Xie, Zhuanlian Ding, Jin Tang 0001
Neurocomputing4
2022 Multi-head collaborative learning for graph neural networks
Haiyun Xu, Bo Jiang 0002, Lili Huang 0006, Jin Tang 0001, Shaojie Zhang 0002
Neurocomputing4
2022 DBRANet: Road Extraction by Dual-Branch Encoder and Regional Attention Decoder
abstract
Although widely exploited in recent decades, road extraction is still a very significant and challenging research in the field of remote sensing image processing due to the complex background and road distribution. Among the existing CNN-based methods, U-shape architectures composed of encoders and decoders have shown their effectiveness. In this letter, we propose an improved encoder–decoder method, named DBRANet, for extracting roads from remote sensing images. In the encoding phase, we present a dual-branch network module (DBNM) to construct more effective features, thus improving the fusion feature maps of different scales. One branch utilizes the residual block, and the other branch utilizes the refined asymmetric block, which effectively increases the feature extraction capability of the backbone. In the decoding phase, considering the sinuous shape and the unbalanced distribution of roads in remote sensing images, we design a novel attention module, named the regional attention network module (RANM), to automatically learn the importance of each channel according to the regional information. Extensive experiments on several public remote sensing road data sets show that our DBRANet achieves higher segmentation [$F1$score and Intersection over Union (IoU)] and connectivity [average path length similarity (APLS)] accuracy, which verifies the effectiveness of our approach.
Sibao Chen 0001, Yu-Xin Ji, Jin Tang 0001, Bin Luo 0001, Weiqiang Wang 0001, Ke Lu 0002
IEEE Geosci. Remote. Sens. Lett.3
2022 BDTNet: Road Extraction by Bi-Direction Transformer From Remote Sensing Images
abstract
The past several years have witnessed the rapid development of the task of road extraction in high-resolution remote sensing images. However, due to the complex background and road distribution, road extraction is still a challenging research in remote sensing images. In convolutional neural networks (CNNs), the U-shaped architecture network has shown its effectiveness. But the global representation cannot be captured effectively by CNNs. While in the transformer, the self-attention (SA) module can capture the long-distance feature dependencies. A hybrid encoder-decoder method called BDTNet is proposed in this letter, which enhance the extraction of global and local information in remote sensing images. Firstly, feature maps of different scales are obtained through the backbone network. And then, on the basis of reducing the computational cost of self-attention, the Bi-Direction Transformer Module (BDTM) is constructed to capture the contextual road information in feature maps of different scales. Finally, the Feature Refinement Module (FRM) is introduced to integrate the features extracted from the backbone network and BDTM, which enhances the semantic information of the feature maps and obtains more detailed segmentation results. The results show that the proposed method achieved a high IoU of 67.09% in the DeepGlobe dataset. Extensive experiments also verify the effectiveness of the proposed method on three public remote sensing road datasets.
Jia-Xin Wang, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Semi-Supervised Semantic Segmentation of Remote Sensing Images With Iterative Contrastive Network
abstract
With the development of deep learning, semantic segmentation of remote sensing images has made great progress. However, segmentation algorithms based on deep learning usually require a huge number of labeled images for model training. For remote sensing images, pixel-level annotation usually consumes expensive resources. To alleviate this problem, this letter proposes a semi-supervised segmentation method of remote sensing images based on an iterative contrastive network. This method combines few labeled images and more unlabeled images to significantly improve the model performance. First, contrastive networks continuously learn more potential information by using better pseudo labels. Then, the iterative training method keeps the differences between models to better improve the segmentation performance. The semi-supervised experiments on different remote sensing datasets prove that this method has a better performance than the related methods. Code is available athttps://github.com/VCISwang/ICNet.
Jia-Xin Wang, Sibao Chen 0001, Chris Ding, Jin Tang 0001, Bin Luo 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 LGLNN: Label Guided Graph Learning-Neural Network for few-shot learning
Kangkang Zhao, Bo Jiang 0002, Jin Tang 0001
Neural Networks4
2022 GeCNs: Graph Elastic Convolutional Networks for Data Representation
abstract
Graph representation and learning is a fundamental problem in machine learning area. Graph Convolutional Networks (GCNs) have been recently studied and demonstrated very powerful for graph representation and learning. Graph convolution (GC) operation in GCNs can be regarded as a composition of feature aggregation and nonlinear transformation step. Existing GCs generally conduct feature aggregation on a full neighborhood set in which each node computes its representation by aggregating the feature information of all its neighbors. However, this full aggregation strategy is not guaranteed to be optimal for GCN learning and also can be affected by some graph structure noises, such as incorrect or undesired edge connections. To address these issues, we propose to integrate elastic net based selection into graph convolution and propose a novel graph elastic convolution (GeC) operation. In GeC, each node can adaptively select the optimal neighbors in its feature aggregation. The key aspect of the proposed GeC operation is that it can be formulated by a regularization framework, based on which we can derive a simple update rule to implement GeC in a self-supervised manner. Using GeC, we then present a novel GeCN for graph learning. Experimental results demonstrate the effectiveness and robustness of GeCN.
Bo Jiang 0002, Beibei Wang 0006, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Pedestrian attribute recognition: A survey
Xiao Wang 0014, Shaofei Zheng, Aihua Zheng, Zhe Chen 0013, Jin Tang 0001, Bin Luo 0001
Pattern Recognit.6
2022 RGTransformer: Region-Graph Transformer for Image Representation and Few-Shot Classification
abstract
The goal of few-shot image classification is to learn a classifier that can be well generalized to the unseen classes with a few available labeled samples. One major challenge for few-shot learning is how to conduct effective image representation for support and query images. Recently, local region-based image representation and metric learning approaches have been demonstrated effectively for few-shot classification problem. However, existing approaches generally conduct representations of image regions individually which thus lack of considering the rich spatial/structural relationships among image regions. In this paper, we propose to bridge the individual regions and exploit the structural contexts among regions via a novel Region-Graph Transformer (RGTransformer). In RGTransformer, each region aggregates the information from its neighboring regions and thus can obtain context-aware feature representations for regions. Using the proposed RGTransformer, we propose an effective metric learning model for few-shot image classification. We evaluate the proposed method on four benchmark datasets and experimental results demonstrate the effectiveness and advantages of the proposed RGTransformer.
Bo Jiang 0002, Kangkang Zhao, Jin Tang 0001
IEEE Signal Process. Lett.3
2022 Multitask Multigranularity Aggregation With Global-Guided Attention for Video Person Re-Identification
abstract
The goal of video-based person re-identification (Re-ID) is to identify the same person across multiple non-overlapping cameras. The key to accomplishing this challenging task is to sufficiently exploit both spatial and temporal cues in video sequences. However, most current methods are incapable of accurately locating semantic regions or efficiently filtering discriminative spatio-temporal features; so it is difficult to handle issues such as spatial misalignment and occlusion. Thus, we propose a novel feature aggregation framework, multi-task and multi-granularity aggregation with global-guided attention (MMA-GGA), which aims to adaptively generate more representative spatio-temporal aggregation features. Specifically, we develop a multi-task multi-granularity aggregation (MMA) module to extract features at different locations and scales to identify key semantic-aware regions that are robust to spatial misalignment. Then, to determine the importance of the multi-granular semantic information, we propose a global-guided attention (GGA) mechanism to learn weights based on the global features of the video sequence, allowing our framework to identify stable local features while ignoring occlusions. Therefore, the MMA-GGA framework can efficiently and effectively capture more robust and representative features. Extensive experiments on four benchmark datasets demonstrate that our MMA-GGA framework outperforms current state-of-the-art methods. In particular, our method achieves a rank-1 accuracy of 91.0% on the MARS dataset, the most widely used database, significantly outperforming existing methods.
Dengdi Sun, Jin Tang 0001, Zhuanlian Ding
IEEE Trans. Circuits Syst. Video Technol.4
2022 RGBT Tracking by Trident Fusion Network
abstract
In recent years, RGBT tracking has become a hot topic in the field of visual tracking, and made great progress. In this paper, we propose a novel Trident Fusion Network (TFNet) to achieve effective fusion of different modalities for robust RGBT tracking. In specific, to deploy the complementarity of features of all convolutional layers, we propose a recursive strategy to densely aggregate these features that yield robust representations of target objects in two modalities. Moreover, we design a trident architecture to integrate the fused features and both modality-specific features for robust target representations. There are three main advantages. First, retaining the classification layer of each modality is beneficial to enhance feature learning of single modality, and compared with aggregate branches, single-modality branches pay more attention to the mining of modal specific information. Second, when some modality is noisy or invalid, the modality-specific branches would capture more discriminative features for RGBT tracking. Finally, the integration of aggregation branches and single-modality branches is beneficial to the complementary learning of different modalities. In addition, we also introduce a feature pruning module in each branch to prune the redundant features and avoid network overfitting. Experimental results on four RGBT tracking benchmark datasets suggest that our tracker achieves superior performance against the state-of-the-art RGBT tracking methods.
Yabin Zhu, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 RanPaste: Paste Consistency and Pseudo Label for Semisupervised Remote Sensing Image Semantic Segmentation
abstract
With the development of deep learning, remote sensing (RS) image segmentation has been applied with marked success. However, in the process of model training, the large number of labeled images required more expensive annotation. A key challenge is how to make full use of extensive unlabeled images available to improve the segmentation model. In this article, we propose a semisupervised remote sensing image semantic segmentation method defined as RanPaste, which combines labeled images with unlabeled images to improve segmentation performance. First, we obtain pseudo label by randomly pasting part of the ground truth label into the predicted segmentation map. Then, we combine the labeled and unlabeled images to generate rough predictions after strong augmentation. Finally, by using the semisupervised loss, we achieve better performance on remote sensing image segmentation. Our method combines consistency regularization and pseudo label and then utilizes thresholds to gradually improve the model performance. RanPaste enables the model to learn more underlying information in the unlabeled data. Experimental results on six datasets show that RanPaste can learn more latent information from unlabeled data to improve segmentation performance. Besides, our approach achieves better segmentation results on different network structures and datasets.
Jia-Xin Wang, Sibao Chen 0001, Chris Ding, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Reliable Contrastive Learning for Semi-Supervised Change Detection in Remote Sensing Images
abstract
With the development of deep learning in remote sensing (RS) image change detection (CD), the dependence of CD models on labeled data has become an important problem. To make better use of the comparatively resource-saving unlabeled data, the CD method based on semi-supervised learning (SSL) is worth further study. This article proposes a reliable contrastive learning (RCL) method for semi-supervised RS image CD. First, according to the task characteristics of CD, we design the contrastive loss based on the changed areas to enhance the model’s feature extraction ability for changed objects. Then, to improve the quality of pseudo labels in SSL, we use the uncertainty of unlabeled data to select reliable pseudo labels for model training. Combining these methods, semi-supervised CD models can make full use of unlabeled data. Extensive experiments on three widely used CD datasets demonstrate the effectiveness of the proposed method. The results show that our semi-supervised approach has a better performance than related methods. The code is available athttps://github.com/VCISwang/RC-Change-Detection.
Jia-Xin Wang, Teng Li 0001, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001, Richard C. Wilson 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Grouped Bidirectional LSTM Network and Multistage Fusion Convolutional Transformer for Hyperspectral Image Classification
abstract
The efficiently and effectively discriminative spectral-spatial feature representation is essential for hyperspectral image (HSI) classification. However, most of the existing methods rely on the patch-based convolutional neural networks (CNNs) whose ability of extracting the global spatial information is very limited. To address this issue, in this paper, we propose a two-branch network consisting of a grouped bidirectional long short-term memory (GBiLSTM) network and multi-stage fusion convolutional transformer (MFCT) for HSI classification. In the proposed GBiLSTM-MFCT, to extract the spectral features of HSI efficiently, a GBiLSTM network is designed by dividing the sequence features and hidden units of BiLSTM network into several separate groups. To simultaneously extract the global and local spatial features of HSI, a MFCT is proposed by fusing the features of different levels obtained from the multiple phases of convolutional vision transformer. Moreover, in the multi-headed attention module of each stage, blueprint separable convolution based self-attention (BSCA) module is designed which is able to model the global and local spatial information effectively. The outputs of GBiLSTM network and MFCT are fused to generate discriminative and robust spectral-spatial features for HSI classification. Experiments on three benchmark data sets of IN, UP and KSC demonstrate that the proposed GBiLSTM-MFCT exhibits higher classification performance with very limited labeled samples than eight state-of-the-art methods.
Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Entropy Guided Adversarial Domain Adaptation for Aerial Image Semantic Segmentation
abstract
Recent advances on aerial image semantic segmentation mainly employ the domain adaption to transfer knowledge from the source domain to the target domain. Despite the remarkable achievement, most methods focus on the global marginal distribution alignment to reduce the domain shift between source and target domains, leading to a wrong mapping of the well-aligned features. In this article, we propose an effective unsupervised domain adaptation approach, which relies on a novel entropy guided adversarial learning algorithm, for aerial image semantic segmentation. In specific, we perform local feature alignment between domains by learning a self-adaptive weight from the target prediction probability map to measure the interdomain discrepancy. To exploit the meaningful structure information among semantic regions, we propose to utilize the graph convolutions for long-range semantic reasoning. Comprehensive experimental results on the benchmark dataset of aerial image semantic segmentation and natural scenes demonstrate the superior performance of the proposed method compared to the state-of-the-art methods.
Aihua Zheng, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Remote Sensing Scene Classification via Multi-Branch Local Attention Network
abstract
Remote sensing scene classification (RSSC) is a hotspot and play very important role in the field of remote sensing image interpretation in recent years. With the recent development of the convolutional neural networks, a significant breakthrough has been made in the classification of remote sensing scenes. Many objects form complex and diverse scenes through spatial combination and association, which makes it difficult to classify remote sensing image scenes. The problem of insufficient differentiation of feature representations extracted by Convolutional Neural Networks (CNNs) still exists, which is mainly due to the characteristics of similarity for inter-class images and diversity for intra-class images. In this paper, we propose a remote sensing image scene classification method via Multi-Branch Local Attention Network (MBLANet), where Convolutional Local Attention Module (CLAM) is embedded into all down-sampling blocks and residual blocks of ResNet backbone. CLAM contains two submodules, Convolutional Channel Attention Module (CCAM) and Local Spatial Attention Module (LSAM). The two submodules are placed in parallel to obtain both channel and spatial attentions, which helps to emphasize the main target in the complex background and improve the ability of feature representation. Extensive experiments on three benchmark datasets show that our method is better than state-of-the-art methods.
Sibao Chen 0001, Qing-Song Wei, Wenzhong Wang, Jin Tang 0001, Bin Luo 0001, Zuyuan Wang
IEEE Trans. Image Process.4
2022 Attribute and State Guided Structural Embedding Network for Vehicle Re-Identification
abstract
Vehicle re-identification (Re-ID) is a crucial task in smart city and intelligent transportation, aiming to match vehicle images across non-overlapping surveillance camera scenarios. However, the images of different vehicles may have small visual discrepancies when they have the same/similar attributes, e.g., the same/similar color, type, and manufacturer. Meanwhile, the images from a vehicle may have large visual discrepancies with different states, e.g., different camera views, vehicle viewpoints, and capture time. In this paper, we propose an attribute and state guided structural embedding network (ASSEN) to achieve discriminative feature learning by attribute-based enhancement and state-based weakening for vehicle Re-ID. First, we propose an attribute-based enhancement and expanding module to enhance the discrimination of vehicle features through identity-related attribute information, and we design an attribute-based expanding loss to increase the feature gap between different vehicles. Second, we design a state-based weakening and shrinking module, which not only weakens the state information that interferes with identification but also reduces the intra-class feature gap by a state-based shrinking loss. Third, we propose a global structural embedding module that exploits the attribute information and state information to explore hierarchical relationships between vehicle features, then we use these relationships for feature embedding to learn more robust vehicle features. Extensive experiments on benchmark datasets VeRi-776, VehicleID, and VERI-Wild demonstrate the superior performance and generalization of the proposed method against state-of-the-art vehicle Re-ID methods. The code is available at https://github.com/ttaalle/fast_assen.
Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Image Process.4
2022 LasHeR: A Large-Scale High-Diversity Benchmark for RGBT Tracking
abstract
RGBT tracking receives a surge of interest in the computer vision community, but this research field lacks a large-scale and high-diversity benchmark dataset, which is essential for both the training of deep RGBT trackers and the comprehensive evaluation of RGBT tracking methods. To this end, we present a La rge- s cale H igh-diversity [Formula: see text]nchmark for short-term R GBT tracking (LasHeR) in this work. LasHeR consists of 1224 visible and thermal infrared video pairs with more than 730K frame pairs in total. Each frame pair is spatially aligned and manually annotated with a bounding box, making the dataset well and densely annotated. LasHeR is highly diverse capturing from a broad range of object categories, camera viewpoints, scene complexities and environmental factors across seasons, weathers, day and night. We conduct a comprehensive performance evaluation of 12 RGBT tracking algorithms on the LasHeR dataset and present detailed analysis. In addition, we release the unaligned version of LasHeR to attract the research interest for alignment-free RGBT tracking, which is a more practical task in real-world applications. The datasets and evaluation protocols are available at: https://github.com/mmic-lcl/Datasets-and-benchmark-code.
Chenglong Li 0002, Wanlin Xue, Yaqing Jia, Zhichen Qu, Bin Luo 0001, Jin Tang 0001, Dengdi Sun
IEEE Trans. Image Process.6
2022 Weakly Alignment-Free RGBT Salient Object Detection With Deep Correlation Network
abstract
RGBT Salient Object Detection (SOD) focuses on common salient regions of a pair of visible and thermal infrared images. Existing methods perform on the well-aligned RGBT image pairs, but the captured image pairs are always unaligned and aligning them requires much labor cost. To handle this problem, we propose a novel deep correlation network (DCNet), which explores the correlations across RGB and thermal modalities, for weakly alignment-free RGBT SOD. In particular, DCNet includes a modality alignment module based on the spatial affine transformation, the feature-wise affine transformation and the dynamic convolution to model the strong correlation of two modalities. Moreover, we propose a novel bi-directional decoder model, which combines the coarse-to-fine and fine-to-coarse processes for better feature enhancement. In particular, we design a modality correlation ConvLSTM by adding the first two components of modality alignment module and a global context reinforcement module into ConvLSTM, which is used to decode hierarchical features in both top-down and button-up manners. Extensive experiments on three public benchmark datasets show the remarkable performance of our method against state-of-the-art methods.
Zhengzheng Tu, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Image Process.4
2022 M5L: Multi-Modal Multi-Margin Metric Learning for RGBT Tracking
abstract
Classifying hard samples in the course of RGBT tracking is a quite challenging problem. Existing methods only focus on enlarging the boundary between positive and negative samples, but ignore the relations of multilevel hard samples, which are crucial for the robustness of hard sample classification. To handle this problem, we propose a novel Multi-Modal Multi-Margin Metric Learning framework named M5L for RGBT tracking. In particular, we divided all samples into four parts including normal positive, normal negative, hard positive and hard negative ones, and aim to leverage their relations to improve the robustness of feature embeddings, e.g., normal positive samples are closer to the ground truth than hard positive ones. To this end, we design a multi-modal multi-margin structural loss to preserve the relations of multilevel hard samples in the training stage. In addition, we introduce an attention-based fusion module to achieve quality-aware integration of different source data. Extensive experiments on large-scale datasets testify that our framework clearly improves the tracking performance and performs favorably the state-of-the-art RGBT trackers.
Zhengzheng Tu, Chun Lin, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Image Process.5
2022 Beyond Greedy Search: Tracking by Multi-Agent Reinforcement Learning-Based Beam Search
abstract
To track the target in a video, current visual trackers usually adopt greedy search for target object localization in each frame, that is, the candidate region with the maximum response score will be selected as the tracking result of each frame. However, we found that this may be not an optimal choice, especially when encountering challenging tracking scenarios such as heavy occlusion and fast motion. In particular, if a tracker drifts, errors will be accumulated and would further make response scores estimated by the tracker unreliable in future frames. To address this issue, we propose to maintain multiple tracking trajectories and apply beam search strategy for visual tracking, so that the trajectory with fewer accumulated errors can be identified. Accordingly, this paper introduces a novel multi-agent reinforcement learning based beam search tracking strategy, termed BeamTracking. It is mainly inspired by the image captioning task, which takes an image as input and generates diverse descriptions using beam search algorithm. Accordingly, we formulate the tracking as a sample selection problem fulfilled by multiple parallel decision-making processes, each of which aims at picking out one sample as their tracking result in each frame. Each maintained trajectory is associated with an agent to perform the decision-making and determine what actions should be taken to update related information. More specifically, using the classification-based tracker as the baseline, we first adopt bi-GRU to encode the target feature, proposal feature, and its response score into a unified state representation. The state feature and greedy search result are then fed into the first agent for independent action selection. Afterwards, the output action and state features are fed into the subsequent agent for diverse results prediction. When all the frames are processed, we select the trajectory with the maximum accumulated score as the tracking result. Extensive experiments on seven popular tracking benchmark datasets validated the effectiveness of the proposed algorithm.
Xiao Wang 0014, Zhe Chen 0013, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Dacheng Tao
IEEE Trans. Image Process.4
2022 MsKAT: Multi-Scale Knowledge-Aware Transformer for Vehicle Re-Identification
abstract
Existing vehicle re-identification (Re-ID) methods usually suffer from intra-instance discrepancy and inter-instance similarity. The key to solving this problem lies in filtering out identity-irrelevant interference and collecting identity-relevant vehicle details. In this paper, we aim to design a robust vehicle Re-ID framework that trains a model guided by knowledge vectors yet is able to disentangle the identity-relevant features and identity-irrelevant features. Toward this end, we propose a novel Multi-scale Knowledge-Aware Transformer (MsKAT) to build a knowledge-guided multi-scale feature alignment framework. First, we construct a Knowledge-Aware Transformer (KAT) to interact with semantic knowledge and visual feature. KAT mainly includes State elimination Transformer (SeT) to eliminate state (camera, viewpoint) interference and Attribute aggregation Transformer (AaT) to gather attribute (color, type) information. Second, to learn the knowledge-guided sample differences, we propose to encourage the separation of identity-relevant features and identity-irrelevant features by a Knowledge-Guided Alignment loss ($\mathcal {L}_{KGA}$). Specifically,$\mathcal {L}_{KGA}$suppresses the difference between knowledge-guided positive pairs and the similarity between knowledge-guided negative pairs. Third, with the multi-scale settings of KAT and$\mathcal {L}_{KGA}$, our model can capture knowledge-guided visual consistency features at different scales. Extensive evidence demonstrates our approach achieves new state-of-the-art on three widely-used vehicle re-identification benchmarks.
Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Intell. Transp. Syst.4
2022 Viewpoint-Aware Progressive Clustering for Unsupervised Vehicle Re-Identification
abstract
Vehicle re-identification (Re-ID) is an active task due to its importance in large-scale intelligent monitoring in smart cities. Despite the rapid progress in recent years, most existing methods handle vehicle Re-ID task in a supervised manner, which is both time and labor-consuming and limits their application to real-life scenarios. Recently, unsupervised person Re-ID methods achieve impressive performance by exploring domain adaption or clustering-based techniques. However, one cannot directly generalize these methods to vehicle Re-ID since vehicle images present huge appearance variations in different viewpoints. To handle this problem, we propose a novel viewpoint-aware clustering algorithm for unsupervised vehicle Re-ID. In particular, we first divide the entire feature space into different subspaces according to the predicted viewpoints and then perform a progressive clustering to mine the accurate relationship among samples. Comprehensive experiments against the state-of-the-art methods on two multi-viewpoint benchmark datasets VeRi-776 and VeRi-Wild validate the promising performance of the proposed method in both with and without domain adaption scenarios while handling unsupervised vehicle Re-ID.
Aihua Zheng, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Intell. Transp. Syst.4
2022 PH-GCN: Person Retrieval With Part-Based Hierarchical Graph Convolutional Network
abstract
Compact feature representation of person image is important for person re-identification (Re-ID) task. Recently, part-based representation models have been widely studied for extracting the more compact and robust feature representation for person image to improve person Re-ID results. However, existing part-based representation models mostly extract the features of different parts independently which ignore the spatial relationship information among different parts. To address this issue, in this paper we propose a novel deep learning framework, named Part-based Hierarchical Graph Convolutional Network (PH-GCN) for person Re-ID problem. Given a person image, PH-GCN first constructs a hierarchical graph to represent the spatial relationships among different parts. Then, both local and global feature learning is achieved by the feature information passing in PH-GCN, which takes the information of other parts into account for part feature representation. Finally, a perceptron layer is adopted for the final person part label prediction and re-identification. The proposed framework provides a general solution that integrateslocal,globalandstructuralfeature learning simultaneously in a unified end-to-end network representation and learning. Extensive experiments on several widely used benchmark datasets demonstrate the effectiveness and benefits of the proposed PH-GCN approach for person Re-ID task.
Bo Jiang 0002, Xixi Wang 0005, Aihua Zheng, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.4
2022 RGBT Tracking via Noise-Robust Cross-Modal Ranking
abstract
Existing RGBT tracking methods usually localize a target object with a bounding box, in which the trackers are often affected by the inclusion of background clutter. To address this issue, this article presents a novel algorithm, called noise-robust cross-modal ranking, to suppress background effects in target bounding boxes for RGBT tracking. In particular, we handle the noise interference in cross-modal fusion and seed labels from the following two aspects. First, the soft cross-modality consistency is proposed to allow the sparse inconsistency in fusing different modalities, aiming to take both collaboration and heterogeneity of different modalities into account for more effective fusion. Second, the optimal seed learning is designed to handle label noises of ranking seeds caused by some problems, such as irregular object shape and occlusion. In addition, to deploy the complementarity and maintain the structural information of different features within each modality, we perform an individual ranking for each feature and employ a cross-feature consistency to pursue their collaboration. A unified optimization framework with an efficient convergence speed is developed to solve the proposed model. Extensive experiments demonstrate the effectiveness and efficiency of the proposed approach comparing with state-of-the-art tracking methods on GTOT and RGBT234 benchmark data sets.
Chenglong Li 0002, Zhiqiang Xiang, Jin Tang 0001, Bin Luo 0001, Futian Wang
IEEE Trans. Neural Networks Learn. Syst.3
2022 Tracking by Joint Local and Global Search: A Target-Aware Attention-Based Approach
abstract
Tracking-by-detection is a very popular framework for single-object tracking that attempts to search the target object within a local search window for each frame. Although such a local search mechanism works well on simple videos, however, it makes the trackers sensitive to extremely challenging scenarios, such as heavy occlusion and fast motion. In this article, we propose a novel and general target-aware attention mechanism (termed TANet) and integrate it with a tracking-by-detection framework to conduct joint local and global search for robust tracking. Specifically, we extract the features of the target object patch and continuous video frames; then, we concatenate and feed them into a decoder network to generate target-aware global attention maps. More importantly, we resort to adversarial training for better attention prediction. The appearance and motion discriminator networks are designed to ensure its consistency in spatial and temporal views. In the tracking procedure, we integrate target-aware attention with multiple trackers by exploring candidate search regions for robust tracking. Extensive experiments on both short- and long-term tracking benchmark datasets all validated the effectiveness of our algorithm.
Xiao Wang 0014, Jin Tang 0001, Bin Luo 0001, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2021 Robust Multi-Modality Person Re-identification
abstract
To avoid the illumination limitation in visible person re-identification (Re-ID) and the heterogeneous issue in cross-modality Re-ID, we propose to utilize complementary advantages of multiple modalities including visible (RGB), near infrared (NI) and thermal infrared (TI) ones for robust person Re-ID. A novel progressive fusion network is designed to learn effective multi-modal features from single to multiple modalities and from local to global views. Our method works well in diversely challenging scenarios even in the presence of missing modalities. Moreover, we contribute a comprehensive benchmark dataset, RGBNT201, including 201 identities captured from various challenging conditions, to facilitate the research of RGB-NI-TI multi-modality person Re-ID. Comprehensive experiments on RGBNT201 dataset comparing to the state-of-the-art methods demonstrate the contribution of multi-modality person Re-ID and the effectiveness of the proposed approach, which launch a new benchmark and a new baseline for multi-modality person Re-ID.
Aihua Zheng, Zi Wang 0013, Zi-Han Chen, Chenglong Li 0002, Jin Tang 0001
AAAI5
2021 Progressive Fusion Network for Safety Protection Detection
Futian Wang, Lugang Wang, Jin Tang 0001, Chenglong Li 0002
ICIG (1)3
2021 GAMnet: Robust Feature Matching via Graph Adversarial-Matching Network
abstract
Recently, deep graph matching (GM) methods have gained increasing attention. These methods integrate graph nodes¡¯s embedding, node/edges¡¯s affinity learning and final correspondence solver together in an end-to-end manner. For deep graph matching problem, one main issue is how to generate consensus node's embeddings for both source and target graphs that best serve graph matching tasks. In addition, it is also challenging to incorporate the discrete one-to-one matching constraints into the differentiable correspondence solver in deep matching network. To address these issues, we propose a novel Graph Adversarial Matching Network (GAMnet) for graph matching problem. GAMnet integrates graph adversarial embedding and graph matching simultaneously in a unified end-to-end network which aims to adaptively learn distribution consistent and domain invariant embeddings for GM tasks. Also, GAMnet exploits sparse GM optimization as correspondence solver which is differentiable and can also incorporate discrete one-to-one matching constraints approximately in natural in the final matching prediction. Experimental results on three public benchmarks demonstrate the effectiveness and benefits of the proposed GAMnet.
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
ACM Multimedia4
2021 Deep Double Center Hashing for Face Image Retrieval
Wenzhong Wang, Jin Tang 0001
PRCV (2)3
2021 Learning spatio-temporal correlation filter for visual tracking
Youmin Yan, Xixian Guo, Jin Tang 0001, Chenglong Li 0002, Xin Wang 0013
Neurocomputing3
2021 RGBT tracking via cross-modality message passing
Xiao Wang 0014, Chenglong Li 0002, Jinmin Hu, Jin Tang 0001
Neurocomputing5
2021 Dual-decoder graph autoencoder for unsupervised graph representation learning
Dengdi Sun, Dashuang Li, Zhuanlian Ding, Xingyi Zhang 0001, Jin Tang 0001
Knowl. Based Syst.5
2021 Reversible data hiding in encrypted images based on pixel prediction and multi-MSB planes rearrangement
abstract
Great concern has arisen in the field of reversible data hiding in encrypted images (RDHEI) due to the development of cloud storage and privacy protection. RDHEI is an effective technology that can embed additional data after image encryption, extract additional data error-free and reconstruct original images losslessly. In this paper, a high-capacity and fully reversible RDHEI method is proposed, which is based on pixel prediction and multi-MSB (most significant bit) planes rearrangement. First, the median edge detector (MED) predictor is used to calculate the predicted value. Next, unlike previous methods, in our proposed method, signs of prediction errors (PEs) are represented by one bit plane and absolute values of PEs are represented by other bit planes. Then, we divide bit planes into uniform blocks and non-uniform blocks, and rearrange these blocks. Finally, according to different pixel prediction schemes, different numbers of additional data are embedded adaptively. The experimental results prove that our method has higher embedding capacity compared with state-of-the-art RDHEI methods.
Zhao-Xia Yin, Xiaomeng She, Jin Tang 0001, Bin Luo 0001
Signal Process.3
2021 Edge-Guided Non-Local Fully Convolutional Network for Salient Object Detection
abstract
Fully Convolutional Neural Network (FCN) has been widely applied to salient object detection recently by virtue of high-level semantic feature extraction, but existing FCN-based methods still suffer from continuous striding and pooling operations leading to loss of spatial structure and blurred edges. To maintain the clear edge structure of salient objects, we propose a novel Edge-guided Non-local FCN (ENFNet) to perform edge-guided feature learning for accurate salient object detection. In a specific, we extract hierarchical global and local information in FCN to incorporate non-local features for effective feature representations. To preserve good boundaries of salient objects, we propose a guidance block to embed edge prior knowledge into hierarchical feature maps. The guidance block not only performs feature-wise manipulation but also spatial-wise transformation for effective edge embeddings. Our model is trained on the MSRA-B dataset and tested on five popular benchmark datasets. Comparing with the state-of-the-art methods, the proposed method performance well on five datasets.
Zhengzheng Tu, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.4
2021 Dynamic Attention Guided Multi-Trajectory Analysis for Single Object Tracking
abstract
Most of the existing single object trackers track the target in a unitary local search window, making them particularly vulnerable to challenging factors such as heavy occlusions and out-of-view movements. Despite the attempts to further incorporate global search, prevailing mechanisms that cooperate local and global search are relatively static, thus are still sub-optimal for improving tracking performance. By further studying the local and global search results, we raise a question: can we allow more dynamics for cooperating both results? In this paper, we propose to introduce more dynamics by devising a dynamic attention-guided multi-trajectory tracking strategy. In particular, we construct dynamic appearance model that contains multiple target templates, each of which provides its own attention for locating the target in the new frame. Guided by different attention, we maintain diversified tracking results for the target to build multi-trajectory tracking history, allowing more candidates to represent the true target trajectory. After spanning the whole sequence, we introduce a multi-trajectory selection network to find the best trajectory that deliver improved tracking performance. Extensive experimental results show that our proposed tracking strategy achieves compelling performance on various large-scale tracking benchmarks. The project page of this paper can be found athttps://sites.google.com/view/mt-track/.
Xiao Wang 0014, Zhe Chen 0013, Jin Tang 0001, Bin Luo 0001, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 RGBT Tracking via Multi-Adapter Network with Hierarchical Divergence Loss
abstract
RGBT tracking has attracted increasing attention since RGB and thermal infrared data have strong complementary advantages, which could make trackers all-day and all-weather work. Existing works usually focus on extracting modality-shared or modality-specific information, but the potentials of these two cues are not well explored and exploited in RGBT tracking. In this paper, we propose a novel multi-adapter network to jointly perform modality-shared, modality-specific and instance-aware target representation learning for RGBT tracking. To this end, we design three kinds of adapters within an end-to-end deep learning framework. In specific, we use the modified VGG-M as the generality adapter to extract the modality-shared target representations. To extract the modality-specific features while reducing the computational complexity, we design a modality adapter, which adds a small block to the generality adapter in each layer and each modality in a parallel manner. Such a design could learn multilevel modality-specific representations with a modest number of parameters as the vast majority of parameters are shared with the generality adapter. We also design instance adapter to capture the appearance properties and temporal variations of a certain target. Moreover, to enhance the shared and specific features, we employ the loss of multiple kernel maximum mean discrepancy to measure the distribution divergence of different modal features and integrate it into each layer for more robust representation learning. Extensive experiments on two RGBT tracking benchmark datasets demonstrate the outstanding performance of the proposed tracker against the state-of-the-art methods.
Andong Lu, Chenglong Li 0002, Yuqing Yan, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Image Process.4
2021 Multi-Interactive Dual-Decoder for RGB-Thermal Salient Object Detection
abstract
RGB-thermal salient object detection (SOD) aims to segment the common prominent regions of visible image and corresponding thermal infrared image that we call it RGBT SOD. Existing methods don't fully explore and exploit the potentials of complementarity of different modalities and multi-type cues of image contents, which play a vital role in achieving accurate results. In this paper, we propose a multi-interactive dual-decoder to mine and model the multi-type interactions for accurate RGBT SOD. In specific, we first encode two modalities into multi-level multi-modal feature representations. Then, we design a novel dual-decoder to conduct the interactions of multi-level features, two modalities and global contexts. With these interactions, our method works well in diversely challenging scenarios even in the presence of invalid modality. Finally, we carry out extensive experiments on public RGBT and RGBD SOD datasets, and the results show that the proposed method achieves the outstanding performance against state-of-the-art algorithms. The source code has been released at: https://github.com/lz118/Multi-interactive-Dual-decoder.
Zhengzheng Tu, Chenglong Li 0002, Yang Lang, Jin Tang 0001
IEEE Trans. Image Process.5
2021 Co-Saliency Detection via a General Optimization Model and Adaptive Graph Learning
abstract
Co-saliency detection is an important research problem, and has been widely used in computer vision area. One main challenge for co-saliency detection problem is how to explore both interactive information among different images and individual salient information within each image simultaneously in co-saliency estimation. In this paper, we propose a novel general optimization framework with adaptive graph learning for co-saliency estimation problem. The proposed model integrates multiple cues including background, and foreground priors, structural information of images, and image feature representation together to obtain a uniform, and accurate co-saliency estimation. One main benefit of the proposed co-saliency method is that it conducts co-saliency propagation, and prediction across different images while maintains the individual salient information of each image, which ensures the consistency, and communication across different images effectively in co-saliency estimation. To improve the accuracy of co-saliency estimation, we adaptively learn a neighborhood, and structured graph to conduct co-saliency propagation among superpixels. An effective optimization algorithm has been designed to seek the optimal solution for the proposed co-saliency optimization model. Experimental results on several widely used datasets show that our method outperforms some other related co-saliency detection methods.
Bo Jiang 0002, Xingyue Jiang, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.3
2021 STGL: Spatial-Temporal Graph Representation and Learning for Visual Tracking
abstract
Tracking-by-detection framework has been normally adopted in visual tracking methods. It aims to localize the visual target object with a bounding box. However, the bounding box is usually difficult to describe the target object accurately and thus easily introduces noisy background information, which usually degrades the final tracking results. Recently, weighted patch representation of the object has been shown very effectively for suppressing the undesirable background information and thus can obviously improve the tracking results. In this paper, we propose a novel Spatial-Temporal Graph representation and Learning (STGL) model to generate a kind of robust target representation for visual tracking problem. The main aspect of STGL is that it aims to exploit both spatial (within each frame) and temporal (between consecutive frames) structure of patches simultaneously in a unified graph representation and semi-supervised learning model. Comparing with existing works, STGL naturally exploits the learned representation of object in previous frame and thus can obtain the representation of object in current frame more accurately and robustly. A new ADMM algorithm is derived to solve the proposed STGL model. Based on the proposed object representation, we then adapt the structured SVM by introducing scale estimation to achieve object tracking. Extensive experiments show that our method outperforms the state-of-the-art patch based tracking methods on two standard benchmark datasets.
Bo Jiang 0002, Bin Luo 0001, Xiaochun Cao, Jin Tang 0001
IEEE Trans. Multim.5
2021 cmSalGAN: RGB-D Salient Object Detection With Cross-View Generative Adversarial Networks
abstract
Image salient object detection (SOD) is an active research topic in computer vision and multimedia area. Fusing complementary information of RGB and depth has been demonstrated to be effective for image salient object detection which is known as RGB-D salient object detection problem. The main challenge for RGB-D salient object detection is how to exploit the salient cues of both intra-modality (RGB, depth) and cross-modality simultaneously which is known as cross-modality detection problem. In this paper, we tackle this challenge by designing a novel cross-modality Saliency Generative Adversarial Network (cmSalGAN). cmSalGAN aims to learn an optimal view-invariant and consistent pixel-level representation for RGB and depth images via a novel adversarial learning framework, which thus incorporates both information of intra-view and correlation information of cross-view images simultaneously for RGB-D saliency detection problem. To further improve the detection results, the attention mechanism and edge detection module are also incorporated into cmSalGAN. The entire cmSalGAN can be trained in an end-to-end manner by using the standard deep neural network framework. Experimental results show that cmSalGAN achieves the new state-of-the-art RGB-D saliency detection performance on several benchmark datasets.
Bo Jiang 0002, Zitai Zhou, Xiao Wang 0014, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.4
2021 Segmenting Objects in Day and Night: Edge-Conditioned CNN for Thermal Image Semantic Segmentation
abstract
Despite much research progress in image semantic segmentation, it remains challenging under adverse environmental conditions caused by imaging limitations of the visible spectrum, while thermal infrared cameras have several advantages over cameras for the visible spectrum, such as operating in total darkness, insensitive to illumination variations, robust to shadow effects, and strong ability to penetrate haze and smog. These advantages of thermal infrared cameras make the segmentation of semantic objects in day and night. In this article, we propose a novel network architecture, called edge-conditioned convolutional neural network (EC-CNN), for thermal image semantic segmentation. Particularly, we elaborately design a gated featurewise transform layer in EC-CNN to adaptively incorporate edge prior knowledge. The whole EC-CNN is end-to-end trained and can generate high-quality segmentation results with edge guidance. Meanwhile, we also introduce a new benchmark data set named "Segmenting Objects in Day And night" (SODA) for comprehensive evaluations in thermal image semantic segmentation. SODA contains over 7168 manually annotated and synthetically generated thermal images with 20 semantic region labels and from a broad range of viewpoints and scene complexities. Extensive experiments on SODA demonstrate the effectiveness of the proposed EC-CNN against state-of-the-art methods.
Chenglong Li 0002, Yan Yan 0002, Bin Luo 0001, Jin Tang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2020 Challenge-Aware RGBT Tracking
Chenglong Li 0002, Lei Liu 0049, Andong Lu, Jin Tang 0001
ECCV (22)5
2020 Information Enhanced Graph Convolutional Networks for Skeleton-based Action Recognition
abstract
Skeleton-based action recognition has recently attracted much attention in computer vision. The latest methods are mostly based on graph convolutional networks (GCNs), which construct the human body as spatial-temporal Skeleton graphs, and has achieved excellent performance. However, previous studies only capture the local and rough information based on the physical dependencies among joints, which may miss implicit joint correlations. In this work, we propose a novel action recognition model, namely Information Enhanced Graph Convolutional Networks (IE-GCN). To improve the accuracy and robustness of recognition, this model capture higher-order dependency in the skeleton-based graph by expanding the joint neighbors, and combine second stage skeleton features (the lengths and directions of bones) to enhance the discriminative information simultaneously. In addition, an training strategy is designed to solve the framework. Extensive experiments on two large-scale public datasets, NTU-RGBD and Kinetics-Skeleton, demonstrate the superior performance of the proposed algorithms over the state-of-the-art methods.
Dengdi Sun, Fanchen Zeng, Bin Luo 0001, Jin Tang 0001, Zhuanlian Ding
IJCNN4
2020 Synthesizing Large-Scale Datasets for License Plate Detection and Recognition in the Wild
Chaochen Wang, Wenzhong Wang, Chenglong Li 0002, Jin Tang 0001
PRCV (3)4
2020 Defense against adversarial attacks by low-level image transformations
abstract
Deep neural networks (DNNs) are vulnerable to adversarial examples, which can fool classifiers by maliciously adding imperceptible perturbations to the original input. Currently, a large number of research on defending adversarial examples pay little attention to the real-world applications, either with high computational complexity or poor defensive effects. Motivated by this observation, we develop an efficient preprocessing module to defend adversarial attacks. Specifically, before an adversarial example is fed into the model, we perform two low-level image transformations, WebP compression and flip operation, on the picture. Then we can get a de-perturbed sample that can be correctly classified by DNNs. WebP compression is utilized to remove the small adversarial noises. Due to the introduction of loop filtering, there will be no square effect like JPEG compression, so the visual quality of the denoised image is higher. And flip operation, which flips the image once along one side of the image, destroys the specific structure of adversarial perturbations. By taking class activation mapping to localize the discriminative image regions, we show that flipping image may mitigate adversarial effects. Extensive experiments demonstrate that the proposed scheme outperforms the state-of-the-art defense methods. It can effectively defend adversarial attacks while ensuring only slight accuracy drops on normal images.
Zhao-Xia Yin, Jie Wang 0050, Jin Tang 0001, Wenzhong Wang
Int. J. Intell. Syst.4
2020 Multi-modal foreground detection via inter- and intra-modality-consistent low-rank separation
Aihua Zheng, Naipeng Ye, Chenglong Li 0002, Xiao Wang 0014, Jin Tang 0001
Neurocomputing5
2020 Multi-scale attention vehicle re-identification
Aihua Zheng, Xianmin Lin, Jiacheng Dong, Wenzhong Wang, Jin Tang 0001, Bin Luo 0001
Neural Comput. Appl.5
2020 RGBT Salient Object Detection: Benchmark and A Novel Cooperative Ranking Approach
abstract
Despite significant progress, image saliency detection still remains a challenging task in complex scenes and environments. Integrating multiple different but complementary cues, like RGB and Thermal infrared (RGBT), may be an effective way for boosting saliency detection performance. This work contributes a RGBT image dataset, which includes 821 spatially aligned RGBT image pairs and their ground truth annotations for saliency detection purpose. Moreover, 11 challenges are annotated on these image pairs for performing the challenge-sensitive analysis and 3 kinds of baseline methods are implemented to provide a comprehensive comparison platform. With this benchmark, we propose a novel approach based on a cooperative ranking algorithm for RGBT saliency detection. In particular, we introduce a weight for each modality to describe the reliability and a ℓ1-based cross-modal consistency in a unified ranking model, and design an efficient solver to iteratively optimize several subproblems with closed-form solutions. Extensive experiments against baseline methods demonstrate the effectiveness of the proposed approach on both our introduced dataset and a public dataset.
Jin Tang 0001, Dongzhe Fan, Xiaoxiao Wang 0003, Zhengzheng Tu, Chenglong Li 0002
IEEE Trans. Circuits Syst. Video Technol.1
2020 Feature Matching With Intra-Group Sparse Model
abstract
Feature matching is a fundamental problem in computer vision area. In many real applications, one can usually obtain some potential (candidate) matches C by using some discriminative feature descriptors, such as SIFT descriptor. Then, the feature matching problem can be formulated as the problem of trying to select the correct matches S from the potential match set C. In this paper, we propose to solve matches selection by developing a novel intra-group sparse matching (IGSM) model. Our IGSM is motivated by a simple observation that the potential match set C can be divided into several non-overlapping groups Ci, among which the correct matches S are uniformly distributed. We thus develop an intra-group selection model to conduct matches selection at the intra-group level to incorporate the one-to-one matching constraint more in matches selection process. Our IGSM model has three main advantages: (1) The selection mechanism is parameter-free; (2) it generates an intra-group sparse solution which better maintains the one-to-one matching constraint in nature; (3) a simple yet effective update algorithm has been derived to solve IGSM model. The optimality and convergence of the algorithm are theoretically guaranteed. Experimental results on several image feature matching datasets show the effectiveness and efficiency of the proposed IGSM matching method.
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.2
2020 RGB-T Image Saliency Detection via Collaborative Graph Learning
abstract
Image saliency detection is an active research topic in the community of computer vision and multimedia. Fusing complementary RGB and thermal infrared data has been proven to be effective for image saliency detection. In this paper, we propose an effective approach for RGB-T image saliency detection. Our approach relies on a novel collaborative graph learning algorithm. In particular, we take superpixels as graph nodes, and collaboratively use hierarchical deep features to jointly learn graph affinity and node saliency in a unified optimization framework. Moreover, we contribute a more challenging dataset for the purpose of RGB-T image saliency detection, which contains 1000 spatially aligned RGB-T image pairs and their ground truth annotations. Extensive experiments on the public dataset and the newly created dataset suggest that the proposed approach performs favorably against the state-of-the-art RGB-T saliency detection methods.
Zhengzheng Tu, Chenglong Li 0002, Xiaoxiao Wang 0003, Jin Tang 0001
IEEE Trans. Multim.6
2020 An Improved Reversible Data Hiding in Encrypted Images Using Parametric Binary Tree Labeling
abstract
This work proposes an improved reversible data hiding scheme in encrypted images using parametric binary tree labeling(IPBTL-RDHEI), which takes advantage of the spatial correlation in the entire original image but not in small image blocks to reserve room for hiding data. Then the original image is encrypted with an encryption key and the parametric binary tree is used to label encrypted pixels into two different categories. Finally, one of the two categories of encrypted pixels can embed secret information by bit replacement. According to the experimental results, compared with several state-of-the-art methods, the proposed IPBTL-RDHEI method achieves higher embedding rate and outperforms the competitors. Due to the reversibility of IPBTL-RDHEI, the original plaintext image and the secret information can be restored and extracted losslessly and separately.
Youqing Wu, Youzhi Xiang, Yutang Guo, Jin Tang 0001, Zhao-Xia Yin
IEEE Trans. Multim.4
2019 Data Representation and Learning With Graph Diffusion-Embedding Networks
abstract
Recently, graph convolutional neural networks have been widely studied for graph-structured data representation and learning. In this paper, we present Graph Diffusion-Embedding networks (GDENs), a new model for graph-structured data representation and learning. GDENs are motivated by our development of graph based feature diffusion. GDENs integrate both feature diffusion and graph node (low-dimensional) embedding simultaneously into a unified network by employing a novel diffusion-embedding architecture. GDENs have two main advantages. First, the equilibrium representation of the diffusion-embedding operation in GDENs can be obtained via a simple closed-form solution, which thus guarantees the compactivity and efficiency of GDENs. Second, the proposed GDENs can be naturally extended to address the data with multiple graph structures. Experiments on various semi-supervised learning tasks on several benchmark datasets demonstrate that the proposed GDENs significantly outperform traditional graph convolutional networks.
Bo Jiang 0002, Doudou Lin, Jin Tang 0001, Bin Luo 0001
CVPR3
2019 Semi-Supervised Learning With Graph Learning-Convolutional Networks
abstract
Graph Convolutional Neural Networks (graph CNNs) have been widely used for graph data representation and semi-supervised learning tasks. However, existing graph CNNs generally use a fixed graph which may not be optimal for semi-supervised learning tasks. In this paper, we propose a novel Graph Learning-Convolutional Network (GLCN) for graph data representation and semi-supervised learning. The aim of GLCN is to learn an optimal graph structure that best serves graph CNNs for semi-supervised learning by integrating both graph learning and graph convolution in a unified network architecture. The main advantage is that in GLCN both given labels and the estimated labels are incorporated and thus can provide useful `weakly' supervised information to refine (or learn) the graph construction and also to facilitate the graph convolution operation for unknown label estimation. Experimental results on seven benchmarks demonstrate that GLCN significantly outperforms the state-of-the-art traditional fixed structure based graph CNNs.
Bo Jiang 0002, Doudou Lin, Jin Tang 0001, Bin Luo 0001
CVPR4
2019 Learning to Detect License Plates Using Synthesized Data
Yanhui Pang, Wenzhong Wang, Aihua Zheng, Jin Tang 0001
ICIG (2)4
2019 Visual Tracking Via Siamese Network With Global Similarity
abstract
Visual tracking is a very important and challenging problem in the field of computer vision. In recent years, Siamese networks have been widely used for visual tracking due to their fast tracking speed, but many trackers based on Siamese network train their networks by utilizing either pairwise loss or triplet loss, which easily leads to over-fitting. In addition, it is difficult to distinguish some hard samples in the training samples. In this paper, we propose a novel global similarity loss to train the network. Specifically, we utilize two Gaussian distributions to simulate and optimize the distribution of positive and negative samples in the train set and add constraint on the hard samples. In experiments, without any other modification, we apply the proposed method to the Siamese network. And the results on several popular tracking benchmarks show our method achieves superior tracking performance than the baseline.
Chao Fan 0001, Chenglong Li 0002, Jin Tang 0001
ICIP4
2019 Learning Target-Oriented Dual Attention for Robust RGB-T Tracking
abstract
RGB-Thermal object tracking attempts to locate target object using complementary visual and thermal infrared data. Existing RGB-T trackers fuse different modalities by robust feature representation learning or adaptive modal weighting. However, how to integrate dual attention mechanism for visual tracking is still a subject that has not been studied yet. In this paper, we propose two visual attention mechanisms for robust RGB-T object tracking. Specifically, the local attention is implemented by exploiting the common visual attention of RGB and thermal data to train deep classifiers. We also introduce the global attention, which is a multimodal target-driven attention estimation network. It can provide global proposals for the classifier together with local proposals extracted from previous tracking result. Extensive experiments on two RGB-T benchmark datasets validated the effectiveness of our proposed algorithm.
Yabin Zhu, Xiao Wang 0014, Chenglong Li 0002, Jin Tang 0001
ICIP5
2019 Multiple Graph Convolutional Networks for Co-Saliency Detection
abstract
Recently, Graph Convolutional Networks (GCNs) have been usually utilized for graph data representation in computer vision area. However, existing graph GCNs generally use a single graph which can be not adapted for the data with multiple graphs. In this paper, we first propose a novel multiple graph convolutional network (MGCN) for multiple graph data representation and learning. MGCN propagates information/knowledge across multiple graphs and obtains a consistent representation and learning by integrating the information of multiple graphs simultaneously. Based on the proposed MGCN, we then propose a new global-local unified graph convolutional learning architecture for image co-saliency detection problem. The main benefits of the proposed co-saliency model are twofold. First, it learns an optimal superpixel feature representation for co-saliency detection problem. Second, it can well exploit both intra-image and inter-image cues for co-saliency detection via a unified network. Promising experiments demonstrate the effectiveness of the proposed MGCN based co-saliency detection approach.
Bo Jiang 0002, Xingyue Jiang, Jin Tang 0001, Bin Luo 0001, Shilei Huang
ICME3
2019 A Unified Multiple Graph Learning and Convolutional Network Model for Co-saliency Estimation
abstract
Co-saliency estimation which aims to identify the common salient object regions contained in an image set is an active problem in computer vision. The main challenge for co-saliency estimation problem is how to exploit the salient cues of both intra-image and inter-image simultaneously. In this paper, we first represent intra-image and inter-image as intra-graph and inter-graph respectively and formulate co-saliency estimation as graph nodes labeling. Then, we propose a novel multiple graph learning and convolutional network (M-GLCN) for image co-saliency estimation. M-GLCN conducts graph convolutional learning and labeling on both inter-graph and intra-graph cooperatively and thus can well exploit the salient cues of both intra-image and inter-image simultaneously for co-saliency estimation. Moreover, M-GLCN employs a new graph learning mechanism to learn both inter-graph and intra-graph adaptively. Experimental results on several benchmark datasets demonstrate the effectiveness of M-GLCN on co-saliency estimation task.
Bo Jiang 0002, Xingyue Jiang, Ajian Zhou, Jin Tang 0001, Bin Luo 0001
ACM Multimedia4
2019 Dense Feature Aggregation and Pruning for RGBT Tracking
abstract
How to perform effective information fusion of different modalities is a core factor in boosting the performance of RGBT tracking. This paper presents a novel deep fusion algorithm based on the representations from an end-to-end trained convolutional neural network. To deploy the complementarity of features of all layers, we propose a recursive strategy to densely aggregate these features that yield robust representations of target objects in each modality. In different modalities, we propose to prune the densely aggregated features of all modalities in a collaborative way. In a specific, we employ the operations of global average pooling and weighted random selection to perform channel scoring and selection, which could remove redundant and noisy features to achieve more robust feature representation. Experimental results on two RGBT tracking benchmark datasets suggest that our tracker achieves clear state-of-the-art against other RGB and RGBT tracking methods.
Yabin Zhu, Chenglong Li 0002, Bin Luo 0001, Jin Tang 0001, Xiao Wang 0014
ACM Multimedia4
2019 A Novel Method for Thermal Image Based Electrical-Equipment Detection
Futian Wang, Songjian Hua, Xiao Wang 0014, Zhengzheng Tu, Cheng Zhang 0010, Jin Tang 0001
PRCV (1)6
2019 Multi-scale Convolutional Capsule Network for Hyperspectral Image Classification
Dongyue Wang, Jin Tang 0001, Bin Luo 0001
PRCV (2)4
2019 Multi-scale Densely 3D CNN for Hyperspectral Image Classification
Dongyue Wang, Jin Tang 0001, Bin Luo 0001
PRCV (2)4
2019 Efficient Feature Matching via Nonnegative Orthogonal Relaxation
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
Int. J. Comput. Vis.2
2019 Robust visual tracking via Laplacian Regularized Random Walk Ranking
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Chenglong Li 0002
Neurocomputing3
2019 Robust pixelwise saliency detection via progressive graph rankings
Bo Jiang 0002, Zhengzheng Tu, Amir Hussain 0001, Jin Tang 0001
Neurocomputing5
2019 Saliency detection via multi-view graph based saliency optimization
abstract
Saliency detection is an important problem in computer vision and pattern recognition area. Many works have been proposed for addressing the saliency detection task. As a popular method, graph based saliency optimization has been widely studied. However, previous works have universally focussed on single graph optimization which fails to consider multi-view feature representation of image content . In this paper, we first provide a general framework for traditional graph based saliency optimization models. Then, we extend the general framework to the multi-view case and propose our general multi-view graph based saliency optimization model. Finally, we present a particular implementation of our general model and derive an effective updating algorithm to solve it. Experimental results using several benchmark datasets demonstrate the effectiveness of our proposed saliency model.
Yun Xiao 0003, Bo Jiang 0002, Aihua Zheng, Aiwu Zhou, Amir Hussain 0001, Jin Tang 0001
Neurocomputing6
2019 Background subtraction with multi-scale structured low-rank and sparse factorization
Aihua Zheng, Tian Zou, Yumiao Zhao, Bo Jiang 0002, Jin Tang 0001, Chenglong Li 0002
Neurocomputing5
2019 FMT: fusing multi-task convolutional neural network for person search
Sulan Zhai, Shunqiang Liu, Xiao Wang 0014, Jin Tang 0001
Multim. Tools Appl.4
2019 Visual Tracking via Dynamic Graph Learning
abstract
Existing visual tracking methods usually localize a target object with a bounding box, in which the performance of the foreground object trackers or detectors is often affected by the inclusion of background clutter. To handle this problem, we learn a patch-based graph representation for visual tracking. The tracked object is modeled by with a graph by taking a set of non-overlapping image patches as nodes, in which the weight of each node indicates how likely it belongs to the foreground and edges are weighted for indicating the appearance compatibility of two neighboring nodes. This graph is dynamically learned and applied in object tracking and model updating. During the tracking process, the proposed algorithm performs three main steps in each frame. First, the graph is initialized by assigning binary weights of some image patches to indicate the object and background patches according to the predicted bounding box. Second, the graph is optimized to refine the patch weights by using a novel alternating direction method of multipliers. Third, the object feature representation is updated by imposing the weights of patches on the extracted image features. The object location is predicted by maximizing the classification score in the structured support vector machine. Extensive experiments show that the proposed tracking algorithm performs well against the state-of-the-art methods on large-scale benchmark datasets.
Chenglong Li 0002, Liang Lin 0004, Wangmeng Zuo, Jin Tang 0001, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 RGB-T object tracking: Benchmark and baseline
Chenglong Li 0002, Xinyan Liang, Yijuan Lu, Jin Tang 0001
Pattern Recognit.5
2019 Quality-aware dual-modal saliency detection via deep reinforcement learning
Xiao Wang 0014, Chenglong Li 0002, Bin Luo 0001, Jin Tang 0001
Signal Process. Image Commun.6
2019 Learning Local-Global Multi-Graph Descriptors for RGB-T Object Tracking
abstract
RGB-thermal (RGB-T) object tracking, which has attracted much recent attention, uses thermal infrared information to assist object tracking with visible light information. However, it still faces many challenging problems, especially the background inclusion in the target bounding box which easily results in model drifting. To handle this problem, we propose a novel and general approach to learn a local-global multi-graph descriptor to suppress background effects for RGB-T tracking. Our approach relies on a novel graph learning algorithm. First, the object is represented with multiple graphs, with a set of multi-modal image patches as nodes, for the robustness to prevent deformation and partial occlusion. Second, we dynamically learn a joint graph over time with both local and global considerations using spatial smoothness and low-rank representation. In particular, we design a single unified alternating direction method of multipliers-based optimization framework to learn graph structure, edge weights, and node weights simultaneously. Third, we combine multi-graph information with corresponding graph node weights to form a robust object descriptor, and tracking is finally carried out by adopting the structured support vector machine. Extensive experiments conducted on the tracking benchmark data sets demonstrate the effectiveness of the proposed approach against the state-of-the-art RGB-T trackers.
Chenglong Li 0002, Chengli Zhu, Justin Jian Zhang, Bin Luo 0001, Xiaohao Wu, Jin Tang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2019 Image Representation and Learning With Graph-Laplacian Tucker Tensor Decomposition
abstract
Tucker tensor decomposition (TD) is widely used for image representation, reconstruction, and learning tasks. Compared to principal component analysis (PCA) models, tensor models retain more 2-D characteristics of images whereas PCA models linearize images. However, traditional TD involves attribute information only and thus does not consider the pairwise similarity information between images. In this paper, we propose a graph-Laplacian tucker tensor decomposition (GLTD) which explores both attributes and pairwise similarity information simultaneously. Generally, GLTD has three main benefits: 1) GLTD reconstruction shows clear robustness against image occlusions/outliers. We provide analysis to show that Laplacian regularization is mainly responsible to this robustness via an out-of-sample GLTD model. To the best of our knowledge, this Laplacian regularization induced robustness of TD has not been studied or emphasized before; 2) GLTD representation performs more regularity, which improves both unsupervised and supervised learning results; and 3) an effective algorithm is derived to solve GLTD problem. Although GLTD is a noncovex problem, the proposed algorithm is shown experimentally to provide a stable/unique solution starting from different random initializations. Experimental results on image reconstruction, data clustering, and classification tasks show the benefits of GLTD.
Bo Jiang 0002, Chris Ding, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Cybern.3
2018 SINT++: Robust Visual Tracking via Adversarial Positive Instance Generation
abstract
Existing visual trackers are easily disturbed by occlusion, blur and large deformation. We think the performance of existing visual trackers may be limited due to the following issues: i) Adopting the dense sampling strategy to generate positive examples will make them less diverse; ii) The training data with different challenging factors are limited, even through collecting large training dataset. Collecting even larger training dataset is the most intuitive paradigm, but it may still can not cover all situations and the positive samples are still monotonous. In this paper, we propose to generate hard positive samples via adversarial learning for visual tracking. Specifically speaking, we assume the target objects all lie on a manifold, hence, we introduce the positive samples generation network (PSGN) to sampling massive diverse training data through traversing over the constructed target object manifold. The generated diverse target object images can enrich the training dataset and enhance the robustness of visual trackers. To make the tracker more robust to occlusion, we adopt the hard positive transformation network (HPTN) which can generate hard samples for tracking algorithm to recognize. We train this network with deep reinforcement learning to automatically occlude the target object with a negative patch. Based on the generated hard positive samples, we train a Siamese network for visual tracking and our experiments validate the effectiveness of the introduced algorithm. The project page of this paper can be found from the website1.
Xiao Wang 0014, Chenglong Li 0002, Bin Luo 0001, Jin Tang 0001
CVPR4
2018 Cross-Modal Ranking with Soft Consistency and Noisy Labels for Robust RGB-T Tracking
Chenglong Li 0002, Chengli Zhu, Yan Huang 0008, Jin Tang 0001, Liang Wang 0001
ECCV (13)4
2018 Exploring Scene Geometry for Scale Adaptive Object Tracking in Surveillance Videos
abstract
Object tracking is a key technology in video surveillance. Reliable tracker must be adaptive to the constantly changing object sizes. Most of the state-of-the-art methods estimate the object scales using their appearances. Those methods are vulnerable to occlusion, object deformation, illumination change and background clutter. In this paper, we propose to use the geometric context of the surveillance site as a strong clue for scale adaptation. With three reasonable assumptions on the video cameras and the surveillance sites, we deduce a simple geometric model for object scales. The parameters of this model are learned without any human intervention. Then we integrate this model into baseline trackers for robust scale adaptive object tracking. Experimental results on challenging surveillance videos indicate that our approach favorably improves the performance of single-scale baselines, and performs better or comparative to the state-of-the-art multi-scale trackers while significantly improve the speed.
Ran Zhong, Wenzhong Wang, Chenglong Li 0002, Aihua Zheng, Jin Tang 0001
ICIP5
2018 OWP: Objectness Weighted Patch Descriptor for Visual Tracking
abstract
Visual object tracking is an active research problem and has been widely used in computer vision and pattern recognition area. Existing visual tracking methods usually localize the visual object with a bounding box which are often disturbed by the introduced background information and partial occlusion because of bounding box representation of visual object. To deal with this problem, in this paper, we propose a novel Objectness Weighted Patch (OWP) descriptor for object feature descriptor in visual tracking. The aim of OWP is to assign different objectness weights to the patches of bounding box to reduce the influences of background information and partial occlusion. We propose to compute the objectness weights of patches in OWP by integrating multiple cues (background, foreground and local spatial consistency) together in a general optimization model. Also, the proposed model has a simple closed-form solution and thus can be computed efficiently. We incorporate our OWP into structured SVM tracking framework and provide a new robust tracking method. Extensive experiments on two standard benchmark datasets OTB100 and Temple-Color demonstrate the effectiveness and benefits of the proposed tracking method.
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
ICPR3
2018 Multi-scale Attributed Graph Kernel for Image Categorization
Duo Hu, Jin Tang 0001, Bin Luo 0001
PRCV (3)3
2018 Multi-scale Cooperative Ranking for Saliency Detection
Bo Jiang 0002, Xingyue Jiang, Aihua Zheng, Yun Xiao 0003, Jin Tang 0001
PRCV (1)5
2018 Learning Soft-Consistent Correlation Filters for RGB-T Object Tracking
Chenglong Li 0002, Jin Tang 0001
PRCV (4)3
2018 Non-negative Dual Graph Regularized Sparse Ranking for Multi-shot Person Re-identification
Aihua Zheng, Bo Jiang 0002, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
PRCV (1)5
2018 Fusing two-stream convolutional neural networks for RGB-T object tracking
Chenglong Li 0002, Xiaohao Wu, Xiaochun Cao, Jin Tang 0001
Neurocomputing5
2018 A prior regularized multi-layer graph ranking model for image saliency computation
Yun Xiao 0003, Bo Jiang 0002, Zhengzheng Tu, Jixin Ma 0001, Jin Tang 0001
Neurocomputing5
2018 Spatial-temporal representatives selection and weighted patch descriptor for person re-identification
Aihua Zheng, Foqin Wang, Amir Hussain 0001, Jin Tang 0001, Bo Jiang 0002
Neurocomputing4
2018 Moving object detection via robust background modeling with recurring patterns voting
Chenglong Li 0002, Zhimin Bao, Xiao Wang 0014, Jin Tang 0001
Multim. Tools Appl.4
2018 Reversible data hiding in encrypted AMBTC images
Zhao-Xia Yin, Xuejing Niu, Xinpeng Zhang 0001, Jin Tang 0001, Bin Luo 0001
Multim. Tools Appl.4
2018 Fast Grayscale-Thermal Foreground Detection With Collaborative Low-Rank Decomposition
abstract
This paper investigates how to perform efficient and robust foreground detection in challenging scenarios by leveraging multiple source data. We propose a novel approach, called collaborative low-rank decomposition (CLoD), for grayscale-thermal foreground detection. Given two data matrices by accumulating sequential frames from the grayscale and the thermal videos, CLoD detects the foreground objects as sparse noises against the backgrounds with collaborative low rank structure, and also incorporates modality weights to achieve adaptive fusion of different source data. For the optimization, CLoD seeks a sub-optimal solution by making the background matrix rank explicitly determined. In particular, the background matrix with the fixed rank can be decomposed into two sub-matrices of low rank, and then, we iteratively optimize them and the modality weights with closed-form solutions. For improving the efficiency, we design a block-based accelerated algorithm to speed up CLoD while employing the edge-preserving algorithm to keep the accuracy. Extensive experiments on the recently public benchmark grayscale-thermal foreground detection suggest that our approach achieves comparable performance in terms of both accuracy and efficiency against other state-of-the-art methods.
Bin Luo 0001, Chenglong Li 0002, Guizhao Wang, Jin Tang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2017 Nonnegative Orthogonal Graph Matching
abstract
Graph matching problem that incorporates pair-wise constraints can be formulated as Quadratic Assignment Problem(QAP). The optimal solution of QAP is discrete and combinational, which makes QAP problem NP-hard. Thus, many algorithms have been proposed to find approximate solutions. In this paper, we propose a new algorithm, called Nonnegative Orthogonal Graph Matching (NOGM), for QAP matching problem. NOGM is motivated by our new observation that the discrete mapping constraint of QAP can be equivalently encoded by a nonnegative orthogonal constraint which is much easier to implement computationally. Based on this observation, we develop an effective multiplicative update algorithm to solve NOGM and thus can find an effective approximate solution for QAP problem. Comparing with many traditional continuous methods which usually obtain continuous solutions and should be further discretized, NOGM can obtain a sparse solution and thus incorporates the desirable discrete constraint naturally in its optimization. Promising experimental results demonstrate benefits of NOGM algorithm.
Bo Jiang 0002, Jin Tang 0001, Chris Ding, Bin Luo 0001
AAAI2
2017 Learning Patch-Based Dynamic Graph for Visual Tracking
abstract
Existing visual tracking methods usually localize the object with a bounding box, in which the foreground object trackers/detectors are often disturbed by the introduced background information. To handle this problem, we aim to learn a more robust object representation for visual tracking. In particular, the tracked object is represented with a graph structure (i.e., a set of non-overlapping image patches), in which the weight of each node (patch) indicates how likely it belongs to the foreground and edges are also weighed for indicating the appearance compatibility of two neighboring nodes. This graph is dynamically learnt (i.e., the nodes and edges received weights) and applied in object tracking and model updating. We constrain the graph learning from two aspects: i) the global low-rank structure over all nodes and ii) the local sparseness of node neighbors. During the tracking process, our method performs the following steps at each frame. First, the graph is initialized by assigning either 1 or 0 to the weights of some image patches according to the predicted bounding box. Second, the graph is optimized through designing a new ALM (Augmented Lagrange Multiplier) based algorithm. Third, the object feature representation is updated by imposing the weights of patches on the extracted image features. The object location is finally predicted by adopting the Struck tracker. Extensive experiments show that our approach outperforms the state-of-the-art tracking methods on two standard benchmarks, i.e., OTB100 and NUS-PRO.
Chenglong Li 0002, Liang Lin 0004, Wangmeng Zuo, Jin Tang 0001
AAAI4
2017 Binary Constraint Preserving Graph Matching
abstract
Graph matching is a fundamental problem in computer vision and pattern recognition area. In general, it can be formulated as an Integer Quadratic Programming (IQP) problem. Since it is NP-hard, approximate relaxations are required. In this paper, a new graph matching method has been proposed. There are three main contributions of the proposed method: (1) we propose a new graph matching relaxation model, called Binary Constraint Preserving Graph Matching (BPGM), which aims to incorporate the discrete binary mapping constraints more in graph matching relaxation. Our BPGM is motivated by a new observation that the discrete binary constraints in IQP matching problem can be represented (or encoded) exactly by a ℓ2-norm constraint. (2) An effective projection algorithm has been derived to solve BPGM model. (3) Using BPGM, we propose a path-following strategy to optimize IQP matching problem and thus obtain a desired discrete solution at convergence. Promising experimental results show the effectiveness of the proposed method.
Bo Jiang 0002, Jin Tang 0001, Chris Ding, Bin Luo 0001
CVPR2
2017 Image Set Representation with L_1 -Norm Optimal Mean Robust Principal Component Analysis
Youxia Cao, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
ICIG (2)3
2017 Selecting attentive frames from visually coherent video chunks for surveillance video summarization
abstract
This paper investigates how to extract key-frames from surveillance video while maximizing their diversity and representational ability. We solve this problem by two steps, i.e., video partition and frame selection. The first step is to partition a surveillance video into visually coherent video chunks, which have high intra-chunk similarity and interchunk dissimilarity. In particular, we propose an object-based frame metric to measure the relevance of two frames, and apply the Normalized Cut algorithm to achieve video partition. The second step is to select the attentive frames from the partitioned video chunks. We propose an attention score based on the content completeness and the visual satisfaction for each frame, and select most attentive frame with highest attention score in each chunk. Extensive experiments on both public and our newly created datasets suggest that our approach significantly outperforms other video summarization methods.
Wenzhong Wang, Qiaoqiao Zhang, Bin Luo 0001, Jin Tang 0001, Rui Ruan, Chenglong Li 0002
ICIP4
2017 ReGLe: Spatially Regularized Graph Learning for Visual Tracking
abstract
Weighted patch representation of the target object has been proven to be effective for suppressing the background effects in visual tracking. In this paper, we propose a novel approach, called spatially Regularized Graph Learning (ReGLe), to automatically explore the intrinsic relationship among patches both with global and local cues for robust object representation. In particular, the target object bounding box is partitioned into a set of non-overlapping image patches, which are taken as graph nodes, and each of them is associated with a weight to represent how likely it belongs to the target object. To improve the accuracy of node weight computation, we dynamically learn the edge weights (i.e., the appearance compatibility of two nodes) according to both global and local relationship among patches. First, we pursue the low-rank representation for capturing the global low-dimensional subspace structure of patches. Second, we encode the local information into the low-rank representation by exploiting the fact that neighboring nodes usually have similar appearance. Finally, we utilize the representations to learn their affinities (i.e., graph edge weights). The node and edge weights are jointly optimized by a designed ADMM (Alternating Direction Method of Multipliers) algorithm, the object feature representation is updated by imposing the weights of patches on the extracted image features. The object location is finally predicted by maximizing the classification score in the structured SVM. Extensive experiments demonstrate the effectiveness of the proposed approach on the tracking benchmark datasets: OTB100 and Temple-Color.
Chenglong Li 0002, Xiaohao Wu, Zhimin Bao, Jin Tang 0001
ACM Multimedia4
2017 Weighted Sparse Representation Regularized Graph Learning for RGB-T Object Tracking
abstract
In this paper, we propose a novel graph model, called weighted sparse representation regularized graph, to learn a robust object representation using multispectral (RGB and thermal) data for visual tracking. In particular, the tracked object is represented with a graph with image patches as nodes. This graph is dynamically learned from two aspects. First, the graph affinity (i.e., graph structure and edge weights) that indicates the appearance compatibility of two neighboring nodes is optimized based on the weighted sparse representation, in which the modality weight is introduced to leverage RGB and thermal information adaptively. Second, each node weight that indicates how likely it belongs to the foreground is propagated from others along with graph affinity. The optimized patch weights are then imposed on the extracted RGB and thermal features, and the target object is finally located by adopting the structured SVM algorithm. Moreover, we also contribute a comprehensive dataset for RGB-T tracking purpose. Comparing with existing ones, the new dataset has the following advantages: 1) Its size is sufficiently large for large-scale performance evaluation (total frame number: 210K, maximum frames per video pair: 8K). 2) The alignment between RGB-T video pairs is highly accurate, which does not need pre- and post-processing. 3) The occlusion levels are annotated for analyzing the occlusion-sensitive performance of different methods. Extensive experiments on both public and newly created datasets demonstrate the effectiveness of the proposed tracker against several state-of-the-art tracking methods.
Chenglong Li 0002, Yijuan Lu, Chengli Zhu, Jin Tang 0001
ACM Multimedia5
2017 Graph Matching via Multiplicative Update Algorithm
abstract
As a fundamental problem in computer vision, graph matching problem can usually be formulated as a Quadratic Programming (QP) problem with doubly stochastic and discrete (integer) constraints. Since it is NP-hard, approximate algorithms are required. In this paper, we present a new algorithm, called Multiplicative Update Graph Matching (MPGM), that develops a multiplicative update technique to solve the QP matching problem. MPGM has three main benefits: (1) theoretically, MPGM solves the general QP problem with doubly stochastic constraint naturally whose convergence and KKT optimality are guaranteed. (2) Em- pirically, MPGM generally returns a sparse solution and thus can also incorporate the discrete constraint approximately. (3) It is efficient and simple to implement. Experimental results show the benefits of MPGM algorithm.
Bo Jiang 0002, Jin Tang 0001, Chris Ding, Yihong Gong, Bin Luo 0001
NIPS2
2017 A new graph ranking model for image saliency detection problem
abstract
Saliency detection is an important problem in many computer vision applications. As a kind of popular method, graph based manifold ranking (GMR) has been successfully used in saliency detection problem. In traditional GMR saliency detection, it involves two main stages, i.e., ranking with background queries and ranking with foreground queries. However, in GMR method, these two stages are conducted separately, which ignores the correlation between background and foreground cues. In this paper, we propose a new graph ranking model, which aims to perform background and foreground ranking simultaneously by exploiting the correlation between background and foreground cues. We derive a closed-form solution for it. Experimental results on four benchmark datasets demonstrate that the proposed method performs better than some other state-of-art methods.
Yuanyuan Guan, Bo Jiang 0002, Yun Xiao 0003, Jin Tang 0001, Bin Luo 0001
SERA4
2017 A server selection strategy about cloud workflow based on QoS constraint
abstract
Cloud computing is an emerging business computing model whose basic idea is to transmit all kinds of resources through the Internet such as storage resources, computing resources, bandwidth and so forth. So users do not need to purchase a large of computing systems to manage their business, on the contrary they only need to purchase the resources according to their needs in order to reduce the cost greatly. Due to the flexibility, convenience and low maintenance cost of cloud computing, more and more service providers choose to deploy their services to the cloud. However, under the influence of uncertain factors such as manufacturing technology, the price and delivery time of virtual machine also been changed. Because of these variations, it is difficult for users to select the appropriate cloud servers at a minimum cost, which will lead to a decline in user experience. To solve the problems above, this paper analyzes the performance of different Ali Cloud servers and proposes a server selection strategy according to the size of the whole instance and the required time of completing the workflow, so as to make the whole workflow instances execution costs as little as possible in the regulations. We also provide experimental results to demonstrate the effectiveness of our selection strategies.
Futian Wang, Jin Tang 0001, Cheng Zhang 0010
SERA3
2017 Manifold ranking weighted local maximal occurrence descriptor for person re-identification
abstract
Person re-identification is an important task of matching pedestrians across non-overlapping camera views. In this paper, we exploit a weighted feature descriptor for person re-identification. We firstly compute the weights on the superpixel level via graph-based manifold ranking algorithm, then integrate the computed weights into a patch-based feature descriptor, named local maximal occurrence. Finally, the weighted descriptors are fed into a top-push distance learning to mitigate the cross-view gaps. We evaluate the proposed method on three benchmark datasets iLIDS-VID, PRID 450S and VIPeR. The promising experimental results demonstrate the effectiveness of the proposed method comparing with the state-of-the-arts.
Foqin Wang, Xuehan Zhang, Jinxin Ma, Jin Tang 0001, Aihua Zheng
SERA4
2017 Image Set Representation and Classification with Attributed Covariate-Relation Graph Model and Graph Sparse Representation Classification
Zhuqiang Chen, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
Neurocomputing3
2017 A global and local consistent ranking model for image saliency computation
Yun Xiao 0003, Bo Jiang 0002, Zhengzheng Tu, Jin Tang 0001
J. Vis. Commun. Image Represent.5
2017 Reversible data hiding in encrypted images based on multi-level encryption and block histogram modification
Zhao-Xia Yin, Andrew Abel, Jin Tang 0001, Xinpeng Zhang 0001, Bin Luo 0001
Multim. Tools Appl.3
2017 Local-to-global background modeling for moving object detection from non-static cameras
Aihua Zheng, Lei Zhang 0074, Wei Zhang 0012, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
Multim. Tools Appl.5
2017 A method for defensing against multi-source Sybil attacks in VANET
abstract
Sybil attack can counterfeit traffic scenario by sending false messages with multiple identities, which often causes traffic jams and even leads to vehicular accidents in vehicular ad hoc network (VANET). It is very difficult to be defended and detected, especially when it is launched by some conspired attackers using their legitimate identities. In this paper, we propose an event based reputation system (EBRS), in which dynamic reputation and trusted value for each event are employed to suppress the spread of false messages. EBRS can detect Sybil attack with fabricated identities and stolen identities in the process of communication, it also defends against the conspired Sybil attack since each event has a unique reputation value and trusted value. Meanwhile, we keep the vehicle identity in privacy. Simulation results show that EBRS is able to defend and detect multi-source Sybil attacks with high performances.
Xia Feng, Chun-yan Li, De-xin Chen, Jin Tang 0001
Peer-to-Peer Netw. Appl.4
2017 Lagrangian relaxation graph matching
Bo Jiang 0002, Jin Tang 0001, Xiaochun Cao, Bin Luo 0001
Pattern Recognit.2
2017 Image representation and matching with geometric-edge random structure graph
Bo Jiang 0002, Jin Tang 0001, Aihua Zheng, Bin Luo 0001
Pattern Recognit. Lett.2
2017 Weighted Low-Rank Decomposition for Robust Grayscale-Thermal Foreground Detection
abstract
This paper investigates how to fuse grayscale and thermal video data for detecting foreground objects in challenging scenarios. To this end, we propose an intuitive yet effective method called weighted low-rank decomposition (WELD), which adaptively pursues the cross-modality low-rank representation. Specifically, we form two data matrices by accumulating sequential frames from the grayscale and the thermal videos, respectively. Within these two observing matrices, WELD detects moving foreground pixels as sparse outliers against the low-rank structure background and incorporates the weight variables to make the models of two modalities complementary to each other. The smoothness constraints of object motion are also introduced in WELD to further improve the robustness to noises. For optimization, we propose an iterative algorithm to efficiently solve the low-rank models with three subproblems. Moreover, we utilize an edge-preserving filtering-based method to substantially speed up WELD while preserving its accuracy. To provide a comprehensive evaluation benchmark of grayscale-thermal foreground detection, we create a new data set including 25 aligned grayscale-thermal video pairs with high diversity. Our extensive experiments on both the newly created data set and the public data set OSU3 suggest that WELD achieves superior performance and comparable efficiency against other state-of-the-art approaches.
Chenglong Li 0002, Xiao Wang 0014, Lei Zhang 0074, Jin Tang 0001, Hejun Wu, Liang Lin 0004
IEEE Trans. Circuits Syst. Video Technol.4
2017 Grayscale-Thermal Object Tracking via Multitask Laplacian Sparse Representation
abstract
This paper studies the problem of object tracking in challenging scenarios by leveraging multimodal visual data. We propose a grayscale-thermal object tracking method in Bayesian filtering framework based on multitask Laplacian sparse representation. Given one bounding box, we extract a set of overlapping local patches within it, and pursue the multitask joint sparse representation for grayscale and thermal modalities. Then, the representation coefficients of the two modalities are concatenated into a vector to represent the feature of the bounding box. Moreover, the similarity between each patch pair is deployed to refine their representation coefficients in the sparse representation, which can be formulated as the Laplacian sparse representation. We also incorporate the modal reliability into the Laplacian sparse representation to achieve an adaptive fusion of different source data. Experiments on two grayscale-thermal datasets suggest that the proposed approach outperforms both grayscale and grayscale-thermal tracking approaches.
Chenglong Li 0002, Xiao Wang 0014, Lei Zhang 0074, Jin Tang 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2016 Reversible Data Hiding in Encrypted AMBTC Compressed Images
Xuejing Niu, Zhao-Xia Yin, Xinpeng Zhang 0001, Jin Tang 0001, Bin Luo 0001
IWDW4
2016 Real-Time Grayscale-Thermal Tracking via Laplacian Sparse Representation
Chenglong Li 0002, Shiyi Hu, Sihan Gao, Jin Tang 0001
MMM (2)4
2016 Inference With Collaborative Model for Interactive Tumor Segmentation in Medical Image Sequences
abstract
Segmenting organisms or tumors from medical data (e.g., computed tomography volumetric images, ultrasound, or magnetic resonance imaging images/image sequences) is one of the fundamental tasks in medical image analysis and diagnosis, and has received long-term attentions. This paper studies a novel computational framework of interactive segmentation for extracting liver tumors from image sequences, and it is suitable for different types of medical data. The main contributions are twofold. First, we propose a collaborative model to jointly formulate the tumor segmentation from two aspects: 1) region partition and 2) boundary presence. The two terms are complementary but simultaneously competing: the former extracts the tumor based on its appearance/texture information, while the latter searches for the palpable tumor boundary. Moreover, in order to adapt the data variations, we allow the model to be discriminatively trained based on both the seed pixels traced by the Lucas-Kanade algorithm and the scribbles placed by the user. Second, we present an effective inference algorithm that iterates to: 1) solve tumor segmentation using the augmented Lagrangian method and 2) propagate the segmentation across the image sequence by searching for distinctive matches between images. We keep the collaborative model updated during the inference in order to well capture the tumor variations over time. We have verified our system for segmenting liver tumors from a number of clinical data, and have achieved very promising results. The software developed with this paper can be found at http://vision.sysu.edu.cn/projects/med-interactive-seg/.
Liang Lin 0004, Wei Yang 0019, Chenglong Li 0002, Jin Tang 0001, Xiaochun Cao
IEEE Trans. Cybern.4
2016 Learning Collaborative Sparse Representation for Grayscale-Thermal Tracking
abstract
Integrating multiple different yet complementary feature representations has been proved to be an effective way for boosting tracking performance. This paper investigates how to perform robust object tracking in challenging scenarios by adaptively incorporating information from grayscale and thermal videos, and proposes a novel collaborative algorithm for online tracking. In particular, an adaptive fusion scheme is proposed based on collaborative sparse representation in Bayesian filtering framework. We jointly optimize sparse codes and the reliable weights of different modalities in an online way. In addition, this paper contributes a comprehensive video benchmark, which includes 50 grayscale-thermal sequences and their ground truth annotations for tracking purpose. The videos are with high diversity and the annotations were finished by one single person to guarantee consistency. Extensive experiments against other state-of-the-art trackers with both grayscale and grayscale-thermal inputs demonstrate the effectiveness of the proposed tracking approach. Through analyzing quantitative results, we also provide basic insights and potential future research directions in grayscale-thermal tracking.
Chenglong Li 0002, Shiyi Hu, Xiaobai Liu, Jin Tang 0001, Liang Lin 0004
IEEE Trans. Image Process.5
2016 An Approach to Streaming Video Segmentation With Sub-Optimal Low-Rank Decomposition
abstract
This paper investigates how to perform robust and efficient video segmentation while suppressing the effects of data noises and/or corruptions, and an effective approach is introduced to this end. First, a general algorithm, called sub-optimal low-rank decomposition (SOLD), is proposed to pursue the low-rank representation for video segmentation. Given the data matrix formed by supervoxel features of an observed video sequence, SOLD seeks a sub-optimal solution by making the matrix rank explicitly determined. In particular, the representation coefficient matrix with the fixed rank can be decomposed into two sub-matrices of low rank, and then we iteratively optimize them with closed-form solutions. Moreover, we incorporate a discriminative replication prior into SOLD based on the observation that small-size video patterns tend to recur frequently within the same object. Second, based on SOLD, we present an efficient inference algorithm to perform streaming video segmentation in both unsupervised and interactive scenarios. More specifically, the constrained normalized-cut algorithm is adopted by incorporating the low-rank representation with other low level cues and temporal consistent constraints for spatio-temporal segmentation. Extensive experiments on two public challenging data sets VSB100 and SegTrack suggest that our approach outperforms other video segmentation approaches in both accuracy and efficiency.
Chenglong Li 0002, Liang Lin 0004, Wangmeng Zuo, Wenzhong Wang, Jin Tang 0001
IEEE Trans. Image Process.5
2015 A Local Sparse Model for Matching Problem
abstract
Feature matching problem that incorporates pairwise constraints is usually formulated as a quadratic assignment problem (QAP). Since it is NP-hard, relaxation models are required. In this paper, we first formulate the QAP from the match selection point of view; and then propose a local sparse model for matching problem. Our local sparse matching (LSM) method has the following advantages: (1) It is parameter-free; (2) It generates a local sparse solution which is closer to a discrete matrix than most other continuous relaxation methods for the matching problem. (3) The one-to-one matching constraints are better maintained in LSM solution. Promising experimental results show the effectiveness of the Proposed LSM method.
Bo Jiang 0002, Jin Tang 0001, Chris Ding, Bin Luo 0001
AAAI2
2015 SOLD: Sub-optimal low-rank decomposition for efficient video segmentation
abstract
This paper investigates how to perform robust and efficient unsupervised video segmentation while suppressing the effects of data noises and/or corruptions. We propose a general algorithm, called Sub-Optimal Low-rank Decomposition (SOLD), which pursues the low-rank representation for video segmentation. Given the supervoxels affinity matrix of an observed video sequence, SOLD seeks a sub-optimal solution by making the matrix rank explicitly determined. In particular, the affinity matrix with the rank fixed can be decomposed into two sub-matrices of low rank, and then we iteratively optimize them with closed-form solutions. Moreover, we incorporate a discriminative replication prior into our framework based on the obervation that small-size video patterns tend to recur frequently within the same object. The video can be segmented into several spatio-temporal regions by applying the Normalized-Cut (NCut) algorithm with the solved low-rank representation. To process the streaming videos, we apply our algorithm sequentially over a batch of frames over time, in which we also develop several temporal consistent constraints improving the robustness. Extensive experiments on the public benchmarks demonstrate superior performance of our framework over other state-of-the-art approaches.
Chenglong Li 0002, Liang Lin 0004, Wangmeng Zuo, Shuicheng Yan, Jin Tang 0001
CVPR5
2015 Person Re-identification with Density-Distance Unsupervised Salience Learning
Baoliang Zhou, Aihua Zheng, Bo Jiang 0002, Chenglong Li 0002, Jin Tang 0001
ICIG (3)5
2015 EBRS: Event Based Reputation System for Defensing Multi-source Sybil Attacks in VANET
Xia Feng, Chun-yan Li, De-xin Chen, Jin Tang 0001
WASA4
2015 Image matching using a local distribution based outlier detection technique
Haifeng Zhao 0001, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
Neurocomputing3
2014 A sparse nonnegative matrix factorization technique for graph matching problems
Bo Jiang 0002, Haifeng Zhao 0001, Jin Tang 0001, Bin Luo 0001
Pattern Recognit.3
2014 Robust Feature Point Matching With Sparse Model
abstract
Feature point matching that incorporates pairwise constraints can be cast as an integer quadratic programming (IQP) problem. Since it is NP-hard, approximate methods are required. The optimal solution for IQP matching problem is discrete, binary, and thus sparse in nature. This motivates us to use sparse model for feature point matching problem. The main advantage of the proposed sparse feature point matching (SPM) method is that it generates sparse solution and thus naturally imposes the discrete mapping constraints approximately in the optimization process. Therefore, it can optimize the IQP matching problem in an approximate discrete domain. In addition, an efficient algorithm can be derived to solve SPM problem. Promising experimental results on both synthetic points sets matching and real-world image feature sets matching tasks show the effectiveness of the proposed feature point matching method.
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Liang Lin 0004
IEEE Trans. Image Process.2
2013 Graph-Laplacian PCA: Closed-Form Solution and Robustness
abstract
Principal Component Analysis (PCA) is a widely used to learn a low-dimensional representation. In many applications, both vector data X and graph data W are available. Laplacian embedding is widely used for embedding graph data. We propose a graph-Laplacian PCA (gLPCA) to learn a low dimensional representation of X that incorporates graph structures encoded in W. This model has several advantages: (1) It is a data representation model. (2) It has a compact closed-form solution and can be efficiently computed. (3) It is capable to remove corruptions. Extensive experiments on 8 datasets show promising results on image reconstruction and significant improvement on clustering and classification.
Bo Jiang 0002, Chris Ding, Bin Luo 0001, Jin Tang 0001
CVPR4
2013 Blind Detection of Region Duplication Forgery by Merging Blur and Affine Moment Invariants
abstract
Region duplication is a simple and effective operation to create digital image forgeries, where a part of the image is copied and pasted on another part of the same image. Most existing region duplication detection methods are based on directly matching blocks of image pixels or transform coefficients, and are not effective when duplicated regions have affined transforms or blur degradations. In this work we propose a new region duplicated detection method to automatically detect and localize duplicated regions in digital images. The method is based on merging blur and affine moment invariants, which allows successful detection of region duplication forgery, even under some simple affine transforms and blur degradations. Our experiments in synthesized forgery images with duplicated and distorted regions show that the proposed method gets effective detection. These demonstrate that our method is an effective way to detect the duplication regions under some simple affine transforms and blur degradations blindly.
Jin Tang 0001, Bin Luo 0001
ICIG2
2012 Matching State-Based Sequences with Rich Temporal Aspects
abstract
A General Similarity Measurement (GSM), which takes into account of both non-temporal and rich temporal aspects including temporal order, temporal duration and temporal gap, is proposed for state-sequence matching. It is believed to be versatile enough to subsume representative existing measurements as its special cases.
Aihua Zheng, Jixin Ma 0001, Jin Tang 0001, Bin Luo 0001
AAAI3
2012 Graph matching based on spectral embedding with missing value
Jin Tang 0001, Bo Jiang 0002, Aihua Zheng, Bin Luo 0001
Pattern Recognit.1
2011 Angular Decomposition
abstract
Dimensionality reduction plays a vital role in pattern recognition. However, for normalized vector data, existing methods do not utilize the fact that the data is normalized. In this paper, we propose to employ an Angular Decomposition of the normalized vector data which corresponds to embedding them on a unit surface. On graph data for similarity/ kernel matrices with constant diagonal elements, we propose the Angular Decomposition of the similarity matrices which corresponds to embedding objects on a unit sphere. In these angular embeddings, the Euclidean distance is equivalent to the cosine similarity. Thus data structures best described in the cosine similarity and data structures best captured by the Euclidean distance can both be effectively detected in our angular embedding. We provide the theoretical analysis, derive the computational algorithm, and evaluate the angular embedding on several datasets. Experiments on data clustering demonstrate that our method can provide a more discriminative subspace.
Dengdi Sun, Chris Ding, Bin Luo 0001, Jin Tang 0001
IJCAI4
2009 Registration of blurred images for image mosaic
abstract
Existing methods for the registration of blurred images are efficient for the artificially blurred images or a planar registration. They are not suitable for image mosaic of the source images from a real camera with an almost fixed optical center. We propose a registration method so that a distortion-free registration on naturally captured images can be obtained. It adopts a multi-resolution and robust feature based inter-layer mosaic together. In each layer, Harris corner detector is chosen to effectively detect features and RANSAC is used to find reliable matches for further calibration as well as an initial homography as the initial motion of next layer. Simplex and subspace trust region methods are used consequently to estimate the stable focal length and rotation matrix through the transformation property of feature matches. Experimental results demonstrate the performance of our proposed method.
Xianyong Fang, Bin Luo 0001, Jin Tang 0001, Haifeng Zhao 0001
CAD/Graphics3
2007 Using Eigen-Decomposition Method for Weighted Graph Matching
Guoxing Zhao, Bin Luo 0001, Jin Tang 0001, Jixin Ma 0001
ICIC (1)3
2006 Shape Representation and Distance Measure Based on Relational Graph
Jin Tang 0001, Bin Luo 0001
HIS1
2006 Semi-Supervised Clustering of Corner-Oriented Attributed Graphs
Jin Tang 0001, Bin Luo 0001
HIS1
2006 Automatic T-Mixture Model Selection via Rival Penalized EM
Jin Tang 0001, Bin Luo 0001
HIS2