Bin Luo 0001

dblp:36/4256-1 · DBLP profile ↗
← Back
279ranked-venue papers
22as first author
149since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 138 · 12 first-author · 61 since 2021Graphics, computer vision, multimedia, augmented reality and games · 120 · 15 first-author · 53 since 2021Applied, interdisciplinary, general and emerging computing · 34 · 32 since 2021Security and privacy · 7 · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Semantic-Driven Visual Progressive Refinement for Aerial-Ground Person ReID: A Challenging Large-Scale Benchmark
abstract
Aerial-Ground Person Re-IDentification (AGPReID) aims to extract identity-discriminative representations from heterogeneous perspectives across different platforms in complex real-world environments. However, existing methods primarily focus on visual appearance modeling and make insufficient use of semantic attribute priors, which limits their ability to bridge the aerial-ground view gap. To address this limitation, we propose a Semantic-driven Visual Progressive Refinement framework for AGPReID (SVPR-ReID), which effectively leverages textual attribute priors to guide the extraction of fine-grained visual cues. Specifically, we design a View-Decoupled Feature Extractor that incorporates view-aware textual prompts to decouple view-invariant identity features. Then, to alleviate inter-class ambiguity, we propose an Attribute-Scattered Mixture-of-Experts module that integrates attribute semantics into the visual space, thereby improving discrimination among visually similar pedestrians. Finally, we design a Context-Vision Progressive Refinement module for progressive refinement of attribute and view-invariant features, obtaining robust cross-view identity representations. In particular, we contribute a comprehensive benchmark for AGPReID, named CP2108, which contains 142,817 images of 2,108 identities annotated with 22 attributes. Notably, it includes 191 identities captured across different times, enabling both short- and long-term ReID evaluation, addressing the limitation of existing datasets that focus only on short-term scenarios. Extensive experimental results validate the effectiveness of our SVPR-ReID on four AGPReID datasets.
Aihua Zheng, Xixi Wan, Zi Wang 0013, Jin Tang 0001, Bin Luo 0001
AAAI7
2026 Neurological disorder detection based on an adaptive self-supervised multi-scale spatiotemporal interaction network
Changxu Dong, Zongyun Gu, Donghua Li, Xinyuan Xi, Bin Luo 0001, Dengdi Sun
Expert Syst. Appl.5
2026 Background noise suppression for advanced fine-grained visual classification
Zhi-Gang Wang, Sibao Chen 0001, Bin Luo 0001
Neurocomputing3
2026 Dynamic adaptive multi-view contrastive learning for unsupervised person re-identification
Zhi-Hua Li, Xue-Yan Wang, Sibao Chen 0001, Chris Ding, Bin Luo 0001
Neural Networks5
2026 Reliable and Compact Graph Fine-Tuning via Graph Sparse Prompting
abstract
Recently, graph prompt learning has garnered increasing attention in adapting pre-trained GNN models for downstream graph learning tasks. However, existing works generally conduct prompting over all graph elements (e.g., nodes, edges, node attributes, etc.), which is suboptimal and obviously redundant. To address this issue, we propose exploiting sparse representation theory for graph prompting and present Graph Sparse Prompting (GSP). GSP aims to adaptively and sparsely select the optimal elements (e.g., certain node attributes) to achieve compact prompting for downstream tasks. Specifically, we propose two kinds of GSP models, termed Graph Sparse Feature Prompting (GSFP) and Graph Sparse multi-Feature Prompting (GSmFP). Both GSFP and GSmFP provide a general scheme for tuning any specific pre-trained GNNs that can select some desired attributes for prompting by employing sparsity-guided prompt learning. A simple yet effective algorithm has been designed for solving GSFP and GSmFP models. Experiments on 16 widely-used benchmark datasets validate the effectiveness and advantages of the proposed GSFPs.
Bo Jiang 0002, Beibei Wang 0006, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Revisiting Deformable Convolution on Graphs: Large-Range Modeling and Robustness
abstract
Graph Convolution Networks (GCNs) have achieved remarkable success in representation of structured graph data. As we know that traditional GCNs are generally defined on the fixed first-order neighborhood receptive field which makes them be incapable to capture the long-range dependencies between distant nodes and also vulnerable to graph attacks and noises. To address these limitations, we revisit deformable convolution on graphs and propose a novel deformable graph convolution, termed Neighborhood-Deformable Graph Convolution (NDGC). The core of NDGC is to explicitly achieve the deformable convolution on graphs by introducing virtual neighbors which encode large-range information via the offsetting and interpolation function. That is, the introduced virtual neighbors can provide a larger receptive field with deformable receptive shape for graph convolution definition. Also, NDGC conducts message aggregation on the deformable virtual neighbors which thus performs more robustly w.r.t. graph attacks and noises. In particular, NDGC provides a general neighborhood deformable scheme, seamlessly integrating with many graph convolution definitions to derive their deformable variants. Experimental results validate the effectiveness and advantages of the proposed NDGC networks on several graph learning tasks.
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Multi-scale feature sharing and collaborative sampling for unsupervised vehicle re-identification
Sibao Chen 0001, Chris Ding, Bin Luo 0001
Pattern Recognit.4
2026 Semantic change detection of roads and bridges: A fine-grained dataset and multimodal frequency-driven detector
Qing-Ling Shu, Sibao Chen 0001, Xiao Wang 0014, Zhi-Hui You, Wei Lu 0032, Jin Tang 0001, Bin Luo 0001
Pattern Recognit.7
2026 Kernel entropy graph isomorphism network for graph classification
Lixiang Xu, Feiping Nie 0001, Enhong Chen, Bin Luo 0001
Pattern Recognit.5
2026 Context-Semantic Quality Awareness Network for fine-grained visual categorization
Sitong Li, Bo Jiang 0002, Bin Luo 0001, Jinhui Tang 0001
Pattern Recognit.5
2026 CLNS: Camera-aware label noise suppression for unsupervised visible-infrared person re-identification
Sicheng Zhao, Wei Lu 0032, Sibao Chen 0001, Chris Ding, Futian Wang, Jin Tang 0001, Bin Luo 0001
Pattern Recognit.8
2026 Entropy calibrated prototype embedding for transductive few-shot learning
Mengfei Guo, Bo Jiang 0002, Bin Luo 0001
Pattern Recognit. Lett.5
2026 Multiview Graph Contrastive Learning Based on Learnable Graph Augmentation
abstract
Graph contrastive learning has recently gained prominence as a key technique in the domain of graph representation learning. Most graph contrastive learning methods generate two graph views by augmenting the input graph and maximizing the consistency of their representations. However, due to a lack of prior knowledge of graph data, existing graph augmentation methods often generate low-quality views, which can destroy the core structural information of the graph and thus affect the model’s learning. The fundamental limitation lies in the semantic-agnostic nature of such augmentations, which fail to adapt to the underlying data distribution and the model’s evolving training state. In addition, existing contrastive learning methods typically employ a single contrastive strategy and rely on numerous similarity calculations, which makes it challenging to fully capture the diverse features in a graph. To address these limitations, we propose a multiview graph contrastive learning method based on learnable graph augmentation (MGLGA). Specifically, this method abandons the traditional random graph augmentation, constructs views through dynamic feature learning, and generates feature views using a graph feature learner and a postprocessing module, thereby focusing on the essential structural features of the graph while minimizing interference with the view generation. This learnable paradigm ensures that the augmented views maintain high semantic fidelity and provide adaptive, curriculum-like learning signals throughout the training process, thereby theoretically promoting better generalization. Moreover, we design a three-branch structure to capture both fine-grained and coarse-grained knowledge of the graph through node-level and graph-level contrasts. We further develop a structure-aware group discrimination loss through contrastive objective optimization, enhancing the model’s capacity for capturing graph structural patterns. Comprehensive evaluations demonstrate state-of-the-art performance across multiple downstream benchmarks.
Dengdi Sun, Weilong Gong, Bin Luo 0001, Zhuanlian Ding
IEEE Trans. Comput. Soc. Syst.4
2026 BHGraphAdapter: Parameter-Efficient VLMs Tuning Meets Hyper-Graph Learning
abstract
Adapter-based fine-tuning methods for Visual-Language Models (VLMs) have shown promising performance for feature adaptation in limited data scenarios. However, existing adapters generallyeitheremploy parameterized transformation for multi-modality feature refiningorexploit pairwise relationships between classes (i.e., GraphAdapter) for text enhancement, which ignore the inherent high-order correlations among data samples in the adaptation process. In this paper, for the first time, we propose to exploit the high-order relationships of visual samples within each mini-batch for fine-tuning VLMs and develop a novel Batch HyperGraph Adapter (BHGraphAdapter) to fine-tune VLMs. The core idea of BHGraphAdapter is to conduct feature adapter learning by capturing the inherent high-order semantic information of different samples within each mini-batch, which thus can fully exploit the complex context information in adaptation. Specifically, we first construct a Batch HyperGraph (BHGraph) to model the high-order correlation of samples within each mini-batch. Then, we introduce a message propagation module on BHGraph to update the node embeddings by aggregating information from their high-order neighbors, thereby capturing semantic relationships to enrich feature representation. Finally, we incorporate the proposed BHGraph learning into the pre-trained CLIP framework to achieve the feature adaptation for the downstream tasks. Extensive experiments on 11 benchmark datasets show that our proposed BHGraphAdapter outperforms the SOTA adapter tuning methods. The source code and data will be released at https://github.com/LiuMeilin7195/BHGraphAdapter.
Xixi Wang 0005, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Vehicle-Centric Perception via Multimodal Structured Pre-Training
abstract
Vehicle-centric perception plays a crucial role in many intelligent systems, including large-scale surveillance systems, intelligent transportation, and autonomous driving. Existing approaches typically employ general pre-trained weights to initialize backbone networks, followed by task-specific fine-tuning. However, these models lack effective learning of vehiclerelated knowledge during pre-training, resulting in poor capability for modeling general vehicle perception representations. To handle this problem, we propose VehicleMAE-V2, a novel vehicle-centric pre-trained large model. By exploring and exploiting vehicle-related multimodal structured priors to guide the masked token reconstruction process, our approach can significantly enhance the model’s capability to learn generalizable representations for vehicle-centric perception. Specifically, we design the Symmetry-guided Mask Module (SMM), Contour-guided Representation Module (CRM) and Semantics-guided Representation Module (SRM) to incorporate three kinds of structured priors into token reconstruction including symmetry, contour and semantics of vehicles respectively. SMM utilizes the vehicle symmetry constraints to avoid retaining symmetric patches and can thus select high-quality masked image patches and reduce information redundancy. CRM minimizes the prob23 ability distribution divergence between contour features and reconstructed features and can thus preserve holistic vehicle structure information during pixel-level reconstruction. SRM aligns image-text features through contrastive learning and cross-modal distillation to address the feature confusion caused by insufficient semantic understanding during masked reconstruction. To support the pre-training of VehicleMAE-V2, we construct Autobot4M, a large-scale dataset comprising approximately 4 million vehicle images and 12,693 text descriptions. Extensive experiments on five downstream tasks demonstrate the superior performance of VehicleMAE-V2. The source code, dataset, and pre-trained large models are available on https://github.com/Vehicle-AHU/VehicleMAE.
Xiao Wang 0014, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Pixel-Level RGBT Fusion Tracking via Heterogeneous Multi-Expert Distillation and Decoupled Representation Learning
abstract
Pixel-level fusion is widely considered a lightweight yet limited strategy in RGB-Thermal (RGBT) tracking due to its shallow representational capacity. However, its actual limitations and potential remain largely unexplored. We systematically analyze fusion location, modality alignment, and tracking performance, revealing that despite lower modality gaps than feature-level fusion, pixel-level fusion lacks task-relevant discrimination, restricting its effectiveness. In this paper, we propose the Task-driven Pixel-level Fusion tracker (TPF), which preserves the efficiency of early fusion while enhancing discriminative capacity. Central to TPF is a lightweight pixel fusion adapter that ensures real-time image fusion with only 14.3KB extra parameters over the baseline at inference. To enhance its limited representational capacity, we propose a task-driven progressive learning framework consisting of two key stages. First, a heterogeneous multi-expert distillation scheme adaptively transfers image fusion knowledge from diverse models under tracking-guided evaluation, mitigating the generalization limitations of single-teacher distillation across varied tracking scenarios. Second, to overcome limited task discrimination caused by sparse, target-focused tracking supervision, we propose a decoupled representation learning strategy that offers dense, complementary guidance to improve target-background separation and fusion quality. A nearest-neighbor dynamic template update further enhances robustness to appearance changes. Extensive experiments on four RGBT tracking benchmarks show that TPF achieves competitive accuracy and speed, outperforming both feature-level and existing pixel-level fusion methods, offering new insights into efficient RGBT tracking.
Andong Lu, Yuanzhi Guo, Kunpeng Wang 0005, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Image Process.6
2026 Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection
abstract
Existing multimodal UAV object detection methods often overlook the impact of semantic gaps between modalities, which makes it difficult to achieve accurate semantic and spatial alignments and ultimately limits detection performance. To address this problem, we propose a Large Language Model (LLM) guided Progressive feature Alignment Network called LPANet, which leverages the semantic features extracted from a large language model to guide the progressive semantic and spatial alignment between modalities for multimodal UAV object detection. To employ the powerful semantic representation of LLM, we generate the fine-grained text descriptions of each object category by ChatGPT and then extract the semantic features using the large language model MPNet, providing high-level semantic priors to guide multimodal alignment. Based on the semantic features, we guide the semantic and spatial alignments in a progressive manner as follows. First, we design the Semantic Alignment Module (SAM) to pull the semantic features and multimodal visual features of each object closer, alleviating the semantic differences of objects between modalities. Second, we design the Explicit Spatial Alignment Module (ESM) by integrating the semantic relations into the estimation of feature-level offsets, alleviating the coarse spatial misalignment between modalities. Finally, we design the Implicit Spatial alignment Module (ISM), which leverages the cross-modal correlations to aggregate key features from neighboring regions to achieve implicit spatial alignment. Comprehensive experiments on two public multimodal UAV object detection datasets demonstrate that our approach outperforms state-of-the-art multimodal UAV object detectors. The source code will be released on https://github.com/Vehicle-AHU/LPANet.
Chenglong Li 0002, Xiao Wang 0014, Bin Luo 0001
IEEE Trans. Image Process.4
2026 ICPL-ReID: Identity-Conditional Prompt Learning for Multi-Spectral Object Re-Identification
Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.5
2026 DEEP: Decoupled Semantic Prompt Learning, Guiding and Embedding for Multi-Spectral Object Re-Identification
abstract
Multi-spectral object re-identification (ReID) captures diverse object semantics to robustly recognize identity in complex environments. However, without explicit semantic guidance (e.g., attributes, masks, and keypoints), existing modal fusion-based methods struggle to comprehensively capture person or vehicle semantics across spectra. Thanks to the large-scale vision-language pre-training, CLIP effectively aligns visual concepts across different image modalities to a unified semantic prompt. In this paper, we proposeDEEP, aDEcoupled sEmanticPrompt Learning, Guiding and Embedding framework for Multi-Spectral Object ReID. Specifically, to address the challenges posed by low-quality modality noise and spectral style discrepancies, we first propose a Decoupled Semantic Prompt (DSP) strategy, which explicitly decouples the semantic alignment into spectral-style learning with spectral-shared prompts and object content learning with instance-specific inversion token. Second, to lead the model focusing on semantically faithful regions, we propose a Semantic-Guided Spectral Fusion (SGSF) module that builds a semantic interaction bridge between spectra to explore complementary semantics across modalities. Finally, to further empower the spectral representation, we propose a Spectral Semantic Embedding (SSE) module constrained by semantic-aware structural consistency to refine the fine-grained identity semantics in each spectrum. Extensive experiments on five public benchmarks, RGBNT201, Market-MM, MSVR310, WMVEID863, and RGBNT100, demonstrate the proposed method outperforms the state-of-the-art methods. The source code is released at this link:https://github.com/lsh-ahu/DEEP-ReID.
Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.5
2026 RCNet: Reliable Co-Training Network for Weakly Supervised Change Detection
abstract
Fully supervised change detection (CD) methods in remote sensing (RS) perform well but depend on costly and time-consuming pixel-level annotations, which are impractical to obtain at scale. Therefore, it is essential to develop annotation-efficient alternatives that can narrow the performance gap with fully supervised methods. To this end, we propose a novel weakly supervised CD framework, named RCNet, which employs dual networks to implement reliable co-training using image-level annotations. Our framework is grounded in multi-view learning of co-training and the localization ability of class activation mapping (CAM). In our approach, two sub-nets with the same architecture perform image-level change classification and pixel-level segmentation from different views. Although CAM roughly localizes changes, ambiguity and noise in its pseudo labels may cause confirmation bias, limiting performance. Our approach mitigates this bias by introducing a feature discrepancy loss to enable cross-supervision between two sub-nets. Meanwhile, CAM tends to highlight a single object, but RS images commonly contain many dense and small changed objects with complexity, resulting in decreased reliability of pseudo labels. Therefore, we present an IoU-based reliable pseudo label screening (RPLS) strategy, which minimizes the likelihood of changed areas being misidentified as unchanged, enhancing the reliability of changed information obtained. Besides, to further improve boundary fineness and internal integrity of changed areas, we incorporate an additional strong perturbation branch for each sub-net and develop a consistency regularization loss. Extensive experiments on three challenging RS image CD datasets demonstrate that our RCNet achieves competitive performance with image-level labels. The source code is available athttps://github.com/Youzhihui/RCNet.
Zhi-Hui You, Sibao Chen 0001, Chris Ding, Lili Huang 0006, Jia-Xin Wang, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.7
2026 Visual-Label Alignment and Attribute-Aware Prompt for Multi-Label Image Recognition with Partial Labels
abstract
The problem of Multi-Label Image Recognition with Partial Labels (MLIR-PL) is a significant challenge in computer vision, primarily due to the scarcity and high cost of complete annotations. Recent advances have leveraged large-scale vision-language models, such as CLIP, to establish rich correspondences between images and their labels, thereby improving the MLIR-PL performance. However, the existing CLIP-based methods have not fully exploited fine-grained local image features to mitigate interference from semantically irrelevant regions. Moreover, many studies have oversimplified the use of prompt contexts, limiting their ability to comprehensively capture the multi-dimensional attributes of categories. To address these limitations, this article proposes a novel MLIR-PL model with Visual–Label Alignment and Attribute-Aware Prompt (VA \({}^{3}\) P), which sufficiently harnesses the capabilities of large-scale pre-trained vision-language models. In the model, we design a Visual–Label Alignment module to establish a mapping between local image features and category text representations, conspicuously reducing the interference from irrelevant regions. Additionally, our Attribute-Aware Prompt module offers diverse contextual information, providing a more comprehensive representation of the category’s attributes. Extensive experimental results on the COCO 2014 and VOC 2007 datasets, compared with multiple state-of-the-art methods, demonstrate that our model achieves the best performance comprehensively, verifying the advantages of the proposed model in the MLIR-PL task.
Dengdi Sun, Hongxing Xie, Zhendong Cai, Leilei Ma 0002, Bin Luo 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2026 Millimeter-wave radar-assisted skeleton-guided video reconstruction for surveillance systems
Wenjie Leng, Bin Luo 0001, Mao Dai
Vis. Comput.2
2025 RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba
abstract
Existing RGBT tracking methods often design various interaction models to perform cross-modal fusion of each layer, but can not execute the feature interactions among all layers, which plays a critical role in robust multimodal representation, due to large computational burden. To address this issue, this paper presents a novel All-layer multimodal Interaction Network, named AINet, which performs efficient and effective feature interactions of all modalities and layers in a progressive fusion Mamba, for robust RGBT tracking. Even though modality features in different layers are known to contain different cues, it is always challenging to build multimodal interactions in each layer due to struggling in balancing interaction capabilities and efficiency. Meanwhile, considering that the feature discrepancy between RGB and thermal modalities reflects their complementary information to some extent, we design a Difference-based Fusion Mamba (DFM) to achieve enhanced fusion of different modalities with linear complexity. When interacting with features from all layers, a huge number of token sequences (3840 tokens in this work) are involved and the computational burden is thus large. To handle this problem, we design an Order-dynamic Fusion Mamba (OFM) to execute efficient and effective feature interactions of all layers by dynamically adjusting the scan order of different layers in Mamba. Extensive experiments on four public RGBT tracking datasets show that AINet achieves leading performance against existing state-of-the-art methods. We will release the code upon acceptance of the paper.
Andong Lu, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
AAAI5
2025 Alignment-Free RGB-T Salient Object Detection: A Large-Scale Dataset and Progressive Correlation Network
abstract
Alignment-free RGB-Thermal (RGB-T) salient object detection (SOD) aims to achieve robust performance in complex scenes by directly leveraging the complementary information from unaligned visible-thermal image pairs, without requiring manual alignment. However, the labor-intensive process of collecting and annotating image pairs limits the scale of existing benchmarks, hindering the advancement of alignment-free RGB-T SOD. In this paper, we construct a large-scale and high-diversity unaligned RGB-T SOD dataset named UVT20K, comprising 20,000 image pairs, 407 scenes, and 1256 object categories. All samples are collected from real-world scenarios with various challenges, such as low illumination, image clutter, complex salient objects, and so on. To support the exploration for further research, each sample in UVT20K is annotated with a comprehensive set of ground truths, including saliency masks, scribbles, boundaries, and challenge attributes. In addition, we propose a Progressive Correlation Network (PCNet), which models inter- and intra-modal correlations on the basis of explicit alignment to achieve accurate predictions in unaligned image pairs. Extensive experiments conducted on two unaligned three weakly aligned three aligned datasets demonstrate the effectiveness of our method.
Kunpeng Wang 0005, Keke Chen, Chenglong Li 0002, Zhengzheng Tu, Bin Luo 0001
AAAI5
2025 CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework
abstract
Event cameras have attracted increasing attention in recent years due to their advantages in high dynamic range, high temporal resolution, low power consumption, and low latency. Some researchers have begun exploring pre-training directly on event data. Nevertheless, these efforts often fail to establish strong connections with RGB frames, limiting their applicability in multi-modal fusion scenarios. To address these issues, we propose a novel CM3AE pre-training framework for the RGB-Event perception. This framework accepts multi-modalities/views of data as input, including RGB images, event images, and event voxels, providing robust support for both event-based and RGB-event fusion based downstream tasks. Specifically, we design a multi-modal fusion reconstruction module that reconstructs the original image from fused multi-modal features, explicitly enhancing the model's ability to aggregate cross-modal complementary information. Additionally, we employ a multi-modal contrastive learning strategy to align cross-modal feature representations in a shared latent space, which effectively enhances the model's capability for multi-modal understanding and capturing global dependencies. We construct a large-scale dataset containing 2,535,759 RGB-Event data pairs for the pre-training. Extensive experiments on five downstream tasks fully demonstrated the effectiveness of CM3AE. Source code and pre-trained models will be released on https://github.com/Event-AHU/CM3AE.
Xiao Wang 0014, Chenglong Li 0002, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Qi Liu 0003
ACM Multimedia6
2025 Multi-view brain network classification based on Adaptive Graph Isomorphic Information Bottleneck Mamba
Changxu Dong, Dengdi Sun, Zhenda Yu, Bin Luo 0001
Expert Syst. Appl.4
2025 Modality-missing RGBT Tracking: Invertible Prompt Learning and High-quality Benchmarks
Andong Lu, Chenglong Li 0002, Jiacong Zhao, Jin Tang 0001, Bin Luo 0001
Int. J. Comput. Vis.5
2025 Learning Dynamic Batch-Graph Representation for Deep Representation Learning
Xixi Wang 0005, Bo Jiang 0002, Xiao Wang 0014, Bin Luo 0001
Int. J. Comput. Vis.4
2025 Lightweight oriented object detection with Dynamic Smooth Feature Fusion Network
Wei Lu 0032, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001
Neurocomputing5
2025 Understanding beyond outputs: A novel knowledge distillation method using Schur decomposition
Chang-Ming Pan, Sibao Chen 0001, Bo Jiang 0002, Bin Luo 0001
Neurocomputing4
2025 Self-supervised spatial-temporal contrastive network for EEG-based brain network classification
Changxu Dong, Dengdi Sun, Bin Luo 0001
Neural Networks3
2025 Multi-level semantic-aware transformer for image captioning
Shan Song, Qihang Wu, Bo Jiang 0002, Bin Luo 0001, Jinhui Tang 0001
Neural Networks5
2025 Dynamic semantic-geometric guidance and structure transfer network for cross-scene hyperspectral image classification
Shuke Wang, Bo Jiang 0002, Zhifu Tao, Bin Luo 0001
Neural Networks6
2025 Graph Spiking Attention Network: Sparsity, Efficiency and Robustness
abstract
Existing Graph Attention Networks (GATs) generally adopt the self-attention mechanism to learn graph edge attention, which usually return dense attention coefficients over all neighbors and thus are prone to be sensitive to graph edge noises. To overcome this problem, sparse GATs are desirable and have garnered increasing interest in recent years. However, existing sparse GATs usually suffer from high training complexity and are also not straightforward for inductive learning tasks. To address these issues, we propose to learn sparse GATs by exploiting spiking neuron (SN) mechanism, termed Graph Spiking Attention (GSAT). Specifically, it is known that spiking neuron can perform inexpensive information processing by transmitting the input data into discrete spike trains and return sparse outputs. Inspired by it, this work attempts to exploit spiking neuron to learn sparse attention coefficients, resulting in edge-sparsified graph for GNNs. Therefore, GSAT can perform message passing on the selective neighbors naturally, which makes GSAT perform compactly and robustly w.r.t graph noises. Moreover, GSAT can be used straightforwardly for inductive learning tasks. Extensive experiments on both transductive and inductive tasks demonstrate the effectiveness, robustness and efficiency of GSAT.
Beibei Wang 0006, Bo Jiang 0002, Jin Tang 0001, Lu Bai 0001, Bin Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Unifying Graph Contrastive Learning via Graph Message Augmentation
abstract
Graph contrastive learning is usually performed by first conducting Graph Data Augmentation (GDA) and then employing a contrastive learning pipeline to train GNNs. As we know that GDA is an important issue for graph contrastive learning. Various GDAs have been developed recently which mainly involve dropping or perturbing edges, nodes, node attributes and edge attributes. However, to our knowledge, it still lacks a universal and effective augmentor that is suitable for different types of graph data. To address this issue, in this paper, we first introduce the graph message representation of graph data. Based on it, we then propose a novel Graph Message Augmentation (GMA), a universal scheme for reformulating many existing GDAs. The proposed unified GMA not only gives a new perspective to understand many existing GDAs but also provides a universal and more effective graph data augmentation for graph self-supervised learning tasks. Moreover, GMA introduces an easy way to implement the mixup augmentor which is natural for images but usually challengeable for graphs. Based on the proposed GMA, we then propose a unified graph contrastive learning, termed Graph Message Contrastive Learning (GMCL), that employs attribution-guided universal GMA for graph contrastive learning. Experiments on many graph learning tasks demonstrate the effectiveness and benefits of the proposed GMA and GMCL approaches.
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Semantic-aware frame-event fusion based pattern recognition via large vision-language models
Jiandong Jin, Yanlin Zhong, Yaoyang Wu, Lan Chen 0003, Xiao Wang 0014, Bin Luo 0001
Pattern Recognit.8
2025 Graph neural network based on graph kernel: A survey
Lixiang Xu, Jiawang Peng, Xiaoyi Jiang 0001, Enhong Chen, Bin Luo 0001
Pattern Recognit.5
2025 Instant pose extraction based on mask transformer for occluded person re-identification
Qing-Ling Shu, Sibao Chen 0001, Lili Huang 0006, Bin Luo 0001
Pattern Recognit.5
2025 BET-BiLSTM Model: A Robust Solution for Automated Requirements Classification
abstract
ABSTRACT Transformer methods have revolutionized software requirements classification by combining advanced natural language processing to accurately understand and categorize requirements. While traditional methods like Doc2Vec and TF‐IDF are useful, they often fail to capture the deep contextual relationships and subtle meanings inherent in textual data. Transformer models possess unique strengths and weaknesses, impacting their ability to capture various aspects of the data. Consequently, relying on a single model can lead to suboptimal feature representations, limiting the overall performance of the classification task. To address this challenge, our study introduces an innovative BET‐BiLSTM (balanced ensemble transformers using Bi‐LSTM) model. This model combines the strengths of five transformer–based models BERT, RoBERTa, XLNet, GPT‐2, and T5 through weighted averaging ensemble, resulting in a sophisticated and resilient feature set. By employing data balancing techniques, we ensure a well‐distributed representation of features, addressing the issue of class imbalance. The BET‐BiLSTM model plays a crucial role in the classification process, achieving an impressive accuracy of 96%. Moreover, the practical applicability of this model is validated through its successful implementation on three publicly available unlabeled datasets and one additional labeled dataset. The model significantly improved the completeness and reliability of these datasets by accurately predicting labels for previously unclassified requirements. This makes our approach a powerful tool for large‐scale requirements analysis and classification tasks, outperforming traditional single‐model methods and showcasing its real‐world effectiveness.
Jalil Abbas, Cheng Zhang 0010, Bin Luo 0001
J. Softw. Evol. Process.3
2025 Unified-Modal Salient Object Detection via Adaptive Prompt Learning
abstract
Existing single-modal and multi-modal salient object detection (SOD) methods focus on designing specific architectures tailored for their respective tasks. However, developing completely different models for different tasks leads to labor and time consumption, as well as high computational and practical deployment costs. In this paper, we attempt to address both single-modal and multi-modal SOD in a unified framework called UniSOD, which fully exploits the overlapping prior knowledge between different tasks. Nevertheless, assigning appropriate strategies to modality variable inputs is challenging. To this end, UniSOD learns modality-aware prompts with task-specific hints through adaptive prompt learning, which are seamlessly plugged into the proposed pre-trained baseline SOD model to handle corresponding tasks, while only requiring few learnable parameters compared to training the entire model from scratch. In particular, each modality-aware prompt is solely generated from a homogeneous switchable prompt generation (SPG) block, which adaptively performs structural switching based on single-modal and multi-modal inputs without manual intervention, ensuring that the framework can effectively handle diverse input cases (e.g., RGB-only, RGB-D, RGB-T) with a unified approach. Through end-to-end joint training, UniSOD achieves ovrall competitive performance on 14 benchmark datasets, demonstrating its ability to efficiently unify single-modal and multi-modal SOD tasks. Code has been available athttps://github.com/Angknpng/UniSOD
Kunpeng Wang 0005, Zhengzheng Tu, Chenglong Li 0002, Zhengyi Liu, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Camera-Proxy Enhanced Identity-Recalibration Learning for Unsupervised Visible-Infrared Person Re-Identification
abstract
Visible-Infrared person Re-Identification (VI-ReID) involves querying images of the same person across visible and infrared modalities. To minimize annotation costs, Unsupervised Visible-Infrared person Re-Identification (UVI-ReID) using pseudo-label contrastive learning has emerged. Traditional UVI-ReID approaches often neglected camera domain information and relied on inadequate update strategies during training, only using cosine distance for testing, which led to incorrect mapping of cross-modal relationships. To address these issues, we propose Camera-proxy Enhanced Identity-recalibration Learning (CEIL). It consists of two main stages: first, it employs intra-modal contrastive learning in conjunction with the camera-proxy, updates the memory bank using our innovative Difficulty-aware Cluster-based Memory Updating (DCMU) strategy, and applies Camera Domain-driven Local correlation (CDL) Loss to enhance the learning process. Then utilizes cross-modal contrastive learning, featuring our Proxy-enhanced Cross-modal Mapping (PCM) module, to recalibrate the identity relationships between different modalities. Graph network-based Camera constraint adjustment Re-ranking (GCR) method is adopted during test, utilizing camera domain information to recalibrate the correspondence between identities. Extensive experiments have demonstrated that CEIL achieving state-of-the-art performance on the SYSU-MM01, RegDB, and LLCM datasets and the GCR, as a general unsupervised re-ranking method, can further enhance performance of model on these datasets. The code will be released athttps://github.com/maybeextra/CEIL.
Run-Sen Xia, Xue-Yan Wang, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Adversarial Examples Detection With Enhanced Image Difference Features Based on Local Histogram Equalization
abstract
Deep Neural Networks (DNNs) have recently made significant strides in various fields; however, they are susceptible to adversarial examples—crafted inputs with imperceptible perturbations that can mislead these networks. Notably, even when adversaries lack access to the complete model parameters, they can still generate adversarial examples targeting a range of DNN-based task systems. Various defense mechanisms have been proposed, such as feature compression and gradient masking. Nevertheless, extensive research indicates that these methods often address only specific attacks, rendering them ineffective against novel and unknown attack strategies. Recent studies have highlighted the efficacy of identifying adversarial examples in the frequency domain; however, these approaches are limited to frequency-based analysis. In this study, we experimentally observe that adversarial examples possess significant characteristics in local regions. Specifically, adversarial perturbations exhibit localized randomness, whereas the high-frequency information in normal examples is both locally coherent and semantically relevant. This critical distinction enables effectively distinguishing adversarial examples from normal ones. To leverage this insight, we aim to enhance the high-frequency features of input examples to amplify their feature disparities. We propose an image enhancement method utilizing local histogram equalization. Our experimental results demonstrate that this method substantially improves detector performance without modifying the existing detection models. Furthermore, this technique can be seamlessly integrated with task models, effectively reducing deployment costs in practical applications.
Zhao-Xia Yin, Hang Su 0006, Jianteng Peng, Bin Luo 0001
IEEE Trans. Dependable Secur. Comput.6
2025 Multiscale Adaptive Decoder and Diversity Selection Network for Road Extraction in Remote Sensing Image
abstract
Road extraction has been a common and challenging task in the field of remote sensing images. Due to factors such as the high resolution of remote sensing images and the subtle visibility of road features, existing methods often miss certain areas during detection and extraction. These methods struggle to capture contextual information effectively and tend to exhibit false positives and false negatives when handling objects of varying sizes. This article proposes a network based on a multi-scale adaptive decoder and diverse selection (MADSNet) to address the issue of inadequate contextual information capture. By leveraging feature diverse selection, the method minimizes errors in distinguishing between road features and background interference. Specifically, the multi-scale feature flexible extraction (MFFE) decoder utilizes the relevance inquiry attention (RIA) module and scope flexible fusion (SFF) module to enhance the ability to capture contextual information with relatively low computational demands. The optimal choice graph attention (OCGA) module aggregates neighboring nodes with similar features in a graph structure, improving focus on the single class of roads. Furthermore, a multi-level feature selection (MFS) module is proposed to activate the features relevant to the current stage while suppressing features from other stages and interfering with noise. Quantitative and qualitative experimental results on three public datasets demonstrate that the proposed MADSNet outperforms currently popular methods in terms of performance. The code will be available at https://github.com/Talent02/MADSNet.
Zhen-Tao Hua, Sibao Chen 0001, Wei Lu 0032, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Multidimensional Remote Sensing Change Detection Based on Siamese Dual-Branch Networks
abstract
Deep learning models, particularly convolutional neural networks (CNNs), have demonstrated outstanding feature learning capabilities, leading to remarkable performance in remote sensing change detection (RSCD) tasks. However, their most critical drawback lies in the lack of effective modeling of global information. This deficiency affects the model’s understanding of the overall context and structure of the entire image, making it difficult to distinguish between background and target areas, thereby leading to the erroneous identification of change regions. Second, features extracted by traditional backbone networks contain a significant amount of noise, resulting in blurred boundaries of changed objects. The challenge of effectively fusing detailed and semantic information to accurately differentiate pseudo changes remains significant. Furthermore, how to fully exploit multiscale information is another issue worth considering. We propose a full-scale multidimensional interaction network called SDSN, which enhances feature representation by leveraging both detail and semantic branches. Initially, bi-temporal images are processed by the encoder to extract coarse multiscale features. The semantic branch guides shallow-scale features, while the detail branch focuses on deep-scale features. Multikernel receptive module (MRM) aggregates global information. The detail branch utilizes a diversity variance module (DVM) and differential operations to generate refined change maps with noise reduction and background suppression. A multidimensional cross-perception module (MCM) guides the fusion of these change maps, establishing multidimensional dependencies to enrich feature representation. Compared with previous methods, SDSN demonstrates greater performance under complex environmental conditions, particularly noteworthy for its fewer parameters (4.03 M) and lower computational costs (7.94 G). The code is publicly available athttps://github.com/dpt000121/dpt.
Li-Rong Shen, Sibao Chen 0001, Lili Huang 0006, Zhi-Hui You, Chris Ding, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.7
2025 Bridging the Style Gap: Style-Guided Distillation Domain Adaptation for Hyperspectral Image Classification
abstract
Land cover in different hyperspectral image (HSI) commonly exhibits style differences and similarities in the same category and distinct categories. However, most of existing cross-scene HSI classification methods overlook this issue and only conduct the feature-level alignment with unsupervised domain adaptation (UDA). To address this limitation, we propose the Style-Guided Distillation Domain Adaptation (SGDDA) for HSI classification. First, a Fourier Transform based Style Transfer (FTST) module is proposed to generate an enhanced source HSI. It transfers stylistic features from the target domain (TD) to the source domain (SD) by substituting low-frequency components, thereby preserving semantic invariance while bridging their style gap. Second, the Dual-Path Knowledge Distillation (DPKD) module is designed to reduce ambiguity in category assignment for the SD. This is achieved through cross-domain consistency learning between original SD samples and their enhanced style-transferred counterparts, ensuring robust feature alignment across domains. Third, unlike existing methods that primarily utilize classification threshold to select pseudo-label for target samples, we propose the Confidence-aware Dual Pseudo-label Consensus (CADPLC) strategy. This strategy dynamically selects the reliable pseudo-labels by leveraging both class prototype matching and teacher-student prediction consensus, eliminating reliance on fixed thresholds and significantly improving adaptation to the target domain. Experiments on three benchmark datasets demonstrate the superiorities of the proposed SGDDA in comparison with several state-of-the-art methods. The code is available at https://github.com/Wei-spvl/SGDDA.
Bo Jiang 0002, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Real-World Remote Sensing Image Dehazing: Benchmark and Baseline
abstract
Remote Sensing Image Dehazing (RSID) poses significant challenges in real-world scenarios due to the complex atmospheric conditions and severe color distortions that degrade image quality. The scarcity of real-world remote sensing hazy image pairs has compelled existing methods to rely primarily on synthetic datasets. However, these methods struggle with real-world applications due to the inherent domain gap between synthetic and real data. To address this, we introduce Real-World Remote Sensing Hazy Image Dataset (RRSHID), the first large-scale dataset featuring real-world hazy and hazy-free image pairs across diverse atmospheric conditions. Based on this, we propose MCAF-Net, a novel framework tailored for real-world RSID. Its effectiveness arises from three innovative components: Multi-branch Feature Integration Block Aggregator (MFIBA), which enables robust feature extraction through cascaded integration blocks and parallel multi-branch processing; Color-Calibrated Self-Supervised Attention Module (CSAM), which mitigates complex color distortions via self-supervised learning and attention-guided refinement; and Multi-Scale Feature Adaptive Fusion Module (MFAFM), which integrates features effectively while preserving local details and global context. Extensive experiments validate that MCAF-Net demonstrates state-of-the-art performance in real-world RSID, while maintaining competitive performance on synthetic datasets. The introduction of RRSHID and MCAF-Net sets new benchmarks for real-world RSID research, advancing practical solutions for this complex task. The code and dataset are publicly available at here.
Zeng-Hui Zhu, Wei Lu 0032, Sibao Chen 0001, Chris Ding, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 Nighttime Person Re-Identification via Collaborative Enhancement Network With Multi-Domain Learning
abstract
Prevalent nighttime person re-identification (ReID) methods typically combine image relighting and ReID networks in a sequential manner. However, their performance (recognition accuracy) is limited by the quality of relighting images and insufficient collaboration between image relighting and ReID tasks. To handle these problems, we propose a novel Collaborative Enhancement Network called CENet, which performs the multilevel feature interactions in a parallel framework, for nighttime person ReID. In particular, the designed parallel structure of CENet can not only avoid the impact of the quality of relighting images on ReID performance, but also allow us to mine the collaborative relations between image relighting and person ReID tasks. To this end, we integrate the multilevel feature interactions in CENet, where we first share the Transformer encoder to build the low-level feature interaction, and then perform the feature distillation that transfers the high-level features from image relighting to ReID, thereby alleviating the severe image degradation issue caused by the nighttime scenario while avoiding the impact of relighting images. In addition, the sizes of existing real-world nighttime person ReID datasets are limited, and large-scale synthetic ones exhibit substantial domain gaps with real-world data. To leverage both small-scale real-world and large-scale synthetic training data, we develop a multi-domain learning algorithm, which alternately utilizes both kinds of data to reduce the inter-domain difference in training procedure. Extensive experiments on two real nighttime datasets,Night600andRGBNT201rgb, and a synthetic nighttime ReID dataset are conducted to validate the effectiveness of CENet. We release the code and synthetic dataset at: https://github.com/Alexadlu/CENet.
Andong Lu, Chenglong Li 0002, Tianrui Zha, Xiaofeng Wang 0009, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Inf. Forensics Secur.6
2025 Prototype-Based Diversity and Integrity Learning for All-Day Multi-Modal Person Re-Identification
abstract
Recent multi-modal person re-identification methods have improved model performance by leveraging complementary information from multiple spectra. However, existing methods cannot ensure feature stability under varying illumination and rely on inflexible paired data, remaining inadequate against real-world cross-time retrieval and modality-missing challenges. To solve these, we first propose diversity representation that augments illumination-sensitive images to simulate diverse lighting conditions via illumination augmentation and enriches instance features using modality-specific prototypes via multiple interaction modules. Secondly, we propose integrity reconstruction that leverages prototypes and available instance features to recover information, the reconstruction module effectively utilizes identity and modality cues to address unpredictable missing problems. In addition, we build a more comprehensive dataset (AllDay843) to alleviate the inadequate dataset diversity, which comprises 91,371 images of 843 identities captured by multi-modal cameras across various periods throughout the day, while incorporating numerous real-world challenges. By integrating diversity representation and integrity reconstruction, the proposed Prototype-Based Diversity and Integrity learning network (PDINet) establishes excellence on the AllDay843 dataset, surpassing existing state-of-the-art approaches. The data and codes are available in https://github.com/ziwang1121/PDINet.
Zi Wang 0013, Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Inf. Forensics Secur.6
2025 AFTER: Attention-Based Fusion Router for RGBT Tracking
abstract
Multi-modal feature fusion as a core investigative component of RGBT tracking emerges numerous fusion studies in recent years. However, existing RGBT tracking methods widely adopt fixed fusion structures to integrate multi-modal feature, which are hard to handle various challenges in dynamic scenarios. To address this problem, this work presents a novel Attention-based Fusion router called AFTER, which optimizes the fusion structure to adapt to the dynamic challenging scenarios, for robust RGBT tracking. In particular, we design a fusion structure space based on the hierarchical attention network, each attention-based fusion unit corresponding to a fusion operation and a combination of these attention units corresponding to a fusion structure. Through optimizing the combination of attention-based fusion units, we can dynamically select the fusion structure to adapt to various challenging scenarios. Unlike complex search of different structures in neural architecture search algorithms, we develop a dynamic routing algorithm, which equips each attention-based fusion unit with a router, to predict the combination weights for efficient optimization of the fusion structure. Extensive experiments on five mainstream RGBT tracking datasets demonstrate the superior performance of the proposed AFTER against state-of-the-art RGBT trackers. We release the code in https://github.com/Alexadlu/AFter.
Andong Lu, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Image Process.5
2025 Transductive Few-shot Learning via Joint Message Passing and Prototype-based Soft-label Propagation
abstract
The transductive Few-shot Learning (FSL) mostly employs either prototype learning or label propagation methods to generalize to new classes by using the information of all query samples. However, existing methods have several main limitations. First, the prototype methods mainly focus on support samples which fail to fully exploit the relationships of query samples. Second, existing label propagation methods are generally not effective for the class-imbalanced problem. Third, existing works usually optimize the learnable parameters during inference which significantly reduces the efficiency of existing methods. To address these limitations, this article proposes an efficient and robust method for transductive FSL problem, termed Prototype-based Soft-label Propagation (PSLP), which combines the prototype learning and label propagation together for FSL problem. In our proposed method, first, the soft-label presentation for each query sample is estimated by leveraging prototypes. Then, the soft-label propagation is conducted on the learned query-support graph and the prototype representation is rectified. Both steps are conducted progressively for boosting the performance. Moreover, to learn effective prototypes for soft-label estimation and the desirable query-support graph for soft-label propagation, we design a new joint message passing scheme to learn the sample presentation and relational graph jointly. The PSLP method is parameter-free and can be implemented very efficiently. The experiments conducted on four popular benchmarks show that our method achieves competitive results on both balanced and imbalanced settings compared to the state-of-the-art methods. The code is released at https://github.com/mobulan/PSLP .
Bo Jiang 0002, Bin Luo 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 An Improved WM Pattern Matching Algorithm Based on Cuckoo Filter
abstract
Pattern matching algorithms are widely used in fields such as traffic classification and management, user behavior analysis, and more. The increasingly large and complex network traffic poses significant challenges to feature matching processes. The cuckoo filter, capable of quickly determining whether an element is in a set, can be combined with pattern matching algorithms to accelerate feature matching. Building upon previous research, we propose an improved Wu-Manber (WM) algorithm that further reduces the size of the hash table generated during preprocessing, decreases the number of hash computations, and incorporates a cuckoo filter to eliminate unmatched text prefixes. This improved WM algorithm considers factors affecting algorithm performance, such as the large scale of the pattern set and the prevalence of repeated pattern string suffixes. Experimental results demonstrate that our proposed algorithm, Fast Parallel Wu-Manber (FSPRWM), significantly enhances matching speed while effectively reducing memory consumption.
Zhiyong Zha, Jiangyi Liu, Bin Luo 0001, Mingyuan Ren, Menglan Hu, Kai Peng 0001
HPCC3
2024 Breaking Modality Gap in RGBT Tracking: Coupled Knowledge Distillation
abstract
Modality gap between RGB and thermal infrared (TIR) images is a crucial issue but often overlooked in existing RGBT tracking methods. It can be observed that modality gap mainly lies in the image style difference. In this work, we propose a novel Coupled Knowledge Distillation framework called CKD, which pursues common styles of different modalities to break modality gap, for high performance RGBT tracking. In particular, we introduce two student networks and employ the style distillation loss to make their style features consistent as much as possible. Through alleviating the style difference of two student networks, we can break modality gap of different modalities well. However, the distillation of style features might harm to the content representations of two modalities in student networks. To handle this issue, we take original RGB and TIR networks as the teachers, and distill their content knowledge into two student networks respectively by the style-content orthogonal feature decoupling scheme. We couple the above two distillation processes in an online optimization framework to form new feature representations of RGB and thermal modalities without modality gap. In addition, we design a masked modeling strategy and a multi-modal candidate token elimination strategy into CKD to improve tracking robustness and efficiency respectively. Extensive experiments on five standard RGBT tracking datasets validate the effectiveness of the proposed method against state-of-the-art methods while achieving the fastest tracking speed of 96.4 FPS.
Andong Lu, Jiacong Zhao, Chenglong Li 0002, Yun Xiao 0003, Bin Luo 0001
ACM Multimedia5
2024 ULDC: Unsupervised Learning-Based Data Cleaning for Malicious Traffic With High Noise
abstract
Abstract Since the traffic of novel attacks exceeds current knowledge, realistic traffic labeling methods are prone to mislabeling, which has a significant impact on machine learning-based intrusion detection systems. Data cleaning typically relies on the ability of supervised deep neural networks to learn correct knowledge. Under high noise conditions, noisy labels can affect a supervised network and render it ineffective. To clean traffic datasets under high noise conditions, we propose an unsupervised learning-based data cleaning framework (called ULDC) that does not rely on labels and powerful supervised networks, hence reducing the impact of noisy labels. ULDC evaluates the confidence of observed labels through the distribution and similarity of samples in low dimensions. Moreover, ULDC maximizes the retention of hard samples through adaptive intra-class threshold evaluation, preserving more hard samples for training and improving generalization. In evaluations of ULDC on the CIRA-CIC-DoHBrw-2020 dataset, the percentage of data correction reached more than 75% under high noise, which is better than that of the state-of-the-art methods. ULDC is applicable to traffic data cleaning in both traditional networks and novel networks such as the Internet of Things and mobile networks, and it has been validated on datasets including CIC-IDS-2017 and IoT-23.
Qingjun Yuan, Yuefei Zhu, Gang Xiong 0001, Yongjuan Wang, Bin Luo 0001, Gaopeng Gou
Comput. J.6
2024 MutualFormer: Multi-modal Representation Learning via Cross-Diffusion Attention
Xixi Wang 0005, Xiao Wang 0014, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
Int. J. Comput. Vis.5
2024 DEGANet: Road Extraction Using Dual-Branch Encoder With Gated Attention Mechanism
abstract
Automatic identification and extraction of roads from high-resolution remote sensing images (RSIs) are important in remote sensing and computer vision. Advancements in remote sensing technology have increased the information in images, making road extraction more challenging. Conventional convolutional methods have limitations, such as loss of spatial details and inadequate fusion of multiscale features. To address these challenges, the letter introduces a novel encoder-decoder architecture called dual-branch encoder with gated attention mechanism network (DEGANet), for extracting road networks in remote sensing image (RSI). First, we propose a multigated informative self-attention (MGSA) module that combines information from dual-branch encoders. By integrating the ResNet and the dynamic snake convolution (DSC) block, which conforms to road shapes, the module emphasizes slender structures similar to roads, thus enhancing the extraction of road features and focusing on capturing more road details. Second, we also introduce the cascade receptive field enhancement (CRFE) module, which optimizes both accuracy and computational complexity. This module combines various receptive field enhancement modules to improve capture long-range dependencies and spatial information perception. Comprehensive experiments conducted on various public remote sensing road datasets demonstrate that our network attains greater segmentation accuracy (intersection over union (IoU) and$F1$score) and connectivity [average path length similarity (APLS)], validating the effectiveness of our proposed method.
Sibao Chen 0001, Lili Huang 0006, Chris Ding, Jin Tang 0001, Bin Luo 0001
IEEE Geosci. Remote. Sens. Lett.6
2024 Few-Shot Object Detection in Remote Sensing Images With Multiscale Spatial Selective Attention
abstract
Few-shot object detection (FSOD) leverages limited labeled data and substantial unlabeled data for detection. However, these approaches mainly target natural images and ignore the spatial relationships and contextual information between objects in remote sensing images (RSIs). To overcome these challenges, this letter introduces a novel method for detecting few-shot objects in RSI. First, we propose a new attention, called multiscale spatial selective attention (MSSSA). This attention spatially selects feature maps from convolution kernels of different scales through spatial selection, focusing the network on the most relevant region of spatial context. Then, our proposed pixel-level feature extractor module (PLFEM) was used in the first stage of FSOD, providing pixel-level object position information to reduce false and missed detection. To evaluate the proposed method, we carry out comprehensive experiments on the DIOR dataset. The results show that the novel class mAP of our method reaches 38.2% in ten shots, an increase of 3.0% compared with the baseline, significantly improving the accuracy of FSOD in RSI.
Yingnan Yu, Sibao Chen 0001, Lili Huang 0006, Jin Tang 0001, Bin Luo 0001
IEEE Geosci. Remote. Sens. Lett.5
2024 DLAReID: double-layer attention network for object re-identification
Sibao Chen 0001, Bin Luo 0001
Multim. Tools Appl.3
2024 SemanticFormer: Hyperspectral image classification via semantic transformer
Xixi Wang 0005, Bo Jiang 0002, Lan Chen 0003, Bin Luo 0001
Pattern Recognit. Lett.5
2024 UAV-Ground Visual Tracking: A Unified Dataset and Collaborative Learning Approach
abstract
Visual tracking from the ground view and the UAV view has received increasing attention due to its wide range of practical applications. These two tasks have strong complementary benefits in the description of the target object, such as detailed appearance in the ground view and global motion information in the UAV view, and their combination has the potential to allow the tracking system to be more robust. However, no work has studied this problem in-depth, and it is challenging to accurately combine the ground view information and the UAV view information. To fill the gap and address the challenge, we propose a new computer vision task called UAV-Ground visual tracking. Considering the lack of relevant data and methods, we first propose a unified video dataset called UGVT, which includes 210 pairs of UAV and ground high-resolution video sequences with a total of more than 204K frames, which can be used as a comprehensive evaluation platform for relevant tracking methods. Secondly, based on the newly constructed dataset, we propose a co-learning method called MvCL to fuse the information of ground and UAV views. It first associates the same tracking target in the two views based on cross-attention operation and then fuses the complementary information of the two views. In particular, as a plug-and-play module based on Transformer structure, this method can be flexibly embedded into different tracking frameworks. Extensive experiments are conducted on the newly created dataset. The results demonstrate the effectiveness of the proposed method in improving the robustness of the tracking system compared with 10 state-of-the-art tracking methods and also indicate the prospect and significance of potential UAV-Ground visual tracking research. The dataset is available at:https://github.com/mmic-lcl/Datasets-and-benchmark-code/.
Dengdi Sun, Leilei Cheng, Chenglong Li 0002, Yun Xiao 0003, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Transformer RGBT Tracking With Spatio-Temporal Multimodal Tokens
abstract
Many RGBT tracking researches primarily focus on modal fusion design, while overlooking the effective handling of target appearance changes. While some approaches have introduced historical frames or fuse and replace initial templates to incorporate temporal information, they have the risk of disrupting the original target appearance and accumulating errors over time. To alleviate these limitations, we propose a novel Transformer RGBT tracking approach, which mixes spatio-temporal multimodal tokens from the static multimodal templates and multimodal search regions in Transformer to handle target appearance changes, for robust RGBT tracking. We introduce independent dynamic template tokens to interact with the search region, embedding temporal information to address appearance changes, while also retaining the involvement of the initial static template tokens in the joint feature extraction process to ensure the preservation of the original reliable target appearance information that prevent deviations from the target appearance caused by traditional temporal updates. We also use attention mechanisms to enhance the target features of multimodal template tokens by incorporating supplementary modal cues, and make the multimodal search region tokens interact with multimodal dynamic template tokens via attention mechanisms, which facilitates the conveyance of multimodal-enhanced target change information. Our module is inserted into the transformer backbone network and inherits joint feature extraction, search-template matching, and cross-modal interaction. Extensive experiments on three RGBT benchmark datasets show that the proposed approach maintains competitive performance compared to other state-of-the-art tracking algorithms while running at 39.1 FPS. The project-related materials are available at:https://github.com/yinghaidada/STMT.
Dengdi Sun, Yajie Pan, Andong Lu, Chenglong Li 0002, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Learning Adaptive Fusion Bank for Multi-Modal Salient Object Detection
abstract
Multi-modal salient object detection (MSOD) aims to boost saliency detection performance by integrating visible sources with depth or thermal infrared ones. Existing methods generally design different fusion schemes to handle certain issues or challenges. Although these fusion schemes are effective at addressing specific issues or challenges, they may struggle to handle multiple complex challenges simultaneously. To solve this problem, we propose a novel adaptive fusion bank that makes full use of the complementary benefits from a set of basic fusion schemes to handle different challenges simultaneously for robust MSOD. We focus on handling five major challenges in MSOD, namely center bias, scale variation, image clutter, low illumination, and thermal crossover or depth ambiguity. The fusion bank proposed consists of five representative fusion schemes, which are specifically designed based on the characteristics of each challenge, respectively. The bank is scalable, and more fusion schemes could be incorporated into the bank for more challenges. To adaptively select the appropriate fusion scheme for multi-modal input, we introduce an adaptive ensemble module that forms the adaptive fusion bank, which is embedded into hierarchical layers for sufficient fusion of different source data. Moreover, we design an indirect interactive guidance module to accurately detect salient hollow objects via the skip integration of high-level semantic information and low-level spatial details. Extensive experiments on three RGBT datasets and seven RGBD datasets demonstrate that the proposed method achieves the outstanding performance compared to the state-of-the-art methods.
Kunpeng Wang 0005, Zhengzheng Tu, Chenglong Li 0002, Cheng Zhang 0010, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 An Oriented Object Detector for Hazy Remote Sensing Images
abstract
Currently, a lot of work is focused on aerial object detection and has achieved good results. Though these methods have achieved promising results on the conventional datasets, it is still challenging to locate objects from the low-quality images captured in adverse weather conditions. Currently, there are limited approaches that combine aerial object detection with hazy conditions, and there are few publicly available datasets for real hazy weather based on aerial images. For this purpose, we propose a dataset HRSI, hazy remote sensing images in the real world, which is mainly divided into three categories: airport, large vehicle, and ship. All images in HRSI are from real hazy conditions. In addition, we propose an object detection model DFENet, a dehazing feature enhancement model for hazy remote sensing images, which is suitable for hazy weather. DFENet consists of a two-branch and a dehazing module. The two-branch structure helps to fully learn hazy and dehazing features. In order to avoid the impact of noise caused by the dehezing module, we also designed a haze-predict module (HPM) to predict the information containing haze in the image. We introduce the cross-fuse module (CFM) to utilize the information of haze to guide the feature fusion of two branches. By utilizing the information of haze, DFENet can dynamically adjust the feature weight in the two-branch to avoid the impact of noise generated by the dehazing module. Compared with traditional object detection methods, DFENet not only has good performance in hazy conditions but also improves performance in clear conditions. We tested DFENet on DOTA, HRSI, and Foggy-DOTA to demonstrate that DFENet performs better under hazy conditions.
Sibao Chen 0001, Jia-Xin Wang, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 DecoupleNet: A Lightweight Backbone Network With Efficient Feature Decoupling for Remote Sensing Visual Tasks
abstract
In the realm of computer vision (CV), balancing speed and accuracy remains a significant challenge. Recent efforts have focused on developing lightweight networks that optimize computational efficiency and feature extraction. However, in remote sensing (RS) imagery, where small and multiscale object detection is critical, these networks often fall short in performance. To address these challenges, DecoupleNet is proposed, an innovative lightweight backbone network specifically designed for RS visual tasks in resource-constrained environments. DecoupleNet incorporates two key modules: the feature integration downsampling (FID) module and the multibranch feature decoupling (MBFD) module. The FID module preserves small object features during downsampling, while the MBFD module enhances small and multiscale object feature representation through a novel decoupling approach. Comprehensive evaluations on three RS visual tasks demonstrate DecoupleNet’s superior balance of accuracy and computational efficiency compared to existing lightweight networks. On the NWPU-RESISC45 classification dataset, DecoupleNet achieves a top-1 accuracy of 95.30%, surpassing FasterNet by 2%, with fewer parameters and lower computational overhead. In object detection tasks using the DOTA 1.0 test set, DecoupleNet records an accuracy of 78.04%, outperforming ARC-R50 by 0.69%. For semantic segmentation on the LoveDA test set, DecoupleNet achieves 53.1% accuracy, surpassing UnetFormer by 0.70%. These findings open new avenues for advancing RS image analysis on resource-constrained devices, addressing a pivotal gap in the field. The code and pretrained models are publicly available athttps://github.com/lwCVer/DecoupleNet.
Wei Lu 0032, Sibao Chen 0001, Qing-Ling Shu, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Prior Guidance and Principal Attention Network for Remote Sensing Image Change Detection
abstract
In the field of remote sensing (RS) image change detection (CD), the conventional encoder-decoder architecture networks often encounter three significant challenges. First, noise in the features extracted from traditional backbone networks leads to blurred boundaries of change objects. Second, upsampling techniques employed in the decoder, such as interpolation or deconvolution, are limited by their finite receptive fields, making it challenging to accurately distinguish pseudo-changes. Furthermore, how to merge encoder and decoder features with possible semantic gaps for the fine-grained details is a topic worth considering. To address these challenges, we introduce a prior guidance (PG) module that effectively aggregates prior high-level features as a semantic guidance map to guide encoder features for the enhancement of boundary detection. In addition, we design a principal attention (PA) module, which aggregates global information from principal regions through sparse operations and adaptively allocates this information to the upsampled and encoder features. This not only addresses the deficiency of global information in the upsampled features but also reduces the semantic gap between the encoder and decoder by establishing channel dependencies. PA does not divert attention to irrelevant regions, demonstrating excellent performance and computational efficiency. By integrating these two modules into our method, a novel PG and PA network (PGPANet) is elaborately designed. A wide range of experiments confirms the validity of our method, showcasing outstanding detection accuracy on three publicly available CD datasets: LEVIR-CD, SYSU-CD, and WHU-CD. The demo code of this work is publicly available athttps://github.com/DaGuangDaGuang/PGPANet.
Qing-Ling Shu, Sibao Chen 0001, Zhi-Hui You, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Diffusion Models and Pseudo-Change: A Transfer Learning-Based Change Detection in Remote Sensing Images
abstract
Remote sensing (RS) image change detection (CD) has been a research hotspot in recent years, which plays an important role in urban planning and disaster assessment. However, since CD labels are difficult to obtain, how to utilize semantic information in RS images to improve the change prediction performance is a problem worth exploring. To solve this problem, we propose a transfer learning-based CD method that utilizes a diffusion generation model to translate high-level semantic information into low-level change information. First, we propose a pseudo-change image pair generation method that utilizes semantic labels to guide the diffusion model to generate change images. Then, the refined loss (RL) is designed to improve the model’s ability to recognize change features based on the difference between pseudo-change image pairs and unlabeled image pairs. Experimental results on WHU-CD, LEVIR-CD, and GoogleGZ-CD datasets show that the proposed method effectively transfers the semantic information into change information and finally improves the model’s feature recognition ability for change objects. Compared with recent CD and transfer learning methods, the proposed transfer learning model (TLM) achieves the best performance. The source code is available athttps://github.com/VCISwang/STCD.
Jia-Xin Wang, Teng Li 0001, Sibao Chen 0001, Cheng-Jie Gu, Zhi-Hui You, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Attention-Aware Sobel Graph Convolutional Network for Remote Sensing Image Change Detection
abstract
In the study of remote sensing images, the problem of change detection (CD) is crucial. Convolutional neural networks (CNNs) are well-liked feature extraction structures that are frequently used in CD. On the other hand, graph convolutional networks (GCNs) are effective in building contextual structure information. Compared with CNN, GCN can make full use of the graph structure information to capture the changing features between different areas in the graph by learning the connections and interactions between nodes. In contrast, traditional pixel-based CNNs may have difficulty modeling semantic relationships and temporal variations among features and are susceptible to noise interference. So in this article, we extract optimization information using a GCN structure. Due to the particularity of remote sensing images, edge information is often ignored, which is useful in the field of CD. In this article, we propose an attention-aware Sobel GCN (ASGCN) for remote sensing image CD. First, we use a Siamese CNN to extract primary multilevel features. Then, a dual-branch attention module (DAM) including coordinate attention and multiscale local attention module (MLAM) is proposed to focus on informative pixels, we use Sobel operator to construct graph, and the graph convolutional module can expand receptive field and extract edge information. Attention fusion module (AFM) is adopted at decoder to perform effective feature fusion. Extensive comparative experiments on three CD datasets, LEVIR-CD, WHU-CD, and DSIFN-CD, verify the effectiveness of the proposed ASGCN.
Lei Wang 0095, Zhi-Hui You, Wei Lu 0032, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Prototype Discriminative Learning for Semi-Supervised Change Detection in Remote Sensing Images
abstract
With the continuous progress of deep learning in remote sensing (RS) visual tasks, considerable advancements have been achieved in RS image change detection (CD). However, prevailing CD methods heavily rely on extensive sets of fully pixelwise hand-annotated training data, a time-consuming and costly process, and they fail to fully harness the potential benefits of deep feature representations within the deep feature domain. To tackle the mentioned issues, we propose a novel semi-supervised CD method called PDLCD, which strategically leverages useful information from massive unlabeled data to complement labeled data with just a few samples. Specifically, changed objects and unchanged backgrounds of bitemporal RS images are various and complex, our approach advocates dividing each category into multiple subclasses in the deep feature domain. In this scheme, the high-level feature of each subclass follows a Gaussian distribution. Then, the prototype discriminative learning (PDL) is introduced to explicitly encourage deep features of samples closer to the nearest prototype within their respective category, and away from all prototypes of other categories. We design feature discriminative loss (FDL) to implement PDL for constructing more pronounced intraclass compactness and interclass variability. Finally, we compute the supervised loss based on a limited set of labeled data, incorporate the unsupervised loss leveraging a substantial volume of unlabeled data, and include FDL within the deep feature domain to collectively optimize the model. Extensive experiments carried out on three challenging RS image CD datasets illustrate that our proposed semi-supervised CD method obtains better CD performance than previous counterparts. The source code is available at:https://github.com/Youzhihui/PDLCD.
Zhi-Hui You, Sibao Chen 0001, Jia-Xin Wang, Chris Ding, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Multi-Granularity Part Sampling Attention for Fine-Grained Visual Classification
abstract
Fine-grained visual classification aims to classify similar sub-categories with the challenges of large variations within the same sub-category and high visual similarities between different sub-categories. Recently, methods that extract semantic parts of the discriminative regions have attracted increasing attention. However, most existing methods extract the part features via rectangular bounding boxes by object detection module or attention mechanism, which makes it difficult to capture the rich shape information of objects. In this paper, we propose a novel Multi-Granularity Part Sampling Attention (MPSA) network for fine-grained visual classification. First, a novel multi-granularity part retrospect block is designed to extract the part information of different scales and enhance the high-level feature representation with discriminative part features of different granularities. Then, to extract part features of various shapes at each granularity, we propose part sampling attention, which can sample the implicit semantic parts on the feature maps comprehensively. The proposed part sampling attention not only considers the importance of sampled parts but also adopts the part dropout to reduce the overfitting issue. In addition, we propose a novel multi-granularity fusion method to highlight the foreground features and suppress the background noises with the assistance of the gradient class activation map. Experimental results demonstrate that the proposed MPSA achieves state-of-the-art performance on four commonly used fine-grained visual classification benchmarks. The source code is publicly available at https://github.com/mobulan/MPSA.
Bo Jiang 0002, Bin Luo 0001, Jinhui Tang 0001
IEEE Trans. Image Process.4
2024 Rethinking Batch Sample Relationships for Data Representation: A Batch-Graph Transformer Based Approach
abstract
Exploring sample relationships within each mini-batch has shown great potential for learning image representations. Existing works generally adopt the regular Transformer to model the visual content relationships, ignoring the cues of semantic/label correlations between samples. Also, they generally adopt the ‘full’ self-attention mechanism which are obviously redundant and also sensitive to the noisy samples. To overcome these issues, in this paper, we design a simple yet flexible Batch-Graph Transformer (BGFormer) for mini-batch sample representations by deeply capturing the relationships of image samples from both visual and semantic perspectives. BGFormer has three main aspects. (1) It employs a flexible graph model, termedBatch Graphto jointly encode both visual and semantic relationships of samples within each mini-batch. (2) It explores the neighborhood relationships of samples by borrowing the idea of sparse graph representation which thus performs robustly, w.r.t., noisy samples. (3) It devises a novel specific Transformer architecture that mainly adoptsdualstructure-constrained self-attention (SSA), together with graph normalization, FFN, etc, to carefully exploit the batch graph information for sample tokens (nodes) representations. As an application, we apply BGFormer to the metric learning tasks. Extensive experiments on four popular datasets demonstrate the effectiveness of the proposed model.
Xixi Wang 0005, Bo Jiang 0002, Xiao Wang 0014, Jinhui Tang 0001, Bin Luo 0001
IEEE Trans. Multim.5
2024 Alignment-Free RGBT Salient Object Detection: Semantics-Guided Asymmetric Correlation Network and a Unified Benchmark
abstract
RGB and Thermal (RGBT) Salient Object Detection (SOD) aims to achieve high-quality saliency prediction by exploiting the complementary information of visible and thermal image pairs, which are initially captured in an unaligned manner. However, existing methods are tailored for manually aligned image pairs, which are labor-intensive, and directly applying these methods to original unaligned image pairs could significantly degrade their performance. In this paper, we make the first attempt to address RGBT SOD for initially captured RGB and thermal image pairs without manual alignment. Specifically, we propose a Semantics-guided Asymmetric Correlation Network (SACNet) that consists of two novel components: 1) an asymmetric correlation module utilizing semantics-guided attention to model cross-modal correlations specific to unaligned salient regions; 2) an associated feature sampling module to sample relevant thermal features according to the corresponding RGB features for multi-modal feature integration. In addition, we construct a unified benchmark dataset called UVT2000, containing 2000 RGB and thermal image pairs directly captured from various real-world scenes without any alignment, to facilitate research on alignment-free RGBT SOD. Extensive experiments on both aligned and unaligned datasets demonstrate the effectiveness and superior performance of our method.
Kunpeng Wang 0005, Danying Lin, Chenglong Li 0002, Zhengzheng Tu, Bin Luo 0001
IEEE Trans. Multim.5
2024 Tiny Object Tracking: A Large-Scale Dataset and a Baseline
abstract
Tiny objects, frequently appearing in practical applications, have weak appearance and features, and receive increasing interests in many vision tasks, such as object detection and segmentation. To promote the research and development of tiny object tracking, we create a large-scale video dataset, which contains 434 sequences with a total of more than 217K frames. Each frame is carefully annotated with a high-quality bounding box. In data creation, we take 12 challenge attributes into account to cover a broad range of viewpoints and scene complexities, and annotate these attributes for facilitating the attribute-based performance analysis. To provide a strong baseline in tiny object tracking, we propose a novel multilevel knowledge distillation network (MKDNet), which pursues three-level knowledge distillations in a unified framework to effectively enhance the feature representation, discrimination, and localization abilities in tracking tiny objects. Extensive experiments are performed on the proposed dataset, and the results prove the superiority and effectiveness of MKDNet compared with state-of-the-art methods. The dataset, the algorithm code, and the evaluation code are available at https://github.com/mmic-lcl/Datasets-and-benchmark-code.
Yabin Zhu, Chenglong Li 0002, Xiao Wang 0014, Jin Tang 0001, Bin Luo 0001, Zhixiang Huang
IEEE Trans. Neural Networks Learn. Syst.6
2023 Semantic-Aware Dual Contrastive Learning for Multi-Label Image Classification
abstract
Extracting image semantics effectively and assigning corresponding labels to multiple objects or attributes for natural images is challenging due to the complex scene contents and confusing label dependencies. Recent works have focused on modeling label relationships with graph and understanding object regions using class activation maps (CAM). However, these methods ignore the complex intra- and inter-category relationships among specific semantic features, and CAM is prone to generate noisy information. To this end, we propose a novel semantic-aware dual contrastive learning framework that incorporates sample-to-sample contrastive learning (SSCL) as well as prototype-to-sample contrastive learning (PSCL). Specifically, we leverage semantic-aware representation learning to extract category-related local discriminative features and construct category prototypes. Then based on SSCL, label-level visual representations of the same category are aggregated together, and features belonging to distinct categories are separated. Meanwhile, we construct a novel PSCL module to narrow the distance between positive samples and category prototypes and push negative samples away from the corresponding category prototypes. Finally, the discriminative label-level features related to the image content are accurately captured by the joint training of the above three parts. Experiments on five challenging large-scale public datasets demonstrate that our proposed method is effective and outperforms the state-of-the-art methods. Code and supplementary materials are released on https://github.com/yu-gi-oh-leilei/SADCL.
Leilei Ma 0002, Dengdi Sun, Lei Wang 0095, Haifeng Zhao 0001, Bin Luo 0001
ECAI5
2023 GENet: Guidance Enhancement Network for 3D Shape Recognition
abstract
Both point cloud-based and view-based deep learning methods for 3D shape recognition have achieved relatively remarkable results in recent years. However, there are few methods to jointly represent 3D shapes from both point cloud and multi-view modal data. Therefore, we propose a guidance enhancement network (GENet) for 3D shape recognition based on multimodal data. On the one hand, the point cloud is encoded with features from both explicit and implicit aspects, and on the other hand, all views are encoded and constructed as a graph. In the multilayer guidance enhancement module, graph convolutional neural network (GCN) enhances each view feature, and then temporary high-level features (initially point cloud global feature) guide multiple low-level view features to obtain correlation coefficients, through which the views with higher importance are filtered as inputs for the next layer of the structure and the view features in the current layer are weighted and aggregated. The aggregated view features are then connected to the high-level features with residuals to form the enhanced high-level features. The 3D shape descriptor is finally obtained after several guidance and enhancements. The proposed GENet achieves state-of-the-art results on the 3D benchmark dataset ModelNet.
Xiaofeng Wang 0009, Qingzhe Cui, Lixiang Xu, Haifeng Liu 0004, Lixin He, Bin Luo 0001, Sibao Chen 0001, Yuan Yan Tang
IJCNN6
2023 GLCNet: Global-Local Complementary Network for 3D Shape Recognition
abstract
Both point cloud-based and multi-view-based methods have achieved remarkable results in 3D shape recognition, yet there are few methods that combine the two types of data. In this paper, a novel Global-Local Complementary Network (GLCNet) based on multimodal data is proposed. The network obtains more powerful shape descriptors by stacking multiple layers of Global-Local Complementary Module (GLC Module). More specifically, the Global-Local Relation Score Module is first used to obtain the relationship between view features and global feature. The relationship is then utilized to facilitate the aggregation of view features and to filter out the more important ones. Finally, the aggregated view features are fused with the global features to form a stronger global feature. GLCNet enables the characteristics of various data to be fully utilized and achieves a true sense of complementarity of strengths and weaknesses. Extensive experiments on the benchmark dataset ModelNet show that GLCNet achieves state-of-the-art results in 3D shape classification and retrieval.
Xiaofeng Wang 0009, Qingzhe Cui, Lixiang Xu, Haifeng Liu 0004, Lixin He, Bin Luo 0001, Sibao Chen 0001, Yuan Yan Tang
IJCNN6
2023 UGTransformer: Unsupervised Graph Transformer Representation Learning
abstract
This paper mainly studies graph representation learning in unsupervised scenarios combined with Transformer models. Transformer network models have been widely used in many fields of machine learning and deep learning, and the application of transformer architectures to graph data has been very popular recently. For graph data, the field of graph representation learning has recently attracted a lot of attention. Graph-level representation is widely used in the real world, such as drug molecule design and disease classification in biochemistry. Traditional graph kernel methods, which design different graph kernels for different substructures, are simple but have poor generalization performance. Recently methods based on language models, such as graph2vec, use a particular substructure as the graph representation, which is also similar to the hand-crafted approach and also leads to poor generalization ability. In this paper, we propose the UGTransformer model, which builds on the standard Transformer architecture. We introduce several simple and effective structural encoding methods in order to encode the structural information of the graph into the model efficiently. The unsupervised representation of graphs is learned through a multi-headed attention mechanism and by using powerful aggregation functions. We conducted experiments on a benchmark date set for graph classification, and the experimental results validate the effectiveness of our proposed model.
Lixiang Xu, Haifeng Liu 0004, Qingzhe Cui, Bin Luo 0001, Yan Chen 0037, Yuan Yan Tang
IJCNN4
2023 Imperceptible Adversarial Attack on S Channel of HSV Colorspace
abstract
Deep neural network models are vulnerable to subtle but adversarial perturbations that alter the model. Adversarial perturbations are typically computed for RGB images and, therefore, are evenly distributed among RGB channels. Compared with RGB images, HSV images can express the Hue, saturation, and brightness more intuitively. We find that the adversarial perturbation in the S-channel ensures a high attack success rate, while the perturbation is small, and the visual quality of the adversarial examples is good. Using this finding, we propose an attack method, SPGD, to improve the visual quality of adversarial examples by generating perturbations on the S-channel. Based on the attack principle of the PGD method, the RGB image was converted into an HSV image. The gradient calculated by the model on the S channel was superimposed on the S channel and then combined with the non-interference H and V channels to convert back to the RGB image. The iteration stops until the attack succeed. We compare the SPGD method with the existing state-of-the-art attack methods. The results show that SPGD minimizes pixel perturbation while maintaining a high attack success rate and achieves the best results in terms of structural similarity, imperceptibility, the minimum number of iterations, and the shortest run time.
Zhao-Xia Yin, Jiefei Zhang, Bin Luo 0001
IJCNN5
2023 Multimodal salient object detection via adversarial learning with collaborative generator
Zhengzheng Tu, Wenfang Yang, Kunpeng Wang 0005, Amir Hussain 0001, Bin Luo 0001, Chenglong Li 0002
Eng. Appl. Artif. Intell.5
2023 Robust image steganography against lossy JPEG compression based on embedding domain selection and adaptive error correction
Xiaolong Duan, Bin Li 0011, Zhao-Xia Yin, Xinpeng Zhang 0001, Bin Luo 0001
Expert Syst. Appl.5
2023 A super-resolution-based license plate recognition method for remote surveillance
Sen Pan, Sibao Chen 0001, Bin Luo 0001
J. Vis. Commun. Image Represent.3
2023 Road Extraction by Multiscale Deformable Transformer From Remote Sensing Images
abstract
Rapid progress has been made in the research of high-resolution remote sensing road extraction tasks in the past years, but due to the diversity of road types and the complexity of road context, extracting the perfect road network is still fraught with difficulties and challenges. Many Convolutional Neural Networks (CNNs) based on encoder-decoder structures have demonstrated their effectiveness. Transformer’s self-attention mechanism shows more powerful performance than CNNs in modeling global feature dependencies. In this paper, we propose a Multi-scale Deformable Transformer Network (MDTNet) based on encoder-decoder structure to extract road networks from remote sensing images. The core of MDTNet is our proposed Multi-scale Deformable Self-Attention (MDSA) mechanism. MDSA can capture more comprehensive features than conventional self-attention. In addition, roads are not present in certain blocks of areas like other objects, but are interwoven throughout the image in such a long, linear fashion that information about certain road segments may be overlooked. To minimize residual errors in road segmentations, our MDSA incorporates a deformable design on feature maps, which effectively enhances the salience of road features relative to their surroundings. Extensive experiments on several public remote sensing road datasets show that our MDTNet achieves higher segmentation [F1 score and Intersection over Union (IoU)] and connectivity [Average Path Length Similarity (APLS)] accuracy, which verifies the effectiveness of our approach.
Pengcheng Hu 0001, Sibao Chen 0001, Lili Huang 0006, Guizhou Wang, Jin Tang 0001, Bin Luo 0001
IEEE Geosci. Remote. Sens. Lett.6
2023 Hypergraph convolutional network for hyperspectral image classification
Bo Jiang 0002, Jinpei Liu, Bin Luo 0001
Neural Comput. Appl.5
2023 DropAGG: Robust Graph Neural Networks via Drop Aggregation
Bo Jiang 0002, Beibei Wang 0006, Haiyun Xu, Bin Luo 0001
Neural Networks5
2023 Graph Neural Network Meets Sparse Representation: Graph Sparse Neural Networks via Exclusive Group Lasso
abstract
Existing GNNs usually conduct the layer-wise message propagation via the 'full' aggregation of all neighborhood information which are usually sensitive to the structural noises existed in the graphs, such as incorrect or undesired redundant edge connections. To overcome this issue, we propose to exploit Sparse Representation (SR) theory into GNNs and propose Graph Sparse Neural Networks (GSNNs) which conduct sparse aggregation to select reliable neighbors for message aggregation. GSNNs problem contains discrete/sparse constraint which is difficult to be optimized. Thus, we then develop a tight continuous relaxation model Exclusive Group Lasso GNNs (EGLassoGNNs) for GSNNs. An effective algorithm is derived to optimize the proposed EGLassoGNNs model. Experimental results on several benchmark datasets demonstrate the better performance and robustness of the proposed EGLassoGNNs model.
Bo Jiang 0002, Beibei Wang 0006, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Generalizing Aggregation Functions in GNNs: Building High Capacity and Robust GNNs via Nonlinear Aggregation
abstract
The main aspect powering GNNs is the multi-layer network architecture to learn the nonlinear representation for graph learning task. The core operation in GNNs is the message propagation in which each node updates its information by aggregating the information from its neighbors. Existing GNNs usually adopt either linear neighborhood aggregation (e.g. mean, sum) or max aggregator in their message propagation. 1) For linear aggregators, the whole nonlinearity and network's capacity of GNNs are generally limited because deeper GNNs usually suffer from the over-smoothing issue due to their inherent information propagation mechanism. Also, linear aggregators are usually vulnerable to the spatial perturbations. 2) For max aggregator, it usually fails to be aware of the detailed information of node representations within neighborhood. To overcome these issues, we re-think the message propagation mechanism in GNNs and develop the new general nonlinear aggregators for neighborhood information aggregation in GNNs. One main aspect of our nonlinear aggregators is that they all provide the optimally balanced aggregator between max and mean/sum aggregators. Thus, they can inherit both i) high nonlinearity that enhances network's capacity, robustness and ii) detail-sensitivity that is aware of the detailed information of node representations in GNNs' message propagation. Promising experiments show the effectiveness, high capacity and robustness of the proposed methods.
Beibei Wang 0006, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Sparse norm regularized attribute selection for graph neural networks
Bo Jiang 0002, Beibei Wang 0006, Bin Luo 0001
Pattern Recognit.3
2023 Reversible attack based on adversarial perturbation and reversible data hiding in YUV colorspace
Zhao-Xia Yin, Bin Luo 0001
Pattern Recognit. Lett.4
2023 Combinatorial online high-order interactive feature selection based on dynamic graph convolution network
Wen-Bin Wu, Jun-Jun Sun, Sibao Chen 0001, Chris Ding, Bin Luo 0001
Signal Process.5
2023 Channel Attention TextCNN with Feature Word Extraction for Chinese Sentiment Analysis
abstract
Chinese short text sentiment analysis can help understand society’s views on various hot topics. Many existing sentiment analysis methods are based on sentiment dictionaries. Still, sentiment dictionaries are easily affected by subjective factors. They require a lot of time to build as well as maintenance to prevent obsolescence. For the aim of extracting rich information within texts more effectively, we propose a Channel Attention TextCNN with Feature Word Extraction model (CAT-FWE). The feature word extraction module helps us choose words that affect the sentiment of reviews. Then, these words are integrated with multi-level semantic information to enhance the information of sentences. In addition, the channel attention textCNN module that is a promotion of traditional TextCNN tends to pay more attention to those meaningful features. It eliminates the impacts of features that do not make any sense effectively. We apply our CAT-FWE model to both fine-grained classification and binary classification tasks for Chinese short texts. Experiment results show that it can improve the performance of emotion recognition.
Jiangwei Liu, Zian Yan, Sibao Chen 0001, Xiao Sun 0003, Bin Luo 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2023 Few-Shot Learning Meets Transformer: Unified Query-Support Transformers for Few-Shot Classification
abstract
The goal of Few-shot classification (FSL) is to identify unseen classes with very limited samples has attracted more and more attention. Usually, it is formulated as a metric learning problem. The core issue of few-shot classification is how to learn (1) consistent representations for images in both support and query sets and (2) effective metric learning for images between support and query sets. In this paper, we show that the two challenges can be well modeled simultaneously via a unified Query-Support TransFormer (QSFormer) model. To be specific, the proposed QSFormer involves global query-support sample Transformer (sampleFormer) branch and local patch Transformer (patchFormer) learning branch. sampleFormer aims to capture the dependence of samples in support and query sets for image representation. It adopts the Encoder, QS-Decoder and Cross-Attention to respectively model the Support, Query (image) representation and Metric learning for few-shot classification task. Also, as a complementary to global learning branch, we adopt a local patch Transformer to extract structural representation for each image sample by capturing the long-range dependence of local image patches. In addition, we introduce a novel Cross-scale Interactive Feature Extractor (CIFE) to extract and fuse different scale CNN features as an effective backbone module for the proposed few-shot learning method. We integrate these into a unified framework and train it in an end-to-end way. A large number of experiments are conducted on four popular datasets to validate the superiority and effectiveness of the proposed QSFormer.
Xixi Wang 0005, Xiao Wang 0014, Bo Jiang 0002, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 VcT: Visual Change Transformer for Remote Sensing Image Change Detection
abstract
Given two remote sensing images, the goal of visual change detection task is to detect significantly changed areas between them. Existing visual change detectors usually adopt CNNs or Transformers for feature representation learning and focus on learning effective representation for the changed regions between images. Although good performance can be obtained by enhancing the features of the change regions, however, these works are still limited mainly due to the ignorance of mining the unchanged background context information. It is known that one main challenge for change detection is how to obtain the consistent representations for two images involving different variations, such as spatial variation, sunlight intensity, etc. In this work, we demonstrate that carefully mining the common background information provides an important cue to learn the consistent representations for the two images which thus obviously facilitates the visual change detection problem. Based on this observation, we propose a novel Visual change Transformer (VcT) model for visual change detection problem. To be specific, a shared backbone network is first used to extract the feature maps for the given image pair. Then, each pixel of feature map is regarded as a graph node and the graph neural network is proposed to model the structured information for coarse change map prediction. Top-K reliable tokens can be mined from the map and refined by using the clustering algorithm. Then, these reliable tokens are enhanced by first utilizing self/cross-attention schemes and then interacting with original features via an anchor-primary attention learning module. Finally, the prediction head is proposed to get a more accurate change map. Extensive experiments on multiple benchmark datasets validated the effectiveness of our proposed VcT model. The source code and pre-trained models are available at https://github.com/Event-AHU/VcT_Remote_Sensing_Change_Detection.
Bo Jiang 0002, Zitian Wang, Xixi Wang 0005, Lan Chen 0003, Xiao Wang 0014, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 A Robust Feature Downsampling Module for Remote-Sensing Visual Tasks
abstract
Remote sensing (RS) images present unique challenges for computer vision due to lower resolution, smaller objects, and fewer features. Mainstream backbone networks show promising results for traditional visual tasks. However, they use convolution to reduce feature map dimensionality, which can result in information loss for small objects in RS images and decreased performance. To address this problem, we propose a new and universal downsampling module named Robust Feature Downsampling (RFD). RFD fuses multiple feature maps extracted by different downsampling techniques, creating a more robust feature map with a complementary set of features. Leveraging this, we overcome the limitations of conventional convolutional downsampling, resulting in more accurate and robust analysis of RS images. We develop two versions of RFD module, Shallow RFD (SRFD) and Deep RFD (DRFD), tailored to adapt to different stages of feature capture and improve feature robustness. We replace the downsampling layers of existing mainstream backbones with RFD module and conduct comparative experiments on several public RS image datasets. The results show significant improvements compared to baseline approaches in RS image classification, object detection, and semantic segmentation. Specifically, our RFD module achieved an average performance gain of 1.5% on NWPU-RESISC45 classification dataset without utilizing any additional pretraining data, resulting in state-of-the-art performance on this dataset. Moreover, in detection and segmentation tasks on DOTA and iSAID datasets, our RFD module outperforms the baseline approaches by 2-7% when utilizing pretraining data from NWPU-RESISC45. These results highlight the value of RFD module in enhancing the performance of RS visual tasks.
Wei Lu 0032, Sibao Chen 0001, Jin Tang 0001, Chris Ding, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Boosting Semantic Segmentation of Aerial Images via Decoupled and Multilevel Compaction and Dispersion
abstract
Semantic segmentation is a valuable task in practical applications for aerial images. Nevertheless, the segmentation performance is unsatisfactory due to aerial images’ huge intra-class variance and inter-class similarity. To solve this problem, we propose an approach to increase the distinction between classes and compact the features of the same class. Specifically, since a single aerial image contains only a small number of categories, which is fatal for previous contrastive learning, we discard InfoNCE loss in contrastive learning and use the simple Mean Square Error (MSE) loss that does not require negative samples to decouple the dispersion and compaction operations. Besides, we set up more representative prototypes for classes and extend the prototypes to the whole dataset level, which we call image- and dataset-level prototypes. Based on the calculated prototypes, we propose Multi-level intra-class Feature Compaction (MFC) and Multi-level inter-class Feature Dispersion (MFD) to compact the features of the same class and disperse the features of different classes in the latent feature space. More importantly, some measures are proposed to ensure the two do not conflict. MFC and MFD can be applied to any existing segmentation network to improve performance significantly without increasing computational complexity during inference. Moreover, we feed the calculated multi-level prototypes directly into the classifier, thus keeping the feature extraction and classifier consistent. Results on four challenging datasets, Deepglobe, iSAID, Potsdam, and Vaihingen, demonstrate the significant effect of our method, and sufficient ablation studies verify the role of each module.
Lianlei Shan, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.4
2023 Crossed Siamese Vision Graph Neural Network for Remote-Sensing Image Change Detection
abstract
The development of deep learning in remote sensing (RS) visual tasks has led to remarkable progress in RS image change detection (CD). However, RS bi-temporal images cover complex and confusing scenes due to natural environmental factors, which presents challenges for CD task. How to effectively exploit long-range dependencies and sensitively discriminate real-changes with various scales from pseudo-changes are urgent problems. It is especially obvious for the changes of building structures man-made. This paper presents a CD approach named CSViG, which utilizes Siamese Vision Graph neural network (SViG) with crossed feature fusion. SViG acts as a feature extractor to capture richer short- and long-range dependencies. Crossed feature fusion consists of a horizontal feature fusion module (HFFM) and a vertical feature fusion module (VFFM). HFFM designs cross-concatenation (CC) way to reveal real-changes from pseudo-change in the same horizontal stage, after which global and local features are extracted by using attention mechanism and multi-scale depth-wise separable convolution. VFFM further fuses complementary content from vertical multiple stages to effectively represent change regions of different sizes (tiny or huge) by using attention mechanism. Extensive comparative experiments conducted on three available building change detection datasets demonstrate that the proposed method achieves better CD performance than previous counterparts.
Zhi-Hui You, Jia-Xin Wang, Sibao Chen 0001, Chris Ding, Guizhou Wang, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 Looking and Hearing Into Details: Dual-Enhanced Siamese Adversarial Network for Audio-Visual Matching
abstract
Audio-visual cross-modal matching aims to explore the intrinsic correspondence between face images and audio clips. Existing methods usually focus on the salient features of identities between visual images and voice clips, while neglecting their subtle differences, which are crucial to distinguishing cross-modal samples. To deal with this problem, we propose a novel Dual-enhanced Siamese Adversarial Network (DSANet), which pursues the adversarial dual enhancement to highlight both salient and subtle features for robust audio-visual cross-modal matching. First, we designed a dual enhancement mechanism to enhance potential subtle features by randomly selecting a region feature for salient feature suppression, while enhancing salient features in the corresponding region to ensure the global discriminative ability. Second, to establish the correlation of subtle features in the process of eliminating cross-modal heterogeneity, we design a siamese adversarial structure to perform modal heterogeneity elimination for both enhanced salient and subtle features in a parallel manner. Moreover, we propose an adaptive masked cross-entropy loss to force the network to focus on the feature differences among hard classes. Experiments on public benchmark datasets validate the effectiveness of the proposed algorithm.
Jiaxiang Wang 0001, Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.5
2023 Fine-Grained Visual Classification via Internal Ensemble Learning Transformer
abstract
Recently, vision transformers (ViTs) have been investigated in fine-grained visual recognition (FGVC) and are now considered state of the art. However, most ViT-based works ignore the different learning performances of the heads in the multi-head self-attention (MHSA) mechanism and its layers. To address these issues, in this paper, we propose a novel internal ensemble learning transformer (IELT) for FGVC. The proposed IELT involves three main modules: multi-head voting (MHV) module, cross-layer refinement (CLR) module, and dynamic selection (DS) module. To solve the problem of the inconsistent performances of multiple heads, we propose the MHV module, which considers all of the heads in each layer as weak learners and votes for tokens of discriminative regions as cross-layer feature based on the attention maps and spatial relationships. To effectively mine the cross-layer feature and suppress the noise, the CLR module is proposed, where the refined feature is extracted and the assist logits operation is developed for the final prediction. In addition, a newly designed DS module adjusts the token selection number at each layer by weighting their contributions of the refined feature. In this way, the idea of ensemble learning is combined with the ViT to improve fine-grained feature representation. The experiments demonstrate that our method achieves competitive results compared with the state of the art on five popular FGVC datasets. Source code has been released and can be found athttps://github.com/mobulan/IELT.
Bo Jiang 0002, Bin Luo 0001
IEEE Trans. Multim.4
2023 GPENs: Graph Data Learning With Graph Propagation-Embedding Networks
abstract
Compact representation of graph data is a fundamental problem in pattern recognition and machine learning area. Recently, graph neural networks (GNNs) have been widely studied for graph-structured data representation and learning tasks, such as graph semi-supervised learning, clustering, and low-dimensional embedding. In this article, we present graph propagation-embedding networks (GPENs), a new model for graph-structured data representation and learning problem. GPENs are mainly motivated by 1) revisiting of traditional graph propagation techniques for graph node context-aware feature representation and 2) recent studies on deeply graph embedding and neural network architecture. GPENs integrate both feature propagation on graph and low-dimensional embedding simultaneously into a unified network using a novel propagation-embedding architecture. GPENs have two main advantages. First, GPENs can be well-motivated and explained from feature propagation and deeply learning architecture. Second, the equilibrium representation of the propagation-embedding operation in GPENs has both exact and approximate formulations, both of which have simple closed-form solutions. This guarantees the compactivity and efficiency of GPENs. Third, GPENs can be naturally extended to multiple GPENs (M-GPENs) to address the data with multiple graph structures. Experiments on various semi-supervised learning tasks on several benchmark datasets demonstrate the effectiveness and benefits of the proposed GPENs and M-GPENs.
Bo Jiang 0002, Leiling Wang, Jian Cheng 0001, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 Model Compression Based on Differentiable Network Channel Pruning
abstract
Although neural networks have achieved great success in various fields, applications on mobile devices are limited by the computational and storage costs required for large models. The model compression (neural network pruning) technology can significantly reduce network parameters and improve computational efficiency. In this article, we propose a differentiable network channel pruning (DNCP) method for model compression. Unlike existing methods that require sampling and evaluation of a large number of substructures, our method can efficiently search for optimal substructure that meets resource constraints (e.g., FLOPs) through gradient descent. Specifically, we assign a learnable probability to each possible number of channels in each layer of the network, relax the selection of a particular number of channels to a softmax over all possible numbers of channels, and optimize the learnable probability in an end-to-end manner through gradient descent. After the network parameters are optimized, we prune the network according to the learnable probability to obtain the optimal substructure. To demonstrate the effectiveness and efficiency of DNCP, experiments are conducted with ResNet and MobileNet V2 on CIFAR, Tiny ImageNet, and ImageNet datasets.
Yu-Jie Zheng, Sibao Chen 0001, Chris Ding, Bin Luo 0001
IEEE Trans. Neural Networks Learn. Syst.4
2022 Progressive Attribute Embedding for Accurate Cross-modality Person Re-ID
abstract
Attributes are important information to bridge the appearance gap across modalities, but have not been well explored in cross-modality person ReID. This paper proposes a progressive attribute embedding module (PAE) to effectively fuse the fine-grained semantic attribute information and the global structural visual information. Through a novel cascade way, we use attribute information to learn the relationship between the person images in different modalities, which significantly relieves the modality heterogeneity. Meanwhile, by embedding attribute information to guide more discriminative image feature generation, it simultaneously reduces the inter-class similarity and the intra-class discrepancy. In addition, we propose an attribute-based auxiliary learning strategy (AAL) to supervise the network to learn modality-invariant and identity-specific local features by joint attribute and identity classification losses. The PAE and AAL are jointly optimized in an end-to-end framework, namely, progressive attribute embedding network (PAENet). One can plug PAE and AAL into current mainstream models, as we implement them in five cross-modality person ReID frameworks to further boost the performance. Extensive experiments on public datasets demonstrate the effectiveness of the proposed method against the state-of-the-art cross-modality person ReID methods.
Aihua Zheng, Chenglong Li 0002, Bin Luo 0001, Ruoran Jia
ACM Multimedia5
2022 Universal adversarial perturbation for remote sensing images
abstract
Recently, with the application of deep learning in the remote sensing image (RSI) field, the classification accuracy of the RSI has been dramatically improved compared with traditional technology. However, even the state-of-the-art object recognition convolutional neural networks are fooled by the universal adversarial perturbation (UAP). The research on UAP is mostly limited to ordinary images, and RSIs have not been studied. To explore the basic characteristics of UAPs of RSIs, this paper proposes a novel method combining an encoder-decoder network with an attention mechanism to generate the UAP of RSIs. Firstly, the former is used to generate the UAP, which can learn the distribution of perturbations better, and then the latter is used to find the sensitive regions concerned by the RSI classification model. Finally, the generated regions are used to fine-tune the perturbation making the model misclassified with fewer perturbations. The experimental results show that the UAP can make the classification model misclassify, and the attack success rate of our proposed method on the RSI data set is as high as 97.09%.
Guorui Feng, Zhao-Xia Yin, Bin Luo 0001
MMSP4
2022 EllipseIoU: A General Metric for Aerial Object Detection
Xinbo Yang, Chenglong Li 0002, Rui Ruan, Lei Liu 0049, Bin Luo 0001
PRCV (3)6
2022 Attributes Based Visible-Infrared Person Re-identification
Aihua Zheng, Mengya Feng, Bo Jiang 0002, Bin Luo 0001
PRCV (1)5
2022 Effects of haze and dehazing on deep learning-based vision models
Haseeb Hassan, Pranshu Mishra, Muhammad Ahmad 0002, Ali Kashif Bashir, Bingding Huang, Bin Luo 0001
Appl. Intell.6
2022 PISA: Pixel skipping-based attentional black-box adversarial attack
Jie Wang 0050, Zhao-Xia Yin, Jing Jiang 0021, Jin Tang 0001, Bin Luo 0001
Comput. Secur.5
2022 SIECP: Neural Network Channel Pruning based on Sequential Interval Estimation
Sibao Chen 0001, Yu-Jie Zheng, Chris Ding, Bin Luo 0001
Neurocomputing4
2022 RGBT tracking based on cooperative low-rank graph model
Longfeng Shen, Xiaoxiao Wang 0003, Lei Liu 0049, Bin Hou, Yulei Jian, Jin Tang 0001, Bin Luo 0001
Neurocomputing7
2022 DBRANet: Road Extraction by Dual-Branch Encoder and Regional Attention Decoder
abstract
Although widely exploited in recent decades, road extraction is still a very significant and challenging research in the field of remote sensing image processing due to the complex background and road distribution. Among the existing CNN-based methods, U-shape architectures composed of encoders and decoders have shown their effectiveness. In this letter, we propose an improved encoder–decoder method, named DBRANet, for extracting roads from remote sensing images. In the encoding phase, we present a dual-branch network module (DBNM) to construct more effective features, thus improving the fusion feature maps of different scales. One branch utilizes the residual block, and the other branch utilizes the refined asymmetric block, which effectively increases the feature extraction capability of the backbone. In the decoding phase, considering the sinuous shape and the unbalanced distribution of roads in remote sensing images, we design a novel attention module, named the regional attention network module (RANM), to automatically learn the importance of each channel according to the regional information. Extensive experiments on several public remote sensing road data sets show that our DBRANet achieves higher segmentation [$F1$score and Intersection over Union (IoU)] and connectivity [average path length similarity (APLS)] accuracy, which verifies the effectiveness of our approach.
Sibao Chen 0001, Yu-Xin Ji, Jin Tang 0001, Bin Luo 0001, Weiqiang Wang 0001, Ke Lu 0002
IEEE Geosci. Remote. Sens. Lett.4
2022 BDTNet: Road Extraction by Bi-Direction Transformer From Remote Sensing Images
abstract
The past several years have witnessed the rapid development of the task of road extraction in high-resolution remote sensing images. However, due to the complex background and road distribution, road extraction is still a challenging research in remote sensing images. In convolutional neural networks (CNNs), the U-shaped architecture network has shown its effectiveness. But the global representation cannot be captured effectively by CNNs. While in the transformer, the self-attention (SA) module can capture the long-distance feature dependencies. A hybrid encoder-decoder method called BDTNet is proposed in this letter, which enhance the extraction of global and local information in remote sensing images. Firstly, feature maps of different scales are obtained through the backbone network. And then, on the basis of reducing the computational cost of self-attention, the Bi-Direction Transformer Module (BDTM) is constructed to capture the contextual road information in feature maps of different scales. Finally, the Feature Refinement Module (FRM) is introduced to integrate the features extracted from the backbone network and BDTM, which enhances the semantic information of the feature maps and obtains more detailed segmentation results. The results show that the proposed method achieved a high IoU of 67.09% in the DeepGlobe dataset. Extensive experiments also verify the effectiveness of the proposed method on three public remote sensing road datasets.
Jia-Xin Wang, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001
IEEE Geosci. Remote. Sens. Lett.5
2022 Semi-Supervised Semantic Segmentation of Remote Sensing Images With Iterative Contrastive Network
abstract
With the development of deep learning, semantic segmentation of remote sensing images has made great progress. However, segmentation algorithms based on deep learning usually require a huge number of labeled images for model training. For remote sensing images, pixel-level annotation usually consumes expensive resources. To alleviate this problem, this letter proposes a semi-supervised segmentation method of remote sensing images based on an iterative contrastive network. This method combines few labeled images and more unlabeled images to significantly improve the model performance. First, contrastive networks continuously learn more potential information by using better pseudo labels. Then, the iterative training method keeps the differences between models to better improve the segmentation performance. The semi-supervised experiments on different remote sensing datasets prove that this method has a better performance than the related methods. Code is available athttps://github.com/VCISwang/ICNet.
Jia-Xin Wang, Sibao Chen 0001, Chris Ding, Jin Tang 0001, Bin Luo 0001
IEEE Geosci. Remote. Sens. Lett.5
2022 MGLNN: Semi-supervised learning via Multiple Graph Cooperative Learning Neural Networks
Bo Jiang 0002, Beibei Wang 0006, Bin Luo 0001
Neural Networks4
2022 GeCNs: Graph Elastic Convolutional Networks for Data Representation
abstract
Graph representation and learning is a fundamental problem in machine learning area. Graph Convolutional Networks (GCNs) have been recently studied and demonstrated very powerful for graph representation and learning. Graph convolution (GC) operation in GCNs can be regarded as a composition of feature aggregation and nonlinear transformation step. Existing GCs generally conduct feature aggregation on a full neighborhood set in which each node computes its representation by aggregating the feature information of all its neighbors. However, this full aggregation strategy is not guaranteed to be optimal for GCN learning and also can be affected by some graph structure noises, such as incorrect or undesired edge connections. To address these issues, we propose to integrate elastic net based selection into graph convolution and propose a novel graph elastic convolution (GeC) operation. In GeC, each node can adaptively select the optimal neighbors in its feature aggregation. The key aspect of the proposed GeC operation is that it can be formulated by a regularization framework, based on which we can derive a simple update rule to implement GeC in a self-supervised manner. Using GeC, we then present a novel GeCN for graph learning. Experimental results demonstrate the effectiveness and robustness of GeCN.
Bo Jiang 0002, Beibei Wang 0006, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Residual objectness for imbalance reduction
Joya Chen, Dong Liu 0002, Bin Luo 0001, Xuezheng Peng, Tong Xu 0001, Enhong Chen
Pattern Recognit.3
2022 GLMNet: Graph learning-matching convolutional networks for feature matching
Bo Jiang 0002, Bin Luo 0001
Pattern Recognit.3
2022 Pedestrian attribute recognition: A survey
Xiao Wang 0014, Shaofei Zheng, Aihua Zheng, Zhe Chen 0013, Jin Tang 0001, Bin Luo 0001
Pattern Recognit.7
2022 RGBT Tracking by Trident Fusion Network
abstract
In recent years, RGBT tracking has become a hot topic in the field of visual tracking, and made great progress. In this paper, we propose a novel Trident Fusion Network (TFNet) to achieve effective fusion of different modalities for robust RGBT tracking. In specific, to deploy the complementarity of features of all convolutional layers, we propose a recursive strategy to densely aggregate these features that yield robust representations of target objects in two modalities. Moreover, we design a trident architecture to integrate the fused features and both modality-specific features for robust target representations. There are three main advantages. First, retaining the classification layer of each modality is beneficial to enhance feature learning of single modality, and compared with aggregate branches, single-modality branches pay more attention to the mining of modal specific information. Second, when some modality is noisy or invalid, the modality-specific branches would capture more discriminative features for RGBT tracking. Finally, the integration of aggregation branches and single-modality branches is beneficial to the complementary learning of different modalities. In addition, we also introduce a feature pruning module in each branch to prune the redundant features and avoid network overfitting. Experimental results on four RGBT tracking benchmark datasets suggest that our tracker achieves superior performance against the state-of-the-art RGBT tracking methods.
Yabin Zhu, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Class-Incremental Learning for Semantic Segmentation in Aerial Imagery via Distillation in All Aspects
abstract
Incremental learning using neural networks achieves great success in semantic segmentation but still suffers from catastrophic forgetting. In this article, we propose an effective class-incremental segmentation method without storing old data. To alleviate the issue of forgetting, we present two important modules, i.e., the deep feature distillation (DFD) module and the label mixed (LM) module. The DFD module is established to learn a good feature representation of old classes by distilling a new compact feature representation from different layers of networks. The proposed LM module first identifies the examples (pixels) of old classes with high confidences utilizing the output of old models, and then, they are combined with examples of new classes to supervise the training of new models, which can achieve a good balance between learning new classing and avoiding forgetting old ones. Our ablation studies show that the DFD module and the LM module can make the learning network obtain 6.2% and 15% performance gains [mean Intersection over Union (mIOU)], respectively. Furthermore, by introducing the supervision of output distillation loss, we compare our method with several state-of-the-art methods in the extensive experiments, and the experimental results all show that our method is significantly superior to them on the dataset of aerial images.
Lianlei Shan, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Class-Incremental Semantic Segmentation of Aerial Images via Pixel-Level Feature Generation and Task-Wise Distillation
abstract
Deep neural networks achieve significant progress in semantic segmentation but still suffer from the catastrophic forgetting problem, i.e., networks will forget old classes as they learn new ones. In this article, we propose an effective class-incremental segmentation framework without storing old data. Specifically, to alleviate the issue of catastrophic forgetting, we present two important modules, i.e., the Pixel-level Feature Generation (PFG) module, and the Task-wise Knowledge distillation (TKD) module. The PFG module is designed to constantly generate any number of features of the old classes to keep the old memory. The PFG module is the first attempt to use the generative method in class-incremental segmentation of aerial images, and it abandons the previous image generation approach but to generate pixel-level features, which is more suitable for the segmentation task. Meanwhile, the proposed TKD module is specially designed for class incremental tasks, and it only compares classes in the same learning step (task), thus avoiding the squeezing of new classes to old classes when the output is normalized (softmax), making distillation more effective. Sufficient experiments show that our method is remarkably effective and achieves more than 4.5% gains compared with state-of-the-art methods, and more than 13% compared with baselines, on all learning conditions. The ablation studies show that the PFG module and the TKD module are both indispensables. Besides, the proposed framework can be well combined with any existing class incremental learning method to achieve better performance.
Lianlei Shan, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 ORSI Salient Object Detection via Multiscale Joint Region and Boundary Model
abstract
Salient object detection (SOD) in optical remote sense images (ORSIs) is a valuable and challenging task. The factors in ORSI, such as background clutter, lighting shadows, imaging blur, and low resolution, significantly degrade the completeness and accuracy of salient objects. To handle this problem, we propose a novel model to learn robust multiscale region features of salient objects by simultaneously optimizing their boundaries. First, we extract multiscale region features of salient objects through a hierarchical attention module. Second, we generate the boundary features by combining the local cues and the global information generated by pyramid pooling. Finally, we embed the boundary features into region features at multiple scales. In particular, we design a joint learning scheme based on a bidirectional feature transformation to optimize boundary and region features simultaneously for accurate ORSI SOD. To provide a comprehensive evaluation platform, we construct a new dataset called ORSI-4199 for ORSI SOD. It contains 4199 finely annotated image pairs with diverse scenes, in which nine attributes (i.e., challenge types) are annotated to facilitate analyzing the strengths and weaknesses of SOD models from different perspectives. Extensive experiments on the public dataset ORSSD, EORRSD, and the newly created dataset ORSI-4199 show that the proposed approach achieves promising results against state-of-the-art methods.https://github.com/wchao1213/ORSI-SOD.
Zhengzheng Tu, Chenglong Li 0002, Minghao Fan, Haifeng Zhao 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.6
2022 RanPaste: Paste Consistency and Pseudo Label for Semisupervised Remote Sensing Image Semantic Segmentation
abstract
With the development of deep learning, remote sensing (RS) image segmentation has been applied with marked success. However, in the process of model training, the large number of labeled images required more expensive annotation. A key challenge is how to make full use of extensive unlabeled images available to improve the segmentation model. In this article, we propose a semisupervised remote sensing image semantic segmentation method defined as RanPaste, which combines labeled images with unlabeled images to improve segmentation performance. First, we obtain pseudo label by randomly pasting part of the ground truth label into the predicted segmentation map. Then, we combine the labeled and unlabeled images to generate rough predictions after strong augmentation. Finally, by using the semisupervised loss, we achieve better performance on remote sensing image segmentation. Our method combines consistency regularization and pseudo label and then utilizes thresholds to gradually improve the model performance. RanPaste enables the model to learn more underlying information in the unlabeled data. Experimental results on six datasets show that RanPaste can learn more latent information from unlabeled data to improve segmentation performance. Besides, our approach achieves better segmentation results on different network structures and datasets.
Jia-Xin Wang, Sibao Chen 0001, Chris Ding, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Reliable Contrastive Learning for Semi-Supervised Change Detection in Remote Sensing Images
abstract
With the development of deep learning in remote sensing (RS) image change detection (CD), the dependence of CD models on labeled data has become an important problem. To make better use of the comparatively resource-saving unlabeled data, the CD method based on semi-supervised learning (SSL) is worth further study. This article proposes a reliable contrastive learning (RCL) method for semi-supervised RS image CD. First, according to the task characteristics of CD, we design the contrastive loss based on the changed areas to enhance the model’s feature extraction ability for changed objects. Then, to improve the quality of pseudo labels in SSL, we use the uncertainty of unlabeled data to select reliable pseudo labels for model training. Combining these methods, semi-supervised CD models can make full use of unlabeled data. Extensive experiments on three widely used CD datasets demonstrate the effectiveness of the proposed method. The results show that our semi-supervised approach has a better performance than related methods. The code is available athttps://github.com/VCISwang/RC-Change-Detection.
Jia-Xin Wang, Teng Li 0001, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001, Richard C. Wilson 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Grouped Bidirectional LSTM Network and Multistage Fusion Convolutional Transformer for Hyperspectral Image Classification
abstract
The efficiently and effectively discriminative spectral-spatial feature representation is essential for hyperspectral image (HSI) classification. However, most of the existing methods rely on the patch-based convolutional neural networks (CNNs) whose ability of extracting the global spatial information is very limited. To address this issue, in this paper, we propose a two-branch network consisting of a grouped bidirectional long short-term memory (GBiLSTM) network and multi-stage fusion convolutional transformer (MFCT) for HSI classification. In the proposed GBiLSTM-MFCT, to extract the spectral features of HSI efficiently, a GBiLSTM network is designed by dividing the sequence features and hidden units of BiLSTM network into several separate groups. To simultaneously extract the global and local spatial features of HSI, a MFCT is proposed by fusing the features of different levels obtained from the multiple phases of convolutional vision transformer. Moreover, in the multi-headed attention module of each stage, blueprint separable convolution based self-attention (BSCA) module is designed which is able to model the global and local spatial information effectively. The outputs of GBiLSTM network and MFCT are fused to generate discriminative and robust spectral-spatial features for HSI classification. Experiments on three benchmark data sets of IN, UP and KSC demonstrate that the proposed GBiLSTM-MFCT exhibits higher classification performance with very limited labeled samples than eight state-of-the-art methods.
Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Category-Wise Fusion and Enhancement Learning for Multimodal Remote Sensing Image Semantic Segmentation
abstract
This paper presents a simple yet effective method called Category-wise Fusion and Enhancement learning (CaFE), which leverages the category priors to achieve effective feature fusion and imbalance learning, for multi-modal remote sensing image semantic segmentation. In particular, we disentangle the feature fusion process via the categories to achieve the category-wise fusion based on the fact that the feature fusion in the same category regions tends to have similar characteristics. The disentangled fusion would also increase the fusion capacity with a small number of parameters while reducing the dependence on large-scale training data. For the sample imbalance problem, we design a simple yet effective category-wise enhancement learning scheme. In particular, we assign the weight for each category region based on the proportion of samples in this region over the whole image. By this way, the learning algorithm would focus more on the regions with smaller proportion. Note that both category-wise feature fusion and imbalance learning are only performed in the training stage, and the segmentation efficiency is thus not affected. Experimental results on two benchmark datasets demonstrate the effectiveness of our CaFE against other state-of-the-art methods.
Aihua Zheng, Jinbo He, Chenglong Li 0002, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Entropy Guided Adversarial Domain Adaptation for Aerial Image Semantic Segmentation
abstract
Recent advances on aerial image semantic segmentation mainly employ the domain adaption to transfer knowledge from the source domain to the target domain. Despite the remarkable achievement, most methods focus on the global marginal distribution alignment to reduce the domain shift between source and target domains, leading to a wrong mapping of the well-aligned features. In this article, we propose an effective unsupervised domain adaptation approach, which relies on a novel entropy guided adversarial learning algorithm, for aerial image semantic segmentation. In specific, we perform local feature alignment between domains by learning a self-adaptive weight from the target prediction probability map to measure the interdomain discrepancy. To exploit the meaningful structure information among semantic regions, we propose to utilize the graph convolutions for long-range semantic reasoning. Comprehensive experimental results on the benchmark dataset of aerial image semantic segmentation and natural scenes demonstrate the superior performance of the proposed method compared to the state-of-the-art methods.
Aihua Zheng, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Remote Sensing Scene Classification via Multi-Branch Local Attention Network
abstract
Remote sensing scene classification (RSSC) is a hotspot and play very important role in the field of remote sensing image interpretation in recent years. With the recent development of the convolutional neural networks, a significant breakthrough has been made in the classification of remote sensing scenes. Many objects form complex and diverse scenes through spatial combination and association, which makes it difficult to classify remote sensing image scenes. The problem of insufficient differentiation of feature representations extracted by Convolutional Neural Networks (CNNs) still exists, which is mainly due to the characteristics of similarity for inter-class images and diversity for intra-class images. In this paper, we propose a remote sensing image scene classification method via Multi-Branch Local Attention Network (MBLANet), where Convolutional Local Attention Module (CLAM) is embedded into all down-sampling blocks and residual blocks of ResNet backbone. CLAM contains two submodules, Convolutional Channel Attention Module (CCAM) and Local Spatial Attention Module (LSAM). The two submodules are placed in parallel to obtain both channel and spatial attentions, which helps to emphasize the main target in the complex background and improve the ability of feature representation. Extensive experiments on three benchmark datasets show that our method is better than state-of-the-art methods.
Sibao Chen 0001, Qing-Song Wei, Wenzhong Wang, Jin Tang 0001, Bin Luo 0001, Zuyuan Wang
IEEE Trans. Image Process.5
2022 Attribute and State Guided Structural Embedding Network for Vehicle Re-Identification
abstract
Vehicle re-identification (Re-ID) is a crucial task in smart city and intelligent transportation, aiming to match vehicle images across non-overlapping surveillance camera scenarios. However, the images of different vehicles may have small visual discrepancies when they have the same/similar attributes, e.g., the same/similar color, type, and manufacturer. Meanwhile, the images from a vehicle may have large visual discrepancies with different states, e.g., different camera views, vehicle viewpoints, and capture time. In this paper, we propose an attribute and state guided structural embedding network (ASSEN) to achieve discriminative feature learning by attribute-based enhancement and state-based weakening for vehicle Re-ID. First, we propose an attribute-based enhancement and expanding module to enhance the discrimination of vehicle features through identity-related attribute information, and we design an attribute-based expanding loss to increase the feature gap between different vehicles. Second, we design a state-based weakening and shrinking module, which not only weakens the state information that interferes with identification but also reduces the intra-class feature gap by a state-based shrinking loss. Third, we propose a global structural embedding module that exploits the attribute information and state information to explore hierarchical relationships between vehicle features, then we use these relationships for feature embedding to learn more robust vehicle features. Extensive experiments on benchmark datasets VeRi-776, VehicleID, and VERI-Wild demonstrate the superior performance and generalization of the proposed method against state-of-the-art vehicle Re-ID methods. The code is available at https://github.com/ttaalle/fast_assen.
Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Image Process.5
2022 LasHeR: A Large-Scale High-Diversity Benchmark for RGBT Tracking
abstract
RGBT tracking receives a surge of interest in the computer vision community, but this research field lacks a large-scale and high-diversity benchmark dataset, which is essential for both the training of deep RGBT trackers and the comprehensive evaluation of RGBT tracking methods. To this end, we present a La rge- s cale H igh-diversity [Formula: see text]nchmark for short-term R GBT tracking (LasHeR) in this work. LasHeR consists of 1224 visible and thermal infrared video pairs with more than 730K frame pairs in total. Each frame pair is spatially aligned and manually annotated with a bounding box, making the dataset well and densely annotated. LasHeR is highly diverse capturing from a broad range of object categories, camera viewpoints, scene complexities and environmental factors across seasons, weathers, day and night. We conduct a comprehensive performance evaluation of 12 RGBT tracking algorithms on the LasHeR dataset and present detailed analysis. In addition, we release the unaligned version of LasHeR to attract the research interest for alignment-free RGBT tracking, which is a more practical task in real-world applications. The datasets and evaluation protocols are available at: https://github.com/mmic-lcl/Datasets-and-benchmark-code.
Chenglong Li 0002, Wanlin Xue, Yaqing Jia, Zhichen Qu, Bin Luo 0001, Jin Tang 0001, Dengdi Sun
IEEE Trans. Image Process.5
2022 Beyond Greedy Search: Tracking by Multi-Agent Reinforcement Learning-Based Beam Search
abstract
To track the target in a video, current visual trackers usually adopt greedy search for target object localization in each frame, that is, the candidate region with the maximum response score will be selected as the tracking result of each frame. However, we found that this may be not an optimal choice, especially when encountering challenging tracking scenarios such as heavy occlusion and fast motion. In particular, if a tracker drifts, errors will be accumulated and would further make response scores estimated by the tracker unreliable in future frames. To address this issue, we propose to maintain multiple tracking trajectories and apply beam search strategy for visual tracking, so that the trajectory with fewer accumulated errors can be identified. Accordingly, this paper introduces a novel multi-agent reinforcement learning based beam search tracking strategy, termed BeamTracking. It is mainly inspired by the image captioning task, which takes an image as input and generates diverse descriptions using beam search algorithm. Accordingly, we formulate the tracking as a sample selection problem fulfilled by multiple parallel decision-making processes, each of which aims at picking out one sample as their tracking result in each frame. Each maintained trajectory is associated with an agent to perform the decision-making and determine what actions should be taken to update related information. More specifically, using the classification-based tracker as the baseline, we first adopt bi-GRU to encode the target feature, proposal feature, and its response score into a unified state representation. The state feature and greedy search result are then fed into the first agent for independent action selection. Afterwards, the output action and state features are fed into the subsequent agent for diverse results prediction. When all the frames are processed, we select the trajectory with the maximum accumulated score as the tracking result. Extensive experiments on seven popular tracking benchmark datasets validated the effectiveness of the proposed algorithm.
Xiao Wang 0014, Zhe Chen 0013, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Dacheng Tao
IEEE Trans. Image Process.5
2022 MsKAT: Multi-Scale Knowledge-Aware Transformer for Vehicle Re-Identification
abstract
Existing vehicle re-identification (Re-ID) methods usually suffer from intra-instance discrepancy and inter-instance similarity. The key to solving this problem lies in filtering out identity-irrelevant interference and collecting identity-relevant vehicle details. In this paper, we aim to design a robust vehicle Re-ID framework that trains a model guided by knowledge vectors yet is able to disentangle the identity-relevant features and identity-irrelevant features. Toward this end, we propose a novel Multi-scale Knowledge-Aware Transformer (MsKAT) to build a knowledge-guided multi-scale feature alignment framework. First, we construct a Knowledge-Aware Transformer (KAT) to interact with semantic knowledge and visual feature. KAT mainly includes State elimination Transformer (SeT) to eliminate state (camera, viewpoint) interference and Attribute aggregation Transformer (AaT) to gather attribute (color, type) information. Second, to learn the knowledge-guided sample differences, we propose to encourage the separation of identity-relevant features and identity-irrelevant features by a Knowledge-Guided Alignment loss ($\mathcal {L}_{KGA}$). Specifically,$\mathcal {L}_{KGA}$suppresses the difference between knowledge-guided positive pairs and the similarity between knowledge-guided negative pairs. Third, with the multi-scale settings of KAT and$\mathcal {L}_{KGA}$, our model can capture knowledge-guided visual consistency features at different scales. Extensive evidence demonstrates our approach achieves new state-of-the-art on three widely-used vehicle re-identification benchmarks.
Chenglong Li 0002, Aihua Zheng, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Intell. Transp. Syst.5
2022 PH-GCN: Person Retrieval With Part-Based Hierarchical Graph Convolutional Network
abstract
Compact feature representation of person image is important for person re-identification (Re-ID) task. Recently, part-based representation models have been widely studied for extracting the more compact and robust feature representation for person image to improve person Re-ID results. However, existing part-based representation models mostly extract the features of different parts independently which ignore the spatial relationship information among different parts. To address this issue, in this paper we propose a novel deep learning framework, named Part-based Hierarchical Graph Convolutional Network (PH-GCN) for person Re-ID problem. Given a person image, PH-GCN first constructs a hierarchical graph to represent the spatial relationships among different parts. Then, both local and global feature learning is achieved by the feature information passing in PH-GCN, which takes the information of other parts into account for part feature representation. Finally, a perceptron layer is adopted for the final person part label prediction and re-identification. The proposed framework provides a general solution that integrateslocal,globalandstructuralfeature learning simultaneously in a unified end-to-end network representation and learning. Extensive experiments on several widely used benchmark datasets demonstrate the effectiveness and benefits of the proposed PH-GCN approach for person Re-ID task.
Bo Jiang 0002, Xixi Wang 0005, Aihua Zheng, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.5
2022 Adversarial-Metric Learning for Audio-Visual Cross-Modal Matching
abstract
Audio-visual matching aims to learn the intrinsic correspondence between image and audio clip. Existing works mainly concentrate on learning discriminative features, while ignore the cross-modal heterogeneous issue between audio and visual modalities. To deal with this issue, we propose a novel Adversarial-Metric Learning (AML) model for audio-visual matching. AML aims to generate a modality-independent representation for each person in each modality via adversarial learning, while simultaneously learns a robust similarity measure for cross-modality matching via metric learning. By integrating the discriminative modality-independent representation and robust cross-modality metric learning into an end-to-end trainable deep network, AML can overcome the heterogeneous issue with promising performance for audio-visual matching. Experiments on the various audio-visual learning tasks, including audio-visual matching, audio-visual verification and audio-visual retrieval on benchmark dataset demonstrate the effectiveness of the proposed AML model. The implementation codes are available onhttps://github.com/MLanHu/AML.
Aihua Zheng, Menglan Hu, Bo Jiang 0002, Yan Yan 0002, Bin Luo 0001
IEEE Trans. Multim.6
2022 RGBT Tracking via Noise-Robust Cross-Modal Ranking
abstract
Existing RGBT tracking methods usually localize a target object with a bounding box, in which the trackers are often affected by the inclusion of background clutter. To address this issue, this article presents a novel algorithm, called noise-robust cross-modal ranking, to suppress background effects in target bounding boxes for RGBT tracking. In particular, we handle the noise interference in cross-modal fusion and seed labels from the following two aspects. First, the soft cross-modality consistency is proposed to allow the sparse inconsistency in fusing different modalities, aiming to take both collaboration and heterogeneity of different modalities into account for more effective fusion. Second, the optimal seed learning is designed to handle label noises of ranking seeds caused by some problems, such as irregular object shape and occlusion. In addition, to deploy the complementarity and maintain the structural information of different features within each modality, we perform an individual ranking for each feature and employ a cross-feature consistency to pursue their collaboration. A unified optimization framework with an efficient convergence speed is developed to solve the proposed model. Extensive experiments demonstrate the effectiveness and efficiency of the proposed approach comparing with state-of-the-art tracking methods on GTOT and RGBT234 benchmark data sets.
Chenglong Li 0002, Zhiqiang Xiang, Jin Tang 0001, Bin Luo 0001, Futian Wang
IEEE Trans. Neural Networks Learn. Syst.4
2022 Tracking by Joint Local and Global Search: A Target-Aware Attention-Based Approach
abstract
Tracking-by-detection is a very popular framework for single-object tracking that attempts to search the target object within a local search window for each frame. Although such a local search mechanism works well on simple videos, however, it makes the trackers sensitive to extremely challenging scenarios, such as heavy occlusion and fast motion. In this article, we propose a novel and general target-aware attention mechanism (termed TANet) and integrate it with a tracking-by-detection framework to conduct joint local and global search for robust tracking. Specifically, we extract the features of the target object patch and continuous video frames; then, we concatenate and feed them into a decoder network to generate target-aware global attention maps. More importantly, we resort to adversarial training for better attention prediction. The appearance and motion discriminator networks are designed to ensure its consistency in spatial and temporal views. In the tracking procedure, we integrate target-aware attention with multiple trackers by exploring candidate search regions for robust tracking. Extensive experiments on both short- and long-term tracking benchmark datasets all validated the effectiveness of our algorithm.
Xiao Wang 0014, Jin Tang 0001, Bin Luo 0001, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2021 MGARL: Multiple Graph Adversarial Regularized Learning
abstract
Graph Convolutional Networks (GCNs) have been commonly studied for graph learning tasks, such as semi-supervised learning, clustering etc. However, many existing GCNs are generally conducted on single graph data and thus can not be applied directly to multi-graph data that consists of various types of edges between nodes. To address this issue, in this paper, we propose a novel multiple Graph Adversarial Regularized Learning (mGARL) framework for multi-graph data representation and learning. mGARL aims to learn an optimal structure invariant/consistent representation for multiple graphs by employing a novel Encoder-Decoder architecture with adversarial learning regularization. It can incorporate the structural information of multiple graphs simultaneously for the node’s representation. We apply the proposed mGARL on the multi-view semi-supervised learning tasks. Experimental results on several datasets demonstrate the effectiveness and benefits of the proposed mGARL model.
Bo Jiang 0002, Bin Luo 0001
ICME3
2021 Intelligent Machine Learning System for Predicting Customer Churn
abstract
Nowadays, customer churn issue is becoming more and more important, which is the key indicator of the business and production success. But how to predict the actual customer churn and take action before customer loss is becoming a difficult issue in the industry. At the same time, how to keep the place of production is the first problem we are facing. After the deep research, we use Artificial Intelligence (AI) and Machine Learning (ML) technology to develop a smart intelligent system and reduce the actual customer churn about the production. This paper will explain the machine learning technology which used in this smart intelligent system and the reader will learn how to use this system to reduce customer loss. In the customer’s churn prediction model aspect, the most popular predictive models have been used, namely, support vector machines, random forests, K-nearest neighbors, and Gradient boosting classifier are applied to check the effect on accuracy, AUC, and F1-score. Through the experiment, it proofs that the Gradient boosting classifier and Random forests give the highest accuracy of 95.32% and 94.29% respectively. The highest AUC score of 91% which achieved by both Gradient boosting classifier and random forests. The highest F1-score of 97.3% is achieved by the Gradient boosting classifier which outperforms over others.
Chenggang He, Chris Ding, Sibao Chen 0001, Bin Luo 0001
ICTAI4
2021 GAMnet: Robust Feature Matching via Graph Adversarial-Matching Network
abstract
Recently, deep graph matching (GM) methods have gained increasing attention. These methods integrate graph nodes¡¯s embedding, node/edges¡¯s affinity learning and final correspondence solver together in an end-to-end manner. For deep graph matching problem, one main issue is how to generate consensus node's embeddings for both source and target graphs that best serve graph matching tasks. In addition, it is also challenging to incorporate the discrete one-to-one matching constraints into the differentiable correspondence solver in deep matching network. To address these issues, we propose a novel Graph Adversarial Matching Network (GAMnet) for graph matching problem. GAMnet integrates graph adversarial embedding and graph matching simultaneously in a unified end-to-end network which aims to adaptively learn distribution consistent and domain invariant embeddings for GM tasks. Also, GAMnet exploits sparse GM optimization as correspondence solver which is differentiable and can also incorporate discrete one-to-one matching constraints approximately in natural in the final matching prediction. Experimental results on three public benchmarks demonstrate the effectiveness and benefits of the proposed GAMnet.
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
ACM Multimedia5
2021 Separable Reversible Data Hiding Based on Integer Mapping and Multi-MSB Prediction for Encrypted 3D Mesh Models
Zhao-Xia Yin, Lulu Cheng, Bin Luo 0001
PRCV (2)5
2021 Regularization graph convolutional networks with data augmentation
Xiu-Zhi Tian, Chris Ding, Sibao Chen 0001, Bin Luo 0001, Xin Wang 0013
Neurocomputing4
2021 A novel domain activation mapping-guided network (DA-GNT) for visual tracking
Zhengzheng Tu, Ajian Zhou, Chuang Gan 0003, Bo Jiang 0002, Amir Hussain 0001, Bin Luo 0001
Neurocomputing6
2021 Multi-ColorGAN: Few-shot vehicle recoloring via memory-augmented networks
Wei Zhou 0102, Sibao Chen 0001, Li-Xiang Xu, Bin Luo 0001
J. Vis. Commun. Image Represent.4
2021 Faster Multiscale Capsule Network With Octave Convolution for Hyperspectral Image Classification
abstract
Recently proposed capsule networks have revealed powerfulness in various visual tasks. However, the traditional CNNs adopted in the capsule layer of the capsule network have the problem of high parameter redundancy. In this letter, we propose a faster multiscale capsule network with octave convolution (MSOctCaps) for hyperspectral image classification. In the proposed MSOctCaps, we design multiple kernels of different sizes with parallel convolution to extract deep multiscale features. To feasibly reduce the redundancy of parameters and achieve high accuracy, the octave convolution is explored in the capsule layer, instead of the traditional convolution, which improves the accuracy of the capsule layer above predicted by the capsule layer below. The comparison experiments with six state-of-the-arts on two challenging contest data sets demonstrate the proposed MSOctCaps is able to produce competitive advantages in terms of both classification accuracy and computational time.
Dongyue Wang, Bin Luo 0001
IEEE Geosci. Remote. Sens. Lett.3
2021 Deep Rényi entropy graph kernel
Lixiang Xu, Lu Bai 0001, Xiaoyi Jiang 0001, Daoqiang Zhang, Bin Luo 0001
Pattern Recognit.6
2021 Reversible data hiding in encrypted images based on pixel prediction and multi-MSB planes rearrangement
abstract
Great concern has arisen in the field of reversible data hiding in encrypted images (RDHEI) due to the development of cloud storage and privacy protection. RDHEI is an effective technology that can embed additional data after image encryption, extract additional data error-free and reconstruct original images losslessly. In this paper, a high-capacity and fully reversible RDHEI method is proposed, which is based on pixel prediction and multi-MSB (most significant bit) planes rearrangement. First, the median edge detector (MED) predictor is used to calculate the predicted value. Next, unlike previous methods, in our proposed method, signs of prediction errors (PEs) are represented by one bit plane and absolute values of PEs are represented by other bit planes. Then, we divide bit planes into uniform blocks and non-uniform blocks, and rearrange these blocks. Finally, according to different pixel prediction schemes, different numbers of additional data are embedded adaptively. The experimental results prove that our method has higher embedding capacity compared with state-of-the-art RDHEI methods.
Zhao-Xia Yin, Xiaomeng She, Jin Tang 0001, Bin Luo 0001
Signal Process.4
2021 Author classification using transfer learning and predicting stars in co-author networks
abstract
Summary The vast amount of data is key challenge to mine a new scholar that is plausible to be star in the upcoming period. The enormous amount of unstructured data raise every year is infeasible for traditional learning; consequently, we need a high quality of preprocessing technique to expand the performance of traditional learning. We have persuaded a novel approach, Authors classification algorithm using Transfer Learning (ACTL) to learn new task on target area to mine the external knowledge from the source domain. Comprehensive experimental outcomes on real‐world networks showed that ACTL, Node‐based Influence Predicting Stars, Corresponding Authors Mutual Influence based on Predicting Stars, and Specific Topic Domain‐based Predicting Stars enhanced the node classification accuracy as well as predicting rising stars to compared with contemporary baseline methods.
Rashid Abbasi, Ali Kashif Bashir, Mohammad Jalil Piran, Farhan Amin, Bin Luo 0001
Softw. Pract. Exp.7
2021 Edge-Guided Non-Local Fully Convolutional Network for Salient Object Detection
abstract
Fully Convolutional Neural Network (FCN) has been widely applied to salient object detection recently by virtue of high-level semantic feature extraction, but existing FCN-based methods still suffer from continuous striding and pooling operations leading to loss of spatial structure and blurred edges. To maintain the clear edge structure of salient objects, we propose a novel Edge-guided Non-local FCN (ENFNet) to perform edge-guided feature learning for accurate salient object detection. In a specific, we extract hierarchical global and local information in FCN to incorporate non-local features for effective feature representations. To preserve good boundaries of salient objects, we propose a guidance block to embed edge prior knowledge into hierarchical feature maps. The guidance block not only performs feature-wise manipulation but also spatial-wise transformation for effective edge embeddings. Our model is trained on the MSRA-B dataset and tested on five popular benchmark datasets. Comparing with the state-of-the-art methods, the proposed method performance well on five datasets.
Zhengzheng Tu, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.5
2021 Dynamic Attention Guided Multi-Trajectory Analysis for Single Object Tracking
abstract
Most of the existing single object trackers track the target in a unitary local search window, making them particularly vulnerable to challenging factors such as heavy occlusions and out-of-view movements. Despite the attempts to further incorporate global search, prevailing mechanisms that cooperate local and global search are relatively static, thus are still sub-optimal for improving tracking performance. By further studying the local and global search results, we raise a question: can we allow more dynamics for cooperating both results? In this paper, we propose to introduce more dynamics by devising a dynamic attention-guided multi-trajectory tracking strategy. In particular, we construct dynamic appearance model that contains multiple target templates, each of which provides its own attention for locating the target in the new frame. Guided by different attention, we maintain diversified tracking results for the target to build multi-trajectory tracking history, allowing more candidates to represent the true target trajectory. After spanning the whole sequence, we introduce a multi-trajectory selection network to find the best trajectory that deliver improved tracking performance. Extensive experimental results show that our proposed tracking strategy achieves compelling performance on various large-scale tracking benchmarks. The project page of this paper can be found athttps://sites.google.com/view/mt-track/.
Xiao Wang 0014, Zhe Chen 0013, Jin Tang 0001, Bin Luo 0001, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2021 RGBT Tracking via Multi-Adapter Network with Hierarchical Divergence Loss
abstract
RGBT tracking has attracted increasing attention since RGB and thermal infrared data have strong complementary advantages, which could make trackers all-day and all-weather work. Existing works usually focus on extracting modality-shared or modality-specific information, but the potentials of these two cues are not well explored and exploited in RGBT tracking. In this paper, we propose a novel multi-adapter network to jointly perform modality-shared, modality-specific and instance-aware target representation learning for RGBT tracking. To this end, we design three kinds of adapters within an end-to-end deep learning framework. In specific, we use the modified VGG-M as the generality adapter to extract the modality-shared target representations. To extract the modality-specific features while reducing the computational complexity, we design a modality adapter, which adds a small block to the generality adapter in each layer and each modality in a parallel manner. Such a design could learn multilevel modality-specific representations with a modest number of parameters as the vast majority of parameters are shared with the generality adapter. We also design instance adapter to capture the appearance properties and temporal variations of a certain target. Moreover, to enhance the shared and specific features, we employ the loss of multiple kernel maximum mean discrepancy to measure the distribution divergence of different modal features and integrate it into each layer for more robust representation learning. Extensive experiments on two RGBT tracking benchmark datasets demonstrate the outstanding performance of the proposed tracker against the state-of-the-art methods.
Andong Lu, Chenglong Li 0002, Yuqing Yan, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Image Process.5
2021 Co-Saliency Detection via a General Optimization Model and Adaptive Graph Learning
abstract
Co-saliency detection is an important research problem, and has been widely used in computer vision area. One main challenge for co-saliency detection problem is how to explore both interactive information among different images and individual salient information within each image simultaneously in co-saliency estimation. In this paper, we propose a novel general optimization framework with adaptive graph learning for co-saliency estimation problem. The proposed model integrates multiple cues including background, and foreground priors, structural information of images, and image feature representation together to obtain a uniform, and accurate co-saliency estimation. One main benefit of the proposed co-saliency method is that it conducts co-saliency propagation, and prediction across different images while maintains the individual salient information of each image, which ensures the consistency, and communication across different images effectively in co-saliency estimation. To improve the accuracy of co-saliency estimation, we adaptively learn a neighborhood, and structured graph to conduct co-saliency propagation among superpixels. An effective optimization algorithm has been designed to seek the optimal solution for the proposed co-saliency optimization model. Experimental results on several widely used datasets show that our method outperforms some other related co-saliency detection methods.
Bo Jiang 0002, Xingyue Jiang, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.4
2021 STGL: Spatial-Temporal Graph Representation and Learning for Visual Tracking
abstract
Tracking-by-detection framework has been normally adopted in visual tracking methods. It aims to localize the visual target object with a bounding box. However, the bounding box is usually difficult to describe the target object accurately and thus easily introduces noisy background information, which usually degrades the final tracking results. Recently, weighted patch representation of the object has been shown very effectively for suppressing the undesirable background information and thus can obviously improve the tracking results. In this paper, we propose a novel Spatial-Temporal Graph representation and Learning (STGL) model to generate a kind of robust target representation for visual tracking problem. The main aspect of STGL is that it aims to exploit both spatial (within each frame) and temporal (between consecutive frames) structure of patches simultaneously in a unified graph representation and semi-supervised learning model. Comparing with existing works, STGL naturally exploits the learned representation of object in previous frame and thus can obtain the representation of object in current frame more accurately and robustly. A new ADMM algorithm is derived to solve the proposed STGL model. Based on the proposed object representation, we then adapt the structured SVM by introducing scale estimation to achieve object tracking. Extensive experiments show that our method outperforms the state-of-the-art patch based tracking methods on two standard benchmark datasets.
Bo Jiang 0002, Bin Luo 0001, Xiaochun Cao, Jin Tang 0001
IEEE Trans. Multim.3
2021 cmSalGAN: RGB-D Salient Object Detection With Cross-View Generative Adversarial Networks
abstract
Image salient object detection (SOD) is an active research topic in computer vision and multimedia area. Fusing complementary information of RGB and depth has been demonstrated to be effective for image salient object detection which is known as RGB-D salient object detection problem. The main challenge for RGB-D salient object detection is how to exploit the salient cues of both intra-modality (RGB, depth) and cross-modality simultaneously which is known as cross-modality detection problem. In this paper, we tackle this challenge by designing a novel cross-modality Saliency Generative Adversarial Network (cmSalGAN). cmSalGAN aims to learn an optimal view-invariant and consistent pixel-level representation for RGB and depth images via a novel adversarial learning framework, which thus incorporates both information of intra-view and correlation information of cross-view images simultaneously for RGB-D saliency detection problem. To further improve the detection results, the attention mechanism and edge detection module are also incorporated into cmSalGAN. The entire cmSalGAN can be trained in an end-to-end manner by using the standard deep neural network framework. Experimental results show that cmSalGAN achieves the new state-of-the-art RGB-D saliency detection performance on several benchmark datasets.
Bo Jiang 0002, Zitai Zhou, Xiao Wang 0014, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.5
2021 Segmenting Objects in Day and Night: Edge-Conditioned CNN for Thermal Image Semantic Segmentation
abstract
Despite much research progress in image semantic segmentation, it remains challenging under adverse environmental conditions caused by imaging limitations of the visible spectrum, while thermal infrared cameras have several advantages over cameras for the visible spectrum, such as operating in total darkness, insensitive to illumination variations, robust to shadow effects, and strong ability to penetrate haze and smog. These advantages of thermal infrared cameras make the segmentation of semantic objects in day and night. In this article, we propose a novel network architecture, called edge-conditioned convolutional neural network (EC-CNN), for thermal image semantic segmentation. Particularly, we elaborately design a gated featurewise transform layer in EC-CNN to adaptively incorporate edge prior knowledge. The whole EC-CNN is end-to-end trained and can generate high-quality segmentation results with edge guidance. Meanwhile, we also introduce a new benchmark data set named "Segmenting Objects in Day And night" (SODA) for comprehensive evaluations in thermal image semantic segmentation. SODA contains over 7168 manually annotated and synthetically generated thermal images with 20 semantic region labels and from a broad range of viewpoints and scene complexities. Extensive experiments on SODA demonstrate the effectiveness of the proposed EC-CNN against state-of-the-art methods.
Chenglong Li 0002, Yan Yan 0002, Bin Luo 0001, Jin Tang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2020 Multi-Spectral Vehicle Re-Identification: A Challenge
abstract
Vehicle re-identification (Re-ID) is a crucial task in smart city and intelligent transportation, aiming to match vehicle images across non-overlapping surveillance camera views. Currently, most works focus on RGB-based vehicle Re-ID, which limits its capability of real-life applications in adverse environments such as dark environments and bad weathers. IR (Infrared) spectrum imaging offers complementary information to relieve the illumination issue in computer vision tasks. Furthermore, vehicle Re-ID suffers a big challenge of the diverse appearance with different views, such as trucks. In this work, we address the RGB and IR vehicle Re-ID problem and contribute a multi-spectral vehicle Re-ID benchmark named RGBN300, including RGB and NIR (Near Infrared) vehicle images of 300 identities from 8 camera views, giving in total 50125 RGB images and 50125 NIR images respectively. In addition, we have acquired additional TIR (Thermal Infrared) data for 100 vehicles from RGBN300 to form another dataset for three-spectral vehicle Re-ID. Furthermore, we propose a Heterogeneity-collaboration Aware Multi-stream convolutional Network (HAMNet) towards automatically fusing different spectrum features in an end-to-end learning framework. Comprehensive experiments on prevalent networks show that our HAMNet can effectively integrate multi-spectral data for robust vehicle Re-ID in day and night. Our work provides a benchmark dataset for RGB-NIR and RGB-NIR-TIR multi-spectral vehicle Re-ID and a baseline network for both research and industrial communities. The dataset and baseline codes are available at: https://github.com/ttaalle/multi-modal-vehicle-Re-ID.
Chenglong Li 0002, Xianpeng Zhu, Aihua Zheng, Bin Luo 0001
AAAI5
2020 Global-Local Attention Network for Semantic Segmentation in Aerial Images
abstract
Errors in semantic segmentation could be classified into two types: the large area misclassification and inaccurate local boundaries. Previously attention-based methods typically capture rich global contextual information, which benefits the large area classification but cannot address the local errors of boundaries. In this paper, we propose a Global-Local Attention Network (GLANet) which can simultaneously consider the global context and local details. Specifically, our GLANet consists of two branches: (1) the global attention branch and (2) local attention branch. Furthermore, three different modules are embedded in GLANet for respectively modelling the semantic interdependencies in spatial, channel and boundary dimension. Lastly, we merge the outputs of different branches to enhance the feature representation further, resulting in more precise segmentation. Overall, the proposed method achieves the competitive segmentation accuracy on two public aerial image datasets, bringing significant improvements over the existing baselines.
Minglong Li, Lianlei Shan, Xiaobin Li 0006, Dengji Zhou, Weiqiang Wang 0001, Ke Lu 0002, Bin Luo 0001, Sibao Chen 0001
ICPR8
2020 UHRSNet: A Semantic Segmentation Network Specifically for Ultra-High-Resolution Images
abstract
Semantic segmentation is a basic task in computer vision, but only limited attention has been devoted to the ultra-high-resolution (UHR) image segmentation. Since UHR images occupy too much memory, they cannot be directly put into GPU for training. Previous methods are cropping images to small patches or downsampling the whole images. Cropping and downsampling cause the loss of contexts and details, which is essential for segmentation accuracy. To solve this problem, we improve and simplify the local and global feature fusion method in previous works. Local features are extracted from patches and global features are from downsampled images. Meanwhile, we propose one new fusion called local feature fusion for the first time, which can make patches get information from surrounding patches. We call the network with these two fusions ultra-high-resolution segmentation network (UHRSNet). These two fusions can effectively and efficiently solve the problem caused by cropping and downsampling. Experiments show a remarkable improvement on Deepglobe dataset [1].
Lianlei Shan, Minglong Li, Xiaobin Li 0006, Ke Lu 0002, Bin Luo 0001, Sibao Chen 0001, Weiqiang Wang 0001
ICPR6
2020 Information Enhanced Graph Convolutional Networks for Skeleton-based Action Recognition
abstract
Skeleton-based action recognition has recently attracted much attention in computer vision. The latest methods are mostly based on graph convolutional networks (GCNs), which construct the human body as spatial-temporal Skeleton graphs, and has achieved excellent performance. However, previous studies only capture the local and rough information based on the physical dependencies among joints, which may miss implicit joint correlations. In this work, we propose a novel action recognition model, namely Information Enhanced Graph Convolutional Networks (IE-GCN). To improve the accuracy and robustness of recognition, this model capture higher-order dependency in the skeleton-based graph by expanding the joint neighbors, and combine second stage skeleton features (the lengths and directions of bones) to enhance the discriminative information simultaneously. In addition, an training strategy is designed to solve the framework. Extensive experiments on two large-scale public datasets, NTU-RGBD and Kinetics-Skeleton, demonstrate the superior performance of the proposed algorithms over the state-of-the-art methods.
Dengdi Sun, Fanchen Zeng, Bin Luo 0001, Jin Tang 0001, Zhuanlian Ding
IJCNN3
2020 LSAM: Local Spatial Attention Module
Miao-Miao Lv, Sibao Chen 0001, Bin Luo 0001
PRCV (3)3
2020 Large-Scale Network Representation Learning Based on Improved Louvain Algorithm and Deep Autoencoder
Shou-Jiu Xiong, Sibao Chen 0001, Chris Ding, Bin Luo 0001
PRCV (3)4
2020 Efficient synthetical clustering validity indexes for hierarchical clustering
Jinpei Liu, Bin Luo 0001
Expert Syst. Appl.4
2020 Multi-scale attention vehicle re-identification
Aihua Zheng, Xianmin Lin, Jiacheng Dong, Wenzhong Wang, Jin Tang 0001, Bin Luo 0001
Neural Comput. Appl.6
2020 Probabilistic SVM classifier ensemble selection based on GMDH-type neural network
Lixiang Xu, Xiaofeng Wang 0009, Lu Bai 0001, Jin Xiao 0003, Qi Liu 0003, Enhong Chen, Xiaoyi Jiang 0001, Bin Luo 0001
Pattern Recognit.8
2020 Joint graph regularized dictionary learning and sparse ranking for multi-modal multi-shot person re-identification
Aihua Zheng, Bo Jiang 0002, Wei-Shi Zheng 0001, Bin Luo 0001
Pattern Recognit.5
2020 Reversible Data Hiding in JPEG Images With Multi-Objective Optimization
abstract
Among various methods of reversible data hiding (RDH) in JPEG images, only rate-distortion, i.e. the image quality with given payload, is taken into consideration during algorithm designing. However, file size expansion is another important evaluation metric for JPEG RDH methods. Based on this situation, we propose a JPEG RDH method considering both the rate-distortion and the file size expansion at the same time while designing the algorithm. The multi-objective optimization strategy is utilized to realize the balance of the two objectives. Specifically, the cover signal is divided into several non-overlapping parts firstly, and after that, the embedding costs of each part are calculated. Next, the optimized combination of parts for embedding data is gained by multi-objective optimization. Experimental results show that the proposed algorithm outperforms the state-of-the-art methods in terms of rate-distortion and file size expansion performance.
Zhao-Xia Yin, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Feature Matching With Intra-Group Sparse Model
abstract
Feature matching is a fundamental problem in computer vision area. In many real applications, one can usually obtain some potential (candidate) matches C by using some discriminative feature descriptors, such as SIFT descriptor. Then, the feature matching problem can be formulated as the problem of trying to select the correct matches S from the potential match set C. In this paper, we propose to solve matches selection by developing a novel intra-group sparse matching (IGSM) model. Our IGSM is motivated by a simple observation that the potential match set C can be divided into several non-overlapping groups Ci, among which the correct matches S are uniformly distributed. We thus develop an intra-group selection model to conduct matches selection at the intra-group level to incorporate the one-to-one matching constraint more in matches selection process. Our IGSM model has three main advantages: (1) The selection mechanism is parameter-free; (2) it generates an intra-group sparse solution which better maintains the one-to-one matching constraint in nature; (3) a simple yet effective update algorithm has been derived to solve IGSM model. The optimality and convergence of the algorithm are theoretically guaranteed. Experimental results on several image feature matching datasets show the effectiveness and efficiency of the proposed IGSM matching method.
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Multim.3
2020 A Subspace Learning Approach to Multishot Person Reidentification
abstract
This paper addresses the challenging problem of multishot person reidentification (Re-ID) in real world uncontrolled surveillance systems. A key issue is how to effectively represent and process the multiple data with various appearance information due to the variations of pose, occlusions, and viewpoints. To this end, this paper develops a novel subspace learning approach, which pursues regularized low-rank and sparse representation for multishot person Re-ID. For the images of a person crossing a certain camera, we assume that the appearances of those subset images with similar viewpoints against a camera draw from the same low-rank subspace, and all the images of a person under a camera lie on a union of low-rank subspaces. Based on this assumption, we propose to learn a nonnegative low-rank and sparse graph to represent the person images. Moreover, the recurring pattern prior is integrated into our model to refine the affinities among images. Extensive experiments on four public benchmark datasets yield impressive performance by improving 22.9% on imagery library for intelligent detection systems video re identification (iLIDS-VID), 42.4% on person RE-ID (PRID) dataset 2011, 39.7% and 30.6% on speech, audio, image, and video technology-SoftBio camera 3/8 and camera 5/8, respectively, and 1.6% on motion analysis and re identification set compared to the state-of-the-art methods.
Aihua Zheng, Xuehan Zhang, Bo Jiang 0002, Bin Luo 0001, Chenglong Li 0002
IEEE Trans. Syst. Man Cybern. Syst.4
2019 Learning Target-aware Attention for Robust Tracking with Conditional Adversarial Network
Xiao Wang 0014, Bin Luo 0001
BMVC4
2019 Data Representation and Learning With Graph Diffusion-Embedding Networks
abstract
Recently, graph convolutional neural networks have been widely studied for graph-structured data representation and learning. In this paper, we present Graph Diffusion-Embedding networks (GDENs), a new model for graph-structured data representation and learning. GDENs are motivated by our development of graph based feature diffusion. GDENs integrate both feature diffusion and graph node (low-dimensional) embedding simultaneously into a unified network by employing a novel diffusion-embedding architecture. GDENs have two main advantages. First, the equilibrium representation of the diffusion-embedding operation in GDENs can be obtained via a simple closed-form solution, which thus guarantees the compactivity and efficiency of GDENs. Second, the proposed GDENs can be naturally extended to address the data with multiple graph structures. Experiments on various semi-supervised learning tasks on several benchmark datasets demonstrate that the proposed GDENs significantly outperform traditional graph convolutional networks.
Bo Jiang 0002, Doudou Lin, Jin Tang 0001, Bin Luo 0001
CVPR4
2019 Semi-Supervised Learning With Graph Learning-Convolutional Networks
abstract
Graph Convolutional Neural Networks (graph CNNs) have been widely used for graph data representation and semi-supervised learning tasks. However, existing graph CNNs generally use a fixed graph which may not be optimal for semi-supervised learning tasks. In this paper, we propose a novel Graph Learning-Convolutional Network (GLCN) for graph data representation and semi-supervised learning. The aim of GLCN is to learn an optimal graph structure that best serves graph CNNs for semi-supervised learning by integrating both graph learning and graph convolution in a unified network architecture. The main advantage is that in GLCN both given labels and the estimated labels are incorporated and thus can provide useful `weakly' supervised information to refine (or learn) the graph construction and also to facilitate the graph convolution operation for unknown label estimation. Experimental results on seven benchmarks demonstrate that GLCN significantly outperforms the state-of-the-art traditional fixed structure based graph CNNs.
Bo Jiang 0002, Doudou Lin, Jin Tang 0001, Bin Luo 0001
CVPR5
2019 Person Re-identification with Patch-Based Local Sparse Matching and Metric Learning
Bo Jiang 0002, Yibing Lv, Aihua Zheng, Bin Luo 0001
ICIG (2)4
2019 MMA: Motion Memory Attention Network for Video Object Detection
Huai Hu, Wenzhong Wang, Aihua Zheng, Bin Luo 0001
ICIG (2)4
2019 Multi-view Similarity Learning of Manifold Data
Rui-Rui Wang, Sibao Chen 0001, Bin Luo 0001, Justin Jian Zhang
ICIG (1)3
2019 Multiple Graph Convolutional Networks for Co-Saliency Detection
abstract
Recently, Graph Convolutional Networks (GCNs) have been usually utilized for graph data representation in computer vision area. However, existing graph GCNs generally use a single graph which can be not adapted for the data with multiple graphs. In this paper, we first propose a novel multiple graph convolutional network (MGCN) for multiple graph data representation and learning. MGCN propagates information/knowledge across multiple graphs and obtains a consistent representation and learning by integrating the information of multiple graphs simultaneously. Based on the proposed MGCN, we then propose a new global-local unified graph convolutional learning architecture for image co-saliency detection problem. The main benefits of the proposed co-saliency model are twofold. First, it learns an optimal superpixel feature representation for co-saliency detection problem. Second, it can well exploit both intra-image and inter-image cues for co-saliency detection via a unified network. Promising experiments demonstrate the effectiveness of the proposed MGCN based co-saliency detection approach.
Bo Jiang 0002, Xingyue Jiang, Jin Tang 0001, Bin Luo 0001, Shilei Huang
ICME4
2019 Pyramid Attention Dense Network for Image Super-Resolution
abstract
Recent deep convolution neural networks has made remarkable progress in single images super-resolution area. They achieved very high Peak Signal to Noise Ratio (PSNR) and structural similarity (SSIM), by improved learning of high-frequency details to enhance visual perception. However, current models usually ignore relations between adjacent pixels. In this work, we propose a network that incorporate gradients of adjacent pixels in addition to per-pixel loss and perceptual loss. In addition, we utilize multi-stage network learning to progressively generate high resolution images, by incorporate a new inter-stage feedback in the Laplacian pyramid network structure. Furthermore, we adopted recently proposed attention mechanism and dense block structure. The proposed Pyramid Attention Dense model for image super-resolution achieved state-of-the-art performance in experiments on four benchmark datasets.
Sibao Chen 0001, Bin Luo 0001, Chris Ding, Shilei Huang
IJCNN3
2019 Multiple Back Propagation Network and Metric Fusion for Person Re-identification
abstract
Person re-identification (Re-ID) is a research focus in pattern recognition, which is to identify a person from another camera view. Many researches have studied feature representations and metric distances of person images, which are robust to changes of view angle and illumination. In this paper, we propose a Multiple Back Propagation (MBP) network and Metric Fusion (MF) for person Re-ID. The proposed MBP network is based on DenseNet or ResNet. Each Dense-conv layer or Conv-ID block is linked by a MBP layer. Each MBP layer is divided into two sub-streams. One sub-stream is connected to softmax loss and the other sub-stream is transferred to a convolution layer followed by triplet loss. A Metric Fusion (MF) method with an optimized weighting scheme is proposed for deep feature fusion. Furthermore, we propose a new metric Re-ranking Euclidean distance joining metric fusion. Experiments on three large-scale person Re-ID benchmark datasets, including Market1501, CUHK03 and DukeMTMC-reID, show that the proposed MBPMF method can achieve state-of-the-art performances.
Sibao Chen 0001, Bin Luo 0001, Chris Ding
IJCNN3
2019 SRAGAN: Generating Colour Landscape Photograph from Sketch
abstract
Generating sketch from colour landscape photograph is very easy while it is hard to generate colour photograph from landscape sketch. In this paper, a new automatic conversion network, named Sparse Residual Attention Generative Adversarial Networks (SRAGAN), is proposed to generate landscape colour photograph from sketch. Besides of generator adversarial loss, we not only adopt L1-regularized per-pix loss, but also combine L1-regularized perceptual loss together into our model. Due to the sparsity of L1-norm, it can preserve boundary edge information very well, which makes our model can handle well the conversion task of sketch-to-photo. In addition, we proposed a ResAttention block to our network structure, which combines the residual learning blocks with attention module. Experiments show that the landscape colour photographes generated by our SRAGAN looks more natural with bright colour and clear edge information. At the same time, we integrate two models so that we can generate winter-style and summer-style photographes from the same landscape sketch. Experiments demonstrate that our method outperforms many state-of-the-arts both in quantitative and in visual performance.
Sibao Chen 0001, Bin Luo 0001, Chris Ding, Justin Jian Zhang
IJCNN3
2019 A Unified Multiple Graph Learning and Convolutional Network Model for Co-saliency Estimation
abstract
Co-saliency estimation which aims to identify the common salient object regions contained in an image set is an active problem in computer vision. The main challenge for co-saliency estimation problem is how to exploit the salient cues of both intra-image and inter-image simultaneously. In this paper, we first represent intra-image and inter-image as intra-graph and inter-graph respectively and formulate co-saliency estimation as graph nodes labeling. Then, we propose a novel multiple graph learning and convolutional network (M-GLCN) for image co-saliency estimation. M-GLCN conducts graph convolutional learning and labeling on both inter-graph and intra-graph cooperatively and thus can well exploit the salient cues of both intra-image and inter-image simultaneously for co-saliency estimation. Moreover, M-GLCN employs a new graph learning mechanism to learn both inter-graph and intra-graph adaptively. Experimental results on several benchmark datasets demonstrate the effectiveness of M-GLCN on co-saliency estimation task.
Bo Jiang 0002, Xingyue Jiang, Ajian Zhou, Jin Tang 0001, Bin Luo 0001
ACM Multimedia5
2019 Dense Feature Aggregation and Pruning for RGBT Tracking
abstract
How to perform effective information fusion of different modalities is a core factor in boosting the performance of RGBT tracking. This paper presents a novel deep fusion algorithm based on the representations from an end-to-end trained convolutional neural network. To deploy the complementarity of features of all layers, we propose a recursive strategy to densely aggregate these features that yield robust representations of target objects in each modality. In different modalities, we propose to prune the densely aggregated features of all modalities in a collaborative way. In a specific, we employ the operations of global average pooling and weighted random selection to perform channel scoring and selection, which could remove redundant and noisy features to achieve more robust feature representation. Experimental results on two RGBT tracking benchmark datasets suggest that our tracker achieves clear state-of-the-art against other RGB and RGBT tracking methods.
Yabin Zhu, Chenglong Li 0002, Bin Luo 0001, Jin Tang 0001, Xiao Wang 0014
ACM Multimedia3
2019 Multi-scale Convolutional Capsule Network for Hyperspectral Image Classification
Dongyue Wang, Jin Tang 0001, Bin Luo 0001
PRCV (2)5
2019 Multi-scale Densely 3D CNN for Hyperspectral Image Classification
Dongyue Wang, Jin Tang 0001, Bin Luo 0001
PRCV (2)5
2019 Efficient Feature Matching via Nonnegative Orthogonal Relaxation
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
Int. J. Comput. Vis.3
2019 Robust visual tracking via Laplacian Regularized Random Walk Ranking
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Chenglong Li 0002
Neurocomputing4
2019 Extended adaptive Lasso for multi-class and multi-label feature selection
Sibao Chen 0001, Yu-Mei Zhang, Chris Ding, Justin Jian Zhang, Bin Luo 0001
Knowl. Based Syst.5
2019 Feature selection based on correlation deflation
Sibao Chen 0001, Chris Ding, Bin Luo 0001
Neural Comput. Appl.4
2019 Efficient Lossless Compression Based Reversible Data Hiding Using Multilayered n-Bit Localization
abstract
We proposed an innovative reversible data hiding technique that is formulated on histogram shifting by using multilayer localized n-bit truncation image (LBPTI), namely, generated form 8-bit plane by means of efficient lossless compression. After selecting the reference point from the block, the neighbor topmost points are used to attain the data embedding without modifying the peak point; in addition, the key information regarding peak point is not mandatory in extraction end to extract the secret information. In order to make the embedded cover-image similar to the histogram of original cover-image, we exploited the localization with efficient lossless compression on lower block level to increase the embedding capacity while controlling extra bit to expand additional embedding capacity on optimum level besides sustaining the quality of cover-image.
Rashid Abbasi, Lixiang Xu, Farhan Amin, Bin Luo 0001
Secur. Commun. Networks4
2019 Quality-aware dual-modal saliency detection via deep reinforcement learning
Xiao Wang 0014, Chenglong Li 0002, Bin Luo 0001, Jin Tang 0001
Signal Process. Image Commun.5
2019 Learning Local-Global Multi-Graph Descriptors for RGB-T Object Tracking
abstract
RGB-thermal (RGB-T) object tracking, which has attracted much recent attention, uses thermal infrared information to assist object tracking with visible light information. However, it still faces many challenging problems, especially the background inclusion in the target bounding box which easily results in model drifting. To handle this problem, we propose a novel and general approach to learn a local-global multi-graph descriptor to suppress background effects for RGB-T tracking. Our approach relies on a novel graph learning algorithm. First, the object is represented with multiple graphs, with a set of multi-modal image patches as nodes, for the robustness to prevent deformation and partial occlusion. Second, we dynamically learn a joint graph over time with both local and global considerations using spatial smoothness and low-rank representation. In particular, we design a single unified alternating direction method of multipliers-based optimization framework to learn graph structure, edge weights, and node weights simultaneously. Third, we combine multi-graph information with corresponding graph node weights to form a robust object descriptor, and tracking is finally carried out by adopting the structured support vector machine. Extensive experiments conducted on the tracking benchmark data sets demonstrate the effectiveness of the proposed approach against the state-of-the-art RGB-T trackers.
Chenglong Li 0002, Chengli Zhu, Justin Jian Zhang, Bin Luo 0001, Xiaohao Wu, Jin Tang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2019 Image Representation and Learning With Graph-Laplacian Tucker Tensor Decomposition
abstract
Tucker tensor decomposition (TD) is widely used for image representation, reconstruction, and learning tasks. Compared to principal component analysis (PCA) models, tensor models retain more 2-D characteristics of images whereas PCA models linearize images. However, traditional TD involves attribute information only and thus does not consider the pairwise similarity information between images. In this paper, we propose a graph-Laplacian tucker tensor decomposition (GLTD) which explores both attributes and pairwise similarity information simultaneously. Generally, GLTD has three main benefits: 1) GLTD reconstruction shows clear robustness against image occlusions/outliers. We provide analysis to show that Laplacian regularization is mainly responsible to this robustness via an out-of-sample GLTD model. To the best of our knowledge, this Laplacian regularization induced robustness of TD has not been studied or emphasized before; 2) GLTD representation performs more regularity, which improves both unsupervised and supervised learning results; and 3) an effective algorithm is derived to solve GLTD problem. Although GLTD is a noncovex problem, the proposed algorithm is shown experimentally to provide a stable/unique solution starting from different random initializations. Experimental results on image reconstruction, data clustering, and classification tasks show the benefits of GLTD.
Bo Jiang 0002, Chris Ding, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Cybern.4
2018 SINT++: Robust Visual Tracking via Adversarial Positive Instance Generation
abstract
Existing visual trackers are easily disturbed by occlusion, blur and large deformation. We think the performance of existing visual trackers may be limited due to the following issues: i) Adopting the dense sampling strategy to generate positive examples will make them less diverse; ii) The training data with different challenging factors are limited, even through collecting large training dataset. Collecting even larger training dataset is the most intuitive paradigm, but it may still can not cover all situations and the positive samples are still monotonous. In this paper, we propose to generate hard positive samples via adversarial learning for visual tracking. Specifically speaking, we assume the target objects all lie on a manifold, hence, we introduce the positive samples generation network (PSGN) to sampling massive diverse training data through traversing over the constructed target object manifold. The generated diverse target object images can enrich the training dataset and enhance the robustness of visual trackers. To make the tracker more robust to occlusion, we adopt the hard positive transformation network (HPTN) which can generate hard samples for tracking algorithm to recognize. We train this network with deep reinforcement learning to automatically occlude the target object with a negative patch. Based on the generated hard positive samples, we train a Siamese network for visual tracking and our experiments validate the effectiveness of the introduced algorithm. The project page of this paper can be found from the website1.
Xiao Wang 0014, Chenglong Li 0002, Bin Luo 0001, Jin Tang 0001
CVPR3
2018 OWP: Objectness Weighted Patch Descriptor for Visual Tracking
abstract
Visual object tracking is an active research problem and has been widely used in computer vision and pattern recognition area. Existing visual tracking methods usually localize the visual object with a bounding box which are often disturbed by the introduced background information and partial occlusion because of bounding box representation of visual object. To deal with this problem, in this paper, we propose a novel Objectness Weighted Patch (OWP) descriptor for object feature descriptor in visual tracking. The aim of OWP is to assign different objectness weights to the patches of bounding box to reduce the influences of background information and partial occlusion. We propose to compute the objectness weights of patches in OWP by integrating multiple cues (background, foreground and local spatial consistency) together in a general optimization model. Also, the proposed model has a simple closed-form solution and thus can be computed efficiently. We incorporate our OWP into structured SVM tracking framework and provide a new robust tracking method. Extensive experiments on two standard benchmark datasets OTB100 and Temple-Color demonstrate the effectiveness and benefits of the proposed tracking method.
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
ICPR4
2018 Multi-scale Attributed Graph Kernel for Image Categorization
Duo Hu, Jin Tang 0001, Bin Luo 0001
PRCV (3)4
2018 Non-negative Dual Graph Regularized Sparse Ranking for Multi-shot Person Re-identification
Aihua Zheng, Bo Jiang 0002, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
PRCV (1)6
2018 A discriminative multi-class feature selection method via weighted l2, 1-norm and Extended Elastic Net
Sibao Chen 0001, Chris Ding, Bin Luo 0001
Neurocomputing5
2018 Robust data representation using locally linear embedding guided PCA
Bo Jiang 0002, Chris Ding, Bin Luo 0001
Neurocomputing3
2018 Saliency detection via a multi-layer graph based diffusion model
Bo Jiang 0002, Zhouqin He, Chris Ding, Bin Luo 0001
Neurocomputing4
2018 Linear regression based projections for dimensionality reduction
Sibao Chen 0001, Chris Ding, Bin Luo 0001
Inf. Sci.3
2018 Low-rank subspace learning based network community detection
Zhuanlian Ding, Xingyi Zhang 0001, Dengdi Sun, Bin Luo 0001
Knowl. Based Syst.4
2018 Reversible data hiding in encrypted AMBTC images
Zhao-Xia Yin, Xuejing Niu, Xinpeng Zhang 0001, Jin Tang 0001, Bin Luo 0001
Multim. Tools Appl.5
2018 A Nonnegative Locally Linear KNN model for image recognition
Sibao Chen 0001, Yu-Lan Xu, Chris Ding, Bin Luo 0001
Pattern Recognit.4
2018 A hybrid reproducing graph kernel based on information entropy
Lixiang Xu, Xiaoyi Jiang 0001, Lu Bai 0001, Jin Xiao 0003, Bin Luo 0001
Pattern Recognit.5
2018 Non-greedy Max-min Large Margin based on L1-norm
Sibao Chen 0001, Chong Zuo, Chris Ding, Bin Luo 0001
Pattern Recognit. Lett.4
2018 Two-stage modality-graphs regularized manifold ranking for RGB-T tracking
Chenglong Li 0002, Chengli Zhu, Shaofei Zheng, Bin Luo 0001, Jing Tang 0001
Signal Process. Image Commun.4
2018 Fast Grayscale-Thermal Foreground Detection With Collaborative Low-Rank Decomposition
abstract
This paper investigates how to perform efficient and robust foreground detection in challenging scenarios by leveraging multiple source data. We propose a novel approach, called collaborative low-rank decomposition (CLoD), for grayscale-thermal foreground detection. Given two data matrices by accumulating sequential frames from the grayscale and the thermal videos, CLoD detects the foreground objects as sparse noises against the backgrounds with collaborative low rank structure, and also incorporates modality weights to achieve adaptive fusion of different source data. For the optimization, CLoD seeks a sub-optimal solution by making the background matrix rank explicitly determined. In particular, the background matrix with the fixed rank can be decomposed into two sub-matrices of low rank, and then, we iteratively optimize them and the modality weights with closed-form solutions. For improving the efficiency, we design a block-based accelerated algorithm to speed up CLoD while employing the edge-preserving algorithm to keep the accuracy. Extensive experiments on the recently public benchmark grayscale-thermal foreground detection suggest that our approach achieves comparable performance in terms of both accuracy and efficiency against other state-of-the-art methods.
Bin Luo 0001, Chenglong Li 0002, Guizhao Wang, Jin Tang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2017 Nonnegative Orthogonal Graph Matching
abstract
Graph matching problem that incorporates pair-wise constraints can be formulated as Quadratic Assignment Problem(QAP). The optimal solution of QAP is discrete and combinational, which makes QAP problem NP-hard. Thus, many algorithms have been proposed to find approximate solutions. In this paper, we propose a new algorithm, called Nonnegative Orthogonal Graph Matching (NOGM), for QAP matching problem. NOGM is motivated by our new observation that the discrete mapping constraint of QAP can be equivalently encoded by a nonnegative orthogonal constraint which is much easier to implement computationally. Based on this observation, we develop an effective multiplicative update algorithm to solve NOGM and thus can find an effective approximate solution for QAP problem. Comparing with many traditional continuous methods which usually obtain continuous solutions and should be further discretized, NOGM can obtain a sparse solution and thus incorporates the desirable discrete constraint naturally in its optimization. Promising experimental results demonstrate benefits of NOGM algorithm.
Bo Jiang 0002, Jin Tang 0001, Chris Ding, Bin Luo 0001
AAAI4
2017 Binary Constraint Preserving Graph Matching
abstract
Graph matching is a fundamental problem in computer vision and pattern recognition area. In general, it can be formulated as an Integer Quadratic Programming (IQP) problem. Since it is NP-hard, approximate relaxations are required. In this paper, a new graph matching method has been proposed. There are three main contributions of the proposed method: (1) we propose a new graph matching relaxation model, called Binary Constraint Preserving Graph Matching (BPGM), which aims to incorporate the discrete binary mapping constraints more in graph matching relaxation. Our BPGM is motivated by a new observation that the discrete binary constraints in IQP matching problem can be represented (or encoded) exactly by a ℓ2-norm constraint. (2) An effective projection algorithm has been derived to solve BPGM model. (3) Using BPGM, we propose a path-following strategy to optimize IQP matching problem and thus obtain a desired discrete solution at convergence. Promising experimental results show the effectiveness of the proposed method.
Bo Jiang 0002, Jin Tang 0001, Chris Ding, Bin Luo 0001
CVPR4
2017 Image Set Representation with L_1 -Norm Optimal Mean Robust Principal Component Analysis
Youxia Cao, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
ICIG (2)4
2017 Selecting attentive frames from visually coherent video chunks for surveillance video summarization
abstract
This paper investigates how to extract key-frames from surveillance video while maximizing their diversity and representational ability. We solve this problem by two steps, i.e., video partition and frame selection. The first step is to partition a surveillance video into visually coherent video chunks, which have high intra-chunk similarity and interchunk dissimilarity. In particular, we propose an object-based frame metric to measure the relevance of two frames, and apply the Normalized Cut algorithm to achieve video partition. The second step is to select the attentive frames from the partitioned video chunks. We propose an attention score based on the content completeness and the visual satisfaction for each frame, and select most attentive frame with highest attention score in each chunk. Extensive experiments on both public and our newly created datasets suggest that our approach significantly outperforms other video summarization methods.
Wenzhong Wang, Qiaoqiao Zhang, Bin Luo 0001, Jin Tang 0001, Rui Ruan, Chenglong Li 0002
ICIP3
2017 Graph Matching via Multiplicative Update Algorithm
abstract
As a fundamental problem in computer vision, graph matching problem can usually be formulated as a Quadratic Programming (QP) problem with doubly stochastic and discrete (integer) constraints. Since it is NP-hard, approximate algorithms are required. In this paper, we present a new algorithm, called Multiplicative Update Graph Matching (MPGM), that develops a multiplicative update technique to solve the QP matching problem. MPGM has three main benefits: (1) theoretically, MPGM solves the general QP problem with doubly stochastic constraint naturally whose convergence and KKT optimality are guaranteed. (2) Em- pirically, MPGM generally returns a sparse solution and thus can also incorporate the discrete constraint approximately. (3) It is efficient and simple to implement. Experimental results show the benefits of MPGM algorithm.
Bo Jiang 0002, Jin Tang 0001, Chris Ding, Yihong Gong, Bin Luo 0001
NIPS5
2017 A new graph ranking model for image saliency detection problem
abstract
Saliency detection is an important problem in many computer vision applications. As a kind of popular method, graph based manifold ranking (GMR) has been successfully used in saliency detection problem. In traditional GMR saliency detection, it involves two main stages, i.e., ranking with background queries and ranking with foreground queries. However, in GMR method, these two stages are conducted separately, which ignores the correlation between background and foreground cues. In this paper, we propose a new graph ranking model, which aims to perform background and foreground ranking simultaneously by exploiting the correlation between background and foreground cues. We derive a closed-form solution for it. Experimental results on four benchmark datasets demonstrate that the proposed method performs better than some other state-of-art methods.
Yuanyuan Guan, Bo Jiang 0002, Yun Xiao 0003, Jin Tang 0001, Bin Luo 0001
SERA5
2017 Image Set Representation and Classification with Attributed Covariate-Relation Graph Model and Graph Sparse Representation Classification
Zhuqiang Chen, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
Neurocomputing4
2017 Data hiding in AMBTC images using quantization level modification and perturbation technique
Wien Hong, Tung-Shou Chen, Zhao-Xia Yin, Bin Luo 0001, Yuan-bo Ma
Multim. Tools Appl.4
2017 Reversible data hiding in encrypted images based on multi-level encryption and block histogram modification
Zhao-Xia Yin, Andrew Abel, Jin Tang 0001, Xinpeng Zhang 0001, Bin Luo 0001
Multim. Tools Appl.5
2017 Local-to-global background modeling for moving object detection from non-static cameras
Aihua Zheng, Lei Zhang 0074, Wei Zhang 0012, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001
Multim. Tools Appl.6
2017 A multiple attributes convolution kernel with reproducing property
Lixiang Xu, Xiu Chen, Cheng Zhang 0010, Bin Luo 0001
Pattern Anal. Appl.5
2017 Lagrangian relaxation graph matching
Bo Jiang 0002, Jin Tang 0001, Xiaochun Cao, Bin Luo 0001
Pattern Recognit.4
2017 Two-Dimensional Discriminant Locality Preserving Projection Based on ℓ1-norm Maximization
Sibao Chen 0001, Cai-Yin Liu, Bin Luo 0001
Pattern Recognit. Lett.4
2017 Image representation and matching with geometric-edge random structure graph
Bo Jiang 0002, Jin Tang 0001, Aihua Zheng, Bin Luo 0001
Pattern Recognit. Lett.4
2017 Special issue "Advances in graph-based pattern recognition"
Cheng-Lin Liu 0001, Bin Luo 0001, Walter G. Kropatsch
Pattern Recognit. Lett.2
2017 Protein functional annotation refinement based on graph regularized ℓ1-norm PCA
Dengdi Sun, Huadong Liang, Meiling Ge, Zhuanlian Ding, Wan-Ting Cai, Bin Luo 0001
Pattern Recognit. Lett.6
2016 Reversible data hiding in encrypted image based on block histogram shifting
abstract
Since there is good potential for practical applications such as encrypted image authentication, content owner identification and privacy protection, reversible data hiding in encrypted image (RDHEI) has attracted increasing attention in recent years. In this paper, we propose and evaluate a new separable RDHEI framework. Additional data can be embedded into a cipher image previously encrypted using Josephus traversal and a stream cipher. A block histogram shifting (BHS) approach using self-hidden peak pixels is adopted to perform reversible data embedding. Depending on the keys held, legal receivers can extract only the embedded data with the data hiding key, or, they can decrypt an image very similar to the original with the decryption key. They can extract both the embedded data and recover the original image error-free if both keys are available. The results demonstrate that higher embedding payload, better quality of decrypted-marked image and error-free image recovery are achieved.
Zhao-Xia Yin, Andrew Abel, Xinpeng Zhang 0001, Bin Luo 0001
ICASSP4
2016 Robust Out-of-Sample Data Recovery
Bo Jiang 0002, Chris Ding, Bin Luo 0001
IJCAI3
2016 Reversible Data Hiding in Encrypted AMBTC Compressed Images
Xuejing Niu, Zhao-Xia Yin, Xinpeng Zhang 0001, Jin Tang 0001, Bin Luo 0001
IWDW5
2016 MDE-based image steganography with large embedding capacity
abstract
The big data era calls for image steganography with large embedding capacity and good image quality. Previous methods described in the literature pay more attention to image quality rather than payload. This paper proposes a large capacity steganographic method based on modification direction exploitation and pixel pair matching. By virtue of a reference matrix with particular properties, one or two 9-ary digits can be embedded into each cover pixel pair depending on different payloads. Experimental results demonstrate high embedding capacity as well as good image quality and security. Copyright © 2015 John Wiley & Sons, Ltd.
Zhao-Xia Yin, Bin Luo 0001
Secur. Commun. Networks2
2015 A Local Sparse Model for Matching Problem
abstract
Feature matching problem that incorporates pairwise constraints is usually formulated as a quadratic assignment problem (QAP). Since it is NP-hard, relaxation models are required. In this paper, we first formulate the QAP from the match selection point of view; and then propose a local sparse model for matching problem. Our local sparse matching (LSM) method has the following advantages: (1) It is parameter-free; (2) It generates a local sparse solution which is closer to a discrete matrix than most other continuous relaxation methods for the matching problem. (3) The one-to-one matching constraints are better maintained in LSM solution. Promising experimental results show the effectiveness of the Proposed LSM method.
Bo Jiang 0002, Jin Tang 0001, Chris Ding, Bin Luo 0001
AAAI4
2015 Second-order steganographic method based on adaptive reference matrix
abstract
A second‐order steganographic method (SOS) based on pixel pair matching and modification direction exploiting (MDE) is proposed in this study. In SOS, each cover pixel pair is used to conceal two secret digits in a B ‐ary notational system. Therefore the maximum embedding rate (ER) is up to log 2 B bit per pixel (bpp). It is different from the previous MDE‐based methods in which only one secret digit in base B can be embedded into each cover pixel pair and the maximum ER is ½log 2 B bpp. The experimental results demonstrate the improvements to the proposed method in terms of capacity, efficiency and detection rate compared to recent MDE‐based methods. Take B = 3 as an example. The ER is 1.585 bpp and the corresponding average peak‐signal‐to‐noise‐rate is 49.89 dB, demonstrating the best image quality with the same embedding rate compared to recent MDE‐based methods.
Zhao-Xia Yin, Chin-Chen Chang 0001, Bin Luo 0001
IET Image Process.4
2015 An algorithm framework of sparse minimization for positive definite quadratic forms
Sibao Chen 0001, Chris Ding, Bin Luo 0001
Neurocomputing3
2015 A local-global mixed kernel with reproducing property
Lixiang Xu, Jin Xie 0004, Andrew Abel, Bin Luo 0001
Neurocomputing5
2015 Image matching using a local distribution based outlier detection technique
Haifeng Zhao 0001, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001
Neurocomputing4
2015 PCA-guided search for K-means
Chris Ding, Jinpei Liu, Bin Luo 0001
Pattern Recognit. Lett.4
2015 Similarity Learning of Manifold Data
abstract
Without constructing adjacency graph for neighborhood, we propose a method to learn similarity among sample points of manifold in Laplacian embedding (LE) based on adding constraints of linear reconstruction and least absolute shrinkage and selection operator type minimization. Two algorithms and corresponding analyses are presented to learn similarity for mix-signed and nonnegative data respectively. The similarity learning method is further extended to kernel spaces. The experiments on both synthetic and real world benchmark data sets demonstrate that the proposed LE with new similarity has better visualization and achieves higher accuracy in classification.
Sibao Chen 0001, Chris Ding, Bin Luo 0001
IEEE Trans. Cybern.3
2014 Robust Non-Negative Dictionary Learning
abstract
Dictionary learning plays an important role in machine learning, where data vectors are modeled as a sparse linear combinations of basis factors (i.e., dictionary). However, how to conduct dictionary learning in noisy environment has not been well studied. Moreover, in practice, the dictionary (i.e., the lower rank approximation of the data matrix) and the sparse representations are required to be nonnegative, such as applications for image annotation, document summarization, microarray analysis. In this paper, we propose a new formulation for non-negative dictionary learning in noisy environment, where structure sparsity is enforced on sparse representation. The proposed new formulation is also robust for data with noises and outliers, due to a robust loss function used. We derive an efficient multiplicative updating algorithm to solve the optimization problem, where dictionary and sparse representation are updated iteratively. We prove the convergence and correctness of proposed algorithm rigorously.We show the differences of dictionary at different level of sparsity constraint.The proposed algorithm can be adapted for clustering and semi-supervised learning.
Qihe Pan, Deguang Kong, Chris Ding, Bin Luo 0001
AAAI4
2014 Classification of Fish Ectoparasite Genus Gyrodactylus SEM Images Using ASM and Complex Network Model
Rozniza Ali, Bo Jiang 0002, Mustafa Man, Amir Hussain 0001, Bin Luo 0001
ICONIP (3)5
2014 Unmanned aerial vehicles (UAV) heading optimal tracking control using online kernel-based HDP algorithm
abstract
UAV can work in places that are dangerous, or not easy to reach for humans. However, due to active control and operating difficulties, it is still a challenge to develop fully autonomous flight in complex environments. This paper applies a novel heuristic dynamic programming for the UAV heading optimal tracking controller design, using kernel-based heuristic dynamic programming (KHDP). Kernel-based HDP is developed by integrating kernel methods and approximately linear dependence (ALD) analysis with the critic learning of HDP algorithm. Compared with conventional HDP where neural networks are widely used and their features were manually designed, the proposed algorithm can obtain better generalization capability and learning efficiency through applying the sparse kernel machine into the critic learning process of HDP algorithm. Simulation and experimental results of UAV heading optimal tracking control problems demonstrate the effectiveness of the proposed kernel-based HDP algorithm.
Fuxiao Tan, Derong Liu 0001, Xin-Ping Guan, Bin Luo 0001
IJCNN4
2014 Covariate-Correlated Lasso for Feature Selection
Bo Jiang 0002, Chris Ding, Bin Luo 0001
ECML/PKDD (1)3
2014 A novel intrusion detection system based on feature generation with visualization strategy
Bin Luo 0001, Jingbo Xia
Expert Syst. Appl.1
2014 Computational power of tissue P systems for generating control languages
Xingyi Zhang 0001, Yanjun Liu 0002, Bin Luo 0001, Linqiang Pan
Inf. Sci.3
2014 Extended linear regression for undersampled face recognition
Sibao Chen 0001, Chris Ding, Bin Luo 0001
J. Vis. Commun. Image Represent.3
2014 On Some Classes of Sequential Spiking Neural P Systems
abstract
Spiking neural P systems (SN P systems) are a class of distributed parallel computing devices inspired by the way neurons communicate by means of spikes; neurons work in parallel in the sense that each neuron that can fire should fire, but the work in each neuron is sequential in the sense that at most one rule can be applied at each computation step. In this work, with biological inspiration, we consider SN P systems with the restriction that at each step, one of the neurons (i.e., sequential mode) or all neurons (i.e., pseudo-sequential mode) with the maximum (or minimum) number of spikes among the neurons that are active (can spike) will fire. If an active neuron has more than one enabled rule, it nondeterministically chooses one of the enabled rules to be applied, and the chosen rule is applied in an exhaustive manner (a kind of local parallelism): the rule is used as many times as possible. This strategy makes the system sequential or pseudo-sequential from the global view of the whole network and locally parallel at the level of neurons. We obtain four types of SN P systems: maximum/minimum spike number induced sequential/pseudo-sequential SN P systems with exhaustive use of rules. We prove that SN P systems of these four types are all Turing universal as number-generating computation devices. These results illustrate that the restriction of sequentiality may have little effect on the computation power of SN P systems.
Xingyi Zhang 0001, Xiangxiang Zeng, Bin Luo 0001, Linqiang Pan
Neural Comput.3
2014 A sparse nonnegative matrix factorization technique for graph matching problems
Bo Jiang 0002, Haifeng Zhao 0001, Jin Tang 0001, Bin Luo 0001
Pattern Recognit.4
2014 Robust Feature Point Matching With Sparse Model
abstract
Feature point matching that incorporates pairwise constraints can be cast as an integer quadratic programming (IQP) problem. Since it is NP-hard, approximate methods are required. The optimal solution for IQP matching problem is discrete, binary, and thus sparse in nature. This motivates us to use sparse model for feature point matching problem. The main advantage of the proposed sparse feature point matching (SPM) method is that it generates sparse solution and thus naturally imposes the discrete mapping constraints approximately in the optimization process. Therefore, it can optimize the IQP matching problem in an approximate discrete domain. In addition, an efficient algorithm can be derived to solve SPM problem. Promising experimental results on both synthetic points sets matching and real-world image feature sets matching tasks show the effectiveness of the proposed feature point matching method.
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Liang Lin 0004
IEEE Trans. Image Process.3
2013 Uncorrelated Lasso
abstract
Lasso-type variable selection has increasingly expanded its machine learning applications. In this paper, uncorrelated Lasso is proposed for variable selection, where variable de-correlation is considered simultaneously with variable selection, so that selected variables are uncorrelated as much as possible. An effective iterative algorithm, with the proof of convergence, is presented to solve the sparse optimization problem. Experiments on benchmark data sets show that the proposed method has better classification performance than many state-of-the-art variable selection methods.
Sibao Chen 0001, Chris Ding, Bin Luo 0001, Ying Xie 0002
AAAI3
2013 Graph-Laplacian PCA: Closed-Form Solution and Robustness
abstract
Principal Component Analysis (PCA) is a widely used to learn a low-dimensional representation. In many applications, both vector data X and graph data W are available. Laplacian embedding is widely used for embedding graph data. We propose a graph-Laplacian PCA (gLPCA) to learn a low dimensional representation of X that incorporates graph structures encoded in W. This model has several advantages: (1) It is a data representation model. (2) It has a compact closed-form solution and can be efficiently computed. (3) It is capable to remove corruptions. Extensive experiments on 8 datasets show promising results on image reconstruction and significant improvement on clustering and classification.
Bo Jiang 0002, Chris Ding, Bin Luo 0001, Jin Tang 0001
CVPR3
2013 Blind Detection of Region Duplication Forgery by Merging Blur and Affine Moment Invariants
abstract
Region duplication is a simple and effective operation to create digital image forgeries, where a part of the image is copied and pasted on another part of the same image. Most existing region duplication detection methods are based on directly matching blocks of image pixels or transform coefficients, and are not effective when duplicated regions have affined transforms or blur degradations. In this work we propose a new region duplicated detection method to automatically detect and localize duplicated regions in digital images. The method is based on merging blur and affine moment invariants, which allows successful detection of region duplication forgery, even under some simple affine transforms and blur degradations. Our experiments in synthesized forgery images with duplicated and distorted regions show that the proposed method gets effective detection. These demonstrate that our method is an effective way to detect the duplication regions under some simple affine transforms and blur degradations blindly.
Jin Tang 0001, Bin Luo 0001
ICIG3
2012 Matching State-Based Sequences with Rich Temporal Aspects
abstract
A General Similarity Measurement (GSM), which takes into account of both non-temporal and rich temporal aspects including temporal order, temporal duration and temporal gap, is proposed for state-sequence matching. It is believed to be versatile enough to subsume representative existing measurements as its special cases.
Aihua Zheng, Jixin Ma 0001, Jin Tang 0001, Bin Luo 0001
AAAI4
2012 A uniform solution to the independent set problem through tissue P systems with cell separation
Xingyi Zhang 0001, Xiangxiang Zeng, Bin Luo 0001, Zheng Zhang 0001
Frontiers Comput. Sci.3
2012 Graph matching based on spectral embedding with missing value
Jin Tang 0001, Bo Jiang 0002, Aihua Zheng, Bin Luo 0001
Pattern Recognit.4
2011 Angular Decomposition
abstract
Dimensionality reduction plays a vital role in pattern recognition. However, for normalized vector data, existing methods do not utilize the fact that the data is normalized. In this paper, we propose to employ an Angular Decomposition of the normalized vector data which corresponds to embedding them on a unit surface. On graph data for similarity/ kernel matrices with constant diagonal elements, we propose the Angular Decomposition of the similarity matrices which corresponds to embedding objects on a unit sphere. In these angular embeddings, the Euclidean distance is equivalent to the cosine similarity. Thus data structures best described in the cosine similarity and data structures best captured by the Euclidean distance can both be effectively detected in our angular embedding. We provide the theoretical analysis, derive the computational algorithm, and evaluate the angular embedding on several datasets. Experiments on data clustering demonstrate that our method can provide a more discriminative subspace.
Dengdi Sun, Chris Ding, Bin Luo 0001, Jin Tang 0001
IJCAI3
2010 On Dynamic Weighting of Data in Clustering with K-Alpha Means
abstract
Although many methods of refining initialization have appeared, the sensitivity of K-Means to initial centers is still an obstacle in applications. In this paper, we investigate a new class of clustering algorithm, K-Alpha Means (KAM), which is insensitive to the initial centers. With K-Harmonic Means as a special case, KAM dynamically weights data points during iteratively updating centers, which deemphasizes data points that are close to centers while emphasizes data points that are not close to any centers. Through replacing minimum operator in K-Means by alpha-mean operator, KAM significantly improves the clustering performances.
Sibao Chen 0001, Haixian Wang, Bin Luo 0001
ICPR3
2009 Registration of blurred images for image mosaic
abstract
Existing methods for the registration of blurred images are efficient for the artificially blurred images or a planar registration. They are not suitable for image mosaic of the source images from a real camera with an almost fixed optical center. We propose a registration method so that a distortion-free registration on naturally captured images can be obtained. It adopts a multi-resolution and robust feature based inter-layer mosaic together. In each layer, Harris corner detector is chosen to effectively detect features and RANSAC is used to find reliable matches for further calibration as well as an initial homography as the initial motion of next layer. Simplex and subspace trust region methods are used consequently to estimate the stable focal length and rotation matrix through the transformation property of feature matches. Experimental results demonstrate the performance of our proposed method.
Xianyong Fang, Bin Luo 0001, Jin Tang 0001, Haifeng Zhao 0001
CAD/Graphics2
2008 Time-Embedding 2D Locality Preserving Projection for Video Summarization
abstract
In this paper we present an effective approach to creating quality video summarization. Considering the video frame sequence and visual similarity, we defined a novel distance formula, which is equivalent to Euclidean distance in respect of norm. A time embedding two dimensional locality preserving projection (TE-2DLPP) is proposed. Experiments show that the new algorithm has better time performance. From the resulting frame cluster, a summary storyboard of the video is created in TE-2DLPP feature subspace, and the obtained experimental results are encouraging.
MaoSheng Fu, Min Kong, Bin Luo 0001
CW4
2008 Heteroscedastic discriminant analysis with two-dimensional constraints
abstract
Heteroscedastic discriminant analysis (HDA) with two-dimensional (2D) constraints is proposed in this paper. HDA suffers from the small sample size problem and instability when lack of training data or feature dimension is high, even when the number of dimension is in a suitable range. Two-dimensional HDA is first proposed, then we show that 2D methods are actually a kind of structure-constrained 1D methods, and lastly, HDA with 2D constraints is proposed. Experiments on TIMIT and WSJ0 show that the proposed method outperforms other methods.
Sibao Chen 0001, Yu Hu 0003, Bin Luo 0001, Renhua Wang
ICASSP3
2008 Probabilistic two-dimensional principal component analysis and its mixture model for face recognition
Haixian Wang, Sibao Chen 0001, Zilan Hu, Bin Luo 0001
Neural Comput. Appl.4
2007 Bilateral Two-Dimensional Locality Preserving Projections
abstract
In this paper, we investigate locality preserving projections (LPP) in two-dimensional sense. Recently, LPP was proposed for dimensionality reduction, which can detect the intrinsic manifold structure of data and preserve the local information. When image data are concerned, they are often vectorized for LPP. However, the dimension of image data is usually very high, LPP can't be implemented due to singularity of matrix. We propose two methods for image dimensionality reduction: two-dimensional LPP (2DLPP) and bilateral two-dimensional LPP (B2DLPP), which are based directly on 2D image matrices rather than 1D vectors as LPP does. Experiments are conducted on the ORL face database, which shows higher recognition performance of the proposed methods.
Sibao Chen 0001, Bin Luo 0001, Renhua Wang
ICASSP (2)2
2007 Using Eigen-Decomposition Method for Weighted Graph Matching
Guoxing Zhao, Bin Luo 0001, Jin Tang 0001, Jixin Ma 0001
ICIC (1)2
2007 2D-LPP: A two-dimensional extension of locality preserving projections
Sibao Chen 0001, Haifeng Zhao 0001, Min Kong, Bin Luo 0001
Neurocomputing4
2006 Shape Representation and Distance Measure Based on Relational Graph
Jin Tang 0001, Bin Luo 0001
HIS3
2006 Semi-Supervised Clustering of Corner-Oriented Attributed Graphs
Jin Tang 0001, Bin Luo 0001
HIS3
2006 Automatic T-Mixture Model Selection via Rival Penalized EM
Jin Tang 0001, Bin Luo 0001
HIS3
2006 LPP and LPP Mixtures for Graph Spectral Clustering
Bin Luo 0001, Sibao Chen 0001
PSIVT1
2006 A spectral approach to learning structural variations in graphs
Bin Luo 0001, Richard C. Wilson 0001, Edwin R. Hancock
Pattern Recognit.1
2005 Pattern Vectors from Algebraic Graph Theory
abstract
Graph structures have proven computationally cumbersome for pattern analysis. The reason for this is that, before graphs can be converted to pattern vectors, correspondences must be established between the nodes of structures which are potentially of different size. To overcome this problem, in this paper, we turn to the spectral decomposition of the Laplacian matrix. We show how the elements of the spectral matrix for the Laplacian can be used to construct symmetric polynomials that are permutation invariants. The coefficients of these polynomials can be used as graph features which can be encoded in a vectorial manner. We extend this representation to graphs in which there are unary attributes on the nodes and binary attributes on the edges by using the spectral decomposition of a Hermitian property matrix that can be viewed as a complex analogue of the Laplacian. To embed the graphs in a pattern space, we explore whether the vectors of invariants can be embedded in a low-dimensional space using a number of alternative strategies, including principal components analysis (PCA), multidimensional scaling (MDS), and locality preserving projection (LPP). Experimentally, we demonstrate that the embeddings result in well-defined graph clusters. Our experiments with the spectral representation involve both synthetic and real-world data. The experiments with synthetic data demonstrate that the distances between spectral feature vectors can be used to discriminate between graphs on the basis of their structure. The real-world experiments show that the method can be used to locate clusters of graphs.
Richard C. Wilson 0001, Edwin R. Hancock, Bin Luo 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Greedy EM algorithm for robust t-mixture modeling
abstract
This paper concerns a greedy EM algorithm for t-mixture modeling, which is more robust than Gaussian mixture modeling when a typical points exist or the set of data has heavy tail. Local Kullback divergence is used to determine how to insert new component. The greedy algorithm obviates the complicated initialization. The results are comparable to that of split-and-merge EM algorithm while the proposed algorithm is faster. Also the by product of a sequence of mixture models is useful for model selection. Experiments of synthetic data clustering and unsupervised color image segmentation are given.
Sibao Chen 0001, Haixian Wang, Bin Luo 0001
ICIG3
2004 Estimation for the number of components in a mixture model using stepwise split-and-merge EM algorithm
Haixian Wang, Bin Luo 0001, Quan bing Zhang, Sui Wei
Pattern Recognit. Lett.2
2004 Robust mixture modelling using multivariate
Haixian Wang, Quan bing Zhang, Bin Luo 0001, Sui Wei
Pattern Recognit. Lett.3
2003 Spectral Clustering of Graphs
Bin Luo 0001, Richard C. Wilson 0001, Edwin R. Hancock
CAIP1
2003 Spectral method for learning structural variations in graphs
abstract
The paper investigates the use of graph-spectral methods for learning the modes of structural variation in sets of graphs. Our approach is as follows. First, we vectorise the adjacency matrices of the graphs. Using a graph-matching method, we establish correspondences between the components of the vectors. Using the correspondences, we cluster the graphs using a Gaussian mixture model. For each cluster we compute the mean and covariance matrix for the vectorised adjacency matrices. We allow the graphs to undergo structural deformation by linearly perturbing the mean adjacency matrix in the direction of the modes of the covariance matrix. We demonstrate the method on sets of corner Delaunay graphs for 3D objects viewed from varying directions.
Bin Luo 0001, Richard C. Wilson 0001, Edwin R. Hancock
ICASSP (3)1
2003 Learning modes of structural variation in graphs
abstract
This paper investigates the use of graph-spectral methods for learning the modes of structural variation in sets of graphs. Our approach is as follows. First, we vectorise the adjacency matrices of the graphs. Using a graph-matching method we establish correspondences between the components of the vectors. Using the correspondences we cluster the graphs using a Gaussian mixture model. For each cluster we compute the mean and covariance matrix for the vectorised adjacency matrices. We allow the graphs to undergo structural deformation by linearly perturbing the mean adjacency matrix in the direction of the modes of the covariance matrix.
Bin Luo 0001, Richard C. Wilson 0001, Edwin R. Hancock
ICIP (2)1
2003 A Spectral Approach to Learning Structural Variations in Graphs
Bin Luo 0001, Richard C. Wilson 0001, Edwin R. Hancock
ICVS1
2003 A unified framework for alignment and correspondence
Bin Luo 0001, Edwin R. Hancock
Comput. Vis. Image Underst.1
2003 Spectral embedding of graphs
Bin Luo 0001, Richard C. Wilson 0001, Edwin R. Hancock
Pattern Recognit.1
2002 Object recognition by clustering spectral features
abstract
We investigate whether vectors of graph spectral features can be used for the purposes of graph clustering. We commence from the eigenvalues and eigenvectors of the adjacency matrix. Each of the leading eigenmodes represents a cluster of nodes and is mapped to a component of a feature vector. The spectral features used as components of the vectors are the eigenvalues and the shared perimeter length. We explore whether these vectors can be used for the purposes of graph clustering. Here we investigate the use of both central and pairwise clustering methods. On a database of view-graphs, both of the features provide good clusters while the eigenvectors perform better.
Bin Luo 0001, Richard C. Wilson 0001, Edwin R. Hancock
ICIP (1)1
2002 Iterative Procrustes alignment with the EM algorithm
Bin Luo 0001, Edwin R. Hancock
Image Vis. Comput.1
2001 Relational Constraints for Point Distribution Models
Bin Luo 0001, Edwin R. Hancock
CAIP1
2001 Discovering Shape Categories by Clustering Shock Trees
Bin Luo 0001, Antonio Robles-Kelly, Andrea Torsello, Richard C. Wilson 0001, Edwin R. Hancock
CAIP1
2001 A Probabilistic Framework for Graph Clustering
abstract
The paper describes a probabilistic framework for graph clustering. We commence from a set of pairwise distances between graph structures. From this set of distances, we use a mixture model to characterize the pairwise affinity of the different graphs. We present an EM-like algorithm for clustering the graphs by iteratively updating the elements of the affinity matrix. In the M-step we apply eigendcomposition to the affinity matrix to locate the principal clusters. In the M-step we update the affinity probabilities. We apply the resulting unsupervised clustering algorithm to two practical problems. The first of these involves locating shape-categories using shock trees extracted from 2D silhouettes. The second problem involves finding the view structure of a polyhedral object using the Delaunay triangulation of corner features.
Bin Luo 0001, Antonio Robles-Kelly, Andrea Torsello, Richard C. Wilson 0001, Edwin R. Hancock
CVPR (1)1
2001 Learning shape categories by clustering shock trees
abstract
This paper investigates whether meaningful shape categories can be identified in an unsupervised way by clustering shock-trees. We commence by computing weighted and unweighted edit distances between shock-trees extracted from the Hamilton-Jacobi skeleton of 2D binary shapes. Next we use an EM-like algorithm to locate pairwise clusters in the pattern of edit-distances. We show that when the tree edit distance is weighted using the geometry of the skeleton, then the clustering method returns meaningful shape categories.
Bin Luo 0001, Richard C. Wilson 0001, Antonio Robles-Kelly, Andrea Torsello, Edwin R. Hancock
ICIP (3)1
2001 Structural Graph Matching Using the EM Algorithm and Singular Value Decomposition
abstract
This paper describes an efficient algorithm for inexact graph matching. The method is purely structural, that is, it uses only the edge or connectivity structure of the graph and does not draw on node or edge attributes. We make two contributions: 1) commencing from a probability distribution for matching errors, we show how the problem of graph matching can be posed as maximum-likelihood estimation using the apparatus of the EM algorithm; and 2) we cast the recovery of correspondence matches between the graph nodes in a matrix framework. This allows one to efficiently recover correspondence matches using the singular value decomposition. We experiment with the method on both real-world and synthetic data. Here, we demonstrate that the method offers comparable performance to more computationally demanding methods.
Bin Luo 0001, Edwin R. Hancock
IEEE Trans. Pattern Anal. Mach. Intell.1
2000 Symbolic Graph Matching Using the EM Algorithm and Singular Value Decomposition
abstract
This paper describes an efficient algorithm for inexact graph-matching. The method is purely structural, i.e., it uses only the edge or connectivity structure of the graph and does not draw on node or edge attributes. We make two contributions: 1) commencing from a probability distribution for matching errors, we show how the problem of graph-matching can be posed as maximum likelihood estimation using the apparatus of the EM algorithm; and 2) casting the recover of correspondences matches between the graph-nodes in a matrix framework. This allows us to efficiently recover correspondence matches using singular value decomposition. We experiment with the method on both real-world and synthetic data. We demonstrate that the method offers comparable performance to more computationally demanding methods.
Bin Luo 0001, Edwin R. Hancock
ICPR1
1999 Matching Point-sets using Procrustes Alignment and the EM Algorithm
abstract
This paper casts the problem of point-set alignment via Procrustes analysis into a maximum likelihood framework using the EM algorithm. The aim is to improve the robustness of the Procrustes alignment to noise and clutter.
Bin Luo 0001, Edwin R. Hancock
BMVC1
1999 Procrustes Alignment with the EM Algorithm
Bin Luo 0001, Edwin R. Hancock
CAIP1
1999 Corner detection via topographic analysis of vector-potential
Bin Luo 0001, Andrew D. J. Cross, Edwin R. Hancock
Pattern Recognit. Lett.1
1998 Corner Detection Via Topographic Analysis of Vector Potential
abstract
Abstract This paper describes how corner detection can be realised using a new feature representation based on a magneto-static analogy. The idea is to compute a vector-potential by appealing to an analogy in which the Canny edge-map is regarded as an elementary current density residing on the image plane. In this paper, we demonstrate that corners are located at the saddle-points of the magnitude of the vector-potential. These points correspond to the intersections of saddle-ridge and saddle-valley structures, i.e. to junctions of the edge and symmetry lines. We describe a template-based method for locating the saddle-points. This involves performing a non-minimum suppression test in the direction of the vector-potential and a non-maximum suppression test in the orthogonal direction. Experimental results using both synthetic and real images are given. We investigate the angle and scale sensitivity of the new corner detector and compare it with a number of alternative corner detectors.
Bin Luo 0001, Andrew D. J. Cross, Edwin R. Hancock
BMVC1
1998 Corner detection using vector potential
abstract
This paper describes how corner detection can be realised using a new feature representation that has recently been successfully exploited for edge and symmetry detection. The feature representation based on an magneto-static analogy. The idea is to compute a vector potential by appealing to an analogy in which the Canny edge-map is regarded as an elementary current density residing on the image plane. In our previous work we demonstrated that edges are the local maxima of the vector potential while points of symmetry correspond to the local minimum. In this paper we demonstrate that corners are located at the saddle points of the magnitude of the vector potential. These points corresponds to the intersections of saddle-ridge and saddle-valley structures, i.e. to junctions of the edge and symmetry lines. We describe a template-based method for locating the saddle-points. This involves performing a nonminimum suppression test in the direction of the vector potential and a nonmaximum suppression test in the orthogonal direction. Experimental results of both synthetic and real images are given.
Bin Luo 0001, Andrew D. J. Cross, Edwin R. Hancock
ICPR1