EDBT 2026 Demo / reviewers in the wild / expert
Bo Jiang 0002
dblp:34/2005-2
· DBLP profile ↗
128ranked-venue papers
44as first author
84since 2021 · last 2026
0000-0002-6238-1596ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 73 · 29 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 58 · 22 first-author · 39 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When Person Re-Identification Meets Event Camera: A Benchmark Dataset and an Attribute-Guided Re-Identification FrameworkabstractRecent researchers have proposed using event cameras for person re-identification (ReID) due to their promising performance and better balance in terms of privacy protection, event camera-based person ReID has attracted significant attention. Currently, mainstream event-based person ReID algorithms primarily focus on fusing visible light and event stream, as well as preserving privacy. Although significant progress has been made, these methods are typically trained and evaluated on small-scale or simulated event camera datasets, making it difficult to assess their real identification performance and generalization ability. To address the issue of data scarcity, this paper introduces a large-scale RGB-event based person ReID dataset, called EvReID. The dataset contains 118,988 image pairs and covers 1200 pedestrian identities, with data collected across multiple seasons, scenes, and lighting conditions. We also evaluate 15 state-of-the-art person ReID algorithms, laying a solid foundation for future research in terms of both data and benchmarking. Based on our newly constructed dataset, this paper further proposes a pedestrian attribute-guided contrastive learning framework to enhance feature learning for person re-identification, termed TriPro-ReID. This framework not only effectively explores the visual features from both RGB frames and event streams, but also fully utilizes pedestrian attributes as mid-level semantic features. Extensive experiments on the EvReID dataset and MARS datasets fully validated the effectiveness of our proposed RGB-Event person ReID framework. Xiao Wang 0014, Shujuan Wu, Bo Jiang 0002, Shiliang Zhang |
AAAI | 4 |
| 2026 | Spatio-temporal side tuning pre-trained foundation models for video-based pedestrian attribute recognition
Xiao Wang 0014, Jiandong Jin, Jun Zhu 0001, Futian Wang, Bo Jiang 0002, Yaowei Wang 0001, Yonghong Tian 0001 |
Comput. Vis. Image Underst. | 6 |
| 2026 | Event Stream based Human Action Recognition: A High-Definition Benchmark Dataset and Algorithms
Xiao Wang 0014, Shiao Wang, Pengpeng Shao, Lin Zhu 0012, Bo Jiang 0002, Yonghong Tian 0001 |
Int. J. Comput. Vis. | 5 |
| 2026 | Reliable and Compact Graph Fine-Tuning via Graph Sparse PromptingabstractRecently, graph prompt learning has garnered increasing attention in adapting pre-trained GNN models for downstream graph learning tasks. However, existing works generally conduct prompting over all graph elements (e.g., nodes, edges, node attributes, etc.), which is suboptimal and obviously redundant. To address this issue, we propose exploiting sparse representation theory for graph prompting and present Graph Sparse Prompting (GSP). GSP aims to adaptively and sparsely select the optimal elements (e.g., certain node attributes) to achieve compact prompting for downstream tasks. Specifically, we propose two kinds of GSP models, termed Graph Sparse Feature Prompting (GSFP) and Graph Sparse multi-Feature Prompting (GSmFP). Both GSFP and GSmFP provide a general scheme for tuning any specific pre-trained GNNs that can select some desired attributes for prompting by employing sparsity-guided prompt learning. A simple yet effective algorithm has been designed for solving GSFP and GSmFP models. Experiments on 16 widely-used benchmark datasets validate the effectiveness and advantages of the proposed GSFPs. Bo Jiang 0002, Beibei Wang 0006, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Revisiting Deformable Convolution on Graphs: Large-Range Modeling and RobustnessabstractGraph Convolution Networks (GCNs) have achieved remarkable success in representation of structured graph data. As we know that traditional GCNs are generally defined on the fixed first-order neighborhood receptive field which makes them be incapable to capture the long-range dependencies between distant nodes and also vulnerable to graph attacks and noises. To address these limitations, we revisit deformable convolution on graphs and propose a novel deformable graph convolution, termed Neighborhood-Deformable Graph Convolution (NDGC). The core of NDGC is to explicitly achieve the deformable convolution on graphs by introducing virtual neighbors which encode large-range information via the offsetting and interpolation function. That is, the introduced virtual neighbors can provide a larger receptive field with deformable receptive shape for graph convolution definition. Also, NDGC conducts message aggregation on the deformable virtual neighbors which thus performs more robustly w.r.t. graph attacks and noises. In particular, NDGC provides a general neighborhood deformable scheme, seamlessly integrating with many graph convolution definitions to derive their deformable variants. Experimental results validate the effectiveness and advantages of the proposed NDGC networks on several graph learning tasks. Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Harmonizing class uniformity and separability for transferability estimation
Yuhe Ding, Bo Jiang 0002, Lijun Sheng, Aihua Zheng, Jian Liang 0001 |
Pattern Recognit. | 2 |
| 2026 | Revisiting color-event based tracking: A unified network, dataset, and metric
Chuanming Tang, Xiao Wang 0014, Ju Huang, Bo Jiang 0002, Lin Zhu 0012, Shifeng Chen, Jianlin Zhang 0001, Yaowei Wang 0001, Yonghong Tian 0001 |
Pattern Recognit. | 4 |
| 2026 | Context-Semantic Quality Awareness Network for fine-grained visual categorization
Sitong Li, Bo Jiang 0002, Bin Luo 0001, Jinhui Tang 0001 |
Pattern Recognit. | 4 |
| 2026 | Entropy calibrated prototype embedding for transductive few-shot learning
Mengfei Guo, Bo Jiang 0002, Bin Luo 0001 |
Pattern Recognit. Lett. | 4 |
| 2026 | BHGraphAdapter: Parameter-Efficient VLMs Tuning Meets Hyper-Graph LearningabstractAdapter-based fine-tuning methods for Visual-Language Models (VLMs) have shown promising performance for feature adaptation in limited data scenarios. However, existing adapters generallyeitheremploy parameterized transformation for multi-modality feature refiningorexploit pairwise relationships between classes (i.e., GraphAdapter) for text enhancement, which ignore the inherent high-order correlations among data samples in the adaptation process. In this paper, for the first time, we propose to exploit the high-order relationships of visual samples within each mini-batch for fine-tuning VLMs and develop a novel Batch HyperGraph Adapter (BHGraphAdapter) to fine-tune VLMs. The core idea of BHGraphAdapter is to conduct feature adapter learning by capturing the inherent high-order semantic information of different samples within each mini-batch, which thus can fully exploit the complex context information in adaptation. Specifically, we first construct a Batch HyperGraph (BHGraph) to model the high-order correlation of samples within each mini-batch. Then, we introduce a message propagation module on BHGraph to update the node embeddings by aggregating information from their high-order neighbors, thereby capturing semantic relationships to enrich feature representation. Finally, we incorporate the proposed BHGraph learning into the pre-trained CLIP framework to achieve the feature adaptation for the downstream tasks. Extensive experiments on 11 benchmark datasets show that our proposed BHGraphAdapter outperforms the SOTA adapter tuning methods. The source code and data will be released at https://github.com/LiuMeilin7195/BHGraphAdapter. Xixi Wang 0005, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | MambaEVT: Event Stream-Based Visual Object Tracking Using State Space ModelabstractEvent camera-based visual tracking has drawn more and more attention in recent years due to the unique imaging principle and advantages of low energy consumption, high dynamic range, and dense temporal resolution. Current event-based tracking algorithms are gradually hitting their performance bottlenecks, due to the utilization of vision Transformer and the static template for target object localization. In this paper, we propose a novel Mamba-based visual tracking framework that adopts the state space model with linear complexity as a backbone network. The search regions and target template are fed into the vision Mamba network for simultaneous feature extraction and interaction. The output tokens of search regions will be fed into the tracking head for target localization. More importantly, we consider introducing a dynamic template update strategy into the tracking framework using the Memory Mamba network. By considering the diversity of samples in the target template library and making appropriate adjustments to the template memory module, a more effective dynamic template can be integrated. The effective combination of dynamic and static templates allows our Mamba-based tracking algorithm to achieve a good balance between accuracy and computational cost on multiple large-scale datasets, including EventVOT, VisEvent, and FE240hz. The source code and checkpoint have been released on https://github.com/Event-AHU/MambaEVT. Xiao Wang 0014, Shiao Wang, Xixi Wang 0005, Zhicheng Zhao 0002, Lin Zhu 0012, Bo Jiang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Reliable Multi-Modal Object Re-Identification via Modality-Aware Graph ReasoningabstractMulti-modal data provides abundant and diverse object information, crucial for effective modal interactions in Re-Identification (ReID) task. However, existing approaches often overlook the quality variations in local features and fail to fully leverage the complementary information across modalities, particularly in cases where features are of low quality. In this paper, we propose to address this issue by leveraging a novel graph reasoning model, termed the Modality-aware Graph Reasoning Network (MGRNet). Specifically, we first construct modality-aware graphs to enhance the extraction of fine-grained local details by effectively capturing and modeling the relationships between patches. Subsequently, the selective graph nodes swap operation is employed to alleviate the adverse effects of low-quality local features by considering both local and global information, enhancing the representation of discriminative information. Finally, the swapped modality-aware graphs are fed into the local-aware graph reasoning module, which propagates multi-modal information to yield a reliable feature representation. Another advantage of the proposed graph reasoning approach is its ability to reconstruct missing modal information by exploiting inherent structural relationships, thereby minimizing disparities between different modalities. Experimental results on four benchmarks (RGBNT201, Market1501-MM, RGBNT100, MSVR310) indicate that the proposed method achieves state-of-the-art performance in multi-modal object ReID. The code for our method will be available upon acceptance. Xixi Wan, Aihua Zheng, Zi Wang 0013, Bo Jiang 0002, Jin Tang 0001, Jixin Ma 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | Unlocking Cross-Domain Synergies for Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptation semantic segmentation (UDASS) aims to perform dense prediction on the unlabeled target domain by training the model on a labeled source domain. In this field, self-training approaches have demonstrated strong competitiveness and advantages. However, existing methods often rely on additional training data (such as reference datasets or depth maps) to rectify the unreliable pseudo-labels, ignoring the cross-domain interaction between the target and source domains. To address this issue, in this paper, we propose a novel method for unsupervised domain adaptation semantic segmentation, termed Unlocking Cross-Domain Synergies (UCDS). Specifically, in the UCDS network, we design a new Dynamic Self-Correction (DSC) module that effectively transfers source domain knowledge and generates high-confidence pseudo-labels without additional training resources. Unlike the existing methods, DSC proposes a Dynamic Noisy Label Detection method for the target domain. To correct the noisy pseudo-labels, we design a Dual Bank mechanism that explores the reliable and unreliable predictions of the source domain, and conducts cross-domain synergy through Weighted Reassignment Self-Correction and Negative Correction Prevention strategies. To enhance the discriminative ability of features and amplify the dissimilarity of different categories, we propose Discrepancy-based Contrastive Learning (DCL). The DCL selects positive and negative samples in the source and target domains based on the semantic discrepancies among different categories, effectively avoiding the numerous false negative samples found in existing methods. Extensive experimental results on three commonly used datasets demonstrate the superiority of the proposed UCDS in comparison with the state-of-the-art methods. The project and code are available at https://github.com/wqh011128/UCDS. Qihang Wu, Bo Jiang 0002, Yuan Chen 0012, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Activating Associative Disease-Aware Vision Token Memory for LLM-Based X-Ray Report GenerationabstractX-ray image based medical report generation achieves significant progress in recent years with the help of large language models, however, these models have not fully exploited the effective information in visual image regions, resulting in reports that are linguistically sound but insufficient in describing key diseases. In this paper, we propose a novel associative memory-enhanced X-ray report generation model that effectively mimics the process of professional doctors writing medical reports. It considers both the mining of global and local visual information and associates historical report information to better complete the writing of the current report. Specifically, given an X-ray image, we first utilize a classification model along with its activation maps to accomplish the mining of visual regions highly associated with diseases and the learning of disease query tokens. Then, we employ a visual Hopfield network to establish memory associations for disease-related tokens, and a report Hopfield network to retrieve report memory information. This process facilitates the generation of high-quality reports based on a large language model and achieves state-of-the-art performance on multiple benchmark datasets, including the IU X-ray, MIMIC-CXR, and Chexpert Plus. The source code and pre-trained models of this work have been released on https://github.com/Event-AHU/Medical_Image_Analysis. Xiao Wang 0014, Fuling Wang, Bo Jiang 0002, Chuanfu Li, Yaowei Wang 0001, Yonghong Tian 0001, Jin Tang 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2026 | Ranking Vision-Language Models in Fully Unlabeled TasksabstractVision language models (VLMs) like CLIP show stellar zero-shot capability on classification benchmarks. However, selecting the VLM with the highest performance on the unlabeled downstream task is non-trivial. Existing VLM selection methods focus on the class-name-only setting, relying on supervised auxiliary datasets and large language models, which may not be accessible or feasible during deployment. This paper introduces the problem ofunsupervised vision-language model selection, where only unsupervised downstream datasets are available, with no additional information provided. To solve this problem, we propose a method termed Visual-tExtual Graph Alignment (VEGA), to select VLMs without any annotations by measuring the alignment of the VLM between the two modalities on the downstream task. VEGA is motivated by the pretraining paradigm of VLMs, which aligns features with the same semantics from the visual and textual modalities, thereby mapping both modalities into a shared representation space. Specifically, we first construct two graphs on the vision and textual features, respectively. VEGA is then defined as the overall similarity between the visual and textual graphs at both node and edge levels. Extensive experiments across three different benchmarks, covering a variety of application scenarios and downstream datasets, demonstrate that VEGA consistently provides reliable and accurate estimates of VLMs' performance on unlabeled downstream tasks. Yuhe Ding, Bo Jiang 0002, Aihua Zheng, Jian Liang 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Object Detection using Event Camera: A MoE Heat Conduction based Detector and A New Benchmark DatasetabstractObject detection in event streams has emerged as a cutting-edge research area, demonstrating superior performance in low-light conditions, scenarios with motion blur, and rapid movements. Current detectors leverage spiking neural networks, Transformers, or convolutional neural networks as their core architectures, each with its own set of limitations including restricted performance, high computational overhead, or limited local receptive fields. This paper introduces a novel MoE (Mixture of Experts) heat conduction-based object detection algorithm that strikingly balances accuracy and computational efficiency. Initially, we employ a stem network for event data embedding, followed by processing through our innovative MoE-HCO blocks. Each block integrates various expert modules to mimic heat conduction within event streams. Subsequently, an IoU-based query selection module is utilized for efficient token extraction, which is then channeled into a detection head for the final object detection process. Furthermore, we are pleased to introduce EvDET200K, a novel benchmark dataset for event-based object detection. Captured with a high-definition Prophesee EVK4-HD event camera, this dataset encompasses 10 distinct categories, 200,000 bounding boxes, and 10,054 samples, each spanning 2 to 5 seconds. We also provide comprehensive results from over 15 state-of-the-art detectors, offering a solid foundation for future research and comparison. The source code has been released on: https://github.com/Event-AHU/OpenEvDET Xiao Wang 0014, Wei Zhang 0161, Lin Zhu 0012, Bo Jiang 0002, Yonghong Tian 0001 |
CVPR | 6 |
| 2025 | CXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus DatasetabstractX-ray image-based medical report generation (MRG) is a pivotal area in artificial intelligence that can significantly reduce diagnostic burdens and patient wait times. Despite significant progress, we believe that the task has reached a bottleneck due to the limited benchmark datasets and the existing large models’ insufficient capability enhancements in this specialized domain. Specifically, the recently released CheXpert Plus dataset lacks comparative evaluation algorithms and their results, providing only the dataset itself. This situation makes the training, evaluation, and comparison of subsequent algorithms challenging. Thus, we conduct a comprehensive benchmarking of existing mainstream X-ray report generation models and large language models (LLMs), on the CheXpert Plus dataset. We believe that the proposed benchmark can provide a solid comparative basis for subsequent algorithms and serve as a guide for researchers to quickly grasp the state-of-the-art models in this field. More importantly, we propose a large model for the X-ray image report generation using a multi-stage pre-training strategy, including self-supervised autoregressive generation and Xray-report contrastive learning, and supervised fine-tuning. Extensive experimental results indicate that the autoregressive pre-training based on Mamba effectively encodes X-ray images, and the image-text contrastive pre-training further aligns the feature spaces, achieving better experimental results. Source code can be found on https://github.com/Event-AHU/Medical_Image_Analysis. Xiao Wang 0014, Fuling Wang, Yuehang Li, Qingchuan Ma, Shiao Wang, Bo Jiang 0002, Jin Tang 0001 |
CVPR | 6 |
| 2025 | Mix-Mask Augmentation and Self-Reconstruction for Cross-Domain Few-Shot Hyperspectral Image ClassificationabstractRecently, the metric-based prototypical methods achieves promising performance in few-shot learning (FSL) for hyperspectral image (HSI) classification. However, the existing models are easily affected by the noisy pixels of different categories around the center pixel of the patch, and tend to focus on the most representative features while ignoring other important ones, which cause the overfitting problem. Moreover, the commonly used dimension reduction operation of the feature of source and target domains inevitably results in the loss of valuable spectral information. To address these issues, we propose the mix-mask augmentation and self-reconstruction for cross-domain HSI classification. The pixel mask augmentation is introduced to enhance the sample diversity of query set and suppress the impact of noisy pixels, thus encouraging the model to discover discriminative features on a wider range. The CutMix augmentation is also adopted to generate the mixed support set and mixed prototypes, mitigating the negative impact of confusing prototypes. Furthermore, we develop the self-reconstruction module which can preserve more useful feature information during the dimension reduction for feature representation of the source and target domains. Extensive experiments on three public HSI datasets demonstrate that the proposed method achieves superior performance with fewer computational costs in comparison with the SOTA methods. Qihang Wu, Xiao Wang 0014, Jinpei Liu, Bo Jiang 0002 |
ICASSP | 7 |
| 2025 | Spatial-Temporal Memory Filtering SAM for Lesion Segmentation in Breast Ultrasound Videos
Zhengzheng Tu, Liang Zong, Bo Jiang 0002, Chaoxue Zhang |
MICCAI (2) | 3 |
| 2025 | SPromptGL: Semantic Prompt Guided Graph Learning for Multi-modal Brain Disease
Xixi Wan, Bo Jiang 0002, Aihua Zheng |
MICCAI (12) | 2 |
| 2025 | CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training FrameworkabstractEvent cameras have attracted increasing attention in recent years due to their advantages in high dynamic range, high temporal resolution, low power consumption, and low latency. Some researchers have begun exploring pre-training directly on event data. Nevertheless, these efforts often fail to establish strong connections with RGB frames, limiting their applicability in multi-modal fusion scenarios. To address these issues, we propose a novel CM3AE pre-training framework for the RGB-Event perception. This framework accepts multi-modalities/views of data as input, including RGB images, event images, and event voxels, providing robust support for both event-based and RGB-event fusion based downstream tasks. Specifically, we design a multi-modal fusion reconstruction module that reconstructs the original image from fused multi-modal features, explicitly enhancing the model's ability to aggregate cross-modal complementary information. Additionally, we employ a multi-modal contrastive learning strategy to align cross-modal feature representations in a shared latent space, which effectively enhances the model's capability for multi-modal understanding and capturing global dependencies. We construct a large-scale dataset containing 2,535,759 RGB-Event data pairs for the pre-training. Extensive experiments on five downstream tasks fully demonstrated the effectiveness of CM3AE. Source code and pre-trained models will be released on https://github.com/Event-AHU/CM3AE. Xiao Wang 0014, Chenglong Li 0002, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Qi Liu 0003 |
ACM Multimedia | 4 |
| 2025 | UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-IdentificationabstractMulti-modal object Re-IDentification (ReID) has gained considerable attention with the goal of retrieving specific targets across cameras using heterogeneous visual data sources. At present, multi-modal object ReID faces two core challenges: (1) learning robust features under fine-grained local noise caused by occlusion, frame loss, and other disruptions; and (2) effectively integrating heterogeneous modalities to enhance multi-modal representation. To address the above challenges, we propose a robust approach named Uncertainty-Guided Graph model for multi-modal object ReID (UGG-ReID). UGG-ReID is designed to mitigate noise interference and facilitate effective multi-modal fusion by estimating both local and sample-level aleatoric uncertainty and explicitly modeling their dependencies. Specifically, we first propose the Gaussian patch-graph representation model that leverages uncertainty to quantify fine-grained local cues and capture their structural relationships. This process boosts the expressiveness of modal-specific information, ensuring that the generated embeddings are both more informative and robust. Subsequently, we design an uncertainty-guided mixture of experts strategy that dynamically routes samples to experts exhibiting low uncertainty. This strategy effectively suppresses noise-induced instability, leading to enhanced robustness. Meanwhile, we design an uncertainty-guided routing to strengthen the multi-modal interaction, improving the performance. UGG-ReID is comprehensively evaluated on five representative multi-modal object ReID datasets, encompassing diverse spectral modalities. Experimental results show that the proposed method achieves excellent performance on all datasets and is significantly better than current methods in terms of noise immunity. Our code is available at https://github.com/wanxixi11/UGG-ReID. Xixi Wan, Aihua Zheng, Bo Jiang 0002, Beibei Wang 0006, Chenglong Li 0002, Jin Tang 0001 |
NeurIPS | 3 |
| 2025 | CroPe: Cross-Modal Semantic Compensation Adaptation for All Adverse Scene UnderstandingabstractScene understanding in adverse conditions, such as fog, snow, and night, is challenging due to the visual appearance degeneration. In this context, we propose a Cross-modal Semantic Compensation Adaptation method (CroPe) for scene understanding. Distinct from the existing methods, which only use the visual information to learn the domain-invariant features, CroPe establishes a visual-textual paradigm which provides textual semantic compensation for visual features, enabling the model to learn more consistent representations. We propose the Complementary Perceptual Text Generation (CPTG) module which generates a set of multi-level complementary-perceptive text embeddings incorporating both generalization and domain awareness. To achieve cross-modal semantic compensation, the Reverse Chain Text-Visual Fusion (RCTVF) module is developed. By the unified attention and reverse decoding chain, compensation information is successively fused to the visual features from the deep (semantic dense) to shallow (semantic sparse) features, maximizing compensation gain. CroPe yields competitive results under all adverse conditions and significantly improves the state-of-the-art performance by 6.5 mIoU for ACDC-Night dataset and 1.2 mIoU for ACDC-All dataset, respectively. Qihang Wu, Hongtao Luo, Xiaoxia Cheng, Bo Jiang 0002 |
NeurIPS | 5 |
| 2025 | Uncertain-GMamba: Graph Mamba with Uncertainty-Guided Node Sorting
Beibei Wang 0006, Bo Jiang 0002 |
PRCV (4) | 3 |
| 2025 | GOBoost: leveraging long-tail gene ontology terms for accurate protein function predictionabstractMOTIVATION: With the advancement of deep learning, researchers have increasingly proposed computational methods based on deep learning techniques to predict protein function. However, many of these methods treat protein function prediction as a multi-label classification problem, often overlooking the long-tail distribution of functional labels (i.e., Gene Ontology Terms) in datasets. To address this issue, we propose the GOBoost method, which incorporates the proposed long-tail optimization ensemble strategy. Besides, GOBoost introduces the proposed global-local label graph module and multi-granularity focal loss function to enhance long-tail functional information, mitigate the long-tail phenomenon, and improve overall prediction accuracy. RESULTS: We evaluate GOBoost and other state-of-the-art (SOTA) protein function prediction methods on the PDB and AF2 datasets. The GOBoost outperformed SOTA methods on all evaluation metrics for both datasets. Notably, in the AUPR evaluation on the PDB test set, GOBoost improved by 10.71%, 35.91%, and 22.71% compared to the SOTA HEAL method in the MF, BP, and CC functions. The experimental results show the necessity and superiority of designing models from the label long-tail distribution perspective. AVAILABILITY AND IMPLEMENTATION: The source code of GOBoost is available at https://github.com/Cao-Labs/GOBoost. Lei Zhang 0060, Jie Hou 0001, Dong Si, Bo Jiang 0002, Hailey Ledenko, Renzhi Cao |
Bioinform. | 7 |
| 2025 | Visual and text prompt learning for multi-modal brain disease diagnosis
Yumiao Zhao, Bo Jiang 0002, Yuhe Ding, Xixi Wan, Jin Tang 0001 |
Sci. China Inf. Sci. | 2 |
| 2025 | Learning Dynamic Batch-Graph Representation for Deep Representation Learning
Xixi Wang 0005, Bo Jiang 0002, Xiao Wang 0014, Bin Luo 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Understanding beyond outputs: A novel knowledge distillation method using Schur decomposition
Chang-Ming Pan, Sibao Chen 0001, Bo Jiang 0002, Bin Luo 0001 |
Neurocomputing | 3 |
| 2025 | Multi-level semantic-aware transformer for image captioning
Shan Song, Qihang Wu, Bo Jiang 0002, Bin Luo 0001, Jinhui Tang 0001 |
Neural Networks | 4 |
| 2025 | Dynamic semantic-geometric guidance and structure transfer network for cross-scene hyperspectral image classification
Shuke Wang, Bo Jiang 0002, Zhifu Tao, Bin Luo 0001 |
Neural Networks | 4 |
| 2025 | Graph Spiking Attention Network: Sparsity, Efficiency and RobustnessabstractExisting Graph Attention Networks (GATs) generally adopt the self-attention mechanism to learn graph edge attention, which usually return dense attention coefficients over all neighbors and thus are prone to be sensitive to graph edge noises. To overcome this problem, sparse GATs are desirable and have garnered increasing interest in recent years. However, existing sparse GATs usually suffer from high training complexity and are also not straightforward for inductive learning tasks. To address these issues, we propose to learn sparse GATs by exploiting spiking neuron (SN) mechanism, termed Graph Spiking Attention (GSAT). Specifically, it is known that spiking neuron can perform inexpensive information processing by transmitting the input data into discrete spike trains and return sparse outputs. Inspired by it, this work attempts to exploit spiking neuron to learn sparse attention coefficients, resulting in edge-sparsified graph for GNNs. Therefore, GSAT can perform message passing on the selective neighbors naturally, which makes GSAT perform compactly and robustly w.r.t graph noises. Moreover, GSAT can be used straightforwardly for inductive learning tasks. Extensive experiments on both transductive and inductive tasks demonstrate the effectiveness, robustness and efficiency of GSAT. Beibei Wang 0006, Bo Jiang 0002, Jin Tang 0001, Lu Bai 0001, Bin Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Unifying Graph Contrastive Learning via Graph Message AugmentationabstractGraph contrastive learning is usually performed by first conducting Graph Data Augmentation (GDA) and then employing a contrastive learning pipeline to train GNNs. As we know that GDA is an important issue for graph contrastive learning. Various GDAs have been developed recently which mainly involve dropping or perturbing edges, nodes, node attributes and edge attributes. However, to our knowledge, it still lacks a universal and effective augmentor that is suitable for different types of graph data. To address this issue, in this paper, we first introduce the graph message representation of graph data. Based on it, we then propose a novel Graph Message Augmentation (GMA), a universal scheme for reformulating many existing GDAs. The proposed unified GMA not only gives a new perspective to understand many existing GDAs but also provides a universal and more effective graph data augmentation for graph self-supervised learning tasks. Moreover, GMA introduces an easy way to implement the mixup augmentor which is natural for images but usually challengeable for graphs. Based on the proposed GMA, we then propose a unified graph contrastive learning, termed Graph Message Contrastive Learning (GMCL), that employs attribution-guided universal GMA for graph contrastive learning. Experiments on many graph learning tasks demonstrate the effectiveness and benefits of the proposed GMA and GMCL approaches. Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Challenge-aware U-net for breast lesion segmentation in ultrasound images
Dengdi Sun, Changxu Dong, Bo Jiang 0002, Yayang Duan, Zhengzheng Tu, Chaoxue Zhang |
Pattern Recognit. | 4 |
| 2025 | UAV Video Vehicle Detection: Benchmark and BaselineabstractWith the increasing application of unmanned aerial vehicles (UAVs) in intelligent transportation systems, vehicle object detection in UAV videos has received increasing attention. Precise categorization and detection for vehicles in UAVs is important in many practical applications. However, existing object detection methods, tailored for natural images, often fall short of accurately identifying vehicle objects. Additionally, high-altitude UAV imaging mainly employs horizontal bounding box annotation, frequently leading to significant obstruction and overlapping. Hence, we propose a new task called UAV video vehicle detection (VVD) to achieve precise detection and categorization of vehicles in high-altitude UAV imaging environments. To facilitate the research and development of UAV VVD, we construct the first large-scale well-annotated benchmark UAV VVD dataset, which includes 70 UAV videos captured at a 500-m altitude, with 361489 vehicle instances annotated by the oriented bounding boxes and vehicle categories. Moreover, we introduce a novel category refinement network (CRNet) approach that extracts and refines vehicle object features from the bounding box of the detection results to classify vehicle categories. This approach effectively eliminates the interference of the background and other vehicle objects in candidate boxes. Notably, the vehicle object features are projected into subspace, enabling the category refinement module (CRM) to focus more on the distinctive characteristics of the vehicle object itself through normalization operations. We conduct extensive experiments on the proposed VVD dataset. Experimental results demonstrate the superiority and effectiveness of the proposed CRNet method. The relevant code and dataset are available athttps://github.com/mmic-lcl. Yun Xiao 0003, Jinfa Wang, Zhicheng Zhao 0002, Bo Jiang 0002, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Multimodal Remote Sensing Image Registration via Modality Perception and Self-Supervised Position EstimationabstractMulti-modal remote sensing images registration ensures that images from different sensors or modalities are spatial and informational consistent for effective comparison and analysis. However, due to the non-linear modality gaps that exist between images, making it difficult to focus only on the spatial position differences of the images and ignore the modality gaps. In this paper, to address this issue, we propose a new framework for Multi-Modal remote sensing image Registration, named MMRNet. The proposed framework comprises the following main aspects. First, a novel self-supervised Positional Misalignment Estimator (PME) is designed for multi-modal image registration. PME is able to efficiently overcome the modality gaps and learn the positional differences between multi-modal images more reliably, optimizing the registration loss by minimizing the positional differences directly. Then, a new paradigm of modality translation, termed Modality Perception Module (MPM), is introduced to effectively learn modality gaps and perform modality translation in the case of positional misalignment. Finally, we further design the modality perception guidance loss to supervise the modality translation task, which can encourage the fidelity of the generated pseudo-modality images. Our registration network integrates both rigid registration model and non-rigid registration model. Experimental results demonstrate that the proposed registration framework can obtain obviously superior performance in both rigid and non-rigid image registration tasks on optical-SAR data, optical-map data and optical-infrared data. The code and relevant dataset will be made publicly available at https://github.com/Ahuer-Lei/MMRNet. Yun Xiao 0003, Bo Jiang 0002, Yuan Chen 0012, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Bridging the Style Gap: Style-Guided Distillation Domain Adaptation for Hyperspectral Image ClassificationabstractLand cover in different hyperspectral image (HSI) commonly exhibits style differences and similarities in the same category and distinct categories. However, most of existing cross-scene HSI classification methods overlook this issue and only conduct the feature-level alignment with unsupervised domain adaptation (UDA). To address this limitation, we propose the Style-Guided Distillation Domain Adaptation (SGDDA) for HSI classification. First, a Fourier Transform based Style Transfer (FTST) module is proposed to generate an enhanced source HSI. It transfers stylistic features from the target domain (TD) to the source domain (SD) by substituting low-frequency components, thereby preserving semantic invariance while bridging their style gap. Second, the Dual-Path Knowledge Distillation (DPKD) module is designed to reduce ambiguity in category assignment for the SD. This is achieved through cross-domain consistency learning between original SD samples and their enhanced style-transferred counterparts, ensuring robust feature alignment across domains. Third, unlike existing methods that primarily utilize classification threshold to select pseudo-label for target samples, we propose the Confidence-aware Dual Pseudo-label Consensus (CADPLC) strategy. This strategy dynamically selects the reliable pseudo-labels by leveraging both class prototype matching and teacher-student prediction consensus, eliminating reliance on fixed thresholds and significantly improving adaptation to the target domain. Experiments on three benchmark datasets demonstrate the superiorities of the proposed SGDDA in comparison with several state-of-the-art methods. The code is available at https://github.com/Wei-spvl/SGDDA. Bo Jiang 0002, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | A Spatial-Temporal Progressive Fusion Network for Breast Lesion Segmentation in Ultrasound VideosabstractUltrasound video-based breast lesion segmentation provides valuable assistance in early breast lesion detection and discrimination. However, this field faces two key challenges: the first is how to simultaneously utilize both intra-frame and inter-frame lesion cues to accurately segment breast lesions, and the second is that the availability of breast ultrasound video datasets is quite limited. In this paper, we propose a novel Spatial-Temporal Progressive Fusion Network (STPFNet) for video-based breast lesion segmentation problem. The proposed STPFNet comprises three main components. First, we propose to adopt a unified network architecture to capture spatial dependencies within each ultrasound frame and temporal correlations between different frames together for feature representation of ultrasound video. Second, we propose a new fusion module called Multi-Granularity Feature Fusion (MGFF) to fuse the extracted information with different granularities for lesion segmentation. MGFF can help improve the issue of lesion boundary blurring. Third, we propose to take the segmentation result of the previous frame as prior knowledge to suppress the noisy background and learn a more robust representation. To further promote the research in this field, we construct a new ultrasound video breast lesion segmentation dataset, called UVBLS200, comprising 200 videos (80 benign and 120 malignant lesions). Experiments on the proposed dataset demonstrate that the proposed STPFNet achieves a better breast lesion detection performance than state-of-the-art methods. Zhengzheng Tu, Zigang Zhu, Yayang Duan, Bo Jiang 0002, Qishun Wang, Chaoxue Zhang |
IEEE Trans. Multim. | 4 |
| 2025 | CRSOT: Cross-Resolution Object Tracking Using Unaligned Frame and Event CamerasabstractExisting datasets for RGB-DVS tracking are collected with DVS346 camera and their resolution ($346 \times 260$) is low for practical applications. Actually, only visible cameras are deployed in many practical systems, and the newly designed neuromorphic cameras may have different resolutions. The latest neuromorphic sensors can output high-definition event streams, but it is very difficult to achieve strict alignment between events and frames on both spatial and temporal views. Therefore, how to achieve accurate tracking with unaligned neuromorphic and visible sensors is a valuable but unresearched problem. In this work, we formally propose the task of object tracking using unaligned neuromorphic and visible cameras. We build the first unaligned frame-event dataset CRSOT collected with a specially built data acquisition system, which contains 1,030 high-definition RGB-Event video pairs, 304,974 video frames. In addition, we propose a novel unaligned object tracking framework that can realize robust tracking even using the loosely aligned RGB-Event data. This proposed method utilizes uncertainty perception techniques, which can effectively reduce the negative impact of noise (especially noise in event data) on tracking performance. Specifically, we extract the template and search regions of RGB and Event data and feed them into a unified ViT backbone for feature embedding. Next, we propose uncertainty perception modules to encode the RGB and Event features, respectively, then, we propose a modality uncertainty fusion module to aggregate the two modalities. These three branches are jointly optimized in the training phase. Extensive experiments demonstrate that our tracker can collaborate the dual modalities for high-performance tracking even without strictly temporal and spatial alignment. Yabin Zhu, Xiao Wang 0014, Chenglong Li 0002, Bo Jiang 0002, Lin Zhu 0012, Zhixiang Huang, Yonghong Tian 0001, Jin Tang 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Transductive Few-shot Learning via Joint Message Passing and Prototype-based Soft-label PropagationabstractThe transductive Few-shot Learning (FSL) mostly employs either prototype learning or label propagation methods to generalize to new classes by using the information of all query samples. However, existing methods have several main limitations. First, the prototype methods mainly focus on support samples which fail to fully exploit the relationships of query samples. Second, existing label propagation methods are generally not effective for the class-imbalanced problem. Third, existing works usually optimize the learnable parameters during inference which significantly reduces the efficiency of existing methods. To address these limitations, this article proposes an efficient and robust method for transductive FSL problem, termed Prototype-based Soft-label Propagation (PSLP), which combines the prototype learning and label propagation together for FSL problem. In our proposed method, first, the soft-label presentation for each query sample is estimated by leveraging prototypes. Then, the soft-label propagation is conducted on the learned query-support graph and the prototype representation is rectified. Both steps are conducted progressively for boosting the performance. Moreover, to learn effective prototypes for soft-label estimation and the desirable query-support graph for soft-label propagation, we design a new joint message passing scheme to learn the sample presentation and relational graph jointly. The PSLP method is parameter-free and can be implemented very efficiently. The experiments conducted on four popular benchmarks show that our method achieves competitive results on both balanced and imbalanced settings compared to the state-of-the-art methods. The code is released at https://github.com/mobulan/PSLP . Bo Jiang 0002, Bin Luo 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | HARDVS: Revisiting Human Activity Recognition with Dynamic Vision SensorsabstractThe main streams of human activity recognition (HAR) algorithms are developed based on RGB cameras which usually suffer from illumination, fast motion, privacy preservation, and large energy consumption. Meanwhile, the biologically inspired event cameras attracted great interest due to their unique features, such as high dynamic range, dense temporal but sparse spatial resolution, low latency, low power, etc. As it is a newly arising sensor, even there is no realistic large-scale dataset for HAR. Considering its great practical value, in this paper, we propose a large-scale benchmark dataset to bridge this gap, termed HARDVS, which contains 300 categories and more than 100K event sequences. We evaluate and report the performance of multiple popular HAR algorithms, which provide extensive baselines for future works to compare. More importantly, we propose a novel spatial-temporal feature learning and fusion framework, termed ESTF, for event stream based human activity recognition. It first projects the event streams into spatial and temporal embeddings using StemNet, then, encodes and fuses the dual-view representations using Transformer networks. Finally, the dual features are concatenated and fed into a classification head for activity prediction. Extensive experiments on multiple datasets fully validated the effectiveness of our model. Both the dataset and source code will be released at https://github.com/Event-AHU/HARDVS. Xiao Wang 0014, Zongzhen Wu, Bo Jiang 0002, Zhimin Bao, Lin Zhu 0012, Guoqi Li 0002, Yaowei Wang 0001, Yonghong Tian 0001 |
AAAI | 3 |
| 2024 | Event Stream-Based Visual Object Tracking: A High-Resolution Benchmark Dataset and A Novel BaselineabstractTracking with bio-inspired event cameras has garnered increasing interest in recent years. Existing works either utilize aligned RGB and event data for accurate tracking or directly learn an event-based tracker. The former incurs higher inference costs while the latter may be susceptible to the impact of noisy events or sparse spatial resolution. In this paper, we propose a novel hierarchical knowledge distillation framework that can fully utilize multimodal / multi-view information during training to facilitate knowledge transfer, enabling us to achieve high-speed and low-latency visual tracking during testing by using only event signals. Specifically, a teacher Transformer-based multimodal tracking framework is first trained by feeding the RGB frame and event stream simultaneously. Then, we design a new hierarchical knowledge distillation strategy which includes pairwise similarity, feature representation, and response maps-based knowledge distillation to guide the learning of the student Transformer network. In particular, since existing event-based tracking datasets are all low-resolution (346 × 260), we propose the first large-scale high-resolution (1280 × 720) dataset named EventVOT. It contains 1141 videos and covers a wide range of categories such as pedestrians, vehicles, UAVs, ping pong, etc. Ex-tensive experiments on both low-resolution (FE240hz, Vi-sEvent, COESOT), and our newly proposed high-resolution EventVOT dataset fully validated the effectiveness of our proposed method. Xiao Wang 0014, Shiao Wang, Chuanming Tang, Lin Zhu 0012, Bo Jiang 0002, Yonghong Tian 0001, Jin Tang 0001 |
CVPR | 5 |
| 2024 | MGDR: Multi-modal Graph Disentangled Representation for Brain Disease Prediction
Bo Jiang 0002, Xixi Wan, Yuan Chen 0012, Zhengzheng Tu, Yumiao Zhao, Jin Tang 0001 |
MICCAI (2) | 1 |
| 2024 | Mamba-FETrack: Frame-Event Tracking via State Space Model
Ju Huang, Shiao Wang, Zhe Wu 0006, Xiao Wang 0014, Bo Jiang 0002 |
PRCV (12) | 6 |
| 2024 | MutualFormer: Multi-modal Representation Learning via Cross-Diffusion Attention
Xixi Wang 0005, Xiao Wang 0014, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001 |
Int. J. Comput. Vis. | 3 |
| 2024 | Learning Graph Attentions via Replicator DynamicsabstractGraph Attention (GA) which aims to learn the attention coefficients for graph edges has achieved impressive performance in GNNs on many graph learning tasks. However, existing GAs are usually learned based on edges' (or connected nodes') features which fail to fully capture the rich structural information of edges. Some recent research attempts to incorporate the structural information into GA learning but how to fully exploit them in GA learning is still a challenging problem. To address this challenge, in this work, we propose to leverage a new Replicator Dynamics model for graph attention learning, termed Graph Replicator Attention (GRA). The core of GRA is our derivation of replicator dynamics based sparse attention diffusion which can explicitly learn context-aware and sparse preserved graph attentions via a simple self-supervised way. Moreover, GRA can be theoretically explained from an energy minimization model. This provides a more theoretical justification for the proposed GRA method. Experiments on several graph learning tasks demonstrate the effectiveness and advantages of the proposed GRA method on ten benchmark datasets. Bo Jiang 0002, Sheng Ge, Beibei Wang 0006, Xiao Wang 0014, Jin Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Robust Audio-Visual Contrastive Learning for Proposal-Based Self-Supervised Sound Source Localization in VideosabstractBy observing a scene and listening to corresponding audio cues, humans can easily recognize where the sound is. To achieve such cross-modal perception on machines, existing methods take advantage of the maps obtained by interpolation operations to localize the sound source. As semantic object-level localization is more attractive for prospective practical applications, we argue that these map-based methods only offer a coarse-grained and indirect description of the sound source. Additionally, these methods utilize a single audio-visual tuple at a time during self-supervised learning, causing the model to lose the crucial chance to reason about the data distribution of large-scale audio-visual samples. Although the introduction of Audio-Visual Contrastive Learning (AVCL) can effectively alleviate this issue, the contrastive set constructed by randomly sampling is based on the assumption that the audio and visual segments from all other videos are not semantically related. Since the resulting contrastive set contains a large number of faulty negatives, we believe that this assumption is rough. In this paper, we advocate a novel proposal-based solution that directly localizes the semantic object-level sound source, without any manual annotations. The Global Response Map (GRM) is incorporated as an unsupervised spatial constraint to filter those instances corresponding to a large number of sound-unrelated regions. As a result, our proposal-based Sound Source Localization (SSL) can be cast into a simpler Multiple Instance Learning (MIL) problem. To overcome the limitation of random sampling in AVCL, we propose a novel Active Contrastive Set Mining (ACSM) to mine the contrastive sets with informative and diverse negatives for robust AVCL. Our approaches achieve state-of-the-art (SOTA) performance when compared to several baselines on multiple SSL datasets with diverse scenarios. Hanyu Xuan, Zhiliang Wu, Jian Yang 0003, Bo Jiang 0002, Lei Luo 0001, Xavier Alameda-Pineda, Yan Yan 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | SemanticFormer: Hyperspectral image classification via semantic transformer
Xixi Wang 0005, Bo Jiang 0002, Lan Chen 0003, Bin Luo 0001 |
Pattern Recognit. Lett. | 3 |
| 2024 | MAPS: A Noise-Robust Progressive Learning Approach for Source-Free Domain Adaptive Keypoint DetectionabstractExisting cross-domain keypoint detection methods always require accessing the source data during adaptation, which may violate the data privacy law and pose serious security concerns. Instead, this paper considers a realistic problem setting called source-free domain adaptive keypoint detection, where only the well-trained source model is provided to the target domain. For the challenging problem, we first construct a teacher-student learning baseline by stabilizing the predictions under data augmentation and network ensembles. Built on this, we further propose a unified approach, Mixup Augmentation and Progressive Selection (MAPS), to fully exploit the noisy pseudo labels of unlabeled target data during training. On the one hand, MAPS regularizes the model to favor simple linear behavior in-between the target samples via self-mixup augmentation, preventing the model from over-fitting to noisy predictions. On the other hand, MAPS employs the self-paced learning paradigm and progressively selects pseudo-labeled samples from ‘easy’ to ‘hard’ into the training process to reduce noise accumulation. Results on four keypoint detection datasets show that MAPS outperforms the baseline and achieves comparable or even better results in comparison to previous non-source-free counterparts. The code is available athttps://github.com/YuheD/MAPS. Yuhe Ding, Jian Liang 0001, Bo Jiang 0002, Aihua Zheng, Ran He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | ADRNet: Affine and Deformable Registration Networks for Multimodal Remote Sensing ImagesabstractMulti-modal remote sensing images registration ensures the consistency of the spatial positions for different images. It can provide the accurate geographic information and supports the fusion of multi-source data for geospatial analyses and applications. Rigid registration method shows high performance in dealing with large-scale deformation, but it is difficult to achieve high-precision image registration. In contrast, non-rigid registration method is suitable for processing local differences, but cannot effectively deal with large-scale deformation differences. Therefore, the combination of rigid and non-rigid registration methods becomes a necessary strategy to address such issues. In this paper, we propose a novel ADRNet method for multi-modal remote sensing images registration. The proposed ADRNet method contains three main modules: affine registration module, deformable registration module, and spatial transformer module that integrates the affine and deformable transformation parameters to obtain the final aligned images. Meanwhile, we design a new feature enhancement module and an attention module with dilated convolutions which have different dilation rates, which are used to alleviate the limitations imposed by receptive fields in the convolution operation. Moreover, we propose a specific symmetric loss function to optimize the whole network from the perspective of inverse consistency. To assess the efficiency and performance of the network, we extend the experimental data, ranging from cross-modal images in a conventional viewpoint to cross-modal images in a remote sensing viewpoint. The experimental results show that our method exhibits excellent performance for the images with different viewpoints and deformation scales. The relevant code will be released at: https://github.com/Ahuer-Lei/ADRNet. Yun Xiao 0003, Yuan Chen 0012, Bo Jiang 0002, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Multi-Granularity Part Sampling Attention for Fine-Grained Visual ClassificationabstractFine-grained visual classification aims to classify similar sub-categories with the challenges of large variations within the same sub-category and high visual similarities between different sub-categories. Recently, methods that extract semantic parts of the discriminative regions have attracted increasing attention. However, most existing methods extract the part features via rectangular bounding boxes by object detection module or attention mechanism, which makes it difficult to capture the rich shape information of objects. In this paper, we propose a novel Multi-Granularity Part Sampling Attention (MPSA) network for fine-grained visual classification. First, a novel multi-granularity part retrospect block is designed to extract the part information of different scales and enhance the high-level feature representation with discriminative part features of different granularities. Then, to extract part features of various shapes at each granularity, we propose part sampling attention, which can sample the implicit semantic parts on the feature maps comprehensively. The proposed part sampling attention not only considers the importance of sampled parts but also adopts the part dropout to reduce the overfitting issue. In addition, we propose a novel multi-granularity fusion method to highlight the foreground features and suppress the background noises with the assistance of the gradient class activation map. Experimental results demonstrate that the proposed MPSA achieves state-of-the-art performance on four commonly used fine-grained visual classification benchmarks. The source code is publicly available at https://github.com/mobulan/MPSA. Bo Jiang 0002, Bin Luo 0001, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | AMatFormer: Efficient Feature Matching via Anchor Matching TransformerabstractLearning based feature matching methods have been commonly studied in recent years. The core issue for learning feature matching is to how to learn (1) discriminative representations for feature points (or regions) within each intra-image and (2) consensus representations for feature points across inter-images. Recently, self- and cross-attention models have been exploited to address this issue. However, in many scenes, features are coming with large-scale, redundant and outliers contaminated. Previous self-/cross-attention models generally conduct message passing on all primal features which thus lead to redundant learning and high computational cost. To mitigate limitations, inspired by recent seed matching methods, in this article, we propose a novel efficient Anchor Matching Transformer (AMatFormer) for the feature matching problem. AMatFormer has two main aspects: First, it mainly conducts self-/cross-attention on some anchor features and leverages these anchor features as message bottleneck to learn the representations for all primal features. Thus, it can be implemented efficiently and compactly. Second, AMatFormer adopts a shared FFN module to further embed the features of two images into the common domain and thus learn the consensus feature representations for the matching problem. Experiments on several benchmarks demonstrate the effectiveness and efficiency of the proposed AMatFormer matching approach. Bo Jiang 0002, Shuxian Luo, Xiao Wang 0014, Chuanfu Li, Jin Tang 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Rethinking Batch Sample Relationships for Data Representation: A Batch-Graph Transformer Based ApproachabstractExploring sample relationships within each mini-batch has shown great potential for learning image representations. Existing works generally adopt the regular Transformer to model the visual content relationships, ignoring the cues of semantic/label correlations between samples. Also, they generally adopt the ‘full’ self-attention mechanism which are obviously redundant and also sensitive to the noisy samples. To overcome these issues, in this paper, we design a simple yet flexible Batch-Graph Transformer (BGFormer) for mini-batch sample representations by deeply capturing the relationships of image samples from both visual and semantic perspectives. BGFormer has three main aspects. (1) It employs a flexible graph model, termedBatch Graphto jointly encode both visual and semantic relationships of samples within each mini-batch. (2) It explores the neighborhood relationships of samples by borrowing the idea of sparse graph representation which thus performs robustly, w.r.t., noisy samples. (3) It devises a novel specific Transformer architecture that mainly adoptsdualstructure-constrained self-attention (SSA), together with graph normalization, FFN, etc, to carefully exploit the batch graph information for sample tokens (nodes) representations. As an application, we apply BGFormer to the metric learning tasks. Extensive experiments on four popular datasets demonstrate the effectiveness of the proposed model. Xixi Wang 0005, Bo Jiang 0002, Xiao Wang 0014, Jinhui Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | GDCNet: Graph Enrichment Learning via Graph Dropping Convolutional NetworksabstractGraph convolutional networks (GCNs) have been widely studied to address graph data representation and learning. In contrast to traditional convolutional neural networks (CNNs) that employ many various (spatial) convolution filters to obtain rich feature descriptors to encode complex patterns of image data, GCNs, however, are defined on the input observed graph G(X,A) and usually adopt the single fixed spatial convolution filter for graph data feature extraction. This limits the capacity of the existing GCNs to encode the complex patterns of graph data. To overcome this issue, inspired by depthwise separable convolution and DropEdge operation, we first propose to generate various graph convolution filters by randomly dropping out some edges from the input graph A . Then, we propose a novel graph-dropping convolution layer (GDCLayer) to produce rich feature descriptors for graph data. Using GDCLayer, we finally design a new end-to-end network architecture, that is, a graph-dropping convolutional network (GDCNet), for graph data learning. Experiments on several datasets demonstrate the effectiveness of the proposed GDCNet. Bo Jiang 0002, Beibei Wang 0006, Haiyun Xu, Jin Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Graph context-attention network via low and high order aggregation
Haiyun Xu, Shaojie Zhang 0002, Bo Jiang 0002, Jin Tang 0001 |
Neurocomputing | 3 |
| 2023 | Multi-granularity cross attention network for person re-identification
Chengmei Han, Bo Jiang 0002, Jin Tang 0001 |
Multim. Tools Appl. | 2 |
| 2023 | Hypergraph convolutional network for hyperspectral image classification
Bo Jiang 0002, Jinpei Liu, Bin Luo 0001 |
Neural Comput. Appl. | 3 |
| 2023 | DropAGG: Robust Graph Neural Networks via Drop Aggregation
Bo Jiang 0002, Beibei Wang 0006, Haiyun Xu, Bin Luo 0001 |
Neural Networks | 1 |
| 2023 | Graph Neural Network Meets Sparse Representation: Graph Sparse Neural Networks via Exclusive Group LassoabstractExisting GNNs usually conduct the layer-wise message propagation via the 'full' aggregation of all neighborhood information which are usually sensitive to the structural noises existed in the graphs, such as incorrect or undesired redundant edge connections. To overcome this issue, we propose to exploit Sparse Representation (SR) theory into GNNs and propose Graph Sparse Neural Networks (GSNNs) which conduct sparse aggregation to select reliable neighbors for message aggregation. GSNNs problem contains discrete/sparse constraint which is difficult to be optimized. Thus, we then develop a tight continuous relaxation model Exclusive Group Lasso GNNs (EGLassoGNNs) for GSNNs. An effective algorithm is derived to optimize the proposed EGLassoGNNs model. Experimental results on several benchmark datasets demonstrate the better performance and robustness of the proposed EGLassoGNNs model. Bo Jiang 0002, Beibei Wang 0006, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Generalizing Aggregation Functions in GNNs: Building High Capacity and Robust GNNs via Nonlinear AggregationabstractThe main aspect powering GNNs is the multi-layer network architecture to learn the nonlinear representation for graph learning task. The core operation in GNNs is the message propagation in which each node updates its information by aggregating the information from its neighbors. Existing GNNs usually adopt either linear neighborhood aggregation (e.g. mean, sum) or max aggregator in their message propagation. 1) For linear aggregators, the whole nonlinearity and network's capacity of GNNs are generally limited because deeper GNNs usually suffer from the over-smoothing issue due to their inherent information propagation mechanism. Also, linear aggregators are usually vulnerable to the spatial perturbations. 2) For max aggregator, it usually fails to be aware of the detailed information of node representations within neighborhood. To overcome these issues, we re-think the message propagation mechanism in GNNs and develop the new general nonlinear aggregators for neighborhood information aggregation in GNNs. One main aspect of our nonlinear aggregators is that they all provide the optimally balanced aggregator between max and mean/sum aggregators. Thus, they can inherit both i) high nonlinearity that enhances network's capacity, robustness and ii) detail-sensitivity that is aware of the detailed information of node representations in GNNs' message propagation. Promising experiments show the effectiveness, high capacity and robustness of the proposed methods. Beibei Wang 0006, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Sparse norm regularized attribute selection for graph neural networks
Bo Jiang 0002, Beibei Wang 0006, Bin Luo 0001 |
Pattern Recognit. | 1 |
| 2023 | Few-Shot Learning Meets Transformer: Unified Query-Support Transformers for Few-Shot ClassificationabstractThe goal of Few-shot classification (FSL) is to identify unseen classes with very limited samples has attracted more and more attention. Usually, it is formulated as a metric learning problem. The core issue of few-shot classification is how to learn (1) consistent representations for images in both support and query sets and (2) effective metric learning for images between support and query sets. In this paper, we show that the two challenges can be well modeled simultaneously via a unified Query-Support TransFormer (QSFormer) model. To be specific, the proposed QSFormer involves global query-support sample Transformer (sampleFormer) branch and local patch Transformer (patchFormer) learning branch. sampleFormer aims to capture the dependence of samples in support and query sets for image representation. It adopts the Encoder, QS-Decoder and Cross-Attention to respectively model the Support, Query (image) representation and Metric learning for few-shot classification task. Also, as a complementary to global learning branch, we adopt a local patch Transformer to extract structural representation for each image sample by capturing the long-range dependence of local image patches. In addition, we introduce a novel Cross-scale Interactive Feature Extractor (CIFE) to extract and fuse different scale CNN features as an effective backbone module for the proposed few-shot learning method. We integrate these into a unified framework and train it in an end-to-end way. A large number of experiments are conducted on four popular datasets to validate the superiority and effectiveness of the proposed QSFormer. Xixi Wang 0005, Xiao Wang 0014, Bo Jiang 0002, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | VcT: Visual Change Transformer for Remote Sensing Image Change DetectionabstractGiven two remote sensing images, the goal of visual change detection task is to detect significantly changed areas between them. Existing visual change detectors usually adopt CNNs or Transformers for feature representation learning and focus on learning effective representation for the changed regions between images. Although good performance can be obtained by enhancing the features of the change regions, however, these works are still limited mainly due to the ignorance of mining the unchanged background context information. It is known that one main challenge for change detection is how to obtain the consistent representations for two images involving different variations, such as spatial variation, sunlight intensity, etc. In this work, we demonstrate that carefully mining the common background information provides an important cue to learn the consistent representations for the two images which thus obviously facilitates the visual change detection problem. Based on this observation, we propose a novel Visual change Transformer (VcT) model for visual change detection problem. To be specific, a shared backbone network is first used to extract the feature maps for the given image pair. Then, each pixel of feature map is regarded as a graph node and the graph neural network is proposed to model the structured information for coarse change map prediction. Top-K reliable tokens can be mined from the map and refined by using the clustering algorithm. Then, these reliable tokens are enhanced by first utilizing self/cross-attention schemes and then interacting with original features via an anchor-primary attention learning module. Finally, the prediction head is proposed to get a more accurate change map. Extensive experiments on multiple benchmark datasets validated the effectiveness of our proposed VcT model. The source code and pre-trained models are available at https://github.com/Event-AHU/VcT_Remote_Sensing_Change_Detection. Bo Jiang 0002, Zitian Wang, Xixi Wang 0005, Lan Chen 0003, Xiao Wang 0014, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | AS3ITransUNet: Spatial-Spectral Interactive Transformer U-Net With Alternating Sampling for Hyperspectral Image Super-ResolutionabstractSingle hyperspectral image (HSI) super-resolution (SR) is an important topic in remote sensing field. However, existing HSI SR methods mainly use the feed-forward upsampling technique and convolutional neural network (CNN) to learn the feature representation, failing to learn the complex mapping relationship between low-resolution (LR) and high-resolution (HR) and long-range joint spectral and spatial features. To address this issue, in this paper, we propose the Spatial-Spectral Interactive Transformer U-Net with Alternating Sampling (AS3ITransUNet) for the HSI SR task. In this method, to mitigate the computational burden resulting from the high spectral dimension of HSI, a group reconstruction strategy is adopted. To effectively explore the hierarchical features of HSI, the U-Net with alternating upsampling and downsampling is designed that allocates the task of learning the complex mapping relationship to each stage of U-Net. To fully extract the spatial-spectral features of HSI, we propose the spatial-spectral interactive transformer (SSIT) block and integrate it into the encoder and decoder of U-Net. The SSIT block contains a cross-branch bidirectional interaction module, which further captures the complementary information between spatial and spectral dimensions. Moreover, the multi-stage complementary information learning (MFEL) is proposed to capture the complementary information in the adjacent HSI groups for recovering the absent details in the current HSI group. The experiments on the three benchmark datasets demonstrate that the proposed AS3ITransUNet can effectively improve the spatial resolution and preserve the spectral information at different scales. Models and code are available at https://github.com/liushiji666/AS3-ITransUNet. Shiji Liu, Bo Jiang 0002, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | MFGNet: Dynamic Modality-Aware Filter Generation for RGB-T TrackingabstractMany RGB-T trackers attempt to attain robust feature representation by utilizing an adaptive weighting scheme (or attention mechanism). Different from these works, we propose a new dynamic modality-aware filter generation module (named MFGNet) to boost the message communication between visible and thermal data by adaptively adjusting the convolutional kernels for various input images in practical tracking. Given the image pairs as input, we first encode their features with the backbone network. Then, we concatenate these feature maps and generate dynamic modality-aware filters with two independent networks. The visible and thermal filters will be used to conduct a dynamic convolutional operation on their corresponding input feature maps respectively. Inspired by residual connection, both the generated visible and thermal feature maps will be summarized with input feature maps. The augmented feature maps will be fed into the RoI align module to generate instance-level features for subsequent classification. To address issues caused by heavy occlusion, fast motion and out-of-view, we propose to conduct a joint local and global search by exploiting a new direction-aware target driven attention mechanism. The spatial and temporal recurrent neural network is used to capture the direction-aware context for accurate global attention prediction. Extensive experiments on three large-scale RGB-T tracking benchmark datasets validated the effectiveness of our proposed algorithm. Xiao Wang 0014, Xiujun Shu, Shiliang Zhang, Bo Jiang 0002, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Fine-Grained Visual Classification via Internal Ensemble Learning TransformerabstractRecently, vision transformers (ViTs) have been investigated in fine-grained visual recognition (FGVC) and are now considered state of the art. However, most ViT-based works ignore the different learning performances of the heads in the multi-head self-attention (MHSA) mechanism and its layers. To address these issues, in this paper, we propose a novel internal ensemble learning transformer (IELT) for FGVC. The proposed IELT involves three main modules: multi-head voting (MHV) module, cross-layer refinement (CLR) module, and dynamic selection (DS) module. To solve the problem of the inconsistent performances of multiple heads, we propose the MHV module, which considers all of the heads in each layer as weak learners and votes for tokens of discriminative regions as cross-layer feature based on the attention maps and spatial relationships. To effectively mine the cross-layer feature and suppress the noise, the CLR module is proposed, where the refined feature is extracted and the assist logits operation is developed for the final prediction. In addition, a newly designed DS module adjusts the token selection number at each layer by weighting their contributions of the refined feature. In this way, the idea of ensemble learning is combined with the ViT to improve fine-grained feature representation. The experiments demonstrate that our method achieves competitive results compared with the state of the art on five popular FGVC datasets. Source code has been released and can be found athttps://github.com/mobulan/IELT. Bo Jiang 0002, Bin Luo 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | GPENs: Graph Data Learning With Graph Propagation-Embedding NetworksabstractCompact representation of graph data is a fundamental problem in pattern recognition and machine learning area. Recently, graph neural networks (GNNs) have been widely studied for graph-structured data representation and learning tasks, such as graph semi-supervised learning, clustering, and low-dimensional embedding. In this article, we present graph propagation-embedding networks (GPENs), a new model for graph-structured data representation and learning problem. GPENs are mainly motivated by 1) revisiting of traditional graph propagation techniques for graph node context-aware feature representation and 2) recent studies on deeply graph embedding and neural network architecture. GPENs integrate both feature propagation on graph and low-dimensional embedding simultaneously into a unified network using a novel propagation-embedding architecture. GPENs have two main advantages. First, GPENs can be well-motivated and explained from feature propagation and deeply learning architecture. Second, the equilibrium representation of the propagation-embedding operation in GPENs has both exact and approximate formulations, both of which have simple closed-form solutions. This guarantees the compactivity and efficiency of GPENs. Third, GPENs can be naturally extended to multiple GPENs (M-GPENs) to address the data with multiple graph structures. Experiments on various semi-supervised learning tasks on several benchmark datasets demonstrate the effectiveness and benefits of the proposed GPENs and M-GPENs. Bo Jiang 0002, Leiling Wang, Jian Cheng 0001, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Semi-supervised Learning via Multiple Layer Graph Regularized PerceptionabstractRecently, Graph Neural Networks (GNNs) have made remarkable achievements in semi-supervised classification tasks. Nevertheless, GNNs usually rely on a specific graph convolution which has high computational complexity. To overcome this issue, recent works attempt to implicitly use adjacency matrix to guide message propagation in multi-layer perception (MLP) via neighboring contrastive loss. However, existing works accomplish implicit message passing only, without considering multi-order graph topology information. In this paper, we propose a novel method called Multiple Layer Graph Regularized Perception (MLGP). The main advantage of MLGP is to incorporate multi-order neighboring information into MLP. Further, inspired by gated mechanism, we design a linear gating to capture important features of nodes. More discriminant features can be obtained to alleviate over-smoothing. MLGP is more effective and more robust than existing works when dealing with large-scale graph data and noisy adjacency information. The comparative experiment results show that our model achieves better performance and strong robustness. Haiyun Xu, Lili Huang 0006, Bo Jiang 0002, Jin Tang 0001, Shaojie Zhang 0002 |
ICPR | 3 |
| 2022 | Attributes Based Visible-Infrared Person Re-identification
Aihua Zheng, Mengya Feng, Bo Jiang 0002, Bin Luo 0001 |
PRCV (1) | 4 |
| 2022 | Multi-head collaborative learning for graph neural networks
Haiyun Xu, Bo Jiang 0002, Lili Huang 0006, Jin Tang 0001, Shaojie Zhang 0002 |
Neurocomputing | 2 |
| 2022 | MGLNN: Semi-supervised learning via Multiple Graph Cooperative Learning Neural Networks
Bo Jiang 0002, Beibei Wang 0006, Bin Luo 0001 |
Neural Networks | 1 |
| 2022 | LGLNN: Label Guided Graph Learning-Neural Network for few-shot learning
Kangkang Zhao, Bo Jiang 0002, Jin Tang 0001 |
Neural Networks | 3 |
| 2022 | GeCNs: Graph Elastic Convolutional Networks for Data RepresentationabstractGraph representation and learning is a fundamental problem in machine learning area. Graph Convolutional Networks (GCNs) have been recently studied and demonstrated very powerful for graph representation and learning. Graph convolution (GC) operation in GCNs can be regarded as a composition of feature aggregation and nonlinear transformation step. Existing GCs generally conduct feature aggregation on a full neighborhood set in which each node computes its representation by aggregating the feature information of all its neighbors. However, this full aggregation strategy is not guaranteed to be optimal for GCN learning and also can be affected by some graph structure noises, such as incorrect or undesired edge connections. To address these issues, we propose to integrate elastic net based selection into graph convolution and propose a novel graph elastic convolution (GeC) operation. In GeC, each node can adaptively select the optimal neighbors in its feature aggregation. The key aspect of the proposed GeC operation is that it can be formulated by a regularization framework, based on which we can derive a simple update rule to implement GeC in a self-supervised manner. Using GeC, we then present a novel GeCN for graph learning. Experimental results demonstrate the effectiveness and robustness of GeCN. Bo Jiang 0002, Beibei Wang 0006, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | GLMNet: Graph learning-matching convolutional networks for feature matching
Bo Jiang 0002, Bin Luo 0001 |
Pattern Recognit. | 1 |
| 2022 | RGTransformer: Region-Graph Transformer for Image Representation and Few-Shot ClassificationabstractThe goal of few-shot image classification is to learn a classifier that can be well generalized to the unseen classes with a few available labeled samples. One major challenge for few-shot learning is how to conduct effective image representation for support and query images. Recently, local region-based image representation and metric learning approaches have been demonstrated effectively for few-shot classification problem. However, existing approaches generally conduct representations of image regions individually which thus lack of considering the rich spatial/structural relationships among image regions. In this paper, we propose to bridge the individual regions and exploit the structural contexts among regions via a novel Region-Graph Transformer (RGTransformer). In RGTransformer, each region aggregates the information from its neighboring regions and thus can obtain context-aware feature representations for regions. Using the proposed RGTransformer, we propose an effective metric learning model for few-shot image classification. We evaluate the proposed method on four benchmark datasets and experimental results demonstrate the effectiveness and advantages of the proposed RGTransformer. Bo Jiang 0002, Kangkang Zhao, Jin Tang 0001 |
IEEE Signal Process. Lett. | 1 |
| 2022 | Beyond Greedy Search: Tracking by Multi-Agent Reinforcement Learning-Based Beam SearchabstractTo track the target in a video, current visual trackers usually adopt greedy search for target object localization in each frame, that is, the candidate region with the maximum response score will be selected as the tracking result of each frame. However, we found that this may be not an optimal choice, especially when encountering challenging tracking scenarios such as heavy occlusion and fast motion. In particular, if a tracker drifts, errors will be accumulated and would further make response scores estimated by the tracker unreliable in future frames. To address this issue, we propose to maintain multiple tracking trajectories and apply beam search strategy for visual tracking, so that the trajectory with fewer accumulated errors can be identified. Accordingly, this paper introduces a novel multi-agent reinforcement learning based beam search tracking strategy, termed BeamTracking. It is mainly inspired by the image captioning task, which takes an image as input and generates diverse descriptions using beam search algorithm. Accordingly, we formulate the tracking as a sample selection problem fulfilled by multiple parallel decision-making processes, each of which aims at picking out one sample as their tracking result in each frame. Each maintained trajectory is associated with an agent to perform the decision-making and determine what actions should be taken to update related information. More specifically, using the classification-based tracker as the baseline, we first adopt bi-GRU to encode the target feature, proposal feature, and its response score into a unified state representation. The state feature and greedy search result are then fed into the first agent for independent action selection. Afterwards, the output action and state features are fed into the subsequent agent for diverse results prediction. When all the frames are processed, we select the trajectory with the maximum accumulated score as the tracking result. Extensive experiments on seven popular tracking benchmark datasets validated the effectiveness of the proposed algorithm. Xiao Wang 0014, Zhe Chen 0013, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Dacheng Tao |
IEEE Trans. Image Process. | 3 |
| 2022 | PH-GCN: Person Retrieval With Part-Based Hierarchical Graph Convolutional NetworkabstractCompact feature representation of person image is important for person re-identification (Re-ID) task. Recently, part-based representation models have been widely studied for extracting the more compact and robust feature representation for person image to improve person Re-ID results. However, existing part-based representation models mostly extract the features of different parts independently which ignore the spatial relationship information among different parts. To address this issue, in this paper we propose a novel deep learning framework, named Part-based Hierarchical Graph Convolutional Network (PH-GCN) for person Re-ID problem. Given a person image, PH-GCN first constructs a hierarchical graph to represent the spatial relationships among different parts. Then, both local and global feature learning is achieved by the feature information passing in PH-GCN, which takes the information of other parts into account for part feature representation. Finally, a perceptron layer is adopted for the final person part label prediction and re-identification. The proposed framework provides a general solution that integrateslocal,globalandstructuralfeature learning simultaneously in a unified end-to-end network representation and learning. Extensive experiments on several widely used benchmark datasets demonstrate the effectiveness and benefits of the proposed PH-GCN approach for person Re-ID task. Bo Jiang 0002, Xixi Wang 0005, Aihua Zheng, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Adversarial-Metric Learning for Audio-Visual Cross-Modal MatchingabstractAudio-visual matching aims to learn the intrinsic correspondence between image and audio clip. Existing works mainly concentrate on learning discriminative features, while ignore the cross-modal heterogeneous issue between audio and visual modalities. To deal with this issue, we propose a novel Adversarial-Metric Learning (AML) model for audio-visual matching. AML aims to generate a modality-independent representation for each person in each modality via adversarial learning, while simultaneously learns a robust similarity measure for cross-modality matching via metric learning. By integrating the discriminative modality-independent representation and robust cross-modality metric learning into an end-to-end trainable deep network, AML can overcome the heterogeneous issue with promising performance for audio-visual matching. Experiments on the various audio-visual learning tasks, including audio-visual matching, audio-visual verification and audio-visual retrieval on benchmark dataset demonstrate the effectiveness of the proposed AML model. The implementation codes are available onhttps://github.com/MLanHu/AML. Aihua Zheng, Menglan Hu, Bo Jiang 0002, Yan Yan 0002, Bin Luo 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | Towards More Flexible and Accurate Object Tracking With Natural Language: Algorithms and BenchmarkabstractTracking by natural language specification is a new rising research topic that aims at locating the target object in the video sequence based on its language description. Compared with traditional bounding box (BBox) based tracking, this setting guides object tracking with high-level semantic information, addresses the ambiguity of BBox, and links local and global search organically together. Those benefits may bring more flexible, robust and accurate tracking performance in practical scenarios. However, existing natural language initialized trackers are developed and compared on benchmark datasets proposed for tracking-by-BBox, which can’t reflect the true power of tracking-by-language. In this work, we propose a new benchmark specifically dedicated to the tracking-by-language, including a large scale dataset, strong and diverse baseline methods. Specifically, we collect 2k video sequences (contains a total of 1,244,340 frames, 663 words) and split 1300/700 for the train/testing respectively. We densely annotate one sentence in English and corresponding bounding boxes of the target object for each video. We also introduce two new challenges into TNL2K for the object tracking task, i.e., adversarial samples and modality switch. A strong baseline method based on an adaptive local-global-search scheme is proposed for future works to compare. We believe this benchmark will greatly boost related researches on natural language guided tracking. Xiao Wang 0014, Xiujun Shu, Bo Jiang 0002, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001 |
CVPR | 4 |
| 2021 | MGARL: Multiple Graph Adversarial Regularized LearningabstractGraph Convolutional Networks (GCNs) have been commonly studied for graph learning tasks, such as semi-supervised learning, clustering etc. However, many existing GCNs are generally conducted on single graph data and thus can not be applied directly to multi-graph data that consists of various types of edges between nodes. To address this issue, in this paper, we propose a novel multiple Graph Adversarial Regularized Learning (mGARL) framework for multi-graph data representation and learning. mGARL aims to learn an optimal structure invariant/consistent representation for multiple graphs by employing a novel Encoder-Decoder architecture with adversarial learning regularization. It can incorporate the structural information of multiple graphs simultaneously for the node’s representation. We apply the proposed mGARL on the multi-view semi-supervised learning tasks. Experimental results on several datasets demonstrate the effectiveness and benefits of the proposed mGARL model. Bo Jiang 0002, Bin Luo 0001 |
ICME | 2 |
| 2021 | GAMnet: Robust Feature Matching via Graph Adversarial-Matching NetworkabstractRecently, deep graph matching (GM) methods have gained increasing attention. These methods integrate graph nodes¡¯s embedding, node/edges¡¯s affinity learning and final correspondence solver together in an end-to-end manner. For deep graph matching problem, one main issue is how to generate consensus node's embeddings for both source and target graphs that best serve graph matching tasks. In addition, it is also challenging to incorporate the discrete one-to-one matching constraints into the differentiable correspondence solver in deep matching network. To address these issues, we propose a novel Graph Adversarial Matching Network (GAMnet) for graph matching problem. GAMnet integrates graph adversarial embedding and graph matching simultaneously in a unified end-to-end network which aims to adaptively learn distribution consistent and domain invariant embeddings for GM tasks. Also, GAMnet exploits sparse GM optimization as correspondence solver which is differentiable and can also incorporate discrete one-to-one matching constraints approximately in natural in the final matching prediction. Experimental results on three public benchmarks demonstrate the effectiveness and benefits of the proposed GAMnet. Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001 |
ACM Multimedia | 1 |
| 2021 | A novel domain activation mapping-guided network (DA-GNT) for visual tracking
Zhengzheng Tu, Ajian Zhou, Chuang Gan 0003, Bo Jiang 0002, Amir Hussain 0001, Bin Luo 0001 |
Neurocomputing | 4 |
| 2021 | Co-Saliency Detection via a General Optimization Model and Adaptive Graph LearningabstractCo-saliency detection is an important research problem, and has been widely used in computer vision area. One main challenge for co-saliency detection problem is how to explore both interactive information among different images and individual salient information within each image simultaneously in co-saliency estimation. In this paper, we propose a novel general optimization framework with adaptive graph learning for co-saliency estimation problem. The proposed model integrates multiple cues including background, and foreground priors, structural information of images, and image feature representation together to obtain a uniform, and accurate co-saliency estimation. One main benefit of the proposed co-saliency method is that it conducts co-saliency propagation, and prediction across different images while maintains the individual salient information of each image, which ensures the consistency, and communication across different images effectively in co-saliency estimation. To improve the accuracy of co-saliency estimation, we adaptively learn a neighborhood, and structured graph to conduct co-saliency propagation among superpixels. An effective optimization algorithm has been designed to seek the optimal solution for the proposed co-saliency optimization model. Experimental results on several widely used datasets show that our method outperforms some other related co-saliency detection methods. Bo Jiang 0002, Xingyue Jiang, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | STGL: Spatial-Temporal Graph Representation and Learning for Visual TrackingabstractTracking-by-detection framework has been normally adopted in visual tracking methods. It aims to localize the visual target object with a bounding box. However, the bounding box is usually difficult to describe the target object accurately and thus easily introduces noisy background information, which usually degrades the final tracking results. Recently, weighted patch representation of the object has been shown very effectively for suppressing the undesirable background information and thus can obviously improve the tracking results. In this paper, we propose a novel Spatial-Temporal Graph representation and Learning (STGL) model to generate a kind of robust target representation for visual tracking problem. The main aspect of STGL is that it aims to exploit both spatial (within each frame) and temporal (between consecutive frames) structure of patches simultaneously in a unified graph representation and semi-supervised learning model. Comparing with existing works, STGL naturally exploits the learned representation of object in previous frame and thus can obtain the representation of object in current frame more accurately and robustly. A new ADMM algorithm is derived to solve the proposed STGL model. Based on the proposed object representation, we then adapt the structured SVM by introducing scale estimation to achieve object tracking. Extensive experiments show that our method outperforms the state-of-the-art patch based tracking methods on two standard benchmark datasets. Bo Jiang 0002, Bin Luo 0001, Xiaochun Cao, Jin Tang 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | cmSalGAN: RGB-D Salient Object Detection With Cross-View Generative Adversarial NetworksabstractImage salient object detection (SOD) is an active research topic in computer vision and multimedia area. Fusing complementary information of RGB and depth has been demonstrated to be effective for image salient object detection which is known as RGB-D salient object detection problem. The main challenge for RGB-D salient object detection is how to exploit the salient cues of both intra-modality (RGB, depth) and cross-modality simultaneously which is known as cross-modality detection problem. In this paper, we tackle this challenge by designing a novel cross-modality Saliency Generative Adversarial Network (cmSalGAN). cmSalGAN aims to learn an optimal view-invariant and consistent pixel-level representation for RGB and depth images via a novel adversarial learning framework, which thus incorporates both information of intra-view and correlation information of cross-view images simultaneously for RGB-D saliency detection problem. To further improve the detection results, the attention mechanism and edge detection module are also incorporated into cmSalGAN. The entire cmSalGAN can be trained in an end-to-end manner by using the standard deep neural network framework. Experimental results show that cmSalGAN achieves the new state-of-the-art RGB-D saliency detection performance on several benchmark datasets. Bo Jiang 0002, Zitai Zhou, Xiao Wang 0014, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | Joint graph regularized dictionary learning and sparse ranking for multi-modal multi-shot person re-identification
Aihua Zheng, Bo Jiang 0002, Wei-Shi Zheng 0001, Bin Luo 0001 |
Pattern Recognit. | 3 |
| 2020 | Feature Matching With Intra-Group Sparse ModelabstractFeature matching is a fundamental problem in computer vision area. In many real applications, one can usually obtain some potential (candidate) matches C by using some discriminative feature descriptors, such as SIFT descriptor. Then, the feature matching problem can be formulated as the problem of trying to select the correct matches S from the potential match set C. In this paper, we propose to solve matches selection by developing a novel intra-group sparse matching (IGSM) model. Our IGSM is motivated by a simple observation that the potential match set C can be divided into several non-overlapping groups Ci, among which the correct matches S are uniformly distributed. We thus develop an intra-group selection model to conduct matches selection at the intra-group level to incorporate the one-to-one matching constraint more in matches selection process. Our IGSM model has three main advantages: (1) The selection mechanism is parameter-free; (2) it generates an intra-group sparse solution which better maintains the one-to-one matching constraint in nature; (3) a simple yet effective update algorithm has been derived to solve IGSM model. The optimality and convergence of the algorithm are theoretically guaranteed. Experimental results on several image feature matching datasets show the effectiveness and efficiency of the proposed IGSM matching method. Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 1 |
| 2020 | Revisiting L2, 1-Norm Robustness With Vector Outlier RegularizationabstractIn many real-world applications, data usually contain outliers. One popular approach is to use the L2,1-norm function as a robust loss/error function. However, the robustness of the L2,1-norm function is not well understood so far. In this brief, we propose a new vector outlier regularization (VOR) framework to understand and analyze the robustness of the L2,1-norm function. Our VOR function defines a data point to be the outlier if it is outside a threshold with respect to a theoretical prediction, and regularizes it, i.e., pull it back to the threshold line. Thus, in the VOR function, how far an outlier lies away from its theoretical predicted value does not affect the final regularization and analysis results. One important aspect of the VOR function is that it has an equivalent continuous formulation, based on which we can prove that the L2,1-norm function is the limiting case of the proposed VOR function. Based on this theoretical result, we thus provide a new and intuitive explanation for the robustness property of the L2,1-norm function. As an example, we use the VOR function to matrix factorization and propose a VOR principal component analysis (PCA) (VORPCA). We show some benefits of VORPCA on data reconstruction and clustering tasks. Bo Jiang 0002, Chris Ding |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | A Subspace Learning Approach to Multishot Person ReidentificationabstractThis paper addresses the challenging problem of multishot person reidentification (Re-ID) in real world uncontrolled surveillance systems. A key issue is how to effectively represent and process the multiple data with various appearance information due to the variations of pose, occlusions, and viewpoints. To this end, this paper develops a novel subspace learning approach, which pursues regularized low-rank and sparse representation for multishot person Re-ID. For the images of a person crossing a certain camera, we assume that the appearances of those subset images with similar viewpoints against a camera draw from the same low-rank subspace, and all the images of a person under a camera lie on a union of low-rank subspaces. Based on this assumption, we propose to learn a nonnegative low-rank and sparse graph to represent the person images. Moreover, the recurring pattern prior is integrated into our model to refine the affinities among images. Extensive experiments on four public benchmark datasets yield impressive performance by improving 22.9% on imagery library for intelligent detection systems video re identification (iLIDS-VID), 42.4% on person RE-ID (PRID) dataset 2011, 39.7% and 30.6% on speech, audio, image, and video technology-SoftBio camera 3/8 and camera 5/8, respectively, and 1.6% on motion analysis and re identification set compared to the state-of-the-art methods. Aihua Zheng, Xuehan Zhang, Bo Jiang 0002, Bin Luo 0001, Chenglong Li 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2019 | Data Representation and Learning With Graph Diffusion-Embedding NetworksabstractRecently, graph convolutional neural networks have been widely studied for graph-structured data representation and learning. In this paper, we present Graph Diffusion-Embedding networks (GDENs), a new model for graph-structured data representation and learning. GDENs are motivated by our development of graph based feature diffusion. GDENs integrate both feature diffusion and graph node (low-dimensional) embedding simultaneously into a unified network by employing a novel diffusion-embedding architecture. GDENs have two main advantages. First, the equilibrium representation of the diffusion-embedding operation in GDENs can be obtained via a simple closed-form solution, which thus guarantees the compactivity and efficiency of GDENs. Second, the proposed GDENs can be naturally extended to address the data with multiple graph structures. Experiments on various semi-supervised learning tasks on several benchmark datasets demonstrate that the proposed GDENs significantly outperform traditional graph convolutional networks. Bo Jiang 0002, Doudou Lin, Jin Tang 0001, Bin Luo 0001 |
CVPR | 1 |
| 2019 | Semi-Supervised Learning With Graph Learning-Convolutional NetworksabstractGraph Convolutional Neural Networks (graph CNNs) have been widely used for graph data representation and semi-supervised learning tasks. However, existing graph CNNs generally use a fixed graph which may not be optimal for semi-supervised learning tasks. In this paper, we propose a novel Graph Learning-Convolutional Network (GLCN) for graph data representation and semi-supervised learning. The aim of GLCN is to learn an optimal graph structure that best serves graph CNNs for semi-supervised learning by integrating both graph learning and graph convolution in a unified network architecture. The main advantage is that in GLCN both given labels and the estimated labels are incorporated and thus can provide useful `weakly' supervised information to refine (or learn) the graph construction and also to facilitate the graph convolution operation for unknown label estimation. Experimental results on seven benchmarks demonstrate that GLCN significantly outperforms the state-of-the-art traditional fixed structure based graph CNNs. Bo Jiang 0002, Doudou Lin, Jin Tang 0001, Bin Luo 0001 |
CVPR | 1 |
| 2019 | Person Re-identification with Patch-Based Local Sparse Matching and Metric Learning
Bo Jiang 0002, Yibing Lv, Aihua Zheng, Bin Luo 0001 |
ICIG (2) | 1 |
| 2019 | Multiple Graph Convolutional Networks for Co-Saliency DetectionabstractRecently, Graph Convolutional Networks (GCNs) have been usually utilized for graph data representation in computer vision area. However, existing graph GCNs generally use a single graph which can be not adapted for the data with multiple graphs. In this paper, we first propose a novel multiple graph convolutional network (MGCN) for multiple graph data representation and learning. MGCN propagates information/knowledge across multiple graphs and obtains a consistent representation and learning by integrating the information of multiple graphs simultaneously. Based on the proposed MGCN, we then propose a new global-local unified graph convolutional learning architecture for image co-saliency detection problem. The main benefits of the proposed co-saliency model are twofold. First, it learns an optimal superpixel feature representation for co-saliency detection problem. Second, it can well exploit both intra-image and inter-image cues for co-saliency detection via a unified network. Promising experiments demonstrate the effectiveness of the proposed MGCN based co-saliency detection approach. Bo Jiang 0002, Xingyue Jiang, Jin Tang 0001, Bin Luo 0001, Shilei Huang |
ICME | 1 |
| 2019 | A Unified Multiple Graph Learning and Convolutional Network Model for Co-saliency EstimationabstractCo-saliency estimation which aims to identify the common salient object regions contained in an image set is an active problem in computer vision. The main challenge for co-saliency estimation problem is how to exploit the salient cues of both intra-image and inter-image simultaneously. In this paper, we first represent intra-image and inter-image as intra-graph and inter-graph respectively and formulate co-saliency estimation as graph nodes labeling. Then, we propose a novel multiple graph learning and convolutional network (M-GLCN) for image co-saliency estimation. M-GLCN conducts graph convolutional learning and labeling on both inter-graph and intra-graph cooperatively and thus can well exploit the salient cues of both intra-image and inter-image simultaneously for co-saliency estimation. Moreover, M-GLCN employs a new graph learning mechanism to learn both inter-graph and intra-graph adaptively. Experimental results on several benchmark datasets demonstrate the effectiveness of M-GLCN on co-saliency estimation task. Bo Jiang 0002, Xingyue Jiang, Ajian Zhou, Jin Tang 0001, Bin Luo 0001 |
ACM Multimedia | 1 |
| 2019 | Efficient Feature Matching via Nonnegative Orthogonal Relaxation
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001 |
Int. J. Comput. Vis. | 1 |
| 2019 | Robust visual tracking via Laplacian Regularized Random Walk Ranking
Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Chenglong Li 0002 |
Neurocomputing | 1 |
| 2019 | Robust pixelwise saliency detection via progressive graph rankings
Bo Jiang 0002, Zhengzheng Tu, Amir Hussain 0001, Jin Tang 0001 |
Neurocomputing | 2 |
| 2019 | Saliency detection via multi-view graph based saliency optimizationabstractSaliency detection is an important problem in computer vision and pattern recognition area. Many works have been proposed for addressing the saliency detection task. As a popular method, graph based saliency optimization has been widely studied. However, previous works have universally focussed on single graph optimization which fails to consider multi-view feature representation of image content . In this paper, we first provide a general framework for traditional graph based saliency optimization models. Then, we extend the general framework to the multi-view case and propose our general multi-view graph based saliency optimization model. Finally, we present a particular implementation of our general model and derive an effective updating algorithm to solve it. Experimental results using several benchmark datasets demonstrate the effectiveness of our proposed saliency model. Yun Xiao 0003, Bo Jiang 0002, Aihua Zheng, Aiwu Zhou, Amir Hussain 0001, Jin Tang 0001 |
Neurocomputing | 2 |
| 2019 | Background subtraction with multi-scale structured low-rank and sparse factorization
Aihua Zheng, Tian Zou, Yumiao Zhao, Bo Jiang 0002, Jin Tang 0001, Chenglong Li 0002 |
Neurocomputing | 4 |
| 2019 | Image Representation and Learning With Graph-Laplacian Tucker Tensor DecompositionabstractTucker tensor decomposition (TD) is widely used for image representation, reconstruction, and learning tasks. Compared to principal component analysis (PCA) models, tensor models retain more 2-D characteristics of images whereas PCA models linearize images. However, traditional TD involves attribute information only and thus does not consider the pairwise similarity information between images. In this paper, we propose a graph-Laplacian tucker tensor decomposition (GLTD) which explores both attributes and pairwise similarity information simultaneously. Generally, GLTD has three main benefits: 1) GLTD reconstruction shows clear robustness against image occlusions/outliers. We provide analysis to show that Laplacian regularization is mainly responsible to this robustness via an out-of-sample GLTD model. To the best of our knowledge, this Laplacian regularization induced robustness of TD has not been studied or emphasized before; 2) GLTD representation performs more regularity, which improves both unsupervised and supervised learning results; and 3) an effective algorithm is derived to solve GLTD problem. Although GLTD is a noncovex problem, the proposed algorithm is shown experimentally to provide a stable/unique solution starting from different random initializations. Experimental results on image reconstruction, data clustering, and classification tasks show the benefits of GLTD. Bo Jiang 0002, Chris Ding, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Cybern. | 1 |
| 2018 | OWP: Objectness Weighted Patch Descriptor for Visual TrackingabstractVisual object tracking is an active research problem and has been widely used in computer vision and pattern recognition area. Existing visual tracking methods usually localize the visual object with a bounding box which are often disturbed by the introduced background information and partial occlusion because of bounding box representation of visual object. To deal with this problem, in this paper, we propose a novel Objectness Weighted Patch (OWP) descriptor for object feature descriptor in visual tracking. The aim of OWP is to assign different objectness weights to the patches of bounding box to reduce the influences of background information and partial occlusion. We propose to compute the objectness weights of patches in OWP by integrating multiple cues (background, foreground and local spatial consistency) together in a general optimization model. Also, the proposed model has a simple closed-form solution and thus can be computed efficiently. We incorporate our OWP into structured SVM tracking framework and provide a new robust tracking method. Extensive experiments on two standard benchmark datasets OTB100 and Temple-Color demonstrate the effectiveness and benefits of the proposed tracking method. Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001 |
ICPR | 1 |
| 2018 | Multi-scale Cooperative Ranking for Saliency Detection
Bo Jiang 0002, Xingyue Jiang, Aihua Zheng, Yun Xiao 0003, Jin Tang 0001 |
PRCV (1) | 1 |
| 2018 | Non-negative Dual Graph Regularized Sparse Ranking for Multi-shot Person Re-identification
Aihua Zheng, Bo Jiang 0002, Chenglong Li 0002, Jin Tang 0001, Bin Luo 0001 |
PRCV (1) | 3 |
| 2018 | Robust data representation using locally linear embedding guided PCA
Bo Jiang 0002, Chris Ding, Bin Luo 0001 |
Neurocomputing | 1 |
| 2018 | Saliency detection via a multi-layer graph based diffusion model
Bo Jiang 0002, Zhouqin He, Chris Ding, Bin Luo 0001 |
Neurocomputing | 1 |
| 2018 | A prior regularized multi-layer graph ranking model for image saliency computation
Yun Xiao 0003, Bo Jiang 0002, Zhengzheng Tu, Jixin Ma 0001, Jin Tang 0001 |
Neurocomputing | 2 |
| 2018 | Spatial-temporal representatives selection and weighted patch descriptor for person re-identification
Aihua Zheng, Foqin Wang, Amir Hussain 0001, Jin Tang 0001, Bo Jiang 0002 |
Neurocomputing | 5 |
| 2017 | Nonnegative Orthogonal Graph MatchingabstractGraph matching problem that incorporates pair-wise constraints can be formulated as Quadratic Assignment Problem(QAP). The optimal solution of QAP is discrete and combinational, which makes QAP problem NP-hard. Thus, many algorithms have been proposed to find approximate solutions. In this paper, we propose a new algorithm, called Nonnegative Orthogonal Graph Matching (NOGM), for QAP matching problem. NOGM is motivated by our new observation that the discrete mapping constraint of QAP can be equivalently encoded by a nonnegative orthogonal constraint which is much easier to implement computationally. Based on this observation, we develop an effective multiplicative update algorithm to solve NOGM and thus can find an effective approximate solution for QAP problem. Comparing with many traditional continuous methods which usually obtain continuous solutions and should be further discretized, NOGM can obtain a sparse solution and thus incorporates the desirable discrete constraint naturally in its optimization. Promising experimental results demonstrate benefits of NOGM algorithm. Bo Jiang 0002, Jin Tang 0001, Chris Ding, Bin Luo 0001 |
AAAI | 1 |
| 2017 | Binary Constraint Preserving Graph MatchingabstractGraph matching is a fundamental problem in computer vision and pattern recognition area. In general, it can be formulated as an Integer Quadratic Programming (IQP) problem. Since it is NP-hard, approximate relaxations are required. In this paper, a new graph matching method has been proposed. There are three main contributions of the proposed method: (1) we propose a new graph matching relaxation model, called Binary Constraint Preserving Graph Matching (BPGM), which aims to incorporate the discrete binary mapping constraints more in graph matching relaxation. Our BPGM is motivated by a new observation that the discrete binary constraints in IQP matching problem can be represented (or encoded) exactly by a ℓ2-norm constraint. (2) An effective projection algorithm has been derived to solve BPGM model. (3) Using BPGM, we propose a path-following strategy to optimize IQP matching problem and thus obtain a desired discrete solution at convergence. Promising experimental results show the effectiveness of the proposed method. Bo Jiang 0002, Jin Tang 0001, Chris Ding, Bin Luo 0001 |
CVPR | 1 |
| 2017 | Image Set Representation with L_1 -Norm Optimal Mean Robust Principal Component Analysis
Youxia Cao, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001 |
ICIG (2) | 2 |
| 2017 | Robust Mapping Learning for Multi-view Multi-label Classification with Missing Labels
Weijieying Ren, Lei Zhang 0060, Bo Jiang 0002, Guangming Guo, Guiquan Liu |
KSEM | 3 |
| 2017 | Groupwise Registration of MR Brain Images Containing Tumors via Spatially Constrained Low-Rank Based Image Recovery
Bo Jiang 0002 |
MICCAI (2) | 3 |
| 2017 | Graph Matching via Multiplicative Update AlgorithmabstractAs a fundamental problem in computer vision, graph matching problem can usually be formulated as a Quadratic Programming (QP) problem with doubly stochastic and discrete (integer) constraints. Since it is NP-hard, approximate algorithms are required. In this paper, we present a new algorithm, called Multiplicative Update Graph Matching (MPGM), that develops a multiplicative update technique to solve the QP matching problem. MPGM has three main benefits: (1) theoretically, MPGM solves the general QP problem with doubly stochastic constraint naturally whose convergence and KKT optimality are guaranteed. (2) Em- pirically, MPGM generally returns a sparse solution and thus can also incorporate the discrete constraint approximately. (3) It is efficient and simple to implement. Experimental results show the benefits of MPGM algorithm. Bo Jiang 0002, Jin Tang 0001, Chris Ding, Yihong Gong, Bin Luo 0001 |
NIPS | 1 |
| 2017 | A new graph ranking model for image saliency detection problemabstractSaliency detection is an important problem in many computer vision applications. As a kind of popular method, graph based manifold ranking (GMR) has been successfully used in saliency detection problem. In traditional GMR saliency detection, it involves two main stages, i.e., ranking with background queries and ranking with foreground queries. However, in GMR method, these two stages are conducted separately, which ignores the correlation between background and foreground cues. In this paper, we propose a new graph ranking model, which aims to perform background and foreground ranking simultaneously by exploiting the correlation between background and foreground cues. We derive a closed-form solution for it. Experimental results on four benchmark datasets demonstrate that the proposed method performs better than some other state-of-art methods. Yuanyuan Guan, Bo Jiang 0002, Yun Xiao 0003, Jin Tang 0001, Bin Luo 0001 |
SERA | 2 |
| 2017 | Image Set Representation and Classification with Attributed Covariate-Relation Graph Model and Graph Sparse Representation Classification
Zhuqiang Chen, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001 |
Neurocomputing | 2 |
| 2017 | A global and local consistent ranking model for image saliency computation
Yun Xiao 0003, Bo Jiang 0002, Zhengzheng Tu, Jin Tang 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Lagrangian relaxation graph matching
Bo Jiang 0002, Jin Tang 0001, Xiaochun Cao, Bin Luo 0001 |
Pattern Recognit. | 1 |
| 2017 | Image representation and matching with geometric-edge random structure graph
Bo Jiang 0002, Jin Tang 0001, Aihua Zheng, Bin Luo 0001 |
Pattern Recognit. Lett. | 1 |
| 2016 | Robust Out-of-Sample Data Recovery
Bo Jiang 0002, Chris Ding, Bin Luo 0001 |
IJCAI | 1 |
| 2015 | A Local Sparse Model for Matching ProblemabstractFeature matching problem that incorporates pairwise constraints is usually formulated as a quadratic assignment problem (QAP). Since it is NP-hard, relaxation models are required. In this paper, we first formulate the QAP from the match selection point of view; and then propose a local sparse model for matching problem. Our local sparse matching (LSM) method has the following advantages: (1) It is parameter-free; (2) It generates a local sparse solution which is closer to a discrete matrix than most other continuous relaxation methods for the matching problem. (3) The one-to-one matching constraints are better maintained in LSM solution. Promising experimental results show the effectiveness of the Proposed LSM method. Bo Jiang 0002, Jin Tang 0001, Chris Ding, Bin Luo 0001 |
AAAI | 1 |
| 2015 | Person Re-identification with Density-Distance Unsupervised Salience Learning
Baoliang Zhou, Aihua Zheng, Bo Jiang 0002, Chenglong Li 0002, Jin Tang 0001 |
ICIG (3) | 3 |
| 2015 | Image matching using a local distribution based outlier detection technique
Haifeng Zhao 0001, Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001 |
Neurocomputing | 2 |
| 2014 | Classification of Fish Ectoparasite Genus Gyrodactylus SEM Images Using ASM and Complex Network Model
Rozniza Ali, Bo Jiang 0002, Mustafa Man, Amir Hussain 0001, Bin Luo 0001 |
ICONIP (3) | 2 |
| 2014 | Covariate-Correlated Lasso for Feature Selection
Bo Jiang 0002, Chris Ding, Bin Luo 0001 |
ECML/PKDD (1) | 1 |
| 2014 | A sparse nonnegative matrix factorization technique for graph matching problems
Bo Jiang 0002, Haifeng Zhao 0001, Jin Tang 0001, Bin Luo 0001 |
Pattern Recognit. | 1 |
| 2014 | Robust Feature Point Matching With Sparse ModelabstractFeature point matching that incorporates pairwise constraints can be cast as an integer quadratic programming (IQP) problem. Since it is NP-hard, approximate methods are required. The optimal solution for IQP matching problem is discrete, binary, and thus sparse in nature. This motivates us to use sparse model for feature point matching problem. The main advantage of the proposed sparse feature point matching (SPM) method is that it generates sparse solution and thus naturally imposes the discrete mapping constraints approximately in the optimization process. Therefore, it can optimize the IQP matching problem in an approximate discrete domain. In addition, an efficient algorithm can be derived to solve SPM problem. Promising experimental results on both synthetic points sets matching and real-world image feature sets matching tasks show the effectiveness of the proposed feature point matching method. Bo Jiang 0002, Jin Tang 0001, Bin Luo 0001, Liang Lin 0004 |
IEEE Trans. Image Process. | 1 |
| 2013 | Graph-Laplacian PCA: Closed-Form Solution and RobustnessabstractPrincipal Component Analysis (PCA) is a widely used to learn a low-dimensional representation. In many applications, both vector data X and graph data W are available. Laplacian embedding is widely used for embedding graph data. We propose a graph-Laplacian PCA (gLPCA) to learn a low dimensional representation of X that incorporates graph structures encoded in W. This model has several advantages: (1) It is a data representation model. (2) It has a compact closed-form solution and can be efficiently computed. (3) It is capable to remove corruptions. Extensive experiments on 8 datasets show promising results on image reconstruction and significant improvement on clustering and classification. Bo Jiang 0002, Chris Ding, Bin Luo 0001, Jin Tang 0001 |
CVPR | 1 |
| 2012 | Object categorization with sketch representation and generalized samples
Liang Lin 0004, Xiaobai Liu, Shaowu Peng, Hongyang Chao, Yongtian Wang, Bo Jiang 0002 |
Pattern Recognit. | 6 |
| 2012 | Graph matching based on spectral embedding with missing value
Jin Tang 0001, Bo Jiang 0002, Aihua Zheng, Bin Luo 0001 |
Pattern Recognit. | 2 |