Jiaqi Zhao 0001

dblp:27/9676-1 · DBLP profile ↗
← Back
103ranked-venue papers
22as first author
78since 2021 · last 2026
0000-0002-3564-5090ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 43 · 11 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 41 · 6 first-author · 37 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 6 first-author · 19 since 2021Computer networks · 11 · 11 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 2 since 2021
YearPublicationVenuePosition
2026 DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal Modeling
abstract
Video shadow detection confronts two entwined difficulties: distinguishing shadows from complex backgrounds and modeling dynamic shadow deformations under varying illumination. To address shadow-background ambiguity, we leverage linguistic priors through the proposed Vision-language Match Module (VMM) and a Dark-aware Semantic Block (DSB), extracting text-guided features to explicitly differentiate shadows from dark objects. Furthermore, we introduce adaptive mask reweighting to downweight penumbra regions during training and apply edge masks at the final decoder stage for better supervision. For temporal modeling of variable shadow shapes, we propose a Tokenized Temporal Block (TTB) that decouples spatiotemporal learning. TTB summarizes cross-frame shadow semantics into learnable temporal tokens, enabling efficient sequence encoding with minimal computation overhead. Comprehensive Experiments on multiple benchmark datasets demonstrate state-of-the-art accuracy and real-time inference efficiency.
Kunyang Sun, Rui Yao 0006, Hancheng Zhu, Fuyuan Hu, Jiaqi Zhao 0001, Zhiwen Shao, Yong Zhou 0003
AAAI6
2026 CLIPDet3D: Vision-Language Collaborative Distillation for 3D Object Detection
abstract
Multi-view 3D object detection plays a vital role in autonomous driving systems due to its ability to perceive complex scenes accurately. However, real-world driving data often exhibits a long-tailed distribution, causing significant drops in detection accuracy for rare categories in existing methods. To mitigate this issue, we propose CLIPDet3D, a novel vision-language collaborative framework for multi-view 3D object detection. First, to tackle the difficulty of capturing the semantic information of rare categories, a Vision-Language Collaborative Learning strategy is proposed to incorporate class-level semantic priors from CLIP. Second, a Depth Feature Contrastive Distillation module is designed to overcome the large depth estimation error for rare categories by aligning depth features between a teacher and a student network. Furthermore, to alleviate the difficulty in focusing on regions of rare categories, a Dual-Stream Prompt Attention mechanism is devised to inject learnable prompts and compute attention along both horizontal and vertical BEV directions. Evaluations on the nuScenes dataset demonstrate that CLIPDet3D achieves state-of-the-art accuracy while maintaining efficient inference.
Jiaqi Zhao 0001, Huanfeng Hu, Yong Zhou 0003, Wen-Liang Du 0002, Kunyang Sun, Rui Yao 0006, Qigong Sun
AAAI1
2026 Unified Representation Causal Prompt Distillation for Re-Inference-Free Lifelong Person Re-Identification
abstract
Lifelong person re-identification (LReID) aims to retrieve the target person from sequentially collected data. Due to significant domain gaps between datasets and the continuous increase of training data from different scenarios, weak inter-domain generalization and catastrophic forgetting issues have remained major challenges for LReID. To tackle these issues, a novel LReID method called Unified Representation Causal Prompt Distillation (URCPD) is proposed. Specifically, to reduce domain gaps among different scene datasets and improve model inter-domain generalization capability, a Feature Decoupling Style Transfer module (FDST) is proposed to map new features into a unified feature space. Furthermore, to reduce the accumulated forgetting of old knowledge during the training stage, a Causal Prompt Distillation module (CPD) is introduced. This module eliminates the re-inference process for distillation and embeds memory prompts to combat catastrophic forgetting. Extensive experiments on five classic LReID seen datasets and seven unseen datasets demonstrate that our method significantly outperforms state-of-the-art methods.
Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006
AAAI1
2026 Causal Decoupling Domain Generalization for Remote Sensing Change Detection
abstract
While current state-of-the-art Remote Sensing Change Detection (RSCD) methods can achieve impressive results on individual datasets, they become unreliable in unseen environments and imaging conditions, with performance metrics declining by as much as 60% to 80%. Simultaneously, variable environments and complex imaging conditions are the main characteristics of remote sensing data, calling for generalizable RSCD methods. To address this issue, we propose a novel RSCD method capable of domain generalization—CDDGNet. This method is based on causal decoupling theory, which progressively decouples invariant change features from variable domain features to extract generalizable characteristics. This enables a network trained on a single domain to accurately identify change regions in other domains. Specifically, firstly, the Causal Feature Adaptation Module is proposed to preliminarily decouple and simplify feature information during the encoding process by using wavelet transformation and feature energy spectralization methods. Secondly, the Causal Feature Fusion Module is presented to fully decouple features and aggregate significant change features during the decoding process through frequency domain processing and feature re-attention mechanisms. Thirdly, the Decoupling Effect Loss Function is proposed to optimize the process by evaluating the effectiveness of causal decoupling. Extensive experiments have shown that our model significantly outperforms existing methods across multiple groups of generalization tasks with varying levels of difficulty.
Jiaqi Zhao 0001, Jianpeng Xie 0001, Yong Zhou 0003, Wen-Liang Du 0002, Hancheng Zhu, Rui Yao 0006
AAAI1
2026 A unified multi-stream diffusion framework for robust video camouflaged object detection
Yuyao Ke, Rui Yao 0006, Kunyang Sun, Hancheng Zhu, Jiaqi Zhao 0001, Bing Liu 0016
Neural Networks6
2026 Dynamic Prompt Memory Network for video shadow detection
Rui Yao 0006, Hancheng Zhu, Kunyang Sun, Jiaqi Zhao 0001, Zhiwen Shao, Abdulmotaleb El Saddik
Pattern Recognit.5
2026 SpaceFormer: Spatial Position Contextual Semantics Embedding for Multi-View 3D Object Detection
abstract
3D object detection aims to accurately localize and recognize objects in 3D space. It serves as a fundamental task for reliable perception in intelligent transportation systems, enabling the monitoring of diverse traffic participants such as vehicles, pedestrians, cyclists, and public transport. Recently, transformer-based methods have gained significant attention in multi-view 3D object detection due to their strong global reasoning capabilities. However, their limited capacity to model spatial positional information hinders accurate object localization, especially in complex and large-scale scenes. To address this limitation, SpaceFormer is proposed as a novel transformer-based multi-view 3D object detector. Specifically, a Contextual Visual Prompts Learning strategy is proposed to enhance the perception of small and sparse traffic participants by incorporating contextual priors. To further suppress background interference, a Semantics-guided Depth Estimation method is proposed to refine depth representations using high-level semantic information. Furthermore, a Spatial Position Embedding mechanism is proposed to improve the spatial localization capability of the transformer by integrating geometric position and polar spatial embedding. Extensive experiments on the nuScenes benchmark demonstrate that SpaceFormer achieves state-of-the-art performance with 55.5% mAP and 62.9% NDS. These improvements indicate not only methodological advances but also practical benefits for intelligent transportation systems, enhancing safety, reliability, and efficiency in real-world deployments.
Jiaqi Zhao 0001, Huanfeng Hu, Wen-Liang Du 0002, Yong Zhou 0003, Kunyang Sun, Rui Yao 0006, Abdulmotaleb El Saddik
IEEE Trans. Intell. Transp. Syst.1
2026 Dual Sparse Long-Short Term Transformer for Video Shadow Detection
abstract
Video Shadow Detection (VSD) is critical yet challenging, primarily due to ambiguous shadow boundaries and the presence of confusing shadow-like non-shadow regions, which existing methods struggle to resolve effectively by limited temporal modeling. We propose the Dual Sparse Long-Short Term Transformer Network (DSLSTT-Net), a novel framework designed to enhance feature learning by integrating robust temporal consistency and detailed local context. DSLSTT-Net utilizes a dual-stream architecture to concurrently process global temporal information and local shadow feature refinement, enabling effective discrimination between true shadows and confusing areas. At its core, the Sparse Long-Short Term Attention Module (Sparse LSTAM) is introduced to efficiently propagate only high-confidence shadow features from memory, significantly enhancing feature discriminability and computational efficiency. Furthermore, an Adaptive Fusion Module (AFM) dynamically merges purified long-term features with short-term details, optimizing final segmentation. Experimental results confirm that DSLSTT-Net significantly outperforms state-of-the-art methods on VSD benchmarks, validating our approach of dual-stream architecture and sparse temporal modeling. The source code is available at https://github.com/rayyao/DSLSTTNet .
Rui Yao 0006, Huili Hao, Hancheng Zhu, Jiaqi Zhao 0001, Yong Zhou 0003
ACM Trans. Multim. Comput. Commun. Appl.6
2025 ReDiffDet: Rotation-equivariant Diffusion Model for Oriented Object Detection
abstract
The diffusion model has been successfully applied to various detection tasks. However, it still faces several challenges when used for oriented object detection: objects that are arbitrarily rotated require the diffusion model to encode their orientation information; uncontrollable random boxes inaccurately locate objects with dense arrangements and extreme aspect ratios; oriented boxes result in the misalignment between them and image features. To overcome these limitations, we propose ReDiffDet, a framework that formulates oriented object detection as a rotation-equivariant denoising diffusion process. First, we represent an oriented box as a 2D Gaussian distribution, forming the basis of the denoising paradigm. The reverse process can be proven to be rotation-equivariant within this representation and model framework. Second, we design a conditional encoder with conditional boxes to prevent boxes from being randomly placed across the entire image. Third, we propose an aligned decoder for alignment between oriented boxes and image features. The extensive experiments demonstrate ReDiffDet achieves promising performance and significantly outperforms the diffusion-based baseline detector. Codes are available at https://github.com/wokaikaixinxin/ReDiffDet.
Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006
CVPR1
2025 GSDet: Gaussian Splatting for Oriented Object Detection
abstract
Oriented object detection has advanced with the development of convolutional neural networks (CNNs) and transformers. However, modern detectors still rely on predefined object candidates, such as anchors in CNN-based methods or queries in transformer-based methods, which struggle to capture spatial information effectively. To address the limitations, we propose GSDet, a novel framework that formulates oriented object detection as Gaussian splatting. Specifically, our approach performs detection within a 3D feature space constructed from image features, where 3D Gaussians are employed to represent oriented objects. These 3D Gaussians are projected onto the image plane to form 2D Gaussians, which are then transformed into oriented boxes. Furthermore, we optimize the mean, anisotropic covariance, and confidence scores of these randomly initialized 3D Gaussians, using a decoder that incorporates 3D Gaussian sampling. Moreover, our method exhibits flexibility, enabling adaptive control and a dynamic number of Gaussians during inference. Experiments on 3 datasets indicate that GSDet achieves AP50 gains of 0.7% on DIOR-R, 0.3% on DOTA-v1.0, and 0.55% on DOTA-v1.5 when evaluated with adaptive control and outperforms mainstream detectors.
Zeyu Ding 0010, Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Hancheng Zhu, Rui Yao 0006
IJCAI2
2025 Beyond Individual and Point: Next POI Recommendation via Region-aware Dynamic Hypergraph with Dual-level Modeling
abstract
Next POI recommendation contributes to the prosperity of various intelligent location-based services. Existing studies focus on exploring sequential patterns and POI interactions using sequential and graph-based methods to enhance recommendation performance. However, they don't effectively exploit geographical information. In addition, methods that focus on modeling mobility patterns using individual limited data may suffer from data sparsity and the information cocoons problem. Moreover, most graph structures focus on adjacent nodes, failing to capture potential high-order associations among POIs. To address these challenges, we propose the Region-aware dynamic Hypergraph learning method with Dual-level interaction Modeling (ReHDM), which exploits users' dynamic mobility beyond individual and point. Specifically, ReHDM utilizes regional encoding to mine the potential spatial relationships among POIs with coarse-grained geographical information. By incorporating POI-level and trajectory-level associations within a hypergraph convolutional network, ReHDM comprehensively captures cross-user collaborative information. Furthermore, ReHDM captures not only dependencies among POIs within each trajectory for a single user, but also the high-order collaborative information across individual user trajectories and associated users' trajectories. Experimental results on three public datasets demonstrate the superiority of ReHDM to the state-of-the-art.
Zhuo Gu, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Wen-Liang Du 0002
IJCAI6
2025 Modality-Guided Dynamic Graph Fusion and Temporal Diffusion for Self-Supervised RGB-T Tracking
abstract
To reduce the reliance on large-scale annotations, self-supervised RGB-T tracking approaches have garnered significant attention. However, the omission of the object region by erroneous pseudo-label or the introduction of background noise affects the efficiency of modality fusion, while pseudo-label noise triggered by similar object noise can further affect the tracking performance. In this paper, we propose GDSTrack, a novel approach that introduces dynamic graph fusion and temporal diffusion to address the above challenges in self-supervised RGB-T tracking. GDSTrack dynamically fuses the modalities of neighboring frames, treats them as distractor noise, and leverages the denoising capability of a generative model. Specifically, by constructing an adjacency matrix via an Adjacency Matrix Generator (AMG), the proposed Modality-guided Dynamic Graph Fusion (MDGF) module uses a dynamic adjacency matrix to guide graph attention, focusing on and fusing the object’s coherent regions. Temporal Graph-Informed Diffusion (TGID) models MDGF features from neighboring frames as interference, and thus improving robustness against similar-object noise. Extensive experiments conducted on four public RGB-T tracking datasets demonstrate that GDSTrack outperforms the existing state-of-the-art methods. The source code is available at https://github.com/LiShenglana/GDSTrack.
Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Kunyang Sun, Bing Liu 0016, Zhiwen Shao, Jiaqi Zhao 0001
IJCAI8
2025 Counterfactual Knowledge Maintenance for Unsupervised Domain Adaptation
abstract
Traditional unsupervised domain adaptation (UDA) struggles to extract rich semantics due to backbone limitations. Recent large-scale pre-trained visual-language models (VLMs) have shown strong zero-shot learning capabilities in UDA tasks. However, directly using VLMs results in a mixture of semantic and domain-specific information, complicating knowledge transfer. Complex scenes with subtle semantic differences are prone to misclassification, which in turn can result in the loss of features that are crucial for distinguishing between classes. To address these challenges, we propose a novel counterfactual knowledge maintenance UDA framework. Specifically, we employ counterfactual disentanglement to separate the representation of semantic information from domain features, thereby reducing domain bias. Furthermore, to clarify ambiguous visual information specific to classes, we maintain the discriminative knowledge of both visual and textual information. This approach synergistically leverages multimodal information to preserve modality-specific distinguishable features. We conducted extensive experimental evaluations on several public datasets to demonstrate the effectiveness of our method. The source code is available at https://github.com/LiYaolab/CMKUDA
Yong Zhou 0003, Jiaqi Zhao 0001, Wen-Liang Du 0002, Rui Yao 0006, Bing Liu 0016
IJCAI3
2025 RQFormer: Rotated Query Transformer for end-to-end oriented object detection
Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
Expert Syst. Appl.1
2025 MoViM: A Hybrid CNN Vision Mamba Network for Lightweight Semantic Segmentation of Multimodal Remote Sensing Images
abstract
The “Others” category in multimodal remote sensing images is characterized by high intra-class variability. Therefore, existing lightweight semantic segmentation models struggle with this category due to limitations in capturing both local details and global dependencies efficiently. We propose MoViM, a lightweight model that integrates a hybrid Vision Mamba (ViM) and CNN backbone to capture global contextual information and local details effectively. In addition, the MoViM also features an Inverted Stem for efficient multimodal fusion, a Global Semantics Extraction (GSE) module for enhanced global feature representation, and a Global-Local Feature Fusion (GLF) module for context-aware feature integration. Extensive experiments on WHU-OPT-SAR and Potsdam datasets demonstrate that MoViM achieves state-of-the-art performance, particularly in the “Others” category, while maintaining low computational complexity. Our codes are available at https://github.com/WenliangDu/MoViM.
Wen-Liang Du 0002, Jiaqi Zhao 0001, Rui Yao 0006, Yong Zhou 0003
IEEE Geosci. Remote. Sens. Lett.3
2025 FA-MSVNet: multi-scale and multi-view feature aggregation methods for stereo 3D reconstruction
Yong Zhou 0003, Jiaqi Zhao 0001, Wen-Liang Du 0002, Rui Yao 0006
Multim. Tools Appl.3
2025 Grid-distance-based selection for fine-grained object detection in aerial images
Jiaqi Zhao 0001, Qingfeng Ou, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006
Pattern Recognit. Lett.1
2025 Hierarchical Relation Learning for Few-Shot Semantic Segmentation in Remote Sensing Images
abstract
Few-shot semantic segmentation (FSS) aims to segment specific semantic classes in a query image using only a few annotated support samples. While FSS has gained significant attention in natural image processing, it remains underexplored in the more challenging domain of remote sensing images (RSIs). Existing FSS approaches for RSIs primarily focus on enhancing feature representations of support or query images through hierarchical/multi-level feature fusion. However, unlike fully supervised segmentation that relies on feature extraction and optimization, FSS requires segmenting the query image based on its relations with annotated support images. To address this need, we propose the concept of Hierarchical Relation Learning (HRL) to explore the intrinsic support-query relations, allowing for the direct refinement of target object appearances in the query image. Specifically, we propose a Hierarchical Relation Network (HRNet), which performs single-scale relation extraction at each network hierarchy and multi-scale relation aggregation across hierarchies. In addition, we construct a Bidirectional Hierarchical Loss (BHLoss) to guide HRNet training, providing targeted supervision at each hierarchy in both top-down and bottom-up directions, thus facilitating robust multi-scale relation learning across hierarchies. Comprehensive experiments on the iSAID-5i, DLRSD-5i, and LoveDA-2i datasets demonstrate the superiority of the proposed HRL. The code will be available at https://github.com/XinnHe/HRL.
Xin He 0024, Yun Liu 0011, Yong Zhou 0003, Henghui Ding, Jiaqi Zhao 0001, Bing Liu 0016, Xudong Jiang 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Hyperspectral Object Tracking With Dual-Stream Prompt
abstract
Hyperspectral images, rich in spectral details, offeradvantages for object tracking across diverse scenarios. Current hyperspectral tracking often fine-tunes parameters using pretrained RGB trackers, but this manner is suboptimal due to redundancy in spectral bands and limited training data. Existing hyperspectral trackers also underuse temporal information. To address these issues, we propose a unified spectral-spatiotemporal multimodal dual-stream prompt hyperspectral object tracking, named HDSP. We design a density clustering-based band selection module (BSM) to preserve spectral prompt information efficiently. Using the generated bands and temporal data as multimodal prompts, a dual-stream visual prompter is proposed. Designed multimodal dual-stream visual prompter (MDVP) transforms the multimodal input into a single modality, enhancing the foundational modality’s representation capabilities for hyperspectral tracking. Experiments on hyperspectral videos (HSVs) tracking datasets demonstrate that the proposed tracker achieves state-of-the-art performance. The source code is available athttps://github.com/rayyao/HDSP.
Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao
IEEE Trans. Geosci. Remote. Sens.5
2025 DDCI: Unsupervised Domain Adaptation for Remote Sensing Images Based on Diffusion Causal Distillation
abstract
The distribution of remote sensing (RS) images can vary significantly due to seasonal changes and lighting conditions, making it difficult for deep learning models to generalize effectively across different RS datasets. This variation leads to a domain gap that hampers model performance when applied to new, unseen data. To tackle this challenge, we introduce DDCI, a novel unsupervised domain adaptation (UDA) framework designed to bridge the domain gap in RS image perception. Our framework consists of two key components, i.e., the adaptation diffusion distillation (ADD) module and the consistent causal intervention (CCI) module. The ADD module addresses the domain gap by aligning the source and target domains. It enhances the representation of the target domain by distilling semantic knowledge from the teacher model of the source domain. This process allows the target domain to benefit from the rich features of the source domain, leading to improved model generalization. The CCI module focuses on removing spurious correlations between domain-agnostic knowledge and domain-specific knowledge. By carefully considering the distinct characteristics of the target domain while preserving the specificity of the source domain, the CCI module ensures that only relevant, causal information is transferred between domains. This prevents overfitting to irrelevant domain-specific features and enhances model robustness. We demonstrate the effectiveness of the DDCI framework on RS scene classification tasks, utilizing four widely recognized RS datasets. Our results show significant performance improvements, underscoring the potential of this approach to boost the adaptability of deep learning models across diverse RS image datasets.
Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
IEEE Trans. Geosci. Remote. Sens.1
2025 GLFRNet: Global-Local Feature Refusion Network for Remote Sensing Image Instance Segmentation
abstract
Instance segmentation is a significant way for remote sensing image (RSI) interpretation. The large number, sharp variation of sizes, and complex background of objects raise higher demands for instance segmentation models. The synergistic usage of global and local features has drawn great attention due to its superior performance but has not been fully explored in mainstream instance segmentation methods. In this work, a global-local feature refusion network (GLFRNet) with two fusion procedures is proposed to fully utilize coarse-grained and fine-grained features for RSI instance segmentation. In this model, the backbone integrates both convolutional neural network (CNN)-based and VMamba-based branches to extract local and global features, respectively. Three novel models are proposed to leverage the features adaptively, i.e., the cross-dim feature fusion (CDFF) module, the semantic complementary feature fusion (SCFF) module, and the guided feature refusion module (GFRM). The CDFF module is designed to aggregate features flexibly by fusing features from two backbones with different attention modules in the first fusion procedure. The GFRM and SCFF module are proposed in the refusion procedure to generate accurate segmentation results. Inspired by agent attention, the GFRM dynamically assembles detailed features for mask generation by refusing local and global features with the guidance of fusion results from CDFF. The SCFF module complements the significant features by enhancing and integrating global, local, and detailed features, and finally generates masks of instances. Extensive experiments demonstrate that GLFRNet outperforms the second-best model by 1.9, 1.3, and 0.3 in mask average precisions (APs) on NWPU VHR-10, WHU Building, and iSAID datasets.
Jiaqi Zhao 0001, Yari Wang, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
IEEE Trans. Geosci. Remote. Sens.1
2025 ST-Mamba: Spatio-Temporal Synergistic Model for Remote Sensing Change Detection
abstract
The advancement of remote sensing and deep learning has spurred interest in high-resolution image change detection (CD). However, pseudo-changes in multi-temporal images, due to complex scenes and variable imaging conditions, often lead to significant misdetection in current methods. To address this problem, we propose a new CD framework: Spatio-Temporal Mamba (ST-Mamba), which consists of three key components. Firstly, a Mamba-based Feature Extraction Module (MFEM) is designed as the encoder to extract essential features from multi-temporal images by leveraging Mamba’s capability to capture inherent information in long data sequences. Secondly, a Spatio-Temporal Synergy Module (STSM) is developed to unify the background features of multi-temporal feature maps into a common domain by employing the state-space model for spatio-temporal modeling. Finally, a Spatio-Temporal Fusion Module (STFM) is created to guide the fusion of image features at different scales and across channels by utilizing a feature map of the unified background features. Experimental results on five widely used change detection datasets show significant improvements over current state-of-the-art methods.
Jiaqi Zhao 0001, Jianpeng Xie 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
IEEE Trans. Geosci. Remote. Sens.1
2025 Adversarial Geometric Attacks for 3D Point Cloud Object Tracking
abstract
3D point cloud object tracking (3D PCOT) plays a vital role in applications such as autonomous driving and robotics. Adversarial attacks offer a promising approach to enhance the robustness and security of tracking models. However, existing adversarial attack methods for 3D PCOT seldom leverage the geometric structure of point clouds and often overlook the transferability of attack strategies. To address these limitations, this paper proposes an adversarial geometric attack method tailored for 3D PCOT, which includes a point perturbation attack module (non-isometric transformation) and a rotation attack module (isometric transformation). First, we introduce a curvature-aware point perturbation attack module that enhances local transformations by applying normal perturbations to critical points identified through geometric features such as curvature and entropy. Second, we design a Thompson sampling-based rotation attack module that applies subtle global rotations to the point cloud, introducing tracking errors while maintaining imperceptibility. Additionally, we design a fused loss function to iteratively optimize the point cloud within the search region, generating adversarially perturbed samples. The proposed method is evaluated on multiple 3D PCOT models and validated through black-box tracking experiments on benchmarks. For P2B, white-box attacks on KITTI reduce the success rate from 53.3% to 29.6% and precision from 68.4% to 37.1%. On NuScenes, the success rate drops from 39.0% to 27.6%, and precision from 39.9 to 26.8%. Black-box attacks show a transferability, with BAT showing a maximum 47.0% drop in success rate and 47.2% in precision on KITTI, and a maximum 22.5% and 27.0% on NuScenes.
Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Bing Liu 0016, Abdulmotaleb El Saddik
IEEE Trans. Multim.4
2025 Similarity Regulation and Calibration Alignment for Weakly Supervised Text-Based Person Re-Identification
abstract
Traditional text-based person re-identification relies on identity labels. However, it is impossible to annotate large datasets, since identity annotation is expensive and time-consuming. Weakly supervised text-based person re-identification, where only text–image pairs are available without annotation of identities, is very practical in real life. While dealing with the weakly supervised person re-identification, two issues should be strengthed, i.e., alignment caused by different modal, and cross-modal matching ambiguity caused by the lack of identity labels. In this article, we propose a similarity regulation and calibration alignment (SRCA) framework, which consists of two unimodal encoders for images and text, respectively, and a multi-modal encoder for the masked language modeling task. First, a similarity regulation (SR) strategy is proposed to relax the strict one-to-one constraints for the local similarities between different pairs by introducing a novel soft objective. The soft objective can adjust hard objectives to achieve soft cross-modal alignment by establishing a many-to-many relationship between two modalities. Second, the calibration alignment (CA) module is proposed to improve intra-class compactness by modeling pseudo-label assignment as optimal transport. The ambiguity of cross-modal matching can be reduced by aligning features and pseudo-labels of different modalities and gradually calibrating the distribution of pseudo-labels. Experimental results show that our method has achieved obvious advantages compared with existing methods and also demonstrated competitive performance compared with fully supervised methods.
Ao Fu, Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Historical Object-Aware Prompt Learning for Universal Hyperspectral Object Tracking
abstract
Hyperspectral Object Tracking (HOT), utilizing rich spectral information from hyperspectral video (HSV), holds significant importance for object tracking. We identify that a major obstacle in improving HOT performance lies in effectively leveraging spectral and historical information. Furthermore, due to the mismatch in band dimensions between hyperspectral and RGB images, state-of-the-art RGB-based trackers struggle to adapt to unified HOT tasks. To address this, we propose a Historical Object-Aware Prompt Learning (HOPL) method for universal hyperspectral object tracking. Initially, we transform hyperspectral image ( \( N \) bands) into multiple sets of three bands with different combinations and feed them into a backbone network to generate base features. Subsequently, we introduce a historical object-aware prompter, where historical object-aware images are input to generate prompt features that enhance the representation of object information when combined with base features. Additionally, we design a band information fusion module to integrate the multiple sets of base features. By introducing historical object-aware prompts, HOPL significantly enhances tracking performance without retraining the backbone network. Experimental results on the HOT2023 dataset (comprising HSV with 25-band, 16-band, and 15-band wavelength ranges) and HOT2022 dataset validate the superiority of HOPL over state-of-the-art methods. The source code is available at https://github.com/rayyao/HOPL .
Rui Yao 0006, Yong Zhou 0003, Fuyuan Hu, Jiaqi Zhao 0001, Zhiwen Shao
ACM Trans. Multim. Comput. Commun. Appl.6
2025 CCFL: Customized Client Federated Learning for Unsupervised Person Re-identification
abstract
Federated learning-based person re-identification (Re-ID) aims to address the issue of data silos in surveillance systems caused by increasingly stringent regulations on sensitive data. However, due to differences in data collection locations, times, and scales, severe non-independent and identically distributed (non-IID) characteristics exist across different Re-ID datasets. Existing federated learning-based Re-ID methods often adopt a unified model structure, which prevents the model from adapting well to diverse data environments, thereby significantly degrading the overall Re-ID performance. To address the challenges of training neural networks on non-IID data across different datasets, we propose a customizable federated learning framework. First, customizable clients allow each organization to freely select suitable neural network training methods and model architectures based on local data scales and prior knowledge, thus improving training outcomes. Second, since traditional federated learning frameworks cannot achieve knowledge fusion through parameter exchange between models with different architectures, we introduce an independent model, referred to as the interaction model, specifically designed for knowledge exchange among clients. The interaction model learns parameters (knowledge) from local models on each client through distillation learning. Subsequently, the interaction model is uploaded to the server, where it undergoes parameter fusion (knowledge exchange) with interaction models from other clients. Finally, the interaction model, enriched with knowledge from other clients, guides local model training through knowledge distillation. It is worth noting that selecting a lightweight interaction model, while potentially impacting Re-ID performance, can significantly reduce communication costs between the server and clients.
Yong Zhou 0003, Fayao Liu, Jiaqi Zhao 0001, Hancheng Zhu, Wen-Liang Du 0002
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Image Cropping with Content and Composition Attribute-aware Global Relation Reasoning
abstract
Image cropping aims to find visually pleasing content in an image, which will enhance its aesthetic quality. Existing image cropping approaches mainly emphasize the geometric properties of images, such as composition and layout, neglecting the rich aesthetic information available from the physical attributes (e.g., content and themes), and background information beyond the foreground in images. Consequently, this article proposes an image cropping method based on the content and composition attribute-aware global relation reasoning, which aims at guiding the generation of cropped sub-images by exploring critical attributes based on content and composition as well as global object correlations that affect aesthetics in images. Particularly, to comprehensively introduce aesthetic information into image cropping, we capture feature representations reinforced by content and composition attributes simultaneously. The feature representations can strengthen the visual aesthetics of cropped sub-images. To make the cropped sub-images amply contain more global information, we introduce a global relation reasoning branch in the proposed cropping module, which can fully exploit the dependency relationship between the foreground and background in images. Extensive experiments on image cropping benchmarks demonstrate that our approach is superior to state-of-the-art image cropping methods.
Hancheng Zhu, Yong Zhou 0003, Rui Yao 0006, Zhiwen Shao, Jiaqi Zhao 0001, Leida Li
ACM Trans. Multim. Comput. Commun. Appl.6
2024 Attribute-Driven Multimodal Hierarchical Prompts for Image Aesthetic Quality Assessment
abstract
Image Aesthetic Quality Assessment (IAQA) aims to simulate users' visual perception to judge the aesthetic quality of images. In social media, users' aesthetic experiences are often reflected in their textual comments regarding the aesthetic attributes of images. To fully explore the attribute information perceived by users for evaluating image aesthetic quality, this paper proposes an image aesthetic quality assessment method based on attribute-driven multimodal hierarchical prompts. Unlike existing IAQA methods that utilize multimodal pre-training or straightforward prompts for model learning, the proposed method leverages attribute comments and quality-level text templates to hierarchically learn the aesthetic attributes and quality of images. Specifically, we first leverage users' aesthetic attribute comments to perform prompt learning on images. The learned attribute-driven multimodal features can comprehensively capture the semantic information of image aesthetic attributes perceived by users. Then, we construct text templates for different aesthetic quality levels to further facilitate prompt learning through semantic information related to the aesthetic quality of images. The proposed method can explicitly simulate users' aesthetic judgment of images to obtain more precise aesthetic quality. Experimental results demonstrate that the proposed IAQA method based on hierarchical prompts outperforms existing methods significantly on multiple IAQA databases. Our source code is public at https://github.com/GitHub-Ju/AMHP.
Hancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Leida Li
ACM Multimedia6
2024 Remote sensing image semantic segmentation via class-guided structural interaction and boundary perception
abstract
Existing remote sensing semantic segmentation methods generally ignore the structural information of objects that is vital in the human visual recognition system. The absence of overall structural information often results in weak perceptions of subtle textures and fragmented predictions, especially for complex and variable ground object scenarios. Besides, they still suffer from the semantic ambiguity caused by the unclear object boundary features in remote sensing images. In this paper, we propose a novel remote sensing semantic segmentation framework, called CSBNet, which aims to enhance the capacity of class-guided structural interaction and boundary perception simultaneously. It consists of a class-guided structure interaction module (CSIM), a Transformer-based context aggregation module (TCAM) and a class-guided boundary supervision module (CBSM). The CSIM has the ability to progressively extract the class-specific structural features, i.e. , refining the structural information of each class by iteratively exchanging information between initial coarse class tokens and contexts. Meanwhile, the TCAM is constructed to provide CSIM with more discriminative multi-scale contexts without losing spatial features. In particular, the CBSM plays an auxiliary role, which applies the boundary information obtained from the class tokens to supervise the segmentation of boundary regions. When tested on the ISPRS dataset, LoveDA dataset, UAVid dataset, our method significantly outperforms the state-of-the-art remote sensing semantic segmentation approaches.
Xin He 0024, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006
Expert Syst. Appl.4
2024 Fine-grained semantic oriented embedding set alignment for text-based person search
Jiaqi Zhao 0001, Ao Fu, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006
Image Vis. Comput.1
2024 A Mamba-Diffusion Framework for Multimodal Remote Sensing Image Semantic Segmentation
abstract
Recent advances in deep learning have made significant progress in multimodal remote sensing semantic segmentation. However, current methods face challenges in maintaining geometric consistency, particularly when dealing with large objects, resulting in fragmented segmentation masks. We propose a Mamba-diffusion framework to preserve geometric consistency in segmentation masks. This framework preserves geometric consistency by introducing a generative diffusion-based semantic segmentation pipeline and developing a Mamba-based multimodal fusion model. The fusion model fuses the multimodal images in multiple scales and scanning mechanisms by a double cross-fusion (DCF) module. Then, the cross-modal information is further integrated by a dual-splitting structured state-space (DS-S4) model. Finally, the diffusion-based segmentation pipeline predicts semantic masks by progressively refining random Gaussian noise, guided by fused multimodal features. Our experimental results, verified on WHU-OPT-SAR and Hunan datasets, demonstrate that the proposed framework surpasses state-of-the-art (SOTA) methods by a considerable margin. Our codes are available athttps://github.com/WenliangDu/MambaDiffusion.
Wen-Liang Du 0002, Yang Gu 0005, Jiaqi Zhao 0001, Hancheng Zhu, Rui Yao 0006, Yong Zhou 0003
IEEE Geosci. Remote. Sens. Lett.3
2024 Filter pruning based on evolutionary algorithms for person re-identification
Jiaqi Zhao 0001, Ying Chen 0005, Yufeng Zhong 0001, Yong Zhou 0003, Rui Yao 0006, Lixu Zhang, Shixiong Xia
Multim. Tools Appl.1
2024 Multi-level self attention for unsupervised learning person re-identification
Jiaqi Zhao 0001, Yong Zhou 0003, Fayao Liu, Rui Yao 0006, Hancheng Zhu, Abdulmotaleb El Saddik
Multim. Tools Appl.2
2024 Efficient convolutional neural networks and network compression methods for object detection: a survey
Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Bing Liu 0016
Multim. Tools Appl.3
2024 Dual-Stream Edge-Target Learning Network for Infrared Small Target Detection
abstract
Infrared small target detection (IRSTD) is crucial in both military and civilian applications. However, challenges such as low contrast, low signal-to-noise ratio (SNR), and lack of shape and texture information limit the effectiveness of existing methods in capturing edge details and representing target areas. To address these issues, we propose the dual-stream edge-target learning network (DETL-Net) for IRSTD. This network enhances feature cross-fusion by learning edge details and target regions through a dual-stream framework, significantly improving detection performance. Specifically, we extract multilevel features of the image based on the encoder-decoder structure of U-Net and then reconstruct the feature map. In the decoder, we propose the dual-guided cross-fusion module (DGCFM) to capture edge details of small targets and global contextual features of the target region, achieving complementary advantages. The multiscale context fusion module (MCFM) within DGCFM uses central difference convolution to enhance local contrast and extract rich contextual details, thereby retaining edge information and enhancing overall target representation. In addition, we introduce the cross-dimension interactive aggregation attention module (CIAAM), which dynamically adjusts feature fusion weights across layers to effectively suppress noise and enhance the discrimination of small targets. These modules are sequentially interconnected to progressively refine edge details, and the acquired target features are subsequently utilized for predicting the final target mask via the segmentation head. Experiments on the NUAA-SIRST and IRSTD-1k datasets demonstrate that DETL-Net outperforms state-of-the-art (SOTA) methods. The source code is available athttps://github.com/rayyao/DETL-Net.
Rui Yao 0006, Yong Zhou 0003, Jinqiu Sun, Zihang Yin, Jiaqi Zhao 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 OrientedFormer: An End-to-End Transformer-Based Oriented Object Detector in Remote Sensing Images
abstract
Oriented object detection in remote sensing images is a challenging task due to objects being distributed in multiorientation. Recently, end-to-end transformer-based methods have achieved success by eliminating the need for post-processing operators compared to traditional convolutional neural network (CNN)-based methods. However, directly extending transformers to oriented object detection presents three main issues: 1) objects rotate arbitrarily, necessitating the encoding of angles along with position and size; 2) the geometric relations of oriented objects are lacking in self-attention, due to the absence of interaction between content and positional queries; and 3) oriented objects cause misalignment, mainly between values and positional queries in cross-attention, making accurate classification and localization difficult. In this article, we propose an end-to-end transformer-based oriented object detector, consisting of three dedicated modules to address these issues. First, Gaussian positional encoding (PE) is proposed to encode the angle, position, and size of oriented boxes using Gaussian distributions. Second, Wasserstein self-attention is proposed to introduce geometric relations and facilitate interaction between content and positional queries by utilizing Gaussian Wasserstein distance scores. Third, oriented cross-attention is proposed to align values and positional queries by rotating sampling points around the positional query according to their angles. Experiments on six datasets DIOR-R, a series of DOTA, HRSC2016, and ICDAR2015 show the effectiveness of our approach. Compared with previous end-to-end detectors, the OrientedFormer gains 1.16 and 1.21 AP50 on DIOR-R and DOTA-v1.0, respectively, while reducing training epochs from$3\times $to$1\times $. The code is available athttps://github.com/wokaikaixinxin/OrientedFormer.
Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
IEEE Trans. Geosci. Remote. Sens.1
2024 Black-box Attack against Self-supervised Video Object Segmentation Models with Contrastive Loss
abstract
Deep learning models have been proven to be susceptible to malicious adversarial attacks, which manipulate input images to deceive the model into making erroneous decisions. Consequently, the threat posed to these models serves as a poignant reminder of the necessity to focus on the model security of object segmentation algorithms based on deep learning. However, the current landscape of research on adversarial attacks primarily centers around static images, resulting in a dearth of studies on adversarial attacks targeting Video Object Segmentation (VOS) models. Given that a majority of self-supervised VOS models rely on affinity matrices to learn feature representations of video sequences and achieve robust pixel correspondence, our investigation has delved into the impact of adversarial attacks on self-supervised VOS models. In response, we propose an innovative black-box attack method incorporating contrastive loss. This method induces segmentation errors in the model through perturbations in the feature space and the application of a pixel-level loss function. Diverging from conventional gradient-based attack techniques, we adopt an iterative black-box attack strategy that incorporates contrastive loss across the current frame, any two consecutive frames, and multiple frames. Through extensive experimentation conducted on the DAVIS 2016 and DAVIS 2017 datasets using three self-supervised VOS models and one unsupervised VOS model, we unequivocally demonstrate the potent attack efficiency of the black-box approach. Remarkably, theJ&Fmetric value experiences a significant decline of up to 50.08% post-attack.
Ying Chen 0005, Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Bing Liu 0016, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Motion-Aware Self-Supervised RGBT Tracking with Multi-Modality Hierarchical Transformers
abstract
Supervised RGBT (SRGBT) tracking tasks need both expensive and time-consuming annotations. Therefore, the implementation of Self-Supervised RGBT (SSRGBT) tracking methods has become increasingly important. Straightforward SSRGBT tracking methods use pseudo-labels for tracking, but inaccurate pseudo-labels can lead to object drift, which severely affects tracking performance. This article proposes a self-supervised RGBT object tracking method (S2OTFormer) to bridge the gap between tracking methods supervised under pseudo-labels and ground truth labels. Firstly, to provide more robust appearance features for motion cues, we introduce a multi-modality hierarchical transformer (MHT) module for feature fusion. This module allocates weights to both modalities and strengthens the expressive capability of the MHT module through multiple nonlinear layers to fully utilize the complementary information of the two modalities. Secondly, in order to solve the problems of motion blur caused by camera motion and inaccurate appearance information caused by pseudo-labels, we introduce a motion-aware mechanism (MAM). The MAM extracts the average motion vectors from the previous multi-frame search frame features and constructs the consistency loss with the motion vectors of the current search frame features. The motion vectors of inter-frame objects are obtained by reusing the inter-frame attention map to predict coordinate positions. Finally, to further reduce the effect of inaccurate pseudo-labels, we propose an Attention-Based Multi-Scale Enhancement Module. By introducing cross-attention to achieve more precise and accurate object tracking, this module overcomes the receptive field limitations of traditional CNN tracking heads. We demonstrate the effectiveness of S2OTFormer on four large-scale public datasets through extensive comparisons as well as numerous ablation experiments. The source code is available at https://github.com/LiShenglana/S2OTFormer .
Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Multi-Modal LiDAR Point Cloud Semantic Segmentation with Salience Refinement and Boundary Perception
abstract
Point cloud segmentation is essential for scene understanding, which provides advanced information for many applications, such as autonomous driving, robots, and virtual reality. To improve the accuracy and robustness of point cloud segmentation, many researchers have attempted to fuze camera images to complement the color and texture information. The common fusion strategy is the combination of convolutional operations with concatenation, element-wise addition or element-wise multiplication. However, conventional convolutional operators tend to confine the fusion of modal features within their receptive fields, which can be incomplete and limited. In addition, the inability of encoder–decoder segmentation networks to explicitly perceive segmentation boundary information results in semantic ambiguity and classification errors at object edges. These errors are further amplified in point cloud segmentation tasks, significantly affecting the accuracy of point cloud segmentation. To address the above issues, we propose a novel self-attention multi-modal fusion semantic segmentation network for point cloud semantic segmentation. Firstly, to effectively fuze different modal features, we propose a self-cross fusion module (SCF), which models long-range modality dependencies and transfers complementary image information to the point cloud to fully leverage the modality-specific advantages. Secondly, we design the salience refinement module (SR), which calculates the importance of channels in the feature maps and global descriptors to enhance the representation capability of salient modal features. Finally, we propose the local-aware anisotropy loss measure the element-level importance in the data and explicitly provide boundary information for the model, which alleviates the inherent semantic ambiguity problem in segmentation networks. Extensive experiments on two benchmark datasets demonstrate that our proposed method surpasses current state-of-the-art methods.
Yong Zhou 0003, Zeming Xie, Jiaqi Zhao 0001, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Info-FPN: An Informative Feature Pyramid Network for object detection in remote sensing images
Silin Chen, Jiaqi Zhao 0001, Yong Zhou 0003, Hanzheng Wang, Rui Yao 0006, Lixu Zhang, Yong Xue
Expert Syst. Appl.2
2023 Context-aware and part alignment for visible-infrared person re-identification
Jiaqi Zhao 0001, Hanzheng Wang, Yong Zhou 0003, Rui Yao 0006, Lixu Zhang, Abdulmotaleb El Saddik
Image Vis. Comput.1
2023 Adversarial learning-based skeleton synthesis with spatial-channel attention for robust gait recognition
Ying Chen 0005, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Dongjun Zhu
Multim. Tools Appl.3
2023 Unsupervised RGB-T object tracking with attentional multi-modal feature fusion
Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Jiaqi Zhao 0001, Zhiwen Shao
Multim. Tools Appl.6
2023 Semi-supervised transformable architecture search for feature distillation
Man Zhang 0006, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006, Zhiwen Shao, Hancheng Zhu
Pattern Anal. Appl.4
2023 Attention-guided Adversarial Attack for Video Object Segmentation
abstract
Video Object Segmentation (VOS) methods have made many breakthroughs with the help of the continuous development and advancement of deep learning. However, the deep learning model is vulnerable to malicious adversarial attacks, which mislead the model to make wrong decisions by adding adversarial perturbation that humans cannot perceive to the input image. Threats to deep learning models remind us that video object segmentation methods are also vulnerable to attacks, thereby threatening their security. Therefore, we study adversarial attacks on the VOS task to better identify the vulnerabilities of the VOS method, which in turn provides an opportunity to improve its robustness. In this paper, we propose an attention-guided adversarial attack method, which uses spatial attention blocks to capture features with global dependencies to construct correlations between consecutive video frames, and performs multipath aggregation to effectively integrate spatial-temporal perturbation, thereby guiding the deconvolution network to generate adversarial examples with strong attack capability. Specifically, the class loss function is designed to enable the deconvolution network to better activate noise in other regions and suppress the activation related to the object class based on the enhanced feature map of the object class. At the same time, attentional feature loss is designed to enhance the transferability against attack. The experimental results on the DAVIS dataset show that the proposed attention-guided adversarial attack method can significantly reduce the segmentation accuracy of OSVOS, and the J & F mean on DAVIS 2016 can reach 73.6% drop rate. The generated adversarial examples are also highly transferable to other video object segmentation models.
Rui Yao 0006, Ying Chen 0005, Yong Zhou 0003, Fuyuan Hu, Jiaqi Zhao 0001, Bing Liu 0016, Zhiwen Shao
ACM Trans. Intell. Syst. Technol.5
2023 Weakly Supervised Few-Shot Semantic Segmentation via Pseudo Mask Enhancement and Meta Learning
abstract
Few shot semantic segmentation has been proposed to enhance the generalization ability of traditional models with limited data. Previous works mainly focus on the supervised tasks, while limited amount of work is explored for the weakly supervised tasks. Weakly supervised semantic segmentation has become an active research area because weakly supervised labels effectively reduce the annotation cost of visual tasks. To this end, we propose a weakly supervised few-shot semantic segmentation model based on the meta learning framework, which utilizes prior knowledge and adjusts itself according to new tasks. Thereupon then, the proposed network is capable of both high efficiency and generalization ability to new tasks. In the pseudo mask generation stage, we develop a WRCAM method with the channel-spatial attention mechanism to refine the coverage size of targets in pseudo masks. In the few-shot semantic segmentation stage, the optimization based meta learning method is used to realize few-shot semantic segmentation by virtue of the refined pseudo masks. The experimental results show that the proposed method not only significantly outperforms weakly supervised SOTA methods, but also could be comparative to some supervised SOTA methods.
Man Zhang 0006, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006, Zhiwen Shao, Hancheng Zhu
IEEE Trans. Multim.4
2023 Spatial-Channel Enhanced Transformer for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) is a challenging task in computer vision, aiming at matching people across images from visible and infrared modalities. The widely used VI-ReID framework consists of a convolution neural backbone network that extracts the visual features, and a feature embedding network to project heterogeneous features to the same feature space. However, many studies based on the existing pre-trained models neglect potential correlations between different locations and channels within a single sample during the feature extraction. Inspired by the success of the Transformer in computer vision, we extend it to enhance feature representation for VI-ReID. In this paper, we propose a discriminative feature learning network based on a visual Transformer (DFLN-ViT) for VI-ReID. Firstly, to capture long-term dependencies between different locations, we propose a spatial feature awareness module (SAM), which utilizes a single-layer Transformer with a novel patch-embedding strategy to encode location information. Secondly, to refine the representation at each channel, we design a channel feature enhancement module (CEM). The CEM treats the features of each channel as a sequence of Transformer inputs, taking advantage of the Transformer's ability to model long-term dependencies. Finally, we propose a Triplet-aided Hetero-Center (THC) loss to learn more discriminative feature representation by balancing the cross-modality distance and intra-modality distance of the center. The experimental results on two datasets show that our method can significantly improve the VI-ReID performance, outperforming most state-of-the-art methods.
Jiaqi Zhao 0001, Hanzheng Wang, Yong Zhou 0003, Rui Yao 0006, Silin Chen, Abdulmotaleb El Saddik
IEEE Trans. Multim.1
2023 Distilled Meta-learning for Multi-Class Incremental Learning
abstract
Meta-learning approaches have recently achieved promising performance in multi-class incremental learning. However, meta-learners still suffer from catastrophic forgetting, i.e., they tend to forget the learned knowledge from the old tasks when they focus on rapidly adapting to the new classes of the current task. To solve this problem, we propose a novel distilled meta-learning (DML) framework for multi-class incremental learning that integrates seamlessly meta-learning with knowledge distillation in each incremental stage. Specifically, during inner-loop training, knowledge distillation is incorporated into the DML to overcome catastrophic forgetting. During outer-loop training, a meta-update rule is designed for the meta-learner to learn across tasks and quickly adapt to new tasks. By virtue of the bilevel optimization, our model is encouraged to reach a balance between the retention of old knowledge and the learning of new knowledge. Experimental results on four benchmark datasets demonstrate the effectiveness of our proposal and show that our method significantly outperforms other state-of-the-art incremental learning methods.
Hao Liu 0065, Zhaoyu Yan, Bing Liu 0016, Jiaqi Zhao 0001, Yong Zhou 0003, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Cyclic Self-attention for Point Cloud Recognition
abstract
Point clouds provide a flexible geometric representation for computer vision research. However, the harsh demands for the number of input points and computer hardware are still significant challenges, which hinder their deployment in real applications. To address these challenges, we design a simple and effective module named cyclic self-attention module (CSAM). Specifically, three attention maps of the same input are obtained by cyclically pairing the feature maps, thus exploring the features sufficiently of the attention space of the original input. CSAM can adequately explore the correlation between points to obtain sufficient feature information despite the multiplicative decrease in inputs. Meanwhile, it can direct the computational power to the more essential features, relieving the burden on the computer hardware. We build a point cloud classification network by simply stacking CSAM called cyclic self-attention network (CSAN). We also propose a novel framework for point cloud semantic segmentation called full cyclic self-attention network (FCSAN). By adaptively fusing the original mapping features and the CSAM extracted features, it can better capture the context information of point clouds. Extensive experiments on several benchmark datasets show that our methods can achieve competitive performance in classification and segmentation tasks.
Guanyu Zhu, Yong Zhou 0003, Rui Yao 0006, Hancheng Zhu, Jiaqi Zhao 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2022 Show, Deconfound and Tell: Image Captioning with Causal Inference
abstract
The transformer-based encoder-decoder framework has shown remarkable performance in image captioning. However, most transformer-based captioning methods ever overlook two kinds of elusive confounders: the visual confounder and the linguistic confounder, which generally lead to harmful bias, induce the spurious correlations during training, and degrade the model generalization. In this paper, we first use Structural Causal Models (SCMs) to show how two confounders damage the image captioning. Then we apply the backdoor adjustment to propose a novel causal inference based image captioning (CIIC) framework, which consists of an interventional object detector (IOD) and an interventional transformer decoder (ITD) to jointly confront both confounders. In the encoding stage, the IOD is able to disentangle the region-based visual features by deconfounding the visual confounder. In the decoding stage, the ITD introduces causal intervention into the transformer decoder and deconfounds the visual and linguistic confounders simultaneously. Two modules collaborate with each other to alleviate the spurious correlations caused by the unobserved confounders. When tested on MSCOCO, our proposal significantly outperforms the state-of-the-art encoder-decoder models on Karpathy split and online test split. Code is published in https://github.com/CUMTGG/CIIC.
Bing Liu 0016, Xu Yang 0021, Yong Zhou 0003, Rui Yao 0006, Zhiwen Shao, Jiaqi Zhao 0001
CVPR7
2022 Spatial hierarchy perception and hard samples metric learning for high-resolution remote sensing image object detection
Dongjun Zhu, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Ying Chen 0005
Appl. Intell.3
2022 Edge-aware and spectral-spatial information aggregation network for multispectral image semantic segmentation
Di Zhang 0020, Jiaqi Zhao 0001, Yong Zhou 0003, Boyu Shi, Rui Yao 0006
Eng. Appl. Artif. Intell.2
2022 Multi-granularity semantic alignment distillation learning for remote sensing image semantic segmentation
Di Zhang 0020, Yong Zhou 0003, Jiaqi Zhao 0001, Zhongyuan Yang, Rui Yao 0006, Huifang Ma
Frontiers Comput. Sci.3
2022 Multi-source collaborative enhanced for remote sensing images semantic segmentation
Jiaqi Zhao 0001, Di Zhang 0020, Boyu Shi, Yong Zhou 0003, Rui Yao 0006, Yong Xue
Neurocomputing1
2022 Point cloud recognition based on lightweight embeddable attention module
Guanyu Zhu, Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Man Zhang 0006
Neurocomputing3
2022 A Semi-Supervised Image-to-Image Translation Framework for SAR-Optical Image Matching
abstract
Synthetic Aperture Radar (SAR) and optical image matching aims to acquire correspondences from a certain pair of SAR and optical images. Recent advances in the image-to-image translation provided a way to simplify the SAR-optical image matching into the SAR-SAR or optical-optical image matchings. Existing image-to-image translations mainly focus on supervised or unsupervised learning. However, gathering sufficient amounts of aligned training data for supervised learning is challenging, while unsupervised learning cannot guarantee enough correct correspondences. In this work, we investigate the applicability of semi-supervised image-to-image translation for SAR-optical image matching such that both aligned and unaligned SAR-optical images could be used. To this end, we combine the benefits of both supervised and unsupervised well-known image-to-image translation methods, i.e., Pix2pix and CycleGAN, and propose a simple yet effective semi-supervised image-to-image translation framework. Through extensive experimental comparisons to baseline methods, we verify the effectiveness of the proposed framework in both semi-supervised and fully-supervised settings. Our codes are available at https://github.com/WenliangDu/Semi-I2I.
Wen-Liang Du 0002, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao, Xiaolin Tian 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Semantic Segmentation of Remote-Sensing Images Based on Multiscale Feature Fusion and Attention Refinement
abstract
In recent years, the automatic extraction of remote-sensing image information has attracted full attention. However, the particularity of remote-sensing images and the scarcity of data sets with label information have brought new challenges to existing methods. Therefore, we develop a lightweight semantic segmentation network based onmultiscale feature fusion (MFF) and attention refinement (MFFANet). Our network relies on three crucial modules for improved performance. The multiscale attention refinement module strengthens the representation ability of feature maps extracted by the deep residual network. The MFF module aggregates the information carried by the high-level and low-level features while restoring the image resolution. Furthermore, the boundary enhancement module captures boundary details to solve the semantic ambiguity problem. We achieve 83.5% mean intersection over union (MIoU) on the Urban Semantic 3-D (US3D) data set and 69.3% MIoU on the Vaihingen data set with only 8.2M parameters.
Xin He 0024, Yong Zhou 0003, Jiaqi Zhao 0001, Man Zhang 0006, Rui Yao 0006, Bing Liu 0016
IEEE Geosci. Remote. Sens. Lett.3
2022 Semisupervised Multiscale Generative Adversarial Network for Semantic Segmentation of Remote Sensing Image
abstract
Semantic segmentation of remote sensing images based on deep neural networks has gained wide attention recently. Although many methods have achieved amazing performance, they need large amounts of labeled images to distinguish the differences in angle, color, size, and other aspects for small targets in remote sensing data sets. However, with a few labeled images, it is difficult to extract the key features of small targets. We propose a semisupervised multiscale generative adversarial network (GAN), which not only utilizes the multipath input and atrous spatial pyramid pooling (ASPP) module but leverages unlabeled images and semisupervised learning strategy to improve the performance of small target segmentation in semantic segmentation when labeled data amount is small. Experimental results show that our model outperforms state-of-the-art methods with insufficient labeled data.
Bing Liu 0016, Yong Zhou 0003, Jiaqi Zhao 0001, Shixiong Xia, Yuancan Yang, Man Zhang 0006, Liu Ming Ming
IEEE Geosci. Remote. Sens. Lett.4
2022 Few-Shot Object Detection via Context-Aware Aggregation for Remote Sensing Images
abstract
Few-shot object detection methods have made prodigious progress in recent years. However, these methods are designed for optical images at a single scale, which leads to significantly degraded detection performance due to object scale variation of remote sensing images. In this letter, we propose a few-shot object detection method for the problem of scale variation in remote sensing images. More specifically, our model contains two main components: a context-aware pixel aggregation (CPA) that allows the model to adapt to objects at different scales through different scale convolution and a context-aware feature aggregation (CFA) that enhances context awareness to obtain more semantic information through a graph convolution network (GCN). Experiments on the DIOR dataset demonstrate that our model can achieve a satisfying detection performance on remote sensing images, and our model performs significantly better than the state-of-the-art model.
Yong Zhou 0003, Jiaqi Zhao 0001, Hancheng Zhu, Rui Yao 0006, Wen-Liang Du 0002
IEEE Geosci. Remote. Sens. Lett.3
2022 Fine-Grained Feature Enhancement for Object Detection in Remote Sensing Images
abstract
Recently, object detection in aerial images has ushered in a new challenge—a new benchmark for fine-grained object recognition in high-resolution remote sensing imagery called FAIR1M has been proposed. Fine-grained categories usually have smaller inter class differences and intra-class similarities, which is more difficult to classify with existing object detectors. To address this problem, we propose two enhanced strategies on the current two-stage object detection algorithm. The first strategy uses attention-based group feature enhancement called group enhance module (GEM). By extending and grouping feature channels, the model can improve the ability to extract various discriminative features. The second strategy is to emphasize the sub-saliency feature learning, avoiding the network only focusing on the most significant part of the feature and ignoring the other parts. Our method is easy to implement and effective, and experiments show that our method can improve the Oriented regions with convolutional neural networks features (R-CNN) by about 1.45 mAP on the FAIR1M benchmark.
Yong Zhou 0003, Sifan Wang, Jiaqi Zhao 0001, Hancheng Zhu, Rui Yao 0006
IEEE Geosci. Remote. Sens. Lett.3
2022 Efficient lightweight video person re-identification with online difference discrimination module
Cunyuan Gao, Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Fuyuan Hu
Multim. Tools Appl.4
2022 Survey for person re-identification based on coarse-to-fine feature learning
Minjie Liu, Jiaqi Zhao 0001, Yong Zhou 0003, Hancheng Zhu, Rui Yao 0006, Ying Chen 0005
Multim. Tools Appl.2
2022 ResT-ReID: Transformer block-based residual learning for person re-identification
Ying Chen 0005, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Dongjun Zhu, Dongjingdian Liu
Pattern Recognit. Lett.3
2022 Spatial-Temporal Based Multihead Self-Attention for Remote Sensing Image Change Detection
abstract
The neural network-based remote sensing image change detection method faces a large amount of imaging interference and severe class imbalance problems under high-resolution conditions, which bring new challenges to the accuracy of the detection network. In this work, to address the imaging interference caused by different imaging angles and times, the siamese strategy and multi-head self-attention mechanism are used to reduce the imaging differences between the dual-temporal images and fully exploit the inter-temporal information. Secondly, a learnable multi-part feature learning module is used to adaptively exploit features from different scales to obtain more comprehensive features. Finally, a mixed loss function strategy is used to ensure that the network converges effectively and excludes the adverse interference of a large number of negative samples to the network. Extensive experiments show that our method outperforms numerous methods on LEVIR-CD, WHU, and DSIFN datasets.
Yong Zhou 0003, Fengkai Wang, Jiaqi Zhao 0001, Rui Yao 0006, Silin Chen, Heping Ma
IEEE Trans. Circuits Syst. Video Technol.3
2022 Swin Transformer Embedding UNet for Remote Sensing Image Semantic Segmentation
abstract
Global context information is essential for the semantic segmentation of remote sensing (RS) images. However, most existing methods rely on a convolutional neural network (CNN), which is challenging to directly obtain the global context due to the locality of the convolution operation. Inspired by the Swin transformer with powerful global modeling capabilities, we propose a novel semantic segmentation framework for RS images called ST-U-shaped network (UNet), which embeds the Swin transformer into the classical CNN-based UNet. ST-UNet constitutes a novel dual encoder structure of the Swin transformer and CNN in parallel. First, we propose a spatial interaction module (SIM), which encodes spatial information in the Swin transformer block by establishing pixel-level correlation to enhance the feature representation ability of occluded objects. Second, we construct a feature compression module (FCM) to reduce the loss of detailed information and condense more small-scale features in patch token downsampling of the Swin transformer, which improves the segmentation accuracy of small-scale ground objects. Finally, as a bridge between dual encoders, a relational aggregation module (RAM) is designed to integrate global dependencies from the Swin transformer into the features from CNN hierarchically. Our ST-UNet brings significant improvement on the ISPRS-Vaihingen and Potsdam datasets, respectively. The code will be available athttps://github.com/XinnHe/ST-UNet.
Xin He 0024, Yong Zhou 0003, Jiaqi Zhao 0001, Di Zhang 0020, Rui Yao 0006, Yong Xue
IEEE Trans. Geosci. Remote. Sens.3
2022 CLT-Det: Correlation Learning Based on Transformer for Detecting Dense Objects in Remote Sensing Images
abstract
Challenges still exist in the task of object detection in remote sensing images with densely distributed objects due to large variation in scale and neglect of the relative position and correlation. To address these issues, a Correlation Learning Detector based on Transformer (CLT-Det) is proposed for detecting dense objects in remote sensing images. A Transformer Attention Module (TAM) is designed to improve the densely packed objects’ model representation ability by learning pixel-wise attention with Transformer. To alleviate the semantic gap caused by variations in scale, a Feature Refinement Module (FRM) is proposed by improving the multi-scale feature pyramid. A Correlation Transformer Module (CTM) is proposed to extract correlation information and encodes position information of dense objects’ features on the classification branch for fully utilizing the position information and correlation among objects. Extensive experiments compared with several state-of-art methods on two challenging remote sensing datasets, namely DOTA and HRSC2016, demonstrate that the proposed CLT-Det achieves promising and competitive performance.
Yong Zhou 0003, Silin Chen, Jiaqi Zhao 0001, Rui Yao 0006, Yong Xue, Abdulmotaleb El Saddik
IEEE Trans. Geosci. Remote. Sens.3
2022 Clustering Matters: Sphere Feature for Fully Unsupervised Person Re-identification
abstract
In person re-identification (Re-ID) , the data annotation cost of supervised learning, is huge and it cannot adapt well to complex situations. Therefore, compared with supervised deep learning methods, unsupervised methods are more in line with actual needs. In unsupervised learning, a key to solving Re-ID is to find a standard that can effectively distinguish the difference (distance) between the features of images belonging to different pedestrian identities. However, there are some differences in the images captured by different cameras (such as brightness, angle, etc.). It is well known that the training of neural networks is mainly based on the distance between features, while in unsupervised learning, especially in unsupervised learning methods based on hierarchical clustering, the distance between features plays a more important role in the clustering phase. We improve the accuracy of a deep learning method based on hierarchical clustering under fully unsupervised conditions, starting from both feature and distance metrics. First, we propose to use spherical features, by normalizing the images in the feature space, to weaken the structural differences (length) between features, while saving the feature differences (direction) between different identities. Then, we use the sum of squared errors (SSE) as a regularization term to balance different cluster states. We evaluate our method on four large-scale Re-ID datasets, and experiments show that our method achieves better results than the state-of-the-art unsupervised methods.
Yong Zhou 0003, Jiaqi Zhao 0001, Ying Chen 0005, Rui Yao 0006, Bing Liu 0016, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Facial action unit detection via hybrid relational reasoning
Zhiwen Shao, Yong Zhou 0003, Bing Liu 0016, Hancheng Zhu, Wen-Liang Du 0002, Jiaqi Zhao 0001
Vis. Comput.6
2021 Path Planning based on Multi-objective Topological Map
abstract
Recently, intelligent robot technology has developed rapidly. As an important topic, path planning has attracted more and more attention. However, the performance of multi-objective path planning is limited by the scale of the problem, and the results of path planning are generally not optimal. In this work, based on 12 problems defined in multimodal multiobjective path planning optimization contest for CEC 2021 [1], we present a multi-objective path planning optimization algorithm, which includes data preprocessing, multi-objective genetic evolution path planning algorithm, and non-dominated sorting algorithm with elite strategy. This algorithm flow can realize multi-objective path planning and generate the optimal solution set (i.e. Pareto solution set). We implement our algorithm flow to calculate the time complexity and space complexity indicators. Experimental results on problems indicate the proposed multi-objective path planning algorithm can solve the optimal solution set in time. Due to the constraints of the problem, the number of optimal solutions is different for different problems. We show the validity of our method with experiments for path planning. Finally, we visually present some experimental results intuitively. The code is released at https://github.com/zhangruihao/pathPlanning.
Jiaqi Zhao 0001, Zhijie Jia, Yong Zhou 0003, Zeming Xie, Di Zhang 0020
CEC1
2021 Multi-Objective Net Architecture Pruning for Remote Sensing Classification
abstract
Remote sensing image scene classification has achieved significant breakthroughs in recent years. However, due to the high complexity and expensive computation most of CNNs used in the field of remote sensing imagery scene classification, it has become a challenging task for extracting effective features at restricted hardware conditions. To solve this problem, we present a model compression method by means of evolutionary algorithms. Specifically, we compress the model by pruning filters and transform the compression of the CNN model into a multi-objective optimization problem based on classification accuracy and compression ratio by using the adaptive-BN-based evaluation method. Furthermore, the prior knowledge of ResNet-50 on ImageNet is introduced to reduce the instability of evolutionary algorithm as a result of random population initialization. Experiments are implemented on three datasets with two evolutionary algorithms, and results demonstrate that our method can achieve state-of-the-art performances.
Jiaqi Zhao 0001, Chengrun Yang, Yong Zhou 0003, Zhujun Jiang, Ying Chen 0005
IGARSS1
2021 Joint Attention Mechanism for Unsupervised Video Object Segmentation
Rui Yao 0006, Xin Xu 0009, Yong Zhou 0003, Jiaqi Zhao 0001
PRCV (1)4
2021 Point cloud classification by dynamic graph CNN with adaptive feature fusion
abstract
Abstract The deep neural network has made the most advanced breakthrough in almost all 2D image tasks, so we consider the application of deep learning in 3D images. Point cloud data, as the most basic and important form of representation of 3D images, can accurately and intuitively show the real world. The authors propose a new network based on feature fusion to improve the point cloud classification and segmentation tasks. Our network mainly consists of three parts: global feature extractor, local feature extractor and adaptive feature fusion module. A multi‐scale transformation network is devised to guarantee the invariance of the transformation of the global feature, and a residual block is introduced to alleviate the problem of gradient disappearance to enhance the global feature extractor. Based on the edge convolution and multi‐layer perceptron, a local feature extractor is constructed. Finally, an adaptive feature‐fusion module is proposed to complete the fusion of global features and local features. Extensive experiments on point cloud classification and segmentation tasks are carried out to verify the effectiveness of the proposed method. The classification accuracy of the ModelNet40 is 93.6%, which is 4.4% higher than that of the PointNet. Similarly, the segmentation accuracy on the ShapeNet is 85.6%, which is higher than other methods.
Yong Zhou 0003, Jiaqi Zhao 0001, Yiyun Man, Minjie Liu, Rui Yao 0006, Bing Liu 0016
IET Comput. Vis.3
2021 AMC-Net: Attentive modality-consistent network for visible-infrared person re-identification
Hanzheng Wang, Jiaqi Zhao 0001, Yong Zhou 0003, Rui Yao 0006, Ying Chen 0005, Silin Chen
Neurocomputing2
2021 Unsupervised cross-domain person re-identification with self-attention and joint-flexible optimization
Haopeng Hou, Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Ying Chen 0005, Abdulmotaleb El Saddik
Image Vis. Comput.3
2021 A siamese pedestrian alignment network for person re-identification
Yong Zhou 0003, Jiaqi Zhao 0001, Meng Jian, Rui Yao 0006, Bing Liu 0016, Ying Chen 0005
Multim. Tools Appl.3
2021 Video-based person re-identification by semi-supervised adaptive stepwise learning
Yong Zhou 0003, Jiaqi Zhao 0001, Ying Chen 0005, Rui Yao 0006
Pattern Anal. Appl.3
2021 Semi-supervised blockwisely architecture search for efficient lightweight generative adversarial network
Man Zhang 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Shixiong Xia, Zizheng Huang
Pattern Recognit.3
2021 Multi-Stage Fusion and Multi-Source Attention Network for Multi-Modal Remote Sensing Image Segmentation
abstract
With the rapid development of sensor technology, lots of remote sensing data have been collected. It effectively obtains good semantic segmentation performance by extracting feature maps based on multi-modal remote sensing images since extra modal data provides more information. How to make full use of multi-model remote sensing data for semantic segmentation is challenging. Toward this end, we propose a new network called Multi-Stage Fusion and Multi-Source Attention Network ((MS) 2 -Net) for multi-modal remote sensing data segmentation. The multi-stage fusion module fuses complementary information after calibrating the deviation information by filtering the noise from the multi-modal data. Besides, similar feature points are aggregated by the proposed multi-source attention for enhancing the discriminability of features with different modalities. The proposed model is evaluated on publicly available multi-modal remote sensing data sets, and results demonstrate the effectiveness of the proposed method.
Jiaqi Zhao 0001, Yong Zhou 0003, Boyu Shi, Jingsong Yang, Di Zhang 0020, Rui Yao 0006
ACM Trans. Intell. Syst. Technol.1
2020 Person image synthesis through siamese generative adversarial network
Ying Chen 0005, Shixiong Xia, Jiaqi Zhao 0001, Meng Jian, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Dongjun Zhu
Neurocomputing3
2020 Multiobjective ResNet pruning by means of EMOAs for remote sensing scene classification
Xuning Liu, Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Bing Liu 0016
Neurocomputing3
2020 Diverse sample generation with multi-branch conditional generative adversarial network for remote sensing objects detection
Dongjun Zhu, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Meng Jian, Qiang Niu, Rui Yao 0006, Ying Chen 0005
Neurocomputing3
2020 Appearance and shape based image synthesis by conditional variational generative adversarial network
Ying Chen 0005, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Dongjun Zhu
Knowl. Based Syst.3
2020 Remote sensing image captioning via Variational Autoencoder and Reinforcement Learning
Xiangqing Shen, Bing Liu 0016, Yong Zhou 0003, Jiaqi Zhao 0001
Knowl. Based Syst.4
2020 Remote sensing image caption generation via transformer and reinforcement learning
Xiangqing Shen, Bing Liu 0016, Yong Zhou 0003, Jiaqi Zhao 0001
Multim. Tools Appl.4
2020 Fusion based feature reinforcement component for remote sensing image object detection
Dongjun Zhu, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Ying Chen 0005
Multim. Tools Appl.3
2020 GAN-based person search via deep complementary classifier with center-constrained Triplet loss
Rui Yao 0006, Cunyuan Gao, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Fuyuan Hu
Pattern Recognit.4
2020 Video Object Segmentation and Tracking: A Survey
abstract
Object segmentation and object tracking are fundamental research areas in the computer vision community. These two topics are difficult to handle some common challenges, such as occlusion, deformation, motion blur, scale variation, and more. The former contains heterogeneous object, interacting object, edge ambiguity, and shape complexity; the latter suffers from difficulties in handling fast motion, out-of-view, and real-time processing. Combining the two problems of Video Object Segmentation and Tracking (VOST) can overcome their respective difficulties and improve their performance. VOST can be widely applied to many practical applications such as video summarization, high definition video compression, human computer interaction, and autonomous vehicles. This survey aims to provide a comprehensive review of the state-of-the-art VOST methods, classify these methods into different categories, and identify new trends. First, we broadly categorize VOST methods into Video Object Segmentation (VOS) and Segmentation-based Object Tracking (SOT). Each category is further classified into various types based on the segmentation and tracking mechanism. Moreover, we present some representative VOS and SOT methods of each time node. Second, we provide a detailed discussion and overview of the technical characteristics of the different methods. Third, we summarize the characteristics of the related video dataset and provide a variety of evaluation metrics. Finally, we point out a set of interesting future works and draw our own conclusions.
Rui Yao 0006, Guosheng Lin, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003
ACM Trans. Intell. Syst. Technol.4
2019 Lightweight Video Object Segmentation Based on ConvGRU
Rui Yao 0006, Yikun Zhang 0001, Cunyuan Gao, Yong Zhou 0003, Jiaqi Zhao 0001, Lina Liang
PRCV (2)5
2019 A Siamese Pedestrian Alignment Network for Person Re-identification
Yong Zhou 0003, Jiaqi Zhao 0001, Meng Jian, Rui Yao 0006, Bing Liu 0016, Xuning Liu
PRCV (1)3
2019 Structure-aware person search with self-attention and online instance aggregation matching
Cunyuan Gao, Rui Yao 0006, Jiaqi Zhao 0001, Yong Zhou 0003, Fuyuan Hu, Leida Li
Neurocomputing3
2019 Siamese Convolutional Neural Networks for Remote Sensing Scene Classification
abstract
The convolutional neural networks (CNNs) have shown powerful feature representation capability, which provides novel avenues to improve scene classification of remote sensing imagery. Although we can acquire large collections of satellite images, the lack of rich label information is still a major concern in the remote sensing field. In addition, remote sensing data sets have their own limitations, such as the small scale of scene classes and lack of image diversity. To mitigate the impact of the existing problems, a Siamese CNN, which combines the identification and verification models of CNNs, is proposed in this letter. A metric learning regularization term is explicitly imposed on the features learned through CNNs, which enforce the Siamese networks to be more robust. We carried out experiments on three widely used remote sensing data sets for performance evaluation. Experimental results show that our proposed method outperforms the existing methods.
Xuning Liu, Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Bing Liu 0016
IEEE Geosci. Remote. Sens. Lett.3
2018 Pareto-Based Many-Objective Convolutional Neural Networks
Hongjian Zhao, Shixiong Xia, Jiaqi Zhao 0001, Dongjun Zhu, Rui Yao 0006, Qiang Niu
WISA3
2018 Multiobjective sparse ensemble learning by means of evolutionary algorithms
Jiaqi Zhao 0001, Licheng Jiao, Shixiong Xia, Vitor Basto-Fernandes, Iryna Yevseyeva, Yong Zhou 0003, Michael T. M. Emmerich
Decis. Support Syst.1
2018 Deep Multiple Instance Learning-Based Spatial-Spectral Classification for PAN and MS Imagery
abstract
Panchromatic (PAN) and multispectral (MS) imagery classification is one of the hottest topics in the field of remote sensing. In recent years, deep learning techniques have been widely applied in many areas of image processing. In this paper, an end-to-end learning framework based on deep multiple instance learning (DMIL) is proposed for MS and PAN images’ classification using the joint spectral and spatial information based on feature fusion. There are two instances in the proposed framework: one instance is used to capture the spatial information of PAN and the other is used to describe the spectral information of MS. The features obtained by the two instances are concatenated directly, which can be treated as simple fusion features. To fully fuse the spatial–spectral information for further classification, the simple fusion features are fed into a fusion network with three fully connected layers to learn the high-level fusion features. Classification experiments carried out on four different airborne MS and PAN images indicate that the classifier provides feasible and efficient solution. It demonstrates that DMIL performs better than using a convolutional neural network and a stacked autoencoder network separately. In addition, this paper shows that the DMIL model can learn and fuse spectral and spatial information effectively, and has huge potential for MS and PAN imagery classification.
Xu Liu 0006, Licheng Jiao, Jiaqi Zhao 0001, Jin Zhao 0002, Fang Liu 0001, Shuyuan Yang 0001, Xu Tang 0004
IEEE Trans. Geosci. Remote. Sens.3
2017 Corrigendum to 'Multiobjective optimization of classifiers by means of 3D convex-hull-based evolutionary algorithms' [Information Sciences volumes 367-368 (2016) 80-104]
Jiaqi Zhao 0001, Vitor Basto-Fernandes, Licheng Jiao, Iryna Yevseyeva, Asep Maulana, Rui Li 0001, Thomas Bäck, Ke Tang 0001, Michael T. M. Emmerich
Inf. Sci.1
2017 Sparse learning based fuzzy c-means clustering
Licheng Jiao, Shuyuan Yang 0001, Jiaqi Zhao 0001
Knowl. Based Syst.4
2017 Semi-supervised double sparse graphs based discriminant analysis for dimensionality reduction
Puhua Chen, Licheng Jiao, Fang Liu 0001, Jiaqi Zhao 0001, Shuai Liu 0016
Pattern Recognit.4
2017 Quantum-behaved discrete multi-objective particle swarm optimization for complex network clustering
Lingling Li 0002, Licheng Jiao, Jiaqi Zhao 0001, Ronghua Shang, Maoguo Gong
Pattern Recognit.3
2017 Discriminant deep belief network for high-resolution SAR image classification
Licheng Jiao, Jiaqi Zhao 0001, Jin Zhao 0002
Pattern Recognit.3
2017 Superpixel-Based Multiple Local CNN for Panchromatic and Multispectral Image Classification
abstract
Recently, very high resolution (VHR) panchromatic and multispectral (MS) remote-sensing images can be acquired easily. However, it is still a challenging task to fuse and classify these VHR images. Generally, there are two ways for the fusion and classification of panchromatic and MS images. One way is to use a panchromatic image to sharpen an MS image, and then classify a pan-sharpened MS image. Another way is to extract features from panchromatic and MS images, respectively, and then combine these features for classification. In this paper, we propose a superpixel-based multiple local convolution neural network (SML-CNN) model for panchromatic and MS images classification. In order to reduce the amount of input data for the CNN, we extend simple linear iterative clustering algorithm for segmenting MS images and generating superpixels. Superpixels are taken as the basic analysis unit instead of pixels. To make full advantage of the spatial-spectral and environment information of superpixels, a superpixel-based multiple local regions joint representation method is proposed. Then, an SML-CNN model is established to extract an efficient joint feature representation. A softmax layer is used to classify these features learned by multiple local CNN into different categories. Finally, in order to eliminate the adverse effects on the classification results within and between superpixels, we propose a multi-information modification strategy that combines the detailed information and semantic information to improve the classification performance. Experiments on the classification of Vancouver and Xi’an panchromatic and MS image data sets have demonstrated the effectiveness of the proposed approach.
Wei Zhao 0014, Licheng Jiao, Wenping Ma 0001, Jiaqi Zhao 0001, Jin Zhao 0002, Hongying Liu 0001, Xianghai Cao, Shuyuan Yang 0001
IEEE Trans. Geosci. Remote. Sens.4
2016 Locality-constraint discriminant feature learning for high-resolution SAR image classification
Licheng Jiao, Biao Hou, Shuang Wang 0001, Jiaqi Zhao 0001, Puhua Chen
Neurocomputing5
2016 Multiobjective optimization of classifiers by means of 3D convex-hull-based evolutionary algorithms
Jiaqi Zhao 0001, Vitor Basto-Fernandes, Licheng Jiao, Iryna Yevseyeva, Asep Maulana, Rui Li 0001, Thomas Bäck, Ke Tang 0001, Michael T. M. Emmerich
Inf. Sci.1
2016 Semisupervised Discriminant Feature Learning for SAR Image Category via Sparse Ensemble
abstract
Terrain scene classification plays an important role in various synthetic aperture radar (SAR) image understanding and interpretation. This paper presents a novel approach to characterize SAR image content by addressing category with a limited number of labeled samples. In the proposed approach, each SAR image patch is characterize by a discriminant feature which is generated in a semisupervised manner by utilizing a spare ensemble learning procedure. In particular, a nonnegative sparse coding procedure is applied on the given SAR image patch set to generate the feature descriptors first. The set is combined with a limited number of labeled SAR image patches and an abundant number of unlabeled ones. Then, a semisupervised sampling approach is proposed to construct a set of weak learners, in which each one is modeled by a logistic regression procedure. The discriminant information can be introduced by projecting SAR image patch on each weak learner. Finally, the features of SAR image patches are produced by a sparse ensemble procedure which can reduce the redundancy of multiple weak learners. Experimental results show that the proposed discriminant feature learning approach can achieve a higher classification accuracy than several state-of-the-art approaches.
Licheng Jiao, Fang Liu 0001, Jiaqi Zhao 0001, Puhua Chen
IEEE Trans. Geosci. Remote. Sens.4