Rui Yao 0006

dblp:72/2221-6 · DBLP profile ↗
← Back
116ranked-venue papers
21as first author
86since 2021 · last 2026
0000-0003-2734-915XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 61 · 11 first-author · 44 since 2021Artificial intelligence and machine learning · 49 · 8 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 16 since 2021Computer networks · 11 · 11 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal Modeling
abstract
Video shadow detection confronts two entwined difficulties: distinguishing shadows from complex backgrounds and modeling dynamic shadow deformations under varying illumination. To address shadow-background ambiguity, we leverage linguistic priors through the proposed Vision-language Match Module (VMM) and a Dark-aware Semantic Block (DSB), extracting text-guided features to explicitly differentiate shadows from dark objects. Furthermore, we introduce adaptive mask reweighting to downweight penumbra regions during training and apply edge masks at the final decoder stage for better supervision. For temporal modeling of variable shadow shapes, we propose a Tokenized Temporal Block (TTB) that decouples spatiotemporal learning. TTB summarizes cross-frame shadow semantics into learnable temporal tokens, enabling efficient sequence encoding with minimal computation overhead. Comprehensive Experiments on multiple benchmark datasets demonstrate state-of-the-art accuracy and real-time inference efficiency.
Kunyang Sun, Rui Yao 0006, Hancheng Zhu, Fuyuan Hu, Jiaqi Zhao 0001, Zhiwen Shao, Yong Zhou 0003
AAAI3
2026 CLIPDet3D: Vision-Language Collaborative Distillation for 3D Object Detection
abstract
Multi-view 3D object detection plays a vital role in autonomous driving systems due to its ability to perceive complex scenes accurately. However, real-world driving data often exhibits a long-tailed distribution, causing significant drops in detection accuracy for rare categories in existing methods. To mitigate this issue, we propose CLIPDet3D, a novel vision-language collaborative framework for multi-view 3D object detection. First, to tackle the difficulty of capturing the semantic information of rare categories, a Vision-Language Collaborative Learning strategy is proposed to incorporate class-level semantic priors from CLIP. Second, a Depth Feature Contrastive Distillation module is designed to overcome the large depth estimation error for rare categories by aligning depth features between a teacher and a student network. Furthermore, to alleviate the difficulty in focusing on regions of rare categories, a Dual-Stream Prompt Attention mechanism is devised to inject learnable prompts and compute attention along both horizontal and vertical BEV directions. Evaluations on the nuScenes dataset demonstrate that CLIPDet3D achieves state-of-the-art accuracy while maintaining efficient inference.
Jiaqi Zhao 0001, Huanfeng Hu, Yong Zhou 0003, Wen-Liang Du 0002, Kunyang Sun, Rui Yao 0006, Qigong Sun
AAAI6
2026 Unified Representation Causal Prompt Distillation for Re-Inference-Free Lifelong Person Re-Identification
abstract
Lifelong person re-identification (LReID) aims to retrieve the target person from sequentially collected data. Due to significant domain gaps between datasets and the continuous increase of training data from different scenarios, weak inter-domain generalization and catastrophic forgetting issues have remained major challenges for LReID. To tackle these issues, a novel LReID method called Unified Representation Causal Prompt Distillation (URCPD) is proposed. Specifically, to reduce domain gaps among different scene datasets and improve model inter-domain generalization capability, a Feature Decoupling Style Transfer module (FDST) is proposed to map new features into a unified feature space. Furthermore, to reduce the accumulated forgetting of old knowledge during the training stage, a Causal Prompt Distillation module (CPD) is introduced. This module eliminates the re-inference process for distillation and embeds memory prompts to combat catastrophic forgetting. Extensive experiments on five classic LReID seen datasets and seven unseen datasets demonstrate that our method significantly outperforms state-of-the-art methods.
Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006
AAAI6
2026 Causal Decoupling Domain Generalization for Remote Sensing Change Detection
abstract
While current state-of-the-art Remote Sensing Change Detection (RSCD) methods can achieve impressive results on individual datasets, they become unreliable in unseen environments and imaging conditions, with performance metrics declining by as much as 60% to 80%. Simultaneously, variable environments and complex imaging conditions are the main characteristics of remote sensing data, calling for generalizable RSCD methods. To address this issue, we propose a novel RSCD method capable of domain generalization—CDDGNet. This method is based on causal decoupling theory, which progressively decouples invariant change features from variable domain features to extract generalizable characteristics. This enables a network trained on a single domain to accurately identify change regions in other domains. Specifically, firstly, the Causal Feature Adaptation Module is proposed to preliminarily decouple and simplify feature information during the encoding process by using wavelet transformation and feature energy spectralization methods. Secondly, the Causal Feature Fusion Module is presented to fully decouple features and aggregate significant change features during the decoding process through frequency domain processing and feature re-attention mechanisms. Thirdly, the Decoupling Effect Loss Function is proposed to optimize the process by evaluating the effectiveness of causal decoupling. Extensive experiments have shown that our model significantly outperforms existing methods across multiple groups of generalization tasks with varying levels of difficulty.
Jiaqi Zhao 0001, Jianpeng Xie 0001, Yong Zhou 0003, Wen-Liang Du 0002, Hancheng Zhu, Rui Yao 0006
AAAI6
2026 Few-Shot Quaternion-valued Correlation Squeeze Network for Document Image Layout Segmentation
Rui Yao 0006, Qiwei Yu, Songhui Zhao, Yong Zhou 0003, Bing Liu 0016
Int. J. Document Anal. Recognit.1
2026 Style-controllable adversarial example generation via image editing and prompt embedding optimization
Yong Zhou 0003, Bing Liu 0016, Rui Yao 0006
Neurocomputing4
2026 A unified multi-stream diffusion framework for robust video camouflaged object detection
Yuyao Ke, Rui Yao 0006, Kunyang Sun, Hancheng Zhu, Jiaqi Zhao 0001, Bing Liu 0016
Neural Networks2
2026 Dynamic Prompt Memory Network for video shadow detection
Rui Yao 0006, Hancheng Zhu, Kunyang Sun, Jiaqi Zhao 0001, Zhiwen Shao, Abdulmotaleb El Saddik
Pattern Recognit.1
2026 SpaceFormer: Spatial Position Contextual Semantics Embedding for Multi-View 3D Object Detection
abstract
3D object detection aims to accurately localize and recognize objects in 3D space. It serves as a fundamental task for reliable perception in intelligent transportation systems, enabling the monitoring of diverse traffic participants such as vehicles, pedestrians, cyclists, and public transport. Recently, transformer-based methods have gained significant attention in multi-view 3D object detection due to their strong global reasoning capabilities. However, their limited capacity to model spatial positional information hinders accurate object localization, especially in complex and large-scale scenes. To address this limitation, SpaceFormer is proposed as a novel transformer-based multi-view 3D object detector. Specifically, a Contextual Visual Prompts Learning strategy is proposed to enhance the perception of small and sparse traffic participants by incorporating contextual priors. To further suppress background interference, a Semantics-guided Depth Estimation method is proposed to refine depth representations using high-level semantic information. Furthermore, a Spatial Position Embedding mechanism is proposed to improve the spatial localization capability of the transformer by integrating geometric position and polar spatial embedding. Extensive experiments on the nuScenes benchmark demonstrate that SpaceFormer achieves state-of-the-art performance with 55.5% mAP and 62.9% NDS. These improvements indicate not only methodological advances but also practical benefits for intelligent transportation systems, enhancing safety, reliability, and efficiency in real-world deployments.
Jiaqi Zhao 0001, Huanfeng Hu, Wen-Liang Du 0002, Yong Zhou 0003, Kunyang Sun, Rui Yao 0006, Abdulmotaleb El Saddik
IEEE Trans. Intell. Transp. Syst.6
2026 Dual Sparse Long-Short Term Transformer for Video Shadow Detection
abstract
Video Shadow Detection (VSD) is critical yet challenging, primarily due to ambiguous shadow boundaries and the presence of confusing shadow-like non-shadow regions, which existing methods struggle to resolve effectively by limited temporal modeling. We propose the Dual Sparse Long-Short Term Transformer Network (DSLSTT-Net), a novel framework designed to enhance feature learning by integrating robust temporal consistency and detailed local context. DSLSTT-Net utilizes a dual-stream architecture to concurrently process global temporal information and local shadow feature refinement, enabling effective discrimination between true shadows and confusing areas. At its core, the Sparse Long-Short Term Attention Module (Sparse LSTAM) is introduced to efficiently propagate only high-confidence shadow features from memory, significantly enhancing feature discriminability and computational efficiency. Furthermore, an Adaptive Fusion Module (AFM) dynamically merges purified long-term features with short-term details, optimizing final segmentation. Experimental results confirm that DSLSTT-Net significantly outperforms state-of-the-art methods on VSD benchmarks, validating our approach of dual-stream architecture and sparse temporal modeling. The source code is available at https://github.com/rayyao/DSLSTTNet .
Rui Yao 0006, Huili Hao, Hancheng Zhu, Jiaqi Zhao 0001, Yong Zhou 0003
ACM Trans. Multim. Comput. Commun. Appl.2
2026 Spatio-Temporal Disentanglement and Constrained Self-Attention for Multi-Modal Deception Detection
abstract
Multi-modal deception detection is a challenging yet important task, having pivotal applications in many fields such as business credibility assessment and multimedia anti-frauds. Previous methods either rely solely on spatial features or overemphasize only temporal information within or across modalities, which may overlook potential critical clues. Motivated by these observations, we propose a Spatio-Temporal Representation Disentanglement (STRD) framework for multi-modal deception detection, which uses a dual-encoder structure to learn spatial and temporal representations for each modality. Specifically, we introduce a pre-trained foundation model to act as the spatial encoder and design a lightweight network as the temporal encoder, extracting spatial semantics and capturing dynamic temporal patterns. Then, we propose a Constrained Self-Attention Block (CSAB), in which self-attention distribution of each head is regarded as spatial distribution and is constrained to attend a certain facial local region. Furthermore, we present a Cross-Modal Correlation Fusion Block (CCFB) to achieve temporal synchronization across modalities by measuring the correlations between visual and audio features. Extensive experiments show that our STRD outperforms the state-of-the-art methods on challenging DOLOS, BOL, BgOL, and RLtrial benchmarks. Particularly, STRD improves by 2.12% and 1.88% over the previous best results in terms of ACC on the DOLOS and BOL datasets, respectively. Additionally, STRD outperforms previous methods in cross-dataset testing, highlighting its superior generalization ability.
Zhiwen Shao, Hancheng Zhu, Rui Yao 0006, Lixin Zou, Mengtian Li 0002, Bin Sheng 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Facial Action Unit Detection with Iterative Rank Reduction Adapter and Directional Attention
Zhiwen Shao, Hancheng Zhu, Rui Yao 0006, Bing Liu 0016
CGI (1)5
2025 ReDiffDet: Rotation-equivariant Diffusion Model for Oriented Object Detection
abstract
The diffusion model has been successfully applied to various detection tasks. However, it still faces several challenges when used for oriented object detection: objects that are arbitrarily rotated require the diffusion model to encode their orientation information; uncontrollable random boxes inaccurately locate objects with dense arrangements and extreme aspect ratios; oriented boxes result in the misalignment between them and image features. To overcome these limitations, we propose ReDiffDet, a framework that formulates oriented object detection as a rotation-equivariant denoising diffusion process. First, we represent an oriented box as a 2D Gaussian distribution, forming the basis of the denoising paradigm. The reverse process can be proven to be rotation-equivariant within this representation and model framework. Second, we design a conditional encoder with conditional boxes to prevent boxes from being randomly placed across the entire image. Third, we propose an aligned decoder for alignment between oriented boxes and image features. The extensive experiments demonstrate ReDiffDet achieves promising performance and significantly outperforms the diffusion-based baseline detector. Codes are available at https://github.com/wokaikaixinxin/ReDiffDet.
Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006
CVPR6
2025 Leveraging Large-Scale Pretrained Vision Foundation Models for Label-Efficient 3D Point Cloud Segmentation
Fayao Liu, Rui Yao 0006, Guosheng Lin
ICIG (2)3
2025 GSDet: Gaussian Splatting for Oriented Object Detection
abstract
Oriented object detection has advanced with the development of convolutional neural networks (CNNs) and transformers. However, modern detectors still rely on predefined object candidates, such as anchors in CNN-based methods or queries in transformer-based methods, which struggle to capture spatial information effectively. To address the limitations, we propose GSDet, a novel framework that formulates oriented object detection as Gaussian splatting. Specifically, our approach performs detection within a 3D feature space constructed from image features, where 3D Gaussians are employed to represent oriented objects. These 3D Gaussians are projected onto the image plane to form 2D Gaussians, which are then transformed into oriented boxes. Furthermore, we optimize the mean, anisotropic covariance, and confidence scores of these randomly initialized 3D Gaussians, using a decoder that incorporates 3D Gaussian sampling. Moreover, our method exhibits flexibility, enabling adaptive control and a dynamic number of Gaussians during inference. Experiments on 3 datasets indicate that GSDet achieves AP50 gains of 0.7% on DIOR-R, 0.3% on DOTA-v1.0, and 0.55% on DOTA-v1.5 when evaluated with adaptive control and outperforms mainstream detectors.
Zeyu Ding 0010, Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Hancheng Zhu, Rui Yao 0006
IJCAI6
2025 Beyond Individual and Point: Next POI Recommendation via Region-aware Dynamic Hypergraph with Dual-level Modeling
abstract
Next POI recommendation contributes to the prosperity of various intelligent location-based services. Existing studies focus on exploring sequential patterns and POI interactions using sequential and graph-based methods to enhance recommendation performance. However, they don't effectively exploit geographical information. In addition, methods that focus on modeling mobility patterns using individual limited data may suffer from data sparsity and the information cocoons problem. Moreover, most graph structures focus on adjacent nodes, failing to capture potential high-order associations among POIs. To address these challenges, we propose the Region-aware dynamic Hypergraph learning method with Dual-level interaction Modeling (ReHDM), which exploits users' dynamic mobility beyond individual and point. Specifically, ReHDM utilizes regional encoding to mine the potential spatial relationships among POIs with coarse-grained geographical information. By incorporating POI-level and trajectory-level associations within a hypergraph convolutional network, ReHDM comprehensively captures cross-user collaborative information. Furthermore, ReHDM captures not only dependencies among POIs within each trajectory for a single user, but also the high-order collaborative information across individual user trajectories and associated users' trajectories. Experimental results on three public datasets demonstrate the superiority of ReHDM to the state-of-the-art.
Zhuo Gu, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Wen-Liang Du 0002
IJCAI3
2025 Modality-Guided Dynamic Graph Fusion and Temporal Diffusion for Self-Supervised RGB-T Tracking
abstract
To reduce the reliance on large-scale annotations, self-supervised RGB-T tracking approaches have garnered significant attention. However, the omission of the object region by erroneous pseudo-label or the introduction of background noise affects the efficiency of modality fusion, while pseudo-label noise triggered by similar object noise can further affect the tracking performance. In this paper, we propose GDSTrack, a novel approach that introduces dynamic graph fusion and temporal diffusion to address the above challenges in self-supervised RGB-T tracking. GDSTrack dynamically fuses the modalities of neighboring frames, treats them as distractor noise, and leverages the denoising capability of a generative model. Specifically, by constructing an adjacency matrix via an Adjacency Matrix Generator (AMG), the proposed Modality-guided Dynamic Graph Fusion (MDGF) module uses a dynamic adjacency matrix to guide graph attention, focusing on and fusing the object’s coherent regions. Temporal Graph-Informed Diffusion (TGID) models MDGF features from neighboring frames as interference, and thus improving robustness against similar-object noise. Extensive experiments conducted on four public RGB-T tracking datasets demonstrate that GDSTrack outperforms the existing state-of-the-art methods. The source code is available at https://github.com/LiShenglana/GDSTrack.
Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Kunyang Sun, Bing Liu 0016, Zhiwen Shao, Jiaqi Zhao 0001
IJCAI2
2025 Counterfactual Knowledge Maintenance for Unsupervised Domain Adaptation
abstract
Traditional unsupervised domain adaptation (UDA) struggles to extract rich semantics due to backbone limitations. Recent large-scale pre-trained visual-language models (VLMs) have shown strong zero-shot learning capabilities in UDA tasks. However, directly using VLMs results in a mixture of semantic and domain-specific information, complicating knowledge transfer. Complex scenes with subtle semantic differences are prone to misclassification, which in turn can result in the loss of features that are crucial for distinguishing between classes. To address these challenges, we propose a novel counterfactual knowledge maintenance UDA framework. Specifically, we employ counterfactual disentanglement to separate the representation of semantic information from domain features, thereby reducing domain bias. Furthermore, to clarify ambiguous visual information specific to classes, we maintain the discriminative knowledge of both visual and textual information. This approach synergistically leverages multimodal information to preserve modality-specific distinguishable features. We conducted extensive experimental evaluations on several public datasets to demonstrate the effectiveness of our method. The source code is available at https://github.com/LiYaolab/CMKUDA
Yong Zhou 0003, Jiaqi Zhao 0001, Wen-Liang Du 0002, Rui Yao 0006, Bing Liu 0016
IJCAI5
2025 RQFormer: Rotated Query Transformer for end-to-end oriented object detection
Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
Expert Syst. Appl.6
2025 Facial Action Unit Detection by Adaptively Constraining Self-Attention and Causally Deconfounding Sample
Zhiwen Shao, Hancheng Zhu, Yong Zhou 0003, Xiang Xiang 0001, Bing Liu 0016, Rui Yao 0006, Lizhuang Ma
Int. J. Comput. Vis.6
2025 MoViM: A Hybrid CNN Vision Mamba Network for Lightweight Semantic Segmentation of Multimodal Remote Sensing Images
abstract
The “Others” category in multimodal remote sensing images is characterized by high intra-class variability. Therefore, existing lightweight semantic segmentation models struggle with this category due to limitations in capturing both local details and global dependencies efficiently. We propose MoViM, a lightweight model that integrates a hybrid Vision Mamba (ViM) and CNN backbone to capture global contextual information and local details effectively. In addition, the MoViM also features an Inverted Stem for efficient multimodal fusion, a Global Semantics Extraction (GSE) module for enhanced global feature representation, and a Global-Local Feature Fusion (GLF) module for context-aware feature integration. Extensive experiments on WHU-OPT-SAR and Potsdam datasets demonstrate that MoViM achieves state-of-the-art performance, particularly in the “Others” category, while maintaining low computational complexity. Our codes are available at https://github.com/WenliangDu/MoViM.
Wen-Liang Du 0002, Jiaqi Zhao 0001, Rui Yao 0006, Yong Zhou 0003
IEEE Geosci. Remote. Sens. Lett.4
2025 FA-MSVNet: multi-scale and multi-view feature aggregation methods for stereo 3D reconstruction
Yong Zhou 0003, Jiaqi Zhao 0001, Wen-Liang Du 0002, Rui Yao 0006
Multim. Tools Appl.5
2025 Grid-distance-based selection for fine-grained object detection in aerial images
Jiaqi Zhao 0001, Qingfeng Ou, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006
Pattern Recognit. Lett.5
2025 Progressively Generated Text-Assisted Image Aesthetic Quality Assessment
abstract
Image Aesthetic Quality Assessment (IAQA) aims to simulate user perceptions to judge the aesthetic quality of images. Due to the high subjectivity of users and the complexity of image aesthetics, modeling IAQA solely at the image level is a compromise. Consequently, existing methods mainly focus on multimodal-based models and achieve effective performance. These methods explore aesthetic comments on images to characterize users and serve as auxiliary text information for multimodal modeling. Unfortunately, this may suffer from two limitations. One limitation is that aesthetic comments are often unavailable for an unknown image in the test phase, and another limitation is that the semantic information of these comments may be uncertain and fuzzy. Therefore, this paper proposes a progressively generated text-assisted image aesthetic quality assessment method, aiming to address the lack of aesthetic comments and the fuzziness of aesthetic judgments in these comments. Specifically, we first adopt a Multimodal Large Language Model (MLLM) to generate aesthetic comments on images by simulating user perceptions and utilize the generated comments to characterize their aesthetic perception to assist in the pre-training of our multimodal-based IAQA model. Then, we design an attribute prediction module to determine the attribute levels of aesthetic judgments and utilize text template construction to further generate explicit descriptions of image aesthetics. Finally, we leverage the generated attribute descriptions to further assist in training our IAQA model. By progressively generating textual auxiliary descriptions of aesthetics for images, the proposed model can gradually determine the aesthetic quality of the images. Massive experimental results indicate that the proposed method outperforms existing mainstream methods on multiple IAQA datasets.
Hancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao 0006, Kunyang Sun, Leida Li
IEEE Trans. Fuzzy Syst.4
2025 Hyperspectral Object Tracking With Dual-Stream Prompt
abstract
Hyperspectral images, rich in spectral details, offeradvantages for object tracking across diverse scenarios. Current hyperspectral tracking often fine-tunes parameters using pretrained RGB trackers, but this manner is suboptimal due to redundancy in spectral bands and limited training data. Existing hyperspectral trackers also underuse temporal information. To address these issues, we propose a unified spectral-spatiotemporal multimodal dual-stream prompt hyperspectral object tracking, named HDSP. We design a density clustering-based band selection module (BSM) to preserve spectral prompt information efficiently. Using the generated bands and temporal data as multimodal prompts, a dual-stream visual prompter is proposed. Designed multimodal dual-stream visual prompter (MDVP) transforms the multimodal input into a single modality, enhancing the foundational modality’s representation capabilities for hyperspectral tracking. Experiments on hyperspectral videos (HSVs) tracking datasets demonstrate that the proposed tracker achieves state-of-the-art performance. The source code is available athttps://github.com/rayyao/HDSP.
Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao
IEEE Trans. Geosci. Remote. Sens.1
2025 DDCI: Unsupervised Domain Adaptation for Remote Sensing Images Based on Diffusion Causal Distillation
abstract
The distribution of remote sensing (RS) images can vary significantly due to seasonal changes and lighting conditions, making it difficult for deep learning models to generalize effectively across different RS datasets. This variation leads to a domain gap that hampers model performance when applied to new, unseen data. To tackle this challenge, we introduce DDCI, a novel unsupervised domain adaptation (UDA) framework designed to bridge the domain gap in RS image perception. Our framework consists of two key components, i.e., the adaptation diffusion distillation (ADD) module and the consistent causal intervention (CCI) module. The ADD module addresses the domain gap by aligning the source and target domains. It enhances the representation of the target domain by distilling semantic knowledge from the teacher model of the source domain. This process allows the target domain to benefit from the rich features of the source domain, leading to improved model generalization. The CCI module focuses on removing spurious correlations between domain-agnostic knowledge and domain-specific knowledge. By carefully considering the distinct characteristics of the target domain while preserving the specificity of the source domain, the CCI module ensures that only relevant, causal information is transferred between domains. This prevents overfitting to irrelevant domain-specific features and enhances model robustness. We demonstrate the effectiveness of the DDCI framework on RS scene classification tasks, utilizing four widely recognized RS datasets. Our results show significant performance improvements, underscoring the potential of this approach to boost the adaptability of deep learning models across diverse RS image datasets.
Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
IEEE Trans. Geosci. Remote. Sens.6
2025 GLFRNet: Global-Local Feature Refusion Network for Remote Sensing Image Instance Segmentation
abstract
Instance segmentation is a significant way for remote sensing image (RSI) interpretation. The large number, sharp variation of sizes, and complex background of objects raise higher demands for instance segmentation models. The synergistic usage of global and local features has drawn great attention due to its superior performance but has not been fully explored in mainstream instance segmentation methods. In this work, a global-local feature refusion network (GLFRNet) with two fusion procedures is proposed to fully utilize coarse-grained and fine-grained features for RSI instance segmentation. In this model, the backbone integrates both convolutional neural network (CNN)-based and VMamba-based branches to extract local and global features, respectively. Three novel models are proposed to leverage the features adaptively, i.e., the cross-dim feature fusion (CDFF) module, the semantic complementary feature fusion (SCFF) module, and the guided feature refusion module (GFRM). The CDFF module is designed to aggregate features flexibly by fusing features from two backbones with different attention modules in the first fusion procedure. The GFRM and SCFF module are proposed in the refusion procedure to generate accurate segmentation results. Inspired by agent attention, the GFRM dynamically assembles detailed features for mask generation by refusing local and global features with the guidance of fusion results from CDFF. The SCFF module complements the significant features by enhancing and integrating global, local, and detailed features, and finally generates masks of instances. Extensive experiments demonstrate that GLFRNet outperforms the second-best model by 1.9, 1.3, and 0.3 in mask average precisions (APs) on NWPU VHR-10, WHU Building, and iSAID datasets.
Jiaqi Zhao 0001, Yari Wang, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
IEEE Trans. Geosci. Remote. Sens.5
2025 ST-Mamba: Spatio-Temporal Synergistic Model for Remote Sensing Change Detection
abstract
The advancement of remote sensing and deep learning has spurred interest in high-resolution image change detection (CD). However, pseudo-changes in multi-temporal images, due to complex scenes and variable imaging conditions, often lead to significant misdetection in current methods. To address this problem, we propose a new CD framework: Spatio-Temporal Mamba (ST-Mamba), which consists of three key components. Firstly, a Mamba-based Feature Extraction Module (MFEM) is designed as the encoder to extract essential features from multi-temporal images by leveraging Mamba’s capability to capture inherent information in long data sequences. Secondly, a Spatio-Temporal Synergy Module (STSM) is developed to unify the background features of multi-temporal feature maps into a common domain by employing the state-space model for spatio-temporal modeling. Finally, a Spatio-Temporal Fusion Module (STFM) is created to guide the fusion of image features at different scales and across channels by utilizing a feature map of the unified background features. Experimental results on five widely used change detection datasets show significant improvements over current state-of-the-art methods.
Jiaqi Zhao 0001, Jianpeng Xie 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
IEEE Trans. Geosci. Remote. Sens.5
2025 Adversarial Geometric Attacks for 3D Point Cloud Object Tracking
abstract
3D point cloud object tracking (3D PCOT) plays a vital role in applications such as autonomous driving and robotics. Adversarial attacks offer a promising approach to enhance the robustness and security of tracking models. However, existing adversarial attack methods for 3D PCOT seldom leverage the geometric structure of point clouds and often overlook the transferability of attack strategies. To address these limitations, this paper proposes an adversarial geometric attack method tailored for 3D PCOT, which includes a point perturbation attack module (non-isometric transformation) and a rotation attack module (isometric transformation). First, we introduce a curvature-aware point perturbation attack module that enhances local transformations by applying normal perturbations to critical points identified through geometric features such as curvature and entropy. Second, we design a Thompson sampling-based rotation attack module that applies subtle global rotations to the point cloud, introducing tracking errors while maintaining imperceptibility. Additionally, we design a fused loss function to iteratively optimize the point cloud within the search region, generating adversarially perturbed samples. The proposed method is evaluated on multiple 3D PCOT models and validated through black-box tracking experiments on benchmarks. For P2B, white-box attacks on KITTI reduce the success rate from 53.3% to 29.6% and precision from 68.4% to 37.1%. On NuScenes, the success rate drops from 39.0% to 27.6%, and precision from 39.9 to 26.8%. Black-box attacks show a transferability, with BAT showing a maximum 47.0% drop in success rate and 47.2% in precision on KITTI, and a maximum 22.5% and 27.0% on NuScenes.
Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Bing Liu 0016, Abdulmotaleb El Saddik
IEEE Trans. Multim.1
2025 Similarity Regulation and Calibration Alignment for Weakly Supervised Text-Based Person Re-Identification
abstract
Traditional text-based person re-identification relies on identity labels. However, it is impossible to annotate large datasets, since identity annotation is expensive and time-consuming. Weakly supervised text-based person re-identification, where only text–image pairs are available without annotation of identities, is very practical in real life. While dealing with the weakly supervised person re-identification, two issues should be strengthed, i.e., alignment caused by different modal, and cross-modal matching ambiguity caused by the lack of identity labels. In this article, we propose a similarity regulation and calibration alignment (SRCA) framework, which consists of two unimodal encoders for images and text, respectively, and a multi-modal encoder for the masked language modeling task. First, a similarity regulation (SR) strategy is proposed to relax the strict one-to-one constraints for the local similarities between different pairs by introducing a novel soft objective. The soft objective can adjust hard objectives to achieve soft cross-modal alignment by establishing a many-to-many relationship between two modalities. Second, the calibration alignment (CA) module is proposed to improve intra-class compactness by modeling pseudo-label assignment as optimal transport. The ambiguity of cross-modal matching can be reduced by aligning features and pseudo-labels of different modalities and gradually calibrating the distribution of pseudo-labels. Experimental results show that our method has achieved obvious advantages compared with existing methods and also demonstrated competitive performance compared with fully supervised methods.
Ao Fu, Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Historical Object-Aware Prompt Learning for Universal Hyperspectral Object Tracking
abstract
Hyperspectral Object Tracking (HOT), utilizing rich spectral information from hyperspectral video (HSV), holds significant importance for object tracking. We identify that a major obstacle in improving HOT performance lies in effectively leveraging spectral and historical information. Furthermore, due to the mismatch in band dimensions between hyperspectral and RGB images, state-of-the-art RGB-based trackers struggle to adapt to unified HOT tasks. To address this, we propose a Historical Object-Aware Prompt Learning (HOPL) method for universal hyperspectral object tracking. Initially, we transform hyperspectral image ( \( N \) bands) into multiple sets of three bands with different combinations and feed them into a backbone network to generate base features. Subsequently, we introduce a historical object-aware prompter, where historical object-aware images are input to generate prompt features that enhance the representation of object information when combined with base features. Additionally, we design a band information fusion module to integrate the multiple sets of base features. By introducing historical object-aware prompts, HOPL significantly enhances tracking performance without retraining the backbone network. Experimental results on the HOT2023 dataset (comprising HSV with 25-band, 16-band, and 15-band wavelength ranges) and HOT2022 dataset validate the superiority of HOPL over state-of-the-art methods. The source code is available at https://github.com/rayyao/HOPL .
Rui Yao 0006, Yong Zhou 0003, Fuyuan Hu, Jiaqi Zhao 0001, Zhiwen Shao
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Image Cropping with Content and Composition Attribute-aware Global Relation Reasoning
abstract
Image cropping aims to find visually pleasing content in an image, which will enhance its aesthetic quality. Existing image cropping approaches mainly emphasize the geometric properties of images, such as composition and layout, neglecting the rich aesthetic information available from the physical attributes (e.g., content and themes), and background information beyond the foreground in images. Consequently, this article proposes an image cropping method based on the content and composition attribute-aware global relation reasoning, which aims at guiding the generation of cropped sub-images by exploring critical attributes based on content and composition as well as global object correlations that affect aesthetics in images. Particularly, to comprehensively introduce aesthetic information into image cropping, we capture feature representations reinforced by content and composition attributes simultaneously. The feature representations can strengthen the visual aesthetics of cropped sub-images. To make the cropped sub-images amply contain more global information, we introduce a global relation reasoning branch in the proposed cropping module, which can fully exploit the dependency relationship between the foreground and background in images. Extensive experiments on image cropping benchmarks demonstrate that our approach is superior to state-of-the-art image cropping methods.
Hancheng Zhu, Yong Zhou 0003, Rui Yao 0006, Zhiwen Shao, Jiaqi Zhao 0001, Leida Li
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Attribute-Driven Multimodal Hierarchical Prompts for Image Aesthetic Quality Assessment
abstract
Image Aesthetic Quality Assessment (IAQA) aims to simulate users' visual perception to judge the aesthetic quality of images. In social media, users' aesthetic experiences are often reflected in their textual comments regarding the aesthetic attributes of images. To fully explore the attribute information perceived by users for evaluating image aesthetic quality, this paper proposes an image aesthetic quality assessment method based on attribute-driven multimodal hierarchical prompts. Unlike existing IAQA methods that utilize multimodal pre-training or straightforward prompts for model learning, the proposed method leverages attribute comments and quality-level text templates to hierarchically learn the aesthetic attributes and quality of images. Specifically, we first leverage users' aesthetic attribute comments to perform prompt learning on images. The learned attribute-driven multimodal features can comprehensively capture the semantic information of image aesthetic attributes perceived by users. Then, we construct text templates for different aesthetic quality levels to further facilitate prompt learning through semantic information related to the aesthetic quality of images. The proposed method can explicitly simulate users' aesthetic judgment of images to obtain more precise aesthetic quality. Experimental results demonstrate that the proposed IAQA method based on hierarchical prompts outperforms existing methods significantly on multiple IAQA databases. Our source code is public at https://github.com/GitHub-Ju/AMHP.
Hancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Leida Li
ACM Multimedia4
2024 Remote sensing image semantic segmentation via class-guided structural interaction and boundary perception
abstract
Existing remote sensing semantic segmentation methods generally ignore the structural information of objects that is vital in the human visual recognition system. The absence of overall structural information often results in weak perceptions of subtle textures and fragmented predictions, especially for complex and variable ground object scenarios. Besides, they still suffer from the semantic ambiguity caused by the unclear object boundary features in remote sensing images. In this paper, we propose a novel remote sensing semantic segmentation framework, called CSBNet, which aims to enhance the capacity of class-guided structural interaction and boundary perception simultaneously. It consists of a class-guided structure interaction module (CSIM), a Transformer-based context aggregation module (TCAM) and a class-guided boundary supervision module (CBSM). The CSIM has the ability to progressively extract the class-specific structural features, i.e. , refining the structural information of each class by iteratively exchanging information between initial coarse class tokens and contexts. Meanwhile, the TCAM is constructed to provide CSIM with more discriminative multi-scale contexts without losing spatial features. In particular, the CBSM plays an auxiliary role, which applies the boundary information obtained from the class tokens to supervise the segmentation of boundary regions. When tested on the ISPRS dataset, LoveDA dataset, UAVid dataset, our method significantly outperforms the state-of-the-art remote sensing semantic segmentation approaches.
Xin He 0024, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006
Expert Syst. Appl.5
2024 Fine-grained semantic oriented embedding set alignment for text-based person search
Jiaqi Zhao 0001, Ao Fu, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006
Image Vis. Comput.5
2024 A Mamba-Diffusion Framework for Multimodal Remote Sensing Image Semantic Segmentation
abstract
Recent advances in deep learning have made significant progress in multimodal remote sensing semantic segmentation. However, current methods face challenges in maintaining geometric consistency, particularly when dealing with large objects, resulting in fragmented segmentation masks. We propose a Mamba-diffusion framework to preserve geometric consistency in segmentation masks. This framework preserves geometric consistency by introducing a generative diffusion-based semantic segmentation pipeline and developing a Mamba-based multimodal fusion model. The fusion model fuses the multimodal images in multiple scales and scanning mechanisms by a double cross-fusion (DCF) module. Then, the cross-modal information is further integrated by a dual-splitting structured state-space (DS-S4) model. Finally, the diffusion-based segmentation pipeline predicts semantic masks by progressively refining random Gaussian noise, guided by fused multimodal features. Our experimental results, verified on WHU-OPT-SAR and Hunan datasets, demonstrate that the proposed framework surpasses state-of-the-art (SOTA) methods by a considerable margin. Our codes are available athttps://github.com/WenliangDu/MambaDiffusion.
Wen-Liang Du 0002, Yang Gu 0005, Jiaqi Zhao 0001, Hancheng Zhu, Rui Yao 0006, Yong Zhou 0003
IEEE Geosci. Remote. Sens. Lett.5
2024 Filter pruning based on evolutionary algorithms for person re-identification
Jiaqi Zhao 0001, Ying Chen 0005, Yufeng Zhong 0001, Yong Zhou 0003, Rui Yao 0006, Lixu Zhang, Shixiong Xia
Multim. Tools Appl.5
2024 Multi-level self attention for unsupervised learning person re-identification
Jiaqi Zhao 0001, Yong Zhou 0003, Fayao Liu, Rui Yao 0006, Hancheng Zhu, Abdulmotaleb El Saddik
Multim. Tools Appl.5
2024 Efficient convolutional neural networks and network compression methods for object detection: a survey
Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Bing Liu 0016
Multim. Tools Appl.4
2024 CT-Net: Arbitrary-Shaped Text Detection via Contour Transformer
abstract
Contour based scene text detection methods have rapidly developed recently, but still suffer from inaccurate front-end contour initialization, multi-stage error accumulation, or deficient local information aggregation. To tackle these limitations, we propose a novel arbitrary-shaped scene text detection framework named CT-Net by progressive contour regression with contour transformers. Specifically, we first employ a contour initialization module that generates coarse text contours without any post-processing. Then, we adopt contour refinement modules to adaptively refine text contours in an iterative manner, which are beneficial for context information capturing and progressive global contour deformation. Besides, we propose an adaptive training strategy to enable the contour transformers to learn more potential deformation paths, and introduce a re-score mechanism that can effectively suppress false positives. Extensive experiments are conducted on four challenging datasets, which demonstrate the accuracy and efficiency of our CT-Net over state-of-the-art methods. Particularly, CT-Net achieves F-measure of 86.1 at 11.2 frames per second (FPS) and F-measure of 87.8 at 10.1 FPS for CTW1500 and Total-Text datasets, respectively.
Zhiwen Shao, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Rui Yao 0006
IEEE Trans. Circuits Syst. Video Technol.7
2024 Information Gap Narrowing for Point Cloud Few-Shot Segmentation
abstract
Point-by-point labeling of point clouds is a very costly task. Previous meta-learning-based few-shot methods predict categories by calculating the distance between unlabeled data (query set) and the prototype calculated by a few of data with the label (support set), which can reduce the dependence of point cloud segmentation algorithms on large amounts of labeled data. But it ignores the category information gap caused by object diversity between the two types of data and forcing information transfer is ineffective. To address this issue, we propose a co-occurrent object mining module for mining co-occurring object information from support and query sets. Specifically, the capture of co-occurrent information is used to activate the feature that co-occurs between the support and query set in the high-dimensional feature space so that the prototype generated by computing the mean of support features is more similar to the query set. By reducing the object diversity within the same category, the information gap problem is gradually improved. In addition, we propose a point-attention module to refine the support set features before mining co-occurrent features. It can be widely embedded in the point cloud backbone network. The experimental results on two semantic segmentation datasets demonstrate that our method obtains an average 19.43% lead over the state-of-the-art methods in 4 different few-shot tasks, while inference is around 45 times faster.
Guanyu Zhu, Yong Zhou 0003, Rui Yao 0006, Hancheng Zhu
IEEE Trans. Circuits Syst. Video Technol.3
2024 Dual-Stream Edge-Target Learning Network for Infrared Small Target Detection
abstract
Infrared small target detection (IRSTD) is crucial in both military and civilian applications. However, challenges such as low contrast, low signal-to-noise ratio (SNR), and lack of shape and texture information limit the effectiveness of existing methods in capturing edge details and representing target areas. To address these issues, we propose the dual-stream edge-target learning network (DETL-Net) for IRSTD. This network enhances feature cross-fusion by learning edge details and target regions through a dual-stream framework, significantly improving detection performance. Specifically, we extract multilevel features of the image based on the encoder-decoder structure of U-Net and then reconstruct the feature map. In the decoder, we propose the dual-guided cross-fusion module (DGCFM) to capture edge details of small targets and global contextual features of the target region, achieving complementary advantages. The multiscale context fusion module (MCFM) within DGCFM uses central difference convolution to enhance local contrast and extract rich contextual details, thereby retaining edge information and enhancing overall target representation. In addition, we introduce the cross-dimension interactive aggregation attention module (CIAAM), which dynamically adjusts feature fusion weights across layers to effectively suppress noise and enhance the discrimination of small targets. These modules are sequentially interconnected to progressively refine edge details, and the acquired target features are subsequently utilized for predicting the final target mask via the segmentation head. Experiments on the NUAA-SIRST and IRSTD-1k datasets demonstrate that DETL-Net outperforms state-of-the-art (SOTA) methods. The source code is available athttps://github.com/rayyao/DETL-Net.
Rui Yao 0006, Yong Zhou 0003, Jinqiu Sun, Zihang Yin, Jiaqi Zhao 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 OrientedFormer: An End-to-End Transformer-Based Oriented Object Detector in Remote Sensing Images
abstract
Oriented object detection in remote sensing images is a challenging task due to objects being distributed in multiorientation. Recently, end-to-end transformer-based methods have achieved success by eliminating the need for post-processing operators compared to traditional convolutional neural network (CNN)-based methods. However, directly extending transformers to oriented object detection presents three main issues: 1) objects rotate arbitrarily, necessitating the encoding of angles along with position and size; 2) the geometric relations of oriented objects are lacking in self-attention, due to the absence of interaction between content and positional queries; and 3) oriented objects cause misalignment, mainly between values and positional queries in cross-attention, making accurate classification and localization difficult. In this article, we propose an end-to-end transformer-based oriented object detector, consisting of three dedicated modules to address these issues. First, Gaussian positional encoding (PE) is proposed to encode the angle, position, and size of oriented boxes using Gaussian distributions. Second, Wasserstein self-attention is proposed to introduce geometric relations and facilitate interaction between content and positional queries by utilizing Gaussian Wasserstein distance scores. Third, oriented cross-attention is proposed to align values and positional queries by rotating sampling points around the positional query according to their angles. Experiments on six datasets DIOR-R, a series of DOTA, HRSC2016, and ICDAR2015 show the effectiveness of our approach. Compared with previous end-to-end detectors, the OrientedFormer gains 1.16 and 1.21 AP50 on DIOR-R and DOTA-v1.0, respectively, while reducing training epochs from$3\times $to$1\times $. The code is available athttps://github.com/wokaikaixinxin/OrientedFormer.
Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
IEEE Trans. Geosci. Remote. Sens.6
2024 Black-box Attack against Self-supervised Video Object Segmentation Models with Contrastive Loss
abstract
Deep learning models have been proven to be susceptible to malicious adversarial attacks, which manipulate input images to deceive the model into making erroneous decisions. Consequently, the threat posed to these models serves as a poignant reminder of the necessity to focus on the model security of object segmentation algorithms based on deep learning. However, the current landscape of research on adversarial attacks primarily centers around static images, resulting in a dearth of studies on adversarial attacks targeting Video Object Segmentation (VOS) models. Given that a majority of self-supervised VOS models rely on affinity matrices to learn feature representations of video sequences and achieve robust pixel correspondence, our investigation has delved into the impact of adversarial attacks on self-supervised VOS models. In response, we propose an innovative black-box attack method incorporating contrastive loss. This method induces segmentation errors in the model through perturbations in the feature space and the application of a pixel-level loss function. Diverging from conventional gradient-based attack techniques, we adopt an iterative black-box attack strategy that incorporates contrastive loss across the current frame, any two consecutive frames, and multiple frames. Through extensive experimentation conducted on the DAVIS 2016 and DAVIS 2017 datasets using three self-supervised VOS models and one unsupervised VOS model, we unequivocally demonstrate the potent attack efficiency of the black-box approach. Remarkably, theJ&Fmetric value experiences a significant decline of up to 50.08% post-attack.
Ying Chen 0005, Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Bing Liu 0016, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Motion-Aware Self-Supervised RGBT Tracking with Multi-Modality Hierarchical Transformers
abstract
Supervised RGBT (SRGBT) tracking tasks need both expensive and time-consuming annotations. Therefore, the implementation of Self-Supervised RGBT (SSRGBT) tracking methods has become increasingly important. Straightforward SSRGBT tracking methods use pseudo-labels for tracking, but inaccurate pseudo-labels can lead to object drift, which severely affects tracking performance. This article proposes a self-supervised RGBT object tracking method (S2OTFormer) to bridge the gap between tracking methods supervised under pseudo-labels and ground truth labels. Firstly, to provide more robust appearance features for motion cues, we introduce a multi-modality hierarchical transformer (MHT) module for feature fusion. This module allocates weights to both modalities and strengthens the expressive capability of the MHT module through multiple nonlinear layers to fully utilize the complementary information of the two modalities. Secondly, in order to solve the problems of motion blur caused by camera motion and inaccurate appearance information caused by pseudo-labels, we introduce a motion-aware mechanism (MAM). The MAM extracts the average motion vectors from the previous multi-frame search frame features and constructs the consistency loss with the motion vectors of the current search frame features. The motion vectors of inter-frame objects are obtained by reusing the inter-frame attention map to predict coordinate positions. Finally, to further reduce the effect of inaccurate pseudo-labels, we propose an Attention-Based Multi-Scale Enhancement Module. By introducing cross-attention to achieve more precise and accurate object tracking, this module overcomes the receptive field limitations of traditional CNN tracking heads. We demonstrate the effectiveness of S2OTFormer on four large-scale public datasets through extensive comparisons as well as numerous ablation experiments. The source code is available at https://github.com/LiShenglana/S2OTFormer .
Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Diverse Image Captioning via Conditional Variational Autoencoder and Dual Contrastive Learning
abstract
Diverse image captioning has achieved substantial progress in recent years. However, the discriminability of generative models and the limitation of cross entropy loss are generally overlooked in the traditional diverse image captioning models, which seriously hurts both the diversity and accuracy of image captioning. In this article, aiming to improve diversity and accuracy simultaneously, we propose a novel Conditional Variational Autoencoder (DCL-CVAE) framework for diverse image captioning by seamlessly integrating sequential variational autoencoder with contrastive learning. In the encoding stage, we first build conditional variational autoencoders to separately learn the sequential latent spaces for a pair of captions. Then, we introduce contrastive learning in the sequential latent spaces to enhance the discriminability of latent representations for both image-caption pairs and mismatched pairs. In the decoding stage, we leverage the captions sampled from the pre-trained Long Short-Term Memory (LSTM), LSTM decoder as the negative examples and perform contrastive learning with the greedily sampled positive examples, which can restrain the generation of common words and phrases induced by the cross entropy loss. By virtue of dual constrastive learning, DCL-CVAE is capable of encouraging the discriminability and facilitating the diversity, while promoting the accuracy of the generated captions. Extensive experiments are conducted on the challenging MSCOCO dataset, showing that our proposed methods can achieve a better balance between accuracy and diversity compared to the state-of-the-art diverse image captioning models.
Bing Liu 0016, Yong Zhou 0003, Rui Yao 0006, Zhiwen Shao
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Multi-Modal LiDAR Point Cloud Semantic Segmentation with Salience Refinement and Boundary Perception
abstract
Point cloud segmentation is essential for scene understanding, which provides advanced information for many applications, such as autonomous driving, robots, and virtual reality. To improve the accuracy and robustness of point cloud segmentation, many researchers have attempted to fuze camera images to complement the color and texture information. The common fusion strategy is the combination of convolutional operations with concatenation, element-wise addition or element-wise multiplication. However, conventional convolutional operators tend to confine the fusion of modal features within their receptive fields, which can be incomplete and limited. In addition, the inability of encoder–decoder segmentation networks to explicitly perceive segmentation boundary information results in semantic ambiguity and classification errors at object edges. These errors are further amplified in point cloud segmentation tasks, significantly affecting the accuracy of point cloud segmentation. To address the above issues, we propose a novel self-attention multi-modal fusion semantic segmentation network for point cloud semantic segmentation. Firstly, to effectively fuze different modal features, we propose a self-cross fusion module (SCF), which models long-range modality dependencies and transfers complementary image information to the point cloud to fully leverage the modality-specific advantages. Secondly, we design the salience refinement module (SR), which calculates the importance of channels in the feature maps and global descriptors to enhance the representation capability of salient modal features. Finally, we propose the local-aware anisotropy loss measure the element-level importance in the data and explicitly provide boundary information for the model, which alleviates the inherent semantic ambiguity problem in segmentation networks. Extensive experiments on two benchmark datasets demonstrate that our proposed method surpasses current state-of-the-art methods.
Yong Zhou 0003, Zeming Xie, Jiaqi Zhao 0001, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Identity-invariant representation and transformer-style relation for micro-expression recognition
Zhiwen Shao, Feiran Li, Yong Zhou 0003, Hancheng Zhu, Rui Yao 0006
Appl. Intell.6
2023 Info-FPN: An Informative Feature Pyramid Network for object detection in remote sensing images
Silin Chen, Jiaqi Zhao 0001, Yong Zhou 0003, Hanzheng Wang, Rui Yao 0006, Lixu Zhang, Yong Xue
Expert Syst. Appl.5
2023 Context-aware and part alignment for visible-infrared person re-identification
Jiaqi Zhao 0001, Hanzheng Wang, Yong Zhou 0003, Rui Yao 0006, Lixu Zhang, Abdulmotaleb El Saddik
Image Vis. Comput.4
2023 Visible and Infrared Object Tracking via Convolution-Transformer Network With Joint Multimodal Feature Learning
abstract
The existing Transformer-based RGBT tracker mainly focus on the enhancement of features extracted by Convolutional Neural Network (CNN). The potential of Transformer in representation learning remains under-explored. In this paper, we propose a Convolution-Transformer network with joint multimodal feature learning, in which both representation learning and feature fusion leverage Transformer. Specifically, we use the multi-branch Convolution-Transformer feature extraction network to process the extraction task of local modality-independent features and global modality-shared features respectively. Several simplified Transformer encoder layers form the Transformer backbone network, which is more suitable for the real-time object tracking. Besides, we found that inter-modality correlation is an important factor for modality interactions and mutual exploitation. Therefore, we propose a Joint Multimodal Feature Learning (JMFL) module, which uses cross-attention to capture the dependencies of cross-modal and enhance multimodal fusion by bidirectional guidance of multimodal information. The proposed method is fully experimented on two large benchmark datasets and compared with some current well-performing methods. The experimental results show that the proposed method performs well in terms of tracking accuracy and speed.
Jiazhu Qiu, Rui Yao 0006, Yong Zhou 0003, Peng Wang 0015, Yanning Zhang 0001, Hancheng Zhu
IEEE Geosci. Remote. Sens. Lett.2
2023 Adversarial learning-based skeleton synthesis with spatial-channel attention for robust gait recognition
Ying Chen 0005, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Dongjun Zhu
Multim. Tools Appl.6
2023 Unsupervised RGB-T object tracking with attentional multi-modal feature fusion
Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Jiaqi Zhao 0001, Zhiwen Shao
Multim. Tools Appl.2
2023 Semi-supervised transformable architecture search for feature distillation
Man Zhang 0006, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006, Zhiwen Shao, Hancheng Zhu
Pattern Anal. Appl.5
2023 Facial Action Unit Detection via Adaptive Attention and Relation
abstract
Facial action unit (AU) detection is challenging due to the difficulty in capturing correlated information from subtle and dynamic AUs. Existing methods often resort to the localization of correlated regions of AUs, in which predefining local AU attentions by correlated facial landmarks often discards essential parts, or learning global attention maps often contains irrelevant areas. Furthermore, existing relational reasoning methods often employ common patterns for all AUs while ignoring the specific way of each AU. To tackle these limitations, we propose a novel adaptive attention and relation (AAR) framework for facial AU detection. Specifically, we propose an adaptive attention regression network to regress the global attention map of each AU under the constraint of attention predefinition and the guidance of AU detection, which is beneficial for capturing both specified dependencies by landmarks in strongly correlated regions and facial globally distributed dependencies in weakly correlated regions. Moreover, considering the diversity and dynamics of AUs, we propose an adaptive spatio-temporal graph convolutional network to simultaneously reason the independent pattern of each AU, the inter-dependencies among AUs, as well as the temporal dependencies. Extensive experiments show that our approach (i) achieves competitive performance on challenging benchmarks including BP4D, DISFA, and GFT in constrained scenarios and Aff-Wild2 in unconstrained scenarios, and (ii) can precisely learn the regional correlation distribution of each AU.
Zhiwen Shao, Yong Zhou 0003, Jianfei Cai 0001, Hancheng Zhu, Rui Yao 0006
IEEE Trans. Image Process.5
2023 Attention-guided Adversarial Attack for Video Object Segmentation
abstract
Video Object Segmentation (VOS) methods have made many breakthroughs with the help of the continuous development and advancement of deep learning. However, the deep learning model is vulnerable to malicious adversarial attacks, which mislead the model to make wrong decisions by adding adversarial perturbation that humans cannot perceive to the input image. Threats to deep learning models remind us that video object segmentation methods are also vulnerable to attacks, thereby threatening their security. Therefore, we study adversarial attacks on the VOS task to better identify the vulnerabilities of the VOS method, which in turn provides an opportunity to improve its robustness. In this paper, we propose an attention-guided adversarial attack method, which uses spatial attention blocks to capture features with global dependencies to construct correlations between consecutive video frames, and performs multipath aggregation to effectively integrate spatial-temporal perturbation, thereby guiding the deconvolution network to generate adversarial examples with strong attack capability. Specifically, the class loss function is designed to enable the deconvolution network to better activate noise in other regions and suppress the activation related to the object class based on the enhanced feature map of the object class. At the same time, attentional feature loss is designed to enhance the transferability against attack. The experimental results on the DAVIS dataset show that the proposed attention-guided adversarial attack method can significantly reduce the segmentation accuracy of OSVOS, and the J & F mean on DAVIS 2016 can reach 73.6% drop rate. The generated adversarial examples are also highly transferable to other video object segmentation models.
Rui Yao 0006, Ying Chen 0005, Yong Zhou 0003, Fuyuan Hu, Jiaqi Zhao 0001, Bing Liu 0016, Zhiwen Shao
ACM Trans. Intell. Syst. Technol.1
2023 TextDCT: Arbitrary-Shaped Text Detection via Discrete Cosine Transform Mask
abstract
Arbitrary-shaped scene text detection is a challenging task due to the variety of text changes in font, size, color, and orientation. Most existing regression based methods resort to regress the masks or contour points of text regions to model the text instances. However, regressing the complete masks requires high training complexity, and contour points are not sufficient to capture the details of highly curved texts. To tackle the above limitations, we propose a novel light-weight anchor-free text detection framework called TextDCT, which adopts the discrete cosine transform (DCT) to encode the text masks as compact vectors. Further, considering the imbalanced number of training samples among pyramid layers, we only employ a single-level head for top-down prediction. To model the multi-scale texts in a single-level head, we introduce a novel positive sampling strategy by treating the shrunk text region as positive samples, and design a feature awareness module (FAM) for spatial-awareness and scale-awareness by fusing rich contextual information and focusing on more significant features. Moreover, we propose a segmented non-maximum suppression (S-NMS) method that can filter low-quality mask regressions. Extensive experiments are conducted on four challenging datasets, which demonstrate our TextDCT obtains competitive performance on both accuracy and efficiency. Specifically, TextDCT achieves F-measure of 85.1 at 17.2 frames per second (FPS) and F-measure of 84.9 at 15.1 FPS for CTW1500 and Total-Text datasets, respectively.
Zhiwen Shao, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Rui Yao 0006
IEEE Trans. Multim.7
2023 Weakly Supervised Few-Shot Semantic Segmentation via Pseudo Mask Enhancement and Meta Learning
abstract
Few shot semantic segmentation has been proposed to enhance the generalization ability of traditional models with limited data. Previous works mainly focus on the supervised tasks, while limited amount of work is explored for the weakly supervised tasks. Weakly supervised semantic segmentation has become an active research area because weakly supervised labels effectively reduce the annotation cost of visual tasks. To this end, we propose a weakly supervised few-shot semantic segmentation model based on the meta learning framework, which utilizes prior knowledge and adjusts itself according to new tasks. Thereupon then, the proposed network is capable of both high efficiency and generalization ability to new tasks. In the pseudo mask generation stage, we develop a WRCAM method with the channel-spatial attention mechanism to refine the coverage size of targets in pseudo masks. In the few-shot semantic segmentation stage, the optimization based meta learning method is used to realize few-shot semantic segmentation by virtue of the refined pseudo masks. The experimental results show that the proposed method not only significantly outperforms weakly supervised SOTA methods, but also could be comparative to some supervised SOTA methods.
Man Zhang 0006, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006, Zhiwen Shao, Hancheng Zhu
IEEE Trans. Multim.5
2023 Spatial-Channel Enhanced Transformer for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) is a challenging task in computer vision, aiming at matching people across images from visible and infrared modalities. The widely used VI-ReID framework consists of a convolution neural backbone network that extracts the visual features, and a feature embedding network to project heterogeneous features to the same feature space. However, many studies based on the existing pre-trained models neglect potential correlations between different locations and channels within a single sample during the feature extraction. Inspired by the success of the Transformer in computer vision, we extend it to enhance feature representation for VI-ReID. In this paper, we propose a discriminative feature learning network based on a visual Transformer (DFLN-ViT) for VI-ReID. Firstly, to capture long-term dependencies between different locations, we propose a spatial feature awareness module (SAM), which utilizes a single-layer Transformer with a novel patch-embedding strategy to encode location information. Secondly, to refine the representation at each channel, we design a channel feature enhancement module (CEM). The CEM treats the features of each channel as a sequence of Transformer inputs, taking advantage of the Transformer's ability to model long-term dependencies. Finally, we propose a Triplet-aided Hetero-Center (THC) loss to learn more discriminative feature representation by balancing the cross-modality distance and intra-modality distance of the center. The experimental results on two datasets show that our method can significantly improve the VI-ReID performance, outperforming most state-of-the-art methods.
Jiaqi Zhao 0001, Hanzheng Wang, Yong Zhou 0003, Rui Yao 0006, Silin Chen, Abdulmotaleb El Saddik
IEEE Trans. Multim.4
2023 Cross-Class Bias Rectification for Point Cloud Few-Shot Segmentation
abstract
The point cloud is a densely distributed 3D (three-dimensional) data, and annotating the point cloud is a time-consuming and labor-intensive work. The existing semantics segmentation work adopts few-shot learning to reduce the dependence on labeling samples while improving the generalization of the model to new categories. Since point clouds are 3D structures with rich geometric features, even objects of the same category have feature differences that cannot be ignored. Therefore, a few samples (support set) used to train the model do not cover all the features of this category. There is a distribution difference between the support samples and the samples used to verify the model performance (query set). In this paper, we propose an efficient point cloud few-shot segmentation method based on prototypes for bias rectification. A prototype is a vector representation of a category in the metric space. To make the prototype representation of the support set closer to the query set features, we define a feature bias term and reduce the distribution distance between the two sets by fusing the support set features and the bias term. On this basis, we design a feature cross-reference module. By mining the co-occurring features of the support and query sets, it can generate a more representative prototype which captures the overall features of the point cloud. Extensive experiments on two challenging datasets demonstrate that our method outperforms the state-of-the-art method by an average of 3.31$\%$in several N-way K-shot tasks, and achieves approximately 200 times faster reasoning speed. Our code is available athttps://github.com/964918993/2CBR.
Guanyu Zhu, Yong Zhou 0003, Rui Yao 0006, Hancheng Zhu
IEEE Trans. Multim.3
2023 Cyclic Self-attention for Point Cloud Recognition
abstract
Point clouds provide a flexible geometric representation for computer vision research. However, the harsh demands for the number of input points and computer hardware are still significant challenges, which hinder their deployment in real applications. To address these challenges, we design a simple and effective module named cyclic self-attention module (CSAM). Specifically, three attention maps of the same input are obtained by cyclically pairing the feature maps, thus exploring the features sufficiently of the attention space of the original input. CSAM can adequately explore the correlation between points to obtain sufficient feature information despite the multiplicative decrease in inputs. Meanwhile, it can direct the computational power to the more essential features, relieving the burden on the computer hardware. We build a point cloud classification network by simply stacking CSAM called cyclic self-attention network (CSAN). We also propose a novel framework for point cloud semantic segmentation called full cyclic self-attention network (FCSAN). By adaptively fusing the original mapping features and the CSAM extracted features, it can better capture the context information of point clouds. Extensive experiments on several benchmark datasets show that our methods can achieve competitive performance in classification and segmentation tasks.
Guanyu Zhu, Yong Zhou 0003, Rui Yao 0006, Hancheng Zhu, Jiaqi Zhao 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Show, Deconfound and Tell: Image Captioning with Causal Inference
abstract
The transformer-based encoder-decoder framework has shown remarkable performance in image captioning. However, most transformer-based captioning methods ever overlook two kinds of elusive confounders: the visual confounder and the linguistic confounder, which generally lead to harmful bias, induce the spurious correlations during training, and degrade the model generalization. In this paper, we first use Structural Causal Models (SCMs) to show how two confounders damage the image captioning. Then we apply the backdoor adjustment to propose a novel causal inference based image captioning (CIIC) framework, which consists of an interventional object detector (IOD) and an interventional transformer decoder (ITD) to jointly confront both confounders. In the encoding stage, the IOD is able to disentangle the region-based visual features by deconfounding the visual confounder. In the decoding stage, the ITD introduces causal intervention into the transformer decoder and deconfounds the visual and linguistic confounders simultaneously. Two modules collaborate with each other to alleviate the spurious correlations caused by the unobserved confounders. When tested on MSCOCO, our proposal significantly outperforms the state-of-the-art encoder-decoder models on Karpathy split and online test split. Code is published in https://github.com/CUMTGG/CIIC.
Bing Liu 0016, Xu Yang 0021, Yong Zhou 0003, Rui Yao 0006, Zhiwen Shao, Jiaqi Zhao 0001
CVPR5
2022 Spatial hierarchy perception and hard samples metric learning for high-resolution remote sensing image object detection
Dongjun Zhu, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Ying Chen 0005
Appl. Intell.6
2022 Edge-aware and spectral-spatial information aggregation network for multispectral image semantic segmentation
Di Zhang 0020, Jiaqi Zhao 0001, Yong Zhou 0003, Boyu Shi, Rui Yao 0006
Eng. Appl. Artif. Intell.6
2022 Multi-granularity semantic alignment distillation learning for remote sensing image semantic segmentation
Di Zhang 0020, Yong Zhou 0003, Jiaqi Zhao 0001, Zhongyuan Yang, Rui Yao 0006, Huifang Ma
Frontiers Comput. Sci.6
2022 Multi-source collaborative enhanced for remote sensing images semantic segmentation
Jiaqi Zhao 0001, Di Zhang 0020, Boyu Shi, Yong Zhou 0003, Rui Yao 0006, Yong Xue
Neurocomputing6
2022 Point cloud recognition based on lightweight embeddable attention module
Guanyu Zhu, Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Man Zhang 0006
Neurocomputing4
2022 Semantic Segmentation of Remote-Sensing Images Based on Multiscale Feature Fusion and Attention Refinement
abstract
In recent years, the automatic extraction of remote-sensing image information has attracted full attention. However, the particularity of remote-sensing images and the scarcity of data sets with label information have brought new challenges to existing methods. Therefore, we develop a lightweight semantic segmentation network based onmultiscale feature fusion (MFF) and attention refinement (MFFANet). Our network relies on three crucial modules for improved performance. The multiscale attention refinement module strengthens the representation ability of feature maps extracted by the deep residual network. The MFF module aggregates the information carried by the high-level and low-level features while restoring the image resolution. Furthermore, the boundary enhancement module captures boundary details to solve the semantic ambiguity problem. We achieve 83.5% mean intersection over union (MIoU) on the Urban Semantic 3-D (US3D) data set and 69.3% MIoU on the Vaihingen data set with only 8.2M parameters.
Xin He 0024, Yong Zhou 0003, Jiaqi Zhao 0001, Man Zhang 0006, Rui Yao 0006, Bing Liu 0016
IEEE Geosci. Remote. Sens. Lett.5
2022 Few-Shot Object Detection via Context-Aware Aggregation for Remote Sensing Images
abstract
Few-shot object detection methods have made prodigious progress in recent years. However, these methods are designed for optical images at a single scale, which leads to significantly degraded detection performance due to object scale variation of remote sensing images. In this letter, we propose a few-shot object detection method for the problem of scale variation in remote sensing images. More specifically, our model contains two main components: a context-aware pixel aggregation (CPA) that allows the model to adapt to objects at different scales through different scale convolution and a context-aware feature aggregation (CFA) that enhances context awareness to obtain more semantic information through a graph convolution network (GCN). Experiments on the DIOR dataset demonstrate that our model can achieve a satisfying detection performance on remote sensing images, and our model performs significantly better than the state-of-the-art model.
Yong Zhou 0003, Jiaqi Zhao 0001, Hancheng Zhu, Rui Yao 0006, Wen-Liang Du 0002
IEEE Geosci. Remote. Sens. Lett.5
2022 Fine-Grained Feature Enhancement for Object Detection in Remote Sensing Images
abstract
Recently, object detection in aerial images has ushered in a new challenge—a new benchmark for fine-grained object recognition in high-resolution remote sensing imagery called FAIR1M has been proposed. Fine-grained categories usually have smaller inter class differences and intra-class similarities, which is more difficult to classify with existing object detectors. To address this problem, we propose two enhanced strategies on the current two-stage object detection algorithm. The first strategy uses attention-based group feature enhancement called group enhance module (GEM). By extending and grouping feature channels, the model can improve the ability to extract various discriminative features. The second strategy is to emphasize the sub-saliency feature learning, avoiding the network only focusing on the most significant part of the feature and ignoring the other parts. Our method is easy to implement and effective, and experiments show that our method can improve the Oriented regions with convolutional neural networks features (R-CNN) by about 1.45 mAP on the FAIR1M benchmark.
Yong Zhou 0003, Sifan Wang, Jiaqi Zhao 0001, Hancheng Zhu, Rui Yao 0006
IEEE Geosci. Remote. Sens. Lett.5
2022 Efficient lightweight video person re-identification with online difference discrimination module
Cunyuan Gao, Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Fuyuan Hu
Multim. Tools Appl.2
2022 Survey for person re-identification based on coarse-to-fine feature learning
Minjie Liu, Jiaqi Zhao 0001, Yong Zhou 0003, Hancheng Zhu, Rui Yao 0006, Ying Chen 0005
Multim. Tools Appl.5
2022 ResT-ReID: Transformer block-based residual learning for person re-identification
Ying Chen 0005, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Dongjun Zhu, Dongjingdian Liu
Pattern Recognit. Lett.6
2022 Learning image aesthetic subjectivity from attribute-aware relational reasoning network
Hancheng Zhu, Yong Zhou 0003, Rui Yao 0006, Guangcheng Wang, Yuzhe Yang 0001
Pattern Recognit. Lett.3
2022 Spatial-Temporal Based Multihead Self-Attention for Remote Sensing Image Change Detection
abstract
The neural network-based remote sensing image change detection method faces a large amount of imaging interference and severe class imbalance problems under high-resolution conditions, which bring new challenges to the accuracy of the detection network. In this work, to address the imaging interference caused by different imaging angles and times, the siamese strategy and multi-head self-attention mechanism are used to reduce the imaging differences between the dual-temporal images and fully exploit the inter-temporal information. Secondly, a learnable multi-part feature learning module is used to adaptively exploit features from different scales to obtain more comprehensive features. Finally, a mixed loss function strategy is used to ensure that the network converges effectively and excludes the adverse interference of a large number of negative samples to the network. Extensive experiments show that our method outperforms numerous methods on LEVIR-CD, WHU, and DSIFN datasets.
Yong Zhou 0003, Fengkai Wang, Jiaqi Zhao 0001, Rui Yao 0006, Silin Chen, Heping Ma
IEEE Trans. Circuits Syst. Video Technol.4
2022 Swin Transformer Embedding UNet for Remote Sensing Image Semantic Segmentation
abstract
Global context information is essential for the semantic segmentation of remote sensing (RS) images. However, most existing methods rely on a convolutional neural network (CNN), which is challenging to directly obtain the global context due to the locality of the convolution operation. Inspired by the Swin transformer with powerful global modeling capabilities, we propose a novel semantic segmentation framework for RS images called ST-U-shaped network (UNet), which embeds the Swin transformer into the classical CNN-based UNet. ST-UNet constitutes a novel dual encoder structure of the Swin transformer and CNN in parallel. First, we propose a spatial interaction module (SIM), which encodes spatial information in the Swin transformer block by establishing pixel-level correlation to enhance the feature representation ability of occluded objects. Second, we construct a feature compression module (FCM) to reduce the loss of detailed information and condense more small-scale features in patch token downsampling of the Swin transformer, which improves the segmentation accuracy of small-scale ground objects. Finally, as a bridge between dual encoders, a relational aggregation module (RAM) is designed to integrate global dependencies from the Swin transformer into the features from CNN hierarchically. Our ST-UNet brings significant improvement on the ISPRS-Vaihingen and Potsdam datasets, respectively. The code will be available athttps://github.com/XinnHe/ST-UNet.
Xin He 0024, Yong Zhou 0003, Jiaqi Zhao 0001, Di Zhang 0020, Rui Yao 0006, Yong Xue
IEEE Trans. Geosci. Remote. Sens.5
2022 CLT-Det: Correlation Learning Based on Transformer for Detecting Dense Objects in Remote Sensing Images
abstract
Challenges still exist in the task of object detection in remote sensing images with densely distributed objects due to large variation in scale and neglect of the relative position and correlation. To address these issues, a Correlation Learning Detector based on Transformer (CLT-Det) is proposed for detecting dense objects in remote sensing images. A Transformer Attention Module (TAM) is designed to improve the densely packed objects’ model representation ability by learning pixel-wise attention with Transformer. To alleviate the semantic gap caused by variations in scale, a Feature Refinement Module (FRM) is proposed by improving the multi-scale feature pyramid. A Correlation Transformer Module (CTM) is proposed to extract correlation information and encodes position information of dense objects’ features on the classification branch for fully utilizing the position information and correlation among objects. Extensive experiments compared with several state-of-art methods on two challenging remote sensing datasets, namely DOTA and HRSC2016, demonstrate that the proposed CLT-Det achieves promising and competitive performance.
Yong Zhou 0003, Silin Chen, Jiaqi Zhao 0001, Rui Yao 0006, Yong Xue, Abdulmotaleb El Saddik
IEEE Trans. Geosci. Remote. Sens.4
2022 Clustering Matters: Sphere Feature for Fully Unsupervised Person Re-identification
abstract
In person re-identification (Re-ID) , the data annotation cost of supervised learning, is huge and it cannot adapt well to complex situations. Therefore, compared with supervised deep learning methods, unsupervised methods are more in line with actual needs. In unsupervised learning, a key to solving Re-ID is to find a standard that can effectively distinguish the difference (distance) between the features of images belonging to different pedestrian identities. However, there are some differences in the images captured by different cameras (such as brightness, angle, etc.). It is well known that the training of neural networks is mainly based on the distance between features, while in unsupervised learning, especially in unsupervised learning methods based on hierarchical clustering, the distance between features plays a more important role in the clustering phase. We improve the accuracy of a deep learning method based on hierarchical clustering under fully unsupervised conditions, starting from both feature and distance metrics. First, we propose to use spherical features, by normalizing the images in the feature space, to weaken the structural differences (length) between features, while saving the feature differences (direction) between different identities. Then, we use the sum of squared errors (SSE) as a regularization term to balance different cluster states. We evaluate our method on four large-scale Re-ID datasets, and experiments show that our method achieves better results than the state-of-the-art unsupervised methods.
Yong Zhou 0003, Jiaqi Zhao 0001, Ying Chen 0005, Rui Yao 0006, Bing Liu 0016, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.5
2021 Joint Attention Mechanism for Unsupervised Video Object Segmentation
Rui Yao 0006, Xin Xu 0009, Yong Zhou 0003, Jiaqi Zhao 0001
PRCV (1)1
2021 Point cloud classification by dynamic graph CNN with adaptive feature fusion
abstract
Abstract The deep neural network has made the most advanced breakthrough in almost all 2D image tasks, so we consider the application of deep learning in 3D images. Point cloud data, as the most basic and important form of representation of 3D images, can accurately and intuitively show the real world. The authors propose a new network based on feature fusion to improve the point cloud classification and segmentation tasks. Our network mainly consists of three parts: global feature extractor, local feature extractor and adaptive feature fusion module. A multi‐scale transformation network is devised to guarantee the invariance of the transformation of the global feature, and a residual block is introduced to alleviate the problem of gradient disappearance to enhance the global feature extractor. Based on the edge convolution and multi‐layer perceptron, a local feature extractor is constructed. Finally, an adaptive feature‐fusion module is proposed to complete the fusion of global features and local features. Extensive experiments on point cloud classification and segmentation tasks are carried out to verify the effectiveness of the proposed method. The classification accuracy of the ModelNet40 is 93.6%, which is 4.4% higher than that of the PointNet. Similarly, the segmentation accuracy on the ShapeNet is 85.6%, which is higher than other methods.
Yong Zhou 0003, Jiaqi Zhao 0001, Yiyun Man, Minjie Liu, Rui Yao 0006, Bing Liu 0016
IET Comput. Vis.6
2021 AMC-Net: Attentive modality-consistent network for visible-infrared person re-identification
Hanzheng Wang, Jiaqi Zhao 0001, Yong Zhou 0003, Rui Yao 0006, Ying Chen 0005, Silin Chen
Neurocomputing4
2021 Unsupervised cross-domain person re-identification with self-attention and joint-flexible optimization
Haopeng Hou, Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Ying Chen 0005, Abdulmotaleb El Saddik
Image Vis. Comput.4
2021 A siamese pedestrian alignment network for person re-identification
Yong Zhou 0003, Jiaqi Zhao 0001, Meng Jian, Rui Yao 0006, Bing Liu 0016, Ying Chen 0005
Multim. Tools Appl.5
2021 Video-based person re-identification by semi-supervised adaptive stepwise learning
Yong Zhou 0003, Jiaqi Zhao 0001, Ying Chen 0005, Rui Yao 0006
Pattern Anal. Appl.5
2021 CycleSegNet: Object Co-Segmentation With Cycle Refinement and Region Correspondence
abstract
Image co-segmentation is an active computer vision task that aims to segment the common objects from a set of images. Recently, researchers design various learning-based algorithms to undertake the co-segmentation task. The main difficulty in this task is how to effectively transfer information between images to make conditional predictions. In this paper, we present CycleSegNet, a novel framework for the co-segmentation task. Our network design has two key components: a region correspondence module which is the basic operation for exchanging information between local image regions, and a cycle refinement module, which utilizes ConvLSTMs to progressively update image representations and exchange information in a cycle and iterative manner. Extensive experiments demonstrate that our proposed method significantly outperforms the state-of-the-art methods on four popular benchmark datasets - PASCAL VOC dataset, MSRC dataset, Internet dataset, and iCoseg dataset, by 2.6%, 7.7%, 2.2%, and 2.9%, respectively.
Chi Zhang 0007, Guankai Li, Guosheng Lin, Qingyao Wu, Rui Yao 0006
IEEE Trans. Image Process.5
2021 Multi-Stage Fusion and Multi-Source Attention Network for Multi-Modal Remote Sensing Image Segmentation
abstract
With the rapid development of sensor technology, lots of remote sensing data have been collected. It effectively obtains good semantic segmentation performance by extracting feature maps based on multi-modal remote sensing images since extra modal data provides more information. How to make full use of multi-model remote sensing data for semantic segmentation is challenging. Toward this end, we propose a new network called Multi-Stage Fusion and Multi-Source Attention Network ((MS) 2 -Net) for multi-modal remote sensing data segmentation. The multi-stage fusion module fuses complementary information after calibrating the deviation information by filtering the noise from the multi-modal data. Besides, similar feature points are aggregated by the proposed multi-source attention for enhancing the discriminability of features with different modalities. The proposed model is evaluated on publicly available multi-modal remote sensing data sets, and results demonstrate the effectiveness of the proposed method.
Jiaqi Zhao 0001, Yong Zhou 0003, Boyu Shi, Jingsong Yang, Di Zhang 0020, Rui Yao 0006
ACM Trans. Intell. Syst. Technol.6
2020 Person image synthesis through siamese generative adversarial network
Ying Chen 0005, Shixiong Xia, Jiaqi Zhao 0001, Meng Jian, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Dongjun Zhu
Neurocomputing7
2020 Multiobjective ResNet pruning by means of EMOAs for remote sensing scene classification
Xuning Liu, Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Bing Liu 0016
Neurocomputing4
2020 Diverse sample generation with multi-branch conditional generative adversarial network for remote sensing objects detection
Dongjun Zhu, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Meng Jian, Qiang Niu, Rui Yao 0006, Ying Chen 0005
Neurocomputing7
2020 Appearance and shape based image synthesis by conditional variational generative adversarial network
Ying Chen 0005, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Dongjun Zhu
Knowl. Based Syst.6
2020 Fusion based feature reinforcement component for remote sensing image object detection
Dongjun Zhu, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Ying Chen 0005
Multim. Tools Appl.6
2020 GAN-based person search via deep complementary classifier with center-constrained Triplet loss
Rui Yao 0006, Cunyuan Gao, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Fuyuan Hu
Pattern Recognit.1
2020 Model-Free Tracker for Multiple Objects Using Joint Appearance and Motion Inference
abstract
Model-free tracking is a widely-accepted approach to track an arbitrary object in a video using a single frame annotation with no further prior knowledge about the object of interest. Extending this problem to track multiple objects is really challenging because: a) the tracker is not aware of the objects' type while trying to distinguish them from background (detection task), and b) The tracker needs to distinguish one object from other potentially similar objects (data association task) to generate stable trajectories. In order to track multiple arbitrary objects, most existing model-free tracking approaches rely on tracking each target individually by updating their appearance model independently. Therefore, in this scenario they often fail to perform well due to confusion between the appearance of similar objects, their sudden appearance changes and occlusion. To tackle this problem, we propose to use both appearance and motion models, and to learn them jointly using graphical models and deep neural networks features. We introduce an indicator variable to predict sudden appearance change and/or occlusion. When these happen, our model does not update the appearance model thus avoiding using the background and/or incorrect object to update the appearance of the object of interest mistakenly, and relies on our motion model to track. Moreover, we consider the correlation among all targets, and seek the joint optimal locations for all targets simultaneously as a graphical model inference problem. We learn the joint parameters for both appearance model and motion model in an online fashion under the framework of LaRank. Experiment results show that our method achieved superior performance compared to the competitive methods.
Chongyu Liu, Rui Yao 0006, Seyed Hamid Rezatofighi, Ian D. Reid 0001, Qinfeng Shi
IEEE Trans. Image Process.2
2020 Video Object Segmentation and Tracking: A Survey
abstract
Object segmentation and object tracking are fundamental research areas in the computer vision community. These two topics are difficult to handle some common challenges, such as occlusion, deformation, motion blur, scale variation, and more. The former contains heterogeneous object, interacting object, edge ambiguity, and shape complexity; the latter suffers from difficulties in handling fast motion, out-of-view, and real-time processing. Combining the two problems of Video Object Segmentation and Tracking (VOST) can overcome their respective difficulties and improve their performance. VOST can be widely applied to many practical applications such as video summarization, high definition video compression, human computer interaction, and autonomous vehicles. This survey aims to provide a comprehensive review of the state-of-the-art VOST methods, classify these methods into different categories, and identify new trends. First, we broadly categorize VOST methods into Video Object Segmentation (VOS) and Segmentation-based Object Tracking (SOT). Each category is further classified into various types based on the segmentation and tracking mechanism. Moreover, we present some representative VOS and SOT methods of each time node. Second, we provide a detailed discussion and overview of the technical characteristics of the different methods. Third, we summarize the characteristics of the related video dataset and provide a variety of evaluation metrics. Finally, we point out a set of interesting future works and draw our own conclusions.
Rui Yao 0006, Guosheng Lin, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003
ACM Trans. Intell. Syst. Technol.1
2019 CANet: Class-Agnostic Segmentation Networks With Iterative Refinement and Attentive Few-Shot Learning
abstract
Recent progress in semantic segmentation is driven by deep Convolutional Neural Networks and large-scale labeled image datasets. However, data labeling for pixel-wise segmentation is tedious and costly. Moreover, a trained model can only make predictions within a set of pre-defined classes. In this paper, we present CANet, a class-agnostic segmentation network that performs few-shot segmentation on new classes with only a few annotated images available. Our network consists of a two-branch dense comparison module which performs multi-level feature comparison between the support image and the query image, and an iterative optimization module which iteratively refines the predicted results. Furthermore, we introduce an attention mechanism to effectively fuse information from multiple support examples under the setting of k-shot learning. Experiments on PASCAL VOC 2012 show that our method achieves a mean Intersection-over-Union score of 55.4% for 1-shot segmentation and 57.1% for 5-shot segmentation, outperforming state-of-the-art methods by a large margin of 14.6% and 13.2%, respectively.
Chi Zhang 0007, Guosheng Lin, Fayao Liu, Rui Yao 0006, Chunhua Shen
CVPR4
2019 Pyramid Graph Networks With Connection Attentions for Region-Based One-Shot Semantic Segmentation
abstract
One-shot image segmentation aims to undertake the segmentation task of a novel class with only one training image available. The difficulty lies in that image segmentation has structured data representations, which yields a many-to-many message passing problem. Previous methods often simplify it to a one-to-many problem by squeezing support data to a global descriptor. However, a mixed global representation drops the data structure and information of individual elements. In this paper, we propose to model structured segmentation data with graphs and apply attentive graph reasoning to propagate label information from support data to query data. The graph attention mechanism could establish the element-to-element correspondence across structured data by learning attention weights between connected graph nodes. To capture correspondence at different semantic levels, we further propose a pyramid-like structure that models different sizes of image regions as graph nodes and undertakes graph reasoning at different levels. Experiments on PASCAL VOC 2012 dataset demonstrate that our proposed network significantly outperforms the baseline method and leads to new state-of-the-art performance on 1-shot and 5-shot segmentation benchmarks.
Chi Zhang 0007, Guosheng Lin, Fayao Liu, Jiushuang Guo, Qingyao Wu, Rui Yao 0006
ICCV6
2019 Lightweight Video Object Segmentation Based on ConvGRU
Rui Yao 0006, Yikun Zhang 0001, Cunyuan Gao, Yong Zhou 0003, Jiaqi Zhao 0001, Lina Liang
PRCV (2)1
2019 A Siamese Pedestrian Alignment Network for Person Re-identification
Yong Zhou 0003, Jiaqi Zhao 0001, Meng Jian, Rui Yao 0006, Bing Liu 0016, Xuning Liu
PRCV (1)5
2019 Structure-aware person search with self-attention and online instance aggregation matching
Cunyuan Gao, Rui Yao 0006, Jiaqi Zhao 0001, Yong Zhou 0003, Fuyuan Hu, Leida Li
Neurocomputing2
2019 Local fusion networks with chained residual pooling for video action recognition
abstract
Action recognition is an important yet challenging problem. We here present a novel method, multistage local fusion networks with residual connections, to boost the performance of video action recognition . In realistic videos, an action instance may have a long time span and some frames may suffer from deteriorated object appearance due to motion blur or video defocus. Our method enhances the per-frame representation by capturing information from neighboring frames. We propose a local fusion block which considers neighboring frames to capture appearance and local motion information for generating per-frame representation. Our local fusion is performed in a multistage manner allowing feature fusion from varying neighborhood sizes in the temporal dimension. We employ residual connections in the fusion blocks to enable effective gradient propagation through the whole network allowing effective end-to-end training. We achieve competitive results on two challenging and public available datasets, namely HMDB51 and UCF101, which shows the effectiveness of the proposed method.
Feixiang He, Fayao Liu, Rui Yao 0006, Guosheng Lin
Image Vis. Comput.3
2019 Siamese Convolutional Neural Networks for Remote Sensing Scene Classification
abstract
The convolutional neural networks (CNNs) have shown powerful feature representation capability, which provides novel avenues to improve scene classification of remote sensing imagery. Although we can acquire large collections of satellite images, the lack of rich label information is still a major concern in the remote sensing field. In addition, remote sensing data sets have their own limitations, such as the small scale of scene classes and lack of image diversity. To mitigate the impact of the existing problems, a Siamese CNN, which combines the identification and verification models of CNNs, is proposed in this letter. A metric learning regularization term is explicitly imposed on the features learned through CNNs, which enforce the Siamese networks to be more robust. We carried out experiments on three widely used remote sensing data sets for performance evaluation. Experimental results show that our proposed method outperforms the existing methods.
Xuning Liu, Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Bing Liu 0016
IEEE Geosci. Remote. Sens. Lett.4
2019 Data augmentation of random grid-hiding for video object segmentation
Rui Yao 0006, Yikun Zhang 0001, Qingnan Jiang, Changbin Zhang
Multim. Tools Appl.2
2019 Semantics-Aware Visual Object Tracking
abstract
In this paper, we propose a semantics-aware visual object tracking method, which introduces semantics into the tracking procedure and extends the model of an object with explicit semantics prior to enhancing the robustness of three key aspects of the tracking framework, i.e., appearance model, search scheme, and scale adaptation. We first present a semantic object proposal generation method for video sequences to generate high-quality category-oriented object proposals. Then, a hybrid semantics-aware tracking algorithm with semantic compatibility is proposed. This algorithm takes full advantages of globally sparse semantic object proposal prediction and locally dense prediction with a template model and semantic distractor-aware color appearance model. Furthermore, we propose to exploit semantics to localize object accurately via an energy minimization framework-based scale adaptation method, which jointly integrates dense location prior, instance-specific color, and category-specific semantic information. Extensive experiments are conducted on two widely used benchmarks, and the results demonstrate that our method achieves the state-of-the-art performance.
Rui Yao 0006, Guosheng Lin, Chunhua Shen, Yanning Zhang 0001, Qinfeng Shi
IEEE Trans. Circuits Syst. Video Technol.1
2018 Pareto-Based Many-Objective Convolutional Neural Networks
Hongjian Zhao, Shixiong Xia, Jiaqi Zhao 0001, Dongjun Zhu, Rui Yao 0006, Qiang Niu
WISA5
2018 Efficient eye typing with 9-direction gaze estimation
Chi Zhang 0007, Rui Yao 0006, Jinpeng Cai
Multim. Tools Appl.2
2018 Efficient dense labelling of human activity sequences from wearables using fully convolutional networks
Rui Yao 0006, Guosheng Lin, Qinfeng Shi, Damith Chinthana Ranasinghe
Pattern Recognit.1
2017 Solving Constrained Combinatorial Optimisation Problems via MAP Inference without High-Order Penalties
abstract
Solving constrained combinatorial optimisation problems via MAP inference is often achieved by introducing extra potential functions for each constraint. This can result in very high order potentials, e.g. a 2nd-order objective with pairwise potentials and a quadratic constraint over all N variables would correspond to an unconstrained objective with an order-N potential. This limits the practicality of such an approach, since inference with high order potentials is tractable only for a few special classes of functions. We propose an approach which is able to solve constrained combinatorial problems using belief propagation without increasing the order. For example, in our scheme the 2nd-order problem above remains order 2 instead of order N. Experiments on applications ranging from foreground detection, image reconstruction, quadratic knapsack, and the M-best solutions problem demonstrate the effectiveness and efficiency of our method. Moreover, we show several situations in which our approach outperforms commercial solvers like CPLEX and others designed for specific constrained MAP inference problems.
Zhen Zhang 0008, Qinfeng Shi, Julian J. McAuley, Wei Wei 0008, Yanning Zhang 0001, Rui Yao 0006, Anton van den Hengel
AAAI6
2017 Video stitching based on iterative hashing and dynamic seam-line with local context
Rui Yao 0006, Jinliang Sun, Yong Zhou 0003, Dai Chen
Multim. Tools Appl.1
2017 Part-Based Robust Tracking Using Online Latent Structured Learning
abstract
Despite many advances made in the area, deformable targets and partial occlusions continue to represent key problems in visual tracking. Structured learning has shown good results when applied to tracking whole targets, but applying this approach to a part-based target model is complicated by the need to model the relationships between parts and to avoid lengthy initialization processes. We thus propose a method that models the unknown parts by using latent variables. In doing so, we extend the online algorithm Pegasos to the structured prediction case (i.e., predicting the location of the bounding boxes) with latent part variables. We also incorporate the very recently proposed spatial constraints to preserve distances between parts. To better estimate the parts, and to avoid overfitting caused by the extra model complexity/capacity introduced by the parts, we propose a two-stage training process based on the primal rather than the dual form. We then show that the method outperforms the state of the arts in extensive experiments.
Rui Yao 0006, Qinfeng Shi, Chunhua Shen, Yanning Zhang 0001, Anton van den Hengel
IEEE Trans. Circuits Syst. Video Technol.1
2017 Real-Time Correlation Filter Tracking by Efficient Dense Belief Propagation With Structure Preserving
abstract
Patch-based models that combine local image features or regions into loose geometric assemblies are a powerful paradigm for visual object tracking, and they present favorable properties such as robustness to partial occlusion, deformation, and the ability to address viewpoint changes. However, effectively exploiting the spatial-temporal confidence scores of each patch to construct a robust tracker while ensuring a low computational cost with dense discrete search remains a challenging problem. In this paper, we propose a unified Markov random field (MRF) model that can effectively capture spatio-temporal intrapatch relations and occlusion priors to enhance the tracking performance, and we derive a highly efficient dense belief propagation for inference of the proposed MRF model. We propose a tracker that models the tracking object with a constellation topology (i.e., a global object and several local patches), where the graph structure model describes the pairwise spatial structure and the image observation model corresponding to the correlation filter with occlusion handling measuring the appearance similarity. Furthermore, these two models are updated online by exploiting the flexibility of local patches. Extensive experimental results on single- and multiple-object tracking show that the proposed algorithm performs favorably against state-of-the-art methods and runs in real time.
Rui Yao 0006, Shixiong Xia, Zhen Zhang 0008, Yanning Zhang 0001
IEEE Trans. Multim.1
2016 Robust lifelong visual tracking using compact binary feature with color attributes
Rui Yao 0006, Shixiong Xia, Yong Zhou 0003, Qiang Niu
Neurocomputing1
2016 Exploiting Spatial Structure from Parts for Adaptive Kernelized Correlation Filter Tracker
abstract
Decomposing target into several parts may improve the capability of tracking algorithm to deal with appearance variations such as occlusion and deformation. In this letter, we propose a part-based appearance model by exploiting spatial structure from parts. The model minimizes appearance and deformation cost simultaneously to predict the new position of object. Then, the optimization problem is divided into two parts. Kernelized correlation filter (KCF) is used for tracking the appearance of parts separately to speed up the proposed tracker. Meanwhile, the deformation cost is minimized by structural learning schema, which can reduce the label noise that caused by inaccuracy bounding box. Finally, minimum spanning tree and dynamic programming are employed to combine the score map of the appearance and deformation of parts, and to detect best new position of target. Experimental results on several challenge sequences show the efficiency and effectiveness of the proposed tracking algorithm.
Rui Yao 0006, Shixiong Xia, Fumin Shen, Yong Zhou 0003, Qiang Niu
IEEE Signal Process. Lett.1
2015 Robust tracking via online Max-Margin structural learning with approximate sparse intersection kernel
Rui Yao 0006, Shixiong Xia, Yong Zhou 0003
Neurocomputing1
2015 Robust Model-Free Multi-Object Tracking with Online Kernelized Structural Learning
abstract
One of the most important issues in robust visual tracking is that the method must be flexible enough to endure the inevitable changes in object appearance over time, which is the main propose of many model-free trackers. Nevertheless, existing online model-free methods typically focus on single object tracking. In this letter, we propose a novel multi-object tracker based on online structured learning which allows us to learn a uniform structural classifier from training samples of all objects. We then derive a novel online updating dual form to facilitate efficient non-linear kernels. By formulating a direct online structured learning method for classifying multiple objects, we build a framework for multi-object tracking, where single object tracking is its special case. Both qualitative and quantitative evaluations demonstrate that the proposed multiple object tracker outperforms most current state-of-the-art methods.
Rui Yao 0006
IEEE Signal Process. Lett.1
2013 Part-Based Visual Tracking with Online Latent Structural Learning
abstract
Despite many advances made in the area, deformable targets and partial occlusions continue to represent key problems in visual tracking. Structured learning has shown good results when applied to tracking whole targets, but applying this approach to a part-based target model is complicated by the need to model the relationships between parts, and to avoid lengthy initialisation processes. We thus propose a method which models the unknown parts using latent variables. In doing so we extend the online algorithm pegasos to the structured prediction case (i.e., predicting the location of the bounding boxes) with latent part variables. To better estimate the parts, and to avoid over-fitting caused by the extra model complexity/capacity introduced by the parts, we propose a two-stage training process, based on the primal rather than the dual form. We then show that the method outperforms the state-of-the-art (linear and non-linear kernel) trackers.
Rui Yao 0006, Qinfeng Shi, Chunhua Shen, Yanning Zhang 0001, Anton van den Hengel
CVPR1
2012 Robust Tracking with Weighted Online Structured Learning
Rui Yao 0006, Qinfeng Shi, Chunhua Shen, Yanning Zhang 0001, Anton van den Hengel
ECCV (3)1