EDBT 2026 Demo / reviewers in the wild / expert
Yong Zhou 0003
dblp:90/5836-3
· DBLP profile ↗
158ranked-venue papers
8as first author
119since 2021 · last 2026
0000-0001-6207-0299ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 1 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 61 · 4 first-author · 54 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 3 first-author · 24 since 2021Computer networks · 23 · 1 first-author · 21 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Security and privacy · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal ModelingabstractVideo shadow detection confronts two entwined difficulties: distinguishing shadows from complex backgrounds and modeling dynamic shadow deformations under varying illumination. To address shadow-background ambiguity, we leverage linguistic priors through the proposed Vision-language Match Module (VMM) and a Dark-aware Semantic Block (DSB), extracting text-guided features to explicitly differentiate shadows from dark objects. Furthermore, we introduce adaptive mask reweighting to downweight penumbra regions during training and apply edge masks at the final decoder stage for better supervision. For temporal modeling of variable shadow shapes, we propose a Tokenized Temporal Block (TTB) that decouples spatiotemporal learning. TTB summarizes cross-frame shadow semantics into learnable temporal tokens, enabling efficient sequence encoding with minimal computation overhead. Comprehensive Experiments on multiple benchmark datasets demonstrate state-of-the-art accuracy and real-time inference efficiency. Kunyang Sun, Rui Yao 0006, Hancheng Zhu, Fuyuan Hu, Jiaqi Zhao 0001, Zhiwen Shao, Yong Zhou 0003 |
AAAI | 8 |
| 2026 | CLIPDet3D: Vision-Language Collaborative Distillation for 3D Object DetectionabstractMulti-view 3D object detection plays a vital role in autonomous driving systems due to its ability to perceive complex scenes accurately. However, real-world driving data often exhibits a long-tailed distribution, causing significant drops in detection accuracy for rare categories in existing methods. To mitigate this issue, we propose CLIPDet3D, a novel vision-language collaborative framework for multi-view 3D object detection. First, to tackle the difficulty of capturing the semantic information of rare categories, a Vision-Language Collaborative Learning strategy is proposed to incorporate class-level semantic priors from CLIP. Second, a Depth Feature Contrastive Distillation module is designed to overcome the large depth estimation error for rare categories by aligning depth features between a teacher and a student network. Furthermore, to alleviate the difficulty in focusing on regions of rare categories, a Dual-Stream Prompt Attention mechanism is devised to inject learnable prompts and compute attention along both horizontal and vertical BEV directions. Evaluations on the nuScenes dataset demonstrate that CLIPDet3D achieves state-of-the-art accuracy while maintaining efficient inference. Jiaqi Zhao 0001, Huanfeng Hu, Yong Zhou 0003, Wen-Liang Du 0002, Kunyang Sun, Rui Yao 0006, Qigong Sun |
AAAI | 3 |
| 2026 | Unified Representation Causal Prompt Distillation for Re-Inference-Free Lifelong Person Re-IdentificationabstractLifelong person re-identification (LReID) aims to retrieve the target person from sequentially collected data. Due to significant domain gaps between datasets and the continuous increase of training data from different scenarios, weak inter-domain generalization and catastrophic forgetting issues have remained major challenges for LReID. To tackle these issues, a novel LReID method called Unified Representation Causal Prompt Distillation (URCPD) is proposed. Specifically, to reduce domain gaps among different scene datasets and improve model inter-domain generalization capability, a Feature Decoupling Style Transfer module (FDST) is proposed to map new features into a unified feature space. Furthermore, to reduce the accumulated forgetting of old knowledge during the training stage, a Causal Prompt Distillation module (CPD) is introduced. This module eliminates the re-inference process for distillation and embeds memory prompts to combat catastrophic forgetting. Extensive experiments on five classic LReID seen datasets and seven unseen datasets demonstrate that our method significantly outperforms state-of-the-art methods. Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006 |
AAAI | 3 |
| 2026 | Causal Decoupling Domain Generalization for Remote Sensing Change DetectionabstractWhile current state-of-the-art Remote Sensing Change Detection (RSCD) methods can achieve impressive results on individual datasets, they become unreliable in unseen environments and imaging conditions, with performance metrics declining by as much as 60% to 80%. Simultaneously, variable environments and complex imaging conditions are the main characteristics of remote sensing data, calling for generalizable RSCD methods. To address this issue, we propose a novel RSCD method capable of domain generalization—CDDGNet. This method is based on causal decoupling theory, which progressively decouples invariant change features from variable domain features to extract generalizable characteristics. This enables a network trained on a single domain to accurately identify change regions in other domains. Specifically, firstly, the Causal Feature Adaptation Module is proposed to preliminarily decouple and simplify feature information during the encoding process by using wavelet transformation and feature energy spectralization methods. Secondly, the Causal Feature Fusion Module is presented to fully decouple features and aggregate significant change features during the decoding process through frequency domain processing and feature re-attention mechanisms. Thirdly, the Decoupling Effect Loss Function is proposed to optimize the process by evaluating the effectiveness of causal decoupling. Extensive experiments have shown that our model significantly outperforms existing methods across multiple groups of generalization tasks with varying levels of difficulty. Jiaqi Zhao 0001, Jianpeng Xie 0001, Yong Zhou 0003, Wen-Liang Du 0002, Hancheng Zhu, Rui Yao 0006 |
AAAI | 3 |
| 2026 | A Hybrid Transformer-GCN Framework for CircRNA-Disease Association Prediction
Mengmeng Wei, Ziyuan Shen, Ziqi Xia, Lei Wang 0121, Yong Zhou 0003 |
ICIC (27) | 6 |
| 2026 | Few-Shot Quaternion-valued Correlation Squeeze Network for Document Image Layout Segmentation
Rui Yao 0006, Qiwei Yu, Songhui Zhao, Yong Zhou 0003, Bing Liu 0016 |
Int. J. Document Anal. Recognit. | 5 |
| 2026 | Style-controllable adversarial example generation via image editing and prompt embedding optimization
Yong Zhou 0003, Bing Liu 0016, Rui Yao 0006 |
Neurocomputing | 2 |
| 2026 | Constrained and directional ensemble attention for facial action unit detection
Zhiwen Shao, Bikuan Chen, Yong Zhou 0003, Xuehuai Shi, Canlin Li, Lizhuang Ma, Dit-Yan Yeung |
Pattern Recognit. | 3 |
| 2026 | Retrieval-Augmented Pseudo-Image Guided Alignment and Text Domain-Aware Memory Recall for Continual Zero-Shot Captioning
Bing Liu 0016, Hao Liu 0065, Peng Liu 0013, Yong Zhou 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | SpaceFormer: Spatial Position Contextual Semantics Embedding for Multi-View 3D Object Detectionabstract3D object detection aims to accurately localize and recognize objects in 3D space. It serves as a fundamental task for reliable perception in intelligent transportation systems, enabling the monitoring of diverse traffic participants such as vehicles, pedestrians, cyclists, and public transport. Recently, transformer-based methods have gained significant attention in multi-view 3D object detection due to their strong global reasoning capabilities. However, their limited capacity to model spatial positional information hinders accurate object localization, especially in complex and large-scale scenes. To address this limitation, SpaceFormer is proposed as a novel transformer-based multi-view 3D object detector. Specifically, a Contextual Visual Prompts Learning strategy is proposed to enhance the perception of small and sparse traffic participants by incorporating contextual priors. To further suppress background interference, a Semantics-guided Depth Estimation method is proposed to refine depth representations using high-level semantic information. Furthermore, a Spatial Position Embedding mechanism is proposed to improve the spatial localization capability of the transformer by integrating geometric position and polar spatial embedding. Extensive experiments on the nuScenes benchmark demonstrate that SpaceFormer achieves state-of-the-art performance with 55.5% mAP and 62.9% NDS. These improvements indicate not only methodological advances but also practical benefits for intelligent transportation systems, enhancing safety, reliability, and efficiency in real-world deployments. Jiaqi Zhao 0001, Huanfeng Hu, Wen-Liang Du 0002, Yong Zhou 0003, Kunyang Sun, Rui Yao 0006, Abdulmotaleb El Saddik |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2026 | Dual Sparse Long-Short Term Transformer for Video Shadow DetectionabstractVideo Shadow Detection (VSD) is critical yet challenging, primarily due to ambiguous shadow boundaries and the presence of confusing shadow-like non-shadow regions, which existing methods struggle to resolve effectively by limited temporal modeling. We propose the Dual Sparse Long-Short Term Transformer Network (DSLSTT-Net), a novel framework designed to enhance feature learning by integrating robust temporal consistency and detailed local context. DSLSTT-Net utilizes a dual-stream architecture to concurrently process global temporal information and local shadow feature refinement, enabling effective discrimination between true shadows and confusing areas. At its core, the Sparse Long-Short Term Attention Module (Sparse LSTAM) is introduced to efficiently propagate only high-confidence shadow features from memory, significantly enhancing feature discriminability and computational efficiency. Furthermore, an Adaptive Fusion Module (AFM) dynamically merges purified long-term features with short-term details, optimizing final segmentation. Experimental results confirm that DSLSTT-Net significantly outperforms state-of-the-art methods on VSD benchmarks, validating our approach of dual-stream architecture and sparse temporal modeling. The source code is available at https://github.com/rayyao/DSLSTTNet . Rui Yao 0006, Huili Hao, Hancheng Zhu, Jiaqi Zhao 0001, Yong Zhou 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | ReDiffDet: Rotation-equivariant Diffusion Model for Oriented Object DetectionabstractThe diffusion model has been successfully applied to various detection tasks. However, it still faces several challenges when used for oriented object detection: objects that are arbitrarily rotated require the diffusion model to encode their orientation information; uncontrollable random boxes inaccurately locate objects with dense arrangements and extreme aspect ratios; oriented boxes result in the misalignment between them and image features. To overcome these limitations, we propose ReDiffDet, a framework that formulates oriented object detection as a rotation-equivariant denoising diffusion process. First, we represent an oriented box as a 2D Gaussian distribution, forming the basis of the denoising paradigm. The reverse process can be proven to be rotation-equivariant within this representation and model framework. Second, we design a conditional encoder with conditional boxes to prevent boxes from being randomly placed across the entire image. Third, we propose an aligned decoder for alignment between oriented boxes and image features. The extensive experiments demonstrate ReDiffDet achieves promising performance and significantly outperforms the diffusion-based baseline detector. Codes are available at https://github.com/wokaikaixinxin/ReDiffDet. Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006 |
CVPR | 3 |
| 2025 | MHTGR: Multi-Modal Hierarchical Temporal Graph Representation Learning for Ethereum Phishing DetectionabstractEthereum's swift development has elevated phishing scams to primary security concerns within blockchain networks. Current detection methods face three key challenges: insufficient hierarchical temporal modeling, inadequate pattern-aware structural recognition, and the lack of effective mechanisms to integrate multi-modal information. This paper presents an innovative approach for phishing detection using Multi-modal Hierarchical Temporal Graph Representation (MHTGR). Our method analyzes phishing behaviors by jointly considering temporal dynamics and structural topology of transaction data. First, we construct Hierarchical Transaction Graph Network (HTGN) to organize raw transaction records into structured graph representations. Then, multiple feature modalities are extracted through a Parallel Feature Extraction (PFE) module. Finally, these features are integrated via a Multi-modal Fusion (MMF) module for comprehensive phishing detection. Empirical evaluations conducted across multiple datasets from Ethereum demonstrate that the proposed method outperforms existing methods, providing effective solutions towards blockchain security. Shunrong Jiang, Yong Zhou 0003 |
ICPADS | 3 |
| 2025 | GSDet: Gaussian Splatting for Oriented Object DetectionabstractOriented object detection has advanced with the development of convolutional neural networks (CNNs) and transformers. However, modern detectors still rely on predefined object candidates, such as anchors in CNN-based methods or queries in transformer-based methods, which struggle to capture spatial information effectively. To address the limitations, we propose GSDet, a novel framework that formulates oriented object detection as Gaussian splatting. Specifically, our approach performs detection within a 3D feature space constructed from image features, where 3D Gaussians are employed to represent oriented objects. These 3D Gaussians are projected onto the image plane to form 2D Gaussians, which are then transformed into oriented boxes. Furthermore, we optimize the mean, anisotropic covariance, and confidence scores of these randomly initialized 3D Gaussians, using a decoder that incorporates 3D Gaussian sampling. Moreover, our method exhibits flexibility, enabling adaptive control and a dynamic number of Gaussians during inference. Experiments on 3 datasets indicate that GSDet achieves AP50 gains of 0.7% on DIOR-R, 0.3% on DOTA-v1.0, and 0.55% on DOTA-v1.5 when evaluated with adaptive control and outperforms mainstream detectors. Zeyu Ding 0010, Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Hancheng Zhu, Rui Yao 0006 |
IJCAI | 3 |
| 2025 | Beyond Individual and Point: Next POI Recommendation via Region-aware Dynamic Hypergraph with Dual-level ModelingabstractNext POI recommendation contributes to the prosperity of various intelligent location-based services. Existing studies focus on exploring sequential patterns and POI interactions using sequential and graph-based methods to enhance recommendation performance. However, they don't effectively exploit geographical information. In addition, methods that focus on modeling mobility patterns using individual limited data may suffer from data sparsity and the information cocoons problem. Moreover, most graph structures focus on adjacent nodes, failing to capture potential high-order associations among POIs. To address these challenges, we propose the Region-aware dynamic Hypergraph learning method with Dual-level interaction Modeling (ReHDM), which exploits users' dynamic mobility beyond individual and point. Specifically, ReHDM utilizes regional encoding to mine the potential spatial relationships among POIs with coarse-grained geographical information. By incorporating POI-level and trajectory-level associations within a hypergraph convolutional network, ReHDM comprehensively captures cross-user collaborative information. Furthermore, ReHDM captures not only dependencies among POIs within each trajectory for a single user, but also the high-order collaborative information across individual user trajectories and associated users' trajectories. Experimental results on three public datasets demonstrate the superiority of ReHDM to the state-of-the-art. Zhuo Gu, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Wen-Liang Du 0002 |
IJCAI | 4 |
| 2025 | Modality-Guided Dynamic Graph Fusion and Temporal Diffusion for Self-Supervised RGB-T TrackingabstractTo reduce the reliance on large-scale annotations, self-supervised RGB-T tracking approaches have garnered significant attention. However, the omission of the object region by erroneous pseudo-label or the introduction of background noise affects the efficiency of modality fusion, while pseudo-label noise triggered by similar object noise can further affect the tracking performance. In this paper, we propose GDSTrack, a novel approach that introduces dynamic graph fusion and temporal diffusion to address the above challenges in self-supervised RGB-T tracking. GDSTrack dynamically fuses the modalities of neighboring frames, treats them as distractor noise, and leverages the denoising capability of a generative model. Specifically, by constructing an adjacency matrix via an Adjacency Matrix Generator (AMG), the proposed Modality-guided Dynamic Graph Fusion (MDGF) module uses a dynamic adjacency matrix to guide graph attention, focusing on and fusing the object’s coherent regions. Temporal Graph-Informed Diffusion (TGID) models MDGF features from neighboring frames as interference, and thus improving robustness against similar-object noise. Extensive experiments conducted on four public RGB-T tracking datasets demonstrate that GDSTrack outperforms the existing state-of-the-art methods. The source code is available at https://github.com/LiShenglana/GDSTrack. Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Kunyang Sun, Bing Liu 0016, Zhiwen Shao, Jiaqi Zhao 0001 |
IJCAI | 3 |
| 2025 | Counterfactual Knowledge Maintenance for Unsupervised Domain AdaptationabstractTraditional unsupervised domain adaptation (UDA) struggles to extract rich semantics due to backbone limitations. Recent large-scale pre-trained visual-language models (VLMs) have shown strong zero-shot learning capabilities in UDA tasks. However, directly using VLMs results in a mixture of semantic and domain-specific information, complicating knowledge transfer. Complex scenes with subtle semantic differences are prone to misclassification, which in turn can result in the loss of features that are crucial for distinguishing between classes. To address these challenges, we propose a novel counterfactual knowledge maintenance UDA framework. Specifically, we employ counterfactual disentanglement to separate the representation of semantic information from domain features, thereby reducing domain bias. Furthermore, to clarify ambiguous visual information specific to classes, we maintain the discriminative knowledge of both visual and textual information. This approach synergistically leverages multimodal information to preserve modality-specific distinguishable features. We conducted extensive experimental evaluations on several public datasets to demonstrate the effectiveness of our method. The source code is available at https://github.com/LiYaolab/CMKUDA Yong Zhou 0003, Jiaqi Zhao 0001, Wen-Liang Du 0002, Rui Yao 0006, Bing Liu 0016 |
IJCAI | 2 |
| 2025 | RQFormer: Rotated Query Transformer for end-to-end oriented object detection
Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
Expert Syst. Appl. | 3 |
| 2025 | Facial Action Unit Detection by Adaptively Constraining Self-Attention and Causally Deconfounding Sample
Zhiwen Shao, Hancheng Zhu, Yong Zhou 0003, Xiang Xiang 0001, Bing Liu 0016, Rui Yao 0006, Lizhuang Ma |
Int. J. Comput. Vis. | 3 |
| 2025 | Automated message selection for robust Heterogeneous Graph Contrastive Learning
Rui Bing, Guan Yuan, Yong Zhou 0003, Qiuyan Yan |
Knowl. Based Syst. | 4 |
| 2025 | MoViM: A Hybrid CNN Vision Mamba Network for Lightweight Semantic Segmentation of Multimodal Remote Sensing ImagesabstractThe “Others” category in multimodal remote sensing images is characterized by high intra-class variability. Therefore, existing lightweight semantic segmentation models struggle with this category due to limitations in capturing both local details and global dependencies efficiently. We propose MoViM, a lightweight model that integrates a hybrid Vision Mamba (ViM) and CNN backbone to capture global contextual information and local details effectively. In addition, the MoViM also features an Inverted Stem for efficient multimodal fusion, a Global Semantics Extraction (GSE) module for enhanced global feature representation, and a Global-Local Feature Fusion (GLF) module for context-aware feature integration. Extensive experiments on WHU-OPT-SAR and Potsdam datasets demonstrate that MoViM achieves state-of-the-art performance, particularly in the “Others” category, while maintaining low computational complexity. Our codes are available at https://github.com/WenliangDu/MoViM. Wen-Liang Du 0002, Jiaqi Zhao 0001, Rui Yao 0006, Yong Zhou 0003 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | FA-MSVNet: multi-scale and multi-view feature aggregation methods for stereo 3D reconstruction
Yong Zhou 0003, Jiaqi Zhao 0001, Wen-Liang Du 0002, Rui Yao 0006 |
Multim. Tools Appl. | 2 |
| 2025 | MOL: Joint Estimation of Micro-Expression, Optical Flow, and Landmark via Transformer-Graph-Style ConvolutionabstractFacial micro-expression recognition (MER) is a challenging problem, due to transient and subtle micro-expression (ME) actions. Most existing methods depend on hand-crafted features, key frames like onset, apex, and offset frames, or deep networks limited by small-scale and low-diversity datasets. In this paper, we propose an end-to-end micro-action-aware deep learning framework with advantages from transformer, graph convolution, and vanilla convolution. In particular, we propose a novel F5C block composed of fully-connected convolution and channel correspondence convolution to directly extract local-global features from a sequence of raw frames, without the prior knowledge of key frames. The transformer-style fully-connected convolution is proposed to extract local features while maintaining global receptive fields, and the graph-style channel correspondence convolution is introduced to model the correlations among feature patterns. Moreover, MER, optical flow estimation, and facial landmark detection are jointly trained by sharing the local-global features. The two latter tasks contribute to capturing facial subtle action information for MER, which can alleviate the impact of insufficient training data. Extensive experiments demonstrate that our framework (i) outperforms the state-of-the-art MER methods on CASME II, SAMM, and SMIC benchmarks, (ii) works well for optical flow estimation and facial landmark detection, and (iii) can capture facial subtle muscle actions in local regions associated with MEs. Zhiwen Shao, Feiran Li, Yong Zhou 0003, Xuequan Lu, Yuan Xie 0006, Lizhuang Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Grid-distance-based selection for fine-grained object detection in aerial images
Jiaqi Zhao 0001, Qingfeng Ou, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006 |
Pattern Recognit. Lett. | 3 |
| 2025 | An Efficient Multi-View Heterogeneous Hypergraph Convolutional Network for Heterogeneous Information Network Representation LearningabstractHeterogeneous hypergraph neural networks are powerful tools to capture complex correlations among various nodes in Heterogeneous Information Networks (HINs). Despite satisfied performances of them, they are still plagued by the following problems: 1) They cannot capture the correlations in structural and semantic view at once, leading to topological information loss. 2) Due to the number of nodes being greater than the number of node types, node-level self-attention they used causes massive parameters and leads to high time consumption. 3) Interactions in meta-paths may be redundant, resulting in the correlations bias. To address the three issues, we propose an efficientMulti-ViewHeterogeneousHypergraphConvolutionalNetwork (MVH$^{2}$GCN). It first constructs relational and semantic hypergraphs based on different types of edges and meta-paths respectively, to represent the complex correlations in structural view and semantic view. Meanwhile, the clean semantic hypergraphs are generated by structure learning network to avoid redundancy. Then, an efficient hypergraph convolutional network is designed to learn node embeddings. By doing so, correlations in the two views are captured. Finally, the learned node embeddings from two views are aggregated via a gated embedding fusion module for downstream tasks. Experiment results demonstrate that MVH$^{2}$GCN is effective and efficient. Rui Bing, Guan Yuan, Senzhang Wang, Bohan Li 0001, Yong Zhou 0003 |
IEEE Trans. Big Data | 6 |
| 2025 | L2A: Learning Affinity From Attention for Weakly Supervised Continual Semantic SegmentationabstractDespite significant advances in continual semantic segmentation (CSS), they still rely on the pixel-level annotation to train models, which is time-consuming and labor-intensive. Continual learning from image-level labels is an emerging scheme in continual semantic segmentation to reduce the annotation cost. However, the incomplete and coarse pseudo-labels are insufficient to train a model to maintain a balance between stability and plasticity. To solve these issues, we propose a novel end-to-end framework based on Transformer, called L2A, for Weakly Supervised Continual Semantic Segmentation (WSCSS). In particular, to generate reliable annotations from the image-level supervision, we introduce a semantic affinity from multi-head self-attention (SA-MHSA) module to capture the semantic relationships among adjacent image coordinates. Subsequently, this acquired semantic affinity is employed to refine the initial pseudo labels of new classes trained with the image-level annotations. Furthermore, to minimize catastrophic forgetting, we propose a semantic drift compensation (SDC) strategy to optimize the pseudo-label generation process, which can effectively improve the alignment of object boundaries across both new and old categories. Comprehensive experiments conducted on the PASCAL VOC 2012 and COCO datasets demonstrate the superiority of our framework in existing WSCSS scenarios and a newly proposed challenge protocol, as well as remains competitive compared to the pixel-level supervised CSS methods. Hao Liu 0065, Yong Zhou 0003, Bing Liu 0016, Ming Yan 0007, Joey Tianyi Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Hierarchical Relation Learning for Few-Shot Semantic Segmentation in Remote Sensing ImagesabstractFew-shot semantic segmentation (FSS) aims to segment specific semantic classes in a query image using only a few annotated support samples. While FSS has gained significant attention in natural image processing, it remains underexplored in the more challenging domain of remote sensing images (RSIs). Existing FSS approaches for RSIs primarily focus on enhancing feature representations of support or query images through hierarchical/multi-level feature fusion. However, unlike fully supervised segmentation that relies on feature extraction and optimization, FSS requires segmenting the query image based on its relations with annotated support images. To address this need, we propose the concept of Hierarchical Relation Learning (HRL) to explore the intrinsic support-query relations, allowing for the direct refinement of target object appearances in the query image. Specifically, we propose a Hierarchical Relation Network (HRNet), which performs single-scale relation extraction at each network hierarchy and multi-scale relation aggregation across hierarchies. In addition, we construct a Bidirectional Hierarchical Loss (BHLoss) to guide HRNet training, providing targeted supervision at each hierarchy in both top-down and bottom-up directions, thus facilitating robust multi-scale relation learning across hierarchies. Comprehensive experiments on the iSAID-5i, DLRSD-5i, and LoveDA-2i datasets demonstrate the superiority of the proposed HRL. The code will be available at https://github.com/XinnHe/HRL. Xin He 0024, Yun Liu 0011, Yong Zhou 0003, Henghui Ding, Jiaqi Zhao 0001, Bing Liu 0016, Xudong Jiang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Hyperspectral Object Tracking With Dual-Stream PromptabstractHyperspectral images, rich in spectral details, offeradvantages for object tracking across diverse scenarios. Current hyperspectral tracking often fine-tunes parameters using pretrained RGB trackers, but this manner is suboptimal due to redundancy in spectral bands and limited training data. Existing hyperspectral trackers also underuse temporal information. To address these issues, we propose a unified spectral-spatiotemporal multimodal dual-stream prompt hyperspectral object tracking, named HDSP. We design a density clustering-based band selection module (BSM) to preserve spectral prompt information efficiently. Using the generated bands and temporal data as multimodal prompts, a dual-stream visual prompter is proposed. Designed multimodal dual-stream visual prompter (MDVP) transforms the multimodal input into a single modality, enhancing the foundational modality’s representation capabilities for hyperspectral tracking. Experiments on hyperspectral videos (HSVs) tracking datasets demonstrate that the proposed tracker achieves state-of-the-art performance. The source code is available athttps://github.com/rayyao/HDSP. Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | DDCI: Unsupervised Domain Adaptation for Remote Sensing Images Based on Diffusion Causal DistillationabstractThe distribution of remote sensing (RS) images can vary significantly due to seasonal changes and lighting conditions, making it difficult for deep learning models to generalize effectively across different RS datasets. This variation leads to a domain gap that hampers model performance when applied to new, unseen data. To tackle this challenge, we introduce DDCI, a novel unsupervised domain adaptation (UDA) framework designed to bridge the domain gap in RS image perception. Our framework consists of two key components, i.e., the adaptation diffusion distillation (ADD) module and the consistent causal intervention (CCI) module. The ADD module addresses the domain gap by aligning the source and target domains. It enhances the representation of the target domain by distilling semantic knowledge from the teacher model of the source domain. This process allows the target domain to benefit from the rich features of the source domain, leading to improved model generalization. The CCI module focuses on removing spurious correlations between domain-agnostic knowledge and domain-specific knowledge. By carefully considering the distinct characteristics of the target domain while preserving the specificity of the source domain, the CCI module ensures that only relevant, causal information is transferred between domains. This prevents overfitting to irrelevant domain-specific features and enhances model robustness. We demonstrate the effectiveness of the DDCI framework on RS scene classification tasks, utilizing four widely recognized RS datasets. Our results show significant performance improvements, underscoring the potential of this approach to boost the adaptability of deep learning models across diverse RS image datasets. Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | GLFRNet: Global-Local Feature Refusion Network for Remote Sensing Image Instance SegmentationabstractInstance segmentation is a significant way for remote sensing image (RSI) interpretation. The large number, sharp variation of sizes, and complex background of objects raise higher demands for instance segmentation models. The synergistic usage of global and local features has drawn great attention due to its superior performance but has not been fully explored in mainstream instance segmentation methods. In this work, a global-local feature refusion network (GLFRNet) with two fusion procedures is proposed to fully utilize coarse-grained and fine-grained features for RSI instance segmentation. In this model, the backbone integrates both convolutional neural network (CNN)-based and VMamba-based branches to extract local and global features, respectively. Three novel models are proposed to leverage the features adaptively, i.e., the cross-dim feature fusion (CDFF) module, the semantic complementary feature fusion (SCFF) module, and the guided feature refusion module (GFRM). The CDFF module is designed to aggregate features flexibly by fusing features from two backbones with different attention modules in the first fusion procedure. The GFRM and SCFF module are proposed in the refusion procedure to generate accurate segmentation results. Inspired by agent attention, the GFRM dynamically assembles detailed features for mask generation by refusing local and global features with the guidance of fusion results from CDFF. The SCFF module complements the significant features by enhancing and integrating global, local, and detailed features, and finally generates masks of instances. Extensive experiments demonstrate that GLFRNet outperforms the second-best model by 1.9, 1.3, and 0.3 in mask average precisions (APs) on NWPU VHR-10, WHU Building, and iSAID datasets. Jiaqi Zhao 0001, Yari Wang, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | ST-Mamba: Spatio-Temporal Synergistic Model for Remote Sensing Change DetectionabstractThe advancement of remote sensing and deep learning has spurred interest in high-resolution image change detection (CD). However, pseudo-changes in multi-temporal images, due to complex scenes and variable imaging conditions, often lead to significant misdetection in current methods. To address this problem, we propose a new CD framework: Spatio-Temporal Mamba (ST-Mamba), which consists of three key components. Firstly, a Mamba-based Feature Extraction Module (MFEM) is designed as the encoder to extract essential features from multi-temporal images by leveraging Mamba’s capability to capture inherent information in long data sequences. Secondly, a Spatio-Temporal Synergy Module (STSM) is developed to unify the background features of multi-temporal feature maps into a common domain by employing the state-space model for spatio-temporal modeling. Finally, a Spatio-Temporal Fusion Module (STFM) is created to guide the fusion of image features at different scales and across channels by utilizing a feature map of the unified background features. Experimental results on five widely used change detection datasets show significant improvements over current state-of-the-art methods. Jiaqi Zhao 0001, Jianpeng Xie 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | ECAKM: Efficient Conditional Anonymous Authentication Scheme With On-Chain Key Management in VANETsabstractConditional anonymous authentication can provide anonymity and traceability to Vehicular Ad-Hoc Networks (VANETs), which protects users’ privacy while resisting malicious users and false messages. However, existing schemes suffer from various disadvantages, such as unavailable batch verification, unrenewable user public keys/certificates, and untimely revocation. In this paper, we propose an efficient conditional anonymous authentication scheme with on-chain key management (ECAKM) in VANETs. To achieve lightweight authentication, we design an efficient Signature of Knowledge (SoK) and a batch verification algorithm. We also employ a Bloom filter on the chain to manage the information about revoked anonymous public keys to further improve the efficiency of our scheme. Moreover, we adopt hash chain technology to update users’ anonymous public keys and protect vehicles against linkage attacks. In addition, based on the blockchain and smart contract (SC), we can manage anonymous public keys of users efficiently and transparently. Security analysis and experimental results demonstrate that our scheme ensures conditional privacy with a reduced authentication overhead. Shunrong Jiang, Xiao Zhang 0047, Guohuai Sang, Haotian Chi, Yong Zhou 0003 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Adversarial Geometric Attacks for 3D Point Cloud Object Trackingabstract3D point cloud object tracking (3D PCOT) plays a vital role in applications such as autonomous driving and robotics. Adversarial attacks offer a promising approach to enhance the robustness and security of tracking models. However, existing adversarial attack methods for 3D PCOT seldom leverage the geometric structure of point clouds and often overlook the transferability of attack strategies. To address these limitations, this paper proposes an adversarial geometric attack method tailored for 3D PCOT, which includes a point perturbation attack module (non-isometric transformation) and a rotation attack module (isometric transformation). First, we introduce a curvature-aware point perturbation attack module that enhances local transformations by applying normal perturbations to critical points identified through geometric features such as curvature and entropy. Second, we design a Thompson sampling-based rotation attack module that applies subtle global rotations to the point cloud, introducing tracking errors while maintaining imperceptibility. Additionally, we design a fused loss function to iteratively optimize the point cloud within the search region, generating adversarially perturbed samples. The proposed method is evaluated on multiple 3D PCOT models and validated through black-box tracking experiments on benchmarks. For P2B, white-box attacks on KITTI reduce the success rate from 53.3% to 29.6% and precision from 68.4% to 37.1%. On NuScenes, the success rate drops from 39.0% to 27.6%, and precision from 39.9 to 26.8%. Black-box attacks show a transferability, with BAT showing a maximum 47.0% drop in success rate and 47.2% in precision on KITTI, and a maximum 22.5% and 27.0% on NuScenes. Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Bing Liu 0016, Abdulmotaleb El Saddik |
IEEE Trans. Multim. | 3 |
| 2025 | Similarity Regulation and Calibration Alignment for Weakly Supervised Text-Based Person Re-IdentificationabstractTraditional text-based person re-identification relies on identity labels. However, it is impossible to annotate large datasets, since identity annotation is expensive and time-consuming. Weakly supervised text-based person re-identification, where only text–image pairs are available without annotation of identities, is very practical in real life. While dealing with the weakly supervised person re-identification, two issues should be strengthed, i.e., alignment caused by different modal, and cross-modal matching ambiguity caused by the lack of identity labels. In this article, we propose a similarity regulation and calibration alignment (SRCA) framework, which consists of two unimodal encoders for images and text, respectively, and a multi-modal encoder for the masked language modeling task. First, a similarity regulation (SR) strategy is proposed to relax the strict one-to-one constraints for the local similarities between different pairs by introducing a novel soft objective. The soft objective can adjust hard objectives to achieve soft cross-modal alignment by establishing a many-to-many relationship between two modalities. Second, the calibration alignment (CA) module is proposed to improve intra-class compactness by modeling pseudo-label assignment as optimal transport. The ambiguity of cross-modal matching can be reduced by aligning features and pseudo-labels of different modalities and gradually calibrating the distribution of pseudo-labels. Experimental results show that our method has achieved obvious advantages compared with existing methods and also demonstrated competitive performance compared with fully supervised methods. Ao Fu, Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Syntactic-Conditional Diffusion Networks for Controllable Image CaptioningabstractCurrent diffusion model-based image captioning methods generally focus on generating descriptions in a non-autoregressive manner. Nevertheless, it is not trivial to employ such generative models to control the generation of discrete words while pursuing the balance between diversity and accuracy. Inspired by the success of continuous diffusions in image captioning, we introduce the Part-of-Speech (POS) information and classifier-free guidance into the diffusion model, and propose a novel controllable image captioning model, namely POS-Conditional Diffusion Networks (POSCD-Net), which consists of a Diffusion-based POS Generator (DPG) and a Diffusion-based Caption Generator (DCG). The DPG is built to produce diverse syntactic structures for each input image. The diverse POS sequences are further regarded as the control signals of the DCG, which produces the output sentences in a conditional diffusion process. In the DCG, a syntactic control module (SCM) is designed to strengthen the alignment progressively between words and the corresponding POS tags in a cascaded manner. Furthermore, to improve the controllability of POSCD-Net, the classifier-free guidance with learnable parameters is exploited to jointly optimize both the DPG and DCG in a non-autoregressive manner. Extensive experiments on the MSCOCO dataset demonstrate that our proposed method outperforms the state-of-the-art non-autoregressive counterparts and achieves promising performance compared with the autoregressive models. Bing Liu 0016, Hao Liu 0065, Yong Zhou 0003, Peng Liu 0013 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Historical Object-Aware Prompt Learning for Universal Hyperspectral Object TrackingabstractHyperspectral Object Tracking (HOT), utilizing rich spectral information from hyperspectral video (HSV), holds significant importance for object tracking. We identify that a major obstacle in improving HOT performance lies in effectively leveraging spectral and historical information. Furthermore, due to the mismatch in band dimensions between hyperspectral and RGB images, state-of-the-art RGB-based trackers struggle to adapt to unified HOT tasks. To address this, we propose a Historical Object-Aware Prompt Learning (HOPL) method for universal hyperspectral object tracking. Initially, we transform hyperspectral image ( \( N \) bands) into multiple sets of three bands with different combinations and feed them into a backbone network to generate base features. Subsequently, we introduce a historical object-aware prompter, where historical object-aware images are input to generate prompt features that enhance the representation of object information when combined with base features. Additionally, we design a band information fusion module to integrate the multiple sets of base features. By introducing historical object-aware prompts, HOPL significantly enhances tracking performance without retraining the backbone network. Experimental results on the HOT2023 dataset (comprising HSV with 25-band, 16-band, and 15-band wavelength ranges) and HOT2022 dataset validate the superiority of HOPL over state-of-the-art methods. The source code is available at https://github.com/rayyao/HOPL . Rui Yao 0006, Yong Zhou 0003, Fuyuan Hu, Jiaqi Zhao 0001, Zhiwen Shao |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | CCFL: Customized Client Federated Learning for Unsupervised Person Re-identificationabstractFederated learning-based person re-identification (Re-ID) aims to address the issue of data silos in surveillance systems caused by increasingly stringent regulations on sensitive data. However, due to differences in data collection locations, times, and scales, severe non-independent and identically distributed (non-IID) characteristics exist across different Re-ID datasets. Existing federated learning-based Re-ID methods often adopt a unified model structure, which prevents the model from adapting well to diverse data environments, thereby significantly degrading the overall Re-ID performance. To address the challenges of training neural networks on non-IID data across different datasets, we propose a customizable federated learning framework. First, customizable clients allow each organization to freely select suitable neural network training methods and model architectures based on local data scales and prior knowledge, thus improving training outcomes. Second, since traditional federated learning frameworks cannot achieve knowledge fusion through parameter exchange between models with different architectures, we introduce an independent model, referred to as the interaction model, specifically designed for knowledge exchange among clients. The interaction model learns parameters (knowledge) from local models on each client through distillation learning. Subsequently, the interaction model is uploaded to the server, where it undergoes parameter fusion (knowledge exchange) with interaction models from other clients. Finally, the interaction model, enriched with knowledge from other clients, guides local model training through knowledge distillation. It is worth noting that selecting a lightweight interaction model, while potentially impacting Re-ID performance, can significantly reduce communication costs between the server and clients. Yong Zhou 0003, Fayao Liu, Jiaqi Zhao 0001, Hancheng Zhu, Wen-Liang Du 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Image Cropping with Content and Composition Attribute-aware Global Relation ReasoningabstractImage cropping aims to find visually pleasing content in an image, which will enhance its aesthetic quality. Existing image cropping approaches mainly emphasize the geometric properties of images, such as composition and layout, neglecting the rich aesthetic information available from the physical attributes (e.g., content and themes), and background information beyond the foreground in images. Consequently, this article proposes an image cropping method based on the content and composition attribute-aware global relation reasoning, which aims at guiding the generation of cropped sub-images by exploring critical attributes based on content and composition as well as global object correlations that affect aesthetics in images. Particularly, to comprehensively introduce aesthetic information into image cropping, we capture feature representations reinforced by content and composition attributes simultaneously. The feature representations can strengthen the visual aesthetics of cropped sub-images. To make the cropped sub-images amply contain more global information, we introduce a global relation reasoning branch in the proposed cropping module, which can fully exploit the dependency relationship between the foreground and background in images. Extensive experiments on image cropping benchmarks demonstrate that our approach is superior to state-of-the-art image cropping methods. Hancheng Zhu, Yong Zhou 0003, Rui Yao 0006, Zhiwen Shao, Jiaqi Zhao 0001, Leida Li |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | High-level LoRA and hierarchical fusion for enhanced micro-expression recognition
Zhiwen Shao, Yong Zhou 0003, Xiang Xiang 0001, Jian Li 0054, Bing Liu 0016, Dit-Yan Yeung |
Vis. Comput. | 3 |
| 2024 | LRANet: Towards Accurate and Efficient Scene Text Detection with Low-Rank Approximation NetworkabstractRecently, regression-based methods, which predict parameterized text shapes for text localization, have gained popularity in scene text detection. However, the existing parameterized text shape methods still have limitations in modeling arbitrary-shaped texts due to ignoring the utilization of text-specific shape information. Moreover, the time consumption of the entire pipeline has been largely overlooked, leading to a suboptimal overall inference speed. To address these issues, we first propose a novel parameterized text shape method based on low-rank approximation. Unlike other shape representation methods that employ data-irrelevant parameterization, our approach utilizes singular value decomposition and reconstructs the text shape using a few eigenvectors learned from labeled text contours. By exploring the shape correlation among different text contours, our method achieves consistency, compactness, simplicity, and robustness in shape representation. Next, we propose a dual assignment scheme for speed acceleration. It adopts a sparse assignment branch to accelerate the inference speed, and meanwhile, provides ample supervised signals for training through a dense assignment branch. Building upon these designs, we implement an accurate and efficient arbitrary-shaped text detector named LRANet. Extensive experiments are conducted on several challenging benchmarks, demonstrating the superior accuracy and efficiency of LRANet compared to state-of-the-art methods. Code is available at: https://github.com/ychensu/LRANet.git Zhineng Chen, Zhiwen Shao, Yuning Du, Zhilong Ji, Jinfeng Bai, Yong Zhou 0003, Yu-Gang Jiang 0001 |
AAAI | 7 |
| 2024 | Isolate and Detect the Untrusted Driver with a Virtual BoxabstractIn kernel, the driver code is much more than the core code, thus having a larger attack surface. Especially for the untrusted drivers without source code, they may come from the hot-plug hardware or the user without security knowledge. Traditional isolation methods require analyzing source code to set checkpoints in the driver for control flow protection, which are not available for closed-source drivers. Evenworse, the existing isolation methods can only prevent the hijacked control flows entering/existing drivers, while they cannot discover the illegal control flows inside drivers. Although the kernel address space location randomization (KASLR) can defend against control flow hijacking, it can be bypassed by code probes. In response to these issues, this paper proposes a novel method Dbox to isolate and detect the untrusted drivers whose source code is unavailable. Dbox creates a light hypervisor to monitor and analyze the untrusted driver's behavior without relying on source code. It isolates the untrusted driver in a private space and dynamically changes its virtual space through a sliding space mechanism. Under the protection of Dbox, all control flows jumping to/from untrusted drivers can be detected. Experiments and analysis show that Dbox has good protection against code probes, kernel rootkits and code reuse attacks, and the overhead introduced to the operating system is less than 3.6% in general scenarios. Shunrong Jiang, Yong Zhou 0003, Yeh-Ching Chung |
CCS | 5 |
| 2024 | Learning from Reduced Labels for Long-Tailed DataabstractLong-tailed data is prevalent in real-world classification tasks and heavily relies on supervised information, which makes the annotation process exceptionally labor-intensive and time-consuming. Unfortunately, despite being a common approach to mitigate labeling costs, existing weakly supervised learning methods struggle to adequately preserve supervised information for tail samples, resulting in a decline in accuracy for the tail classes. To alleviate this problem, we introduce a novel weakly supervised labeling setting called Reduced Label. The proposed labeling setting not only avoids the decline of supervised information for the tail samples, but also decreases the labeling costs associated with long-tailed data. Additionally, we propose an straightforward and highly efficient unbiased framework with strong theoretical guarantees to learn from these Reduced Labels. Extensive experiments conducted on benchmark datasets including ImageNet validate the effectiveness of our approach, surpassing the performance of state-of-the-art weakly supervised methods. Source code is available at \hrefhttps://github.com/WilsonMqz/LTRL https://github.com/WilsonMqz/LTRL Meng Wei 0006, Zhongnian Li, Yong Zhou 0003, Xinzheng Xu |
ICMR | 3 |
| 2024 | Attribute-Driven Multimodal Hierarchical Prompts for Image Aesthetic Quality AssessmentabstractImage Aesthetic Quality Assessment (IAQA) aims to simulate users' visual perception to judge the aesthetic quality of images. In social media, users' aesthetic experiences are often reflected in their textual comments regarding the aesthetic attributes of images. To fully explore the attribute information perceived by users for evaluating image aesthetic quality, this paper proposes an image aesthetic quality assessment method based on attribute-driven multimodal hierarchical prompts. Unlike existing IAQA methods that utilize multimodal pre-training or straightforward prompts for model learning, the proposed method leverages attribute comments and quality-level text templates to hierarchically learn the aesthetic attributes and quality of images. Specifically, we first leverage users' aesthetic attribute comments to perform prompt learning on images. The learned attribute-driven multimodal features can comprehensively capture the semantic information of image aesthetic attributes perceived by users. Then, we construct text templates for different aesthetic quality levels to further facilitate prompt learning through semantic information related to the aesthetic quality of images. The proposed method can explicitly simulate users' aesthetic judgment of images to obtain more precise aesthetic quality. Experimental results demonstrate that the proposed IAQA method based on hierarchical prompts outperforms existing methods significantly on multiple IAQA databases. Our source code is public at https://github.com/GitHub-Ju/AMHP. Hancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Leida Li |
ACM Multimedia | 5 |
| 2024 | Remote sensing image semantic segmentation via class-guided structural interaction and boundary perceptionabstractExisting remote sensing semantic segmentation methods generally ignore the structural information of objects that is vital in the human visual recognition system. The absence of overall structural information often results in weak perceptions of subtle textures and fragmented predictions, especially for complex and variable ground object scenarios. Besides, they still suffer from the semantic ambiguity caused by the unclear object boundary features in remote sensing images. In this paper, we propose a novel remote sensing semantic segmentation framework, called CSBNet, which aims to enhance the capacity of class-guided structural interaction and boundary perception simultaneously. It consists of a class-guided structure interaction module (CSIM), a Transformer-based context aggregation module (TCAM) and a class-guided boundary supervision module (CBSM). The CSIM has the ability to progressively extract the class-specific structural features, i.e. , refining the structural information of each class by iteratively exchanging information between initial coarse class tokens and contexts. Meanwhile, the TCAM is constructed to provide CSIM with more discriminative multi-scale contexts without losing spatial features. In particular, the CBSM plays an auxiliary role, which applies the boundary information obtained from the class tokens to supervise the segmentation of boundary regions. When tested on the ISPRS dataset, LoveDA dataset, UAVid dataset, our method significantly outperforms the state-of-the-art remote sensing semantic segmentation approaches. Xin He 0024, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006 |
Expert Syst. Appl. | 2 |
| 2024 | P²SimiDedup: Privacy-Preserving and Similarity-Based Deduplication Scheme for Fog-Assisted Vehicular Crowdsensing SystemabstractThe rapid development of fog-assisted vehicular crowdsensing systems (FVCSs) enables real-time vehicular data sharing, but redundant and similar data in report results in unnecessary costs. However, previous studies only focus on duplicate reports and neglect deduplication of similar data. Besides, transmitting crowdsensing data in Internet of Vehicles (IoV) exposes vulnerabilities to offline brute-force and fake report attacks. In this article, we present P2SimiDedup, a scheme for secure deduplication of similar crowdsensing reports. Specifically, we develop cryptographic primitives and introduce an improved generalized deduplication technique (GreedyGD) to achieve secure deduplication over similar crowdsensing data. Then, we construct a two-level deduplication framework that can perform secure and efficient similar-based deduplication at fog nodes and cloud server. Besides, P2SimiDedup can ensure that only data requesters can decrypt and recover crowdsensing data. The security analysis and evaluation results demonstrate that P2SimiDedup can achieve privacy-preserving deduplication for similar crowdsensing reports with moderate computational, communication, and storage costs. Qiliang Zhang, Tom H. Luan, Yiliang Liu, Shunrong Jiang, Yong Zhou 0003 |
IEEE Internet Things J. | 6 |
| 2024 | Fine-grained semantic oriented embedding set alignment for text-based person search
Jiaqi Zhao 0001, Ao Fu, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006 |
Image Vis. Comput. | 3 |
| 2024 | A Mamba-Diffusion Framework for Multimodal Remote Sensing Image Semantic SegmentationabstractRecent advances in deep learning have made significant progress in multimodal remote sensing semantic segmentation. However, current methods face challenges in maintaining geometric consistency, particularly when dealing with large objects, resulting in fragmented segmentation masks. We propose a Mamba-diffusion framework to preserve geometric consistency in segmentation masks. This framework preserves geometric consistency by introducing a generative diffusion-based semantic segmentation pipeline and developing a Mamba-based multimodal fusion model. The fusion model fuses the multimodal images in multiple scales and scanning mechanisms by a double cross-fusion (DCF) module. Then, the cross-modal information is further integrated by a dual-splitting structured state-space (DS-S4) model. Finally, the diffusion-based segmentation pipeline predicts semantic masks by progressively refining random Gaussian noise, guided by fused multimodal features. Our experimental results, verified on WHU-OPT-SAR and Hunan datasets, demonstrate that the proposed framework surpasses state-of-the-art (SOTA) methods by a considerable margin. Our codes are available athttps://github.com/WenliangDu/MambaDiffusion. Wen-Liang Du 0002, Yang Gu 0005, Jiaqi Zhao 0001, Hancheng Zhu, Rui Yao 0006, Yong Zhou 0003 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | Filter pruning based on evolutionary algorithms for person re-identification
Jiaqi Zhao 0001, Ying Chen 0005, Yufeng Zhong 0001, Yong Zhou 0003, Rui Yao 0006, Lixu Zhang, Shixiong Xia |
Multim. Tools Appl. | 4 |
| 2024 | Multi-level self attention for unsupervised learning person re-identification
Jiaqi Zhao 0001, Yong Zhou 0003, Fayao Liu, Rui Yao 0006, Hancheng Zhu, Abdulmotaleb El Saddik |
Multim. Tools Appl. | 3 |
| 2024 | Efficient convolutional neural networks and network compression methods for object detection: a survey
Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Bing Liu 0016 |
Multim. Tools Appl. | 1 |
| 2024 | Joint facial action unit recognition and self-supervised optical flow estimation
Zhiwen Shao, Yong Zhou 0003, Feiran Li, Hancheng Zhu, Bing Liu 0016 |
Pattern Recognit. Lett. | 2 |
| 2024 | CT-Net: Arbitrary-Shaped Text Detection via Contour TransformerabstractContour based scene text detection methods have rapidly developed recently, but still suffer from inaccurate front-end contour initialization, multi-stage error accumulation, or deficient local information aggregation. To tackle these limitations, we propose a novel arbitrary-shaped scene text detection framework named CT-Net by progressive contour regression with contour transformers. Specifically, we first employ a contour initialization module that generates coarse text contours without any post-processing. Then, we adopt contour refinement modules to adaptively refine text contours in an iterative manner, which are beneficial for context information capturing and progressive global contour deformation. Besides, we propose an adaptive training strategy to enable the contour transformers to learn more potential deformation paths, and introduce a re-score mechanism that can effectively suppress false positives. Extensive experiments are conducted on four challenging datasets, which demonstrate the accuracy and efficiency of our CT-Net over state-of-the-art methods. Particularly, CT-Net achieves F-measure of 86.1 at 11.2 frames per second (FPS) and F-measure of 87.8 at 10.1 FPS for CTW1500 and Total-Text datasets, respectively. Zhiwen Shao, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Rui Yao 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Information Gap Narrowing for Point Cloud Few-Shot SegmentationabstractPoint-by-point labeling of point clouds is a very costly task. Previous meta-learning-based few-shot methods predict categories by calculating the distance between unlabeled data (query set) and the prototype calculated by a few of data with the label (support set), which can reduce the dependence of point cloud segmentation algorithms on large amounts of labeled data. But it ignores the category information gap caused by object diversity between the two types of data and forcing information transfer is ineffective. To address this issue, we propose a co-occurrent object mining module for mining co-occurring object information from support and query sets. Specifically, the capture of co-occurrent information is used to activate the feature that co-occurs between the support and query set in the high-dimensional feature space so that the prototype generated by computing the mean of support features is more similar to the query set. By reducing the object diversity within the same category, the information gap problem is gradually improved. In addition, we propose a point-attention module to refine the support set features before mining co-occurrent features. It can be widely embedded in the point cloud backbone network. The experimental results on two semantic segmentation datasets demonstrate that our method obtains an average 19.43% lead over the state-of-the-art methods in 4 different few-shot tasks, while inference is around 45 times faster. Guanyu Zhu, Yong Zhou 0003, Rui Yao 0006, Hancheng Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Dual-Stream Edge-Target Learning Network for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) is crucial in both military and civilian applications. However, challenges such as low contrast, low signal-to-noise ratio (SNR), and lack of shape and texture information limit the effectiveness of existing methods in capturing edge details and representing target areas. To address these issues, we propose the dual-stream edge-target learning network (DETL-Net) for IRSTD. This network enhances feature cross-fusion by learning edge details and target regions through a dual-stream framework, significantly improving detection performance. Specifically, we extract multilevel features of the image based on the encoder-decoder structure of U-Net and then reconstruct the feature map. In the decoder, we propose the dual-guided cross-fusion module (DGCFM) to capture edge details of small targets and global contextual features of the target region, achieving complementary advantages. The multiscale context fusion module (MCFM) within DGCFM uses central difference convolution to enhance local contrast and extract rich contextual details, thereby retaining edge information and enhancing overall target representation. In addition, we introduce the cross-dimension interactive aggregation attention module (CIAAM), which dynamically adjusts feature fusion weights across layers to effectively suppress noise and enhance the discrimination of small targets. These modules are sequentially interconnected to progressively refine edge details, and the acquired target features are subsequently utilized for predicting the final target mask via the segmentation head. Experiments on the NUAA-SIRST and IRSTD-1k datasets demonstrate that DETL-Net outperforms state-of-the-art (SOTA) methods. The source code is available athttps://github.com/rayyao/DETL-Net. Rui Yao 0006, Yong Zhou 0003, Jinqiu Sun, Zihang Yin, Jiaqi Zhao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | OrientedFormer: An End-to-End Transformer-Based Oriented Object Detector in Remote Sensing ImagesabstractOriented object detection in remote sensing images is a challenging task due to objects being distributed in multiorientation. Recently, end-to-end transformer-based methods have achieved success by eliminating the need for post-processing operators compared to traditional convolutional neural network (CNN)-based methods. However, directly extending transformers to oriented object detection presents three main issues: 1) objects rotate arbitrarily, necessitating the encoding of angles along with position and size; 2) the geometric relations of oriented objects are lacking in self-attention, due to the absence of interaction between content and positional queries; and 3) oriented objects cause misalignment, mainly between values and positional queries in cross-attention, making accurate classification and localization difficult. In this article, we propose an end-to-end transformer-based oriented object detector, consisting of three dedicated modules to address these issues. First, Gaussian positional encoding (PE) is proposed to encode the angle, position, and size of oriented boxes using Gaussian distributions. Second, Wasserstein self-attention is proposed to introduce geometric relations and facilitate interaction between content and positional queries by utilizing Gaussian Wasserstein distance scores. Third, oriented cross-attention is proposed to align values and positional queries by rotating sampling points around the positional query according to their angles. Experiments on six datasets DIOR-R, a series of DOTA, HRSC2016, and ICDAR2015 show the effectiveness of our approach. Compared with previous end-to-end detectors, the OrientedFormer gains 1.16 and 1.21 AP50 on DIOR-R and DOTA-v1.0, respectively, while reducing training epochs from$3\times $to$1\times $. The code is available athttps://github.com/wokaikaixinxin/OrientedFormer. Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Vehicular Edge Computing Meets Cache: An Access Control Scheme With Fair Incentives for Privacy-Aware Content DeliveryabstractVehicular Edge Computing (VEC) integrates mobile edge computing with traditional vehicular networks, which shifts the majority of computation and storage workload of resource-constrained vehicles to the edge nodes. The high mobility of vehicles usually leads to frequent network changes and connection interruptions, making data sharing more challenging in such dynamic and unstable environments. To address this issue, cache-based content delivery is considered a promising solution for efficient data sharing in VEC. However, access control and fair incentive distribution in privacy-aware data sharing are rarely taken into account in prior VEC-oriented studies. In this paper, we propose RFIP-VEC, a Revocable access control scheme with Fair Incentive for Privacy-aware content delivery in VEC. Specifically, to enable anonymous authentication and conditional revocation, we construct a secure group signature scheme with formally proved security guarantees. Subsequently, based on our group signature scheme, we design a two-layer access control framework by employing proxy re-encryption. We also establish an evolutionary game theory model to analyze the effectiveness and fairness of the fair incentive in our scheme. Thus, our scheme can achieve flexible access control and fair incentive distribution with the assistance of edge nodes. Security analysis and experimental results demonstrate that the proposed scheme can achieve security goals with affordable cost in terms of network performance in VEC. Shunrong Jiang, Guohuai Sang, Haiqin Wu, Yong Zhou 0003 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Privacy-Preserving and Fair Crowdsourcing Framework With Fine-Grained Reuse Based on BlockchainabstractCrowdsourcing has gained many developments and wide applications in our daily life. Traditional centralized crowdsourcing systems suffer from high management costs and low efficiency. The recent advance in the blockchain technology has enabled the construction of decentralized crowdsourcing systems, which can overcome the limitations of centralized systems and make crowdsourcing solution reuse possible. However, such systems also bring new security and privacy challenges. For instance, transactions on blockchain are publicly visible which can lead to privacy leakage of crowdsourcing users. Moreover, unfair exchange is a critical issue on these platforms. In this paper, we propose a privacy-preserving and fair crowdsourcing framework with fine-grained reuse based on blockchain to meet the security requirements for decentralized crowdsourcing. Specifically, we construct one-address-only (OAO) authentication to ensure the uniqueness of the participant’s address in the crowdsourcing process. Additionally, We design a submit-then-open method with commitments to resist the “free-riding" and “false-reporting" attacks. Thus, fair exchange between entities can be guaranteed. To ensure data confidentiality and fine-grained solution item reuse, we employ pairing-based cryptography to generate an encryption key and ensure flexible authorization reuse. We also adopt stealth authorization techniques to ensure privacy-preserving access authorization during the reuse phase. Finally, security analysis and implementation results have shown that the proposed framework can effectively achieve privacy-preserving and fair crowdsourcing as well as fine-grained crowdsourcing reuse. Specifically, the gas consumption in the reuse phase is reduced by approximately 49% to 81% compared to the normal operation, which significantly improves the efficiency of blockchain applications. Shunrong Jiang, Xiao Zhang 0047, Haiqin Wu, Yiliang Liu, Yong Zhou 0003 |
IEEE Trans. Netw. Serv. Manag. | 7 |
| 2024 | Black-box Attack against Self-supervised Video Object Segmentation Models with Contrastive LossabstractDeep learning models have been proven to be susceptible to malicious adversarial attacks, which manipulate input images to deceive the model into making erroneous decisions. Consequently, the threat posed to these models serves as a poignant reminder of the necessity to focus on the model security of object segmentation algorithms based on deep learning. However, the current landscape of research on adversarial attacks primarily centers around static images, resulting in a dearth of studies on adversarial attacks targeting Video Object Segmentation (VOS) models. Given that a majority of self-supervised VOS models rely on affinity matrices to learn feature representations of video sequences and achieve robust pixel correspondence, our investigation has delved into the impact of adversarial attacks on self-supervised VOS models. In response, we propose an innovative black-box attack method incorporating contrastive loss. This method induces segmentation errors in the model through perturbations in the feature space and the application of a pixel-level loss function. Diverging from conventional gradient-based attack techniques, we adopt an iterative black-box attack strategy that incorporates contrastive loss across the current frame, any two consecutive frames, and multiple frames. Through extensive experimentation conducted on the DAVIS 2016 and DAVIS 2017 datasets using three self-supervised VOS models and one unsupervised VOS model, we unequivocally demonstrate the potent attack efficiency of the black-box approach. Remarkably, theJ&Fmetric value experiences a significant decline of up to 50.08% post-attack. Ying Chen 0005, Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Bing Liu 0016, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Motion-Aware Self-Supervised RGBT Tracking with Multi-Modality Hierarchical TransformersabstractSupervised RGBT (SRGBT) tracking tasks need both expensive and time-consuming annotations. Therefore, the implementation of Self-Supervised RGBT (SSRGBT) tracking methods has become increasingly important. Straightforward SSRGBT tracking methods use pseudo-labels for tracking, but inaccurate pseudo-labels can lead to object drift, which severely affects tracking performance. This article proposes a self-supervised RGBT object tracking method (S2OTFormer) to bridge the gap between tracking methods supervised under pseudo-labels and ground truth labels. Firstly, to provide more robust appearance features for motion cues, we introduce a multi-modality hierarchical transformer (MHT) module for feature fusion. This module allocates weights to both modalities and strengthens the expressive capability of the MHT module through multiple nonlinear layers to fully utilize the complementary information of the two modalities. Secondly, in order to solve the problems of motion blur caused by camera motion and inaccurate appearance information caused by pseudo-labels, we introduce a motion-aware mechanism (MAM). The MAM extracts the average motion vectors from the previous multi-frame search frame features and constructs the consistency loss with the motion vectors of the current search frame features. The motion vectors of inter-frame objects are obtained by reusing the inter-frame attention map to predict coordinate positions. Finally, to further reduce the effect of inaccurate pseudo-labels, we propose an Attention-Based Multi-Scale Enhancement Module. By introducing cross-attention to achieve more precise and accurate object tracking, this module overcomes the receptive field limitations of traditional CNN tracking heads. We demonstrate the effectiveness of S2OTFormer on four large-scale public datasets through extensive comparisons as well as numerous ablation experiments. The source code is available at https://github.com/LiShenglana/S2OTFormer . Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Diverse Image Captioning via Panoptic Segmentation and Sequential Conditional Variational TransformerabstractRecently, transformer-based image captioning models have achieved significant performance improvement. However, due to the limitations of region visual features and deterministic projections between image space and caption space, existing methods still suffer from disentangled visual features and rigid sentences. To address these issues, we first introduce panoptic segmentation to extract the segmentation region features, which can effectively alleviate the visual confusion caused by the widely-adopted region visual features. Then, we propose a panoptic segmentation based sequential conditional variational transformer (PS-SCVT) framework for diverse image captioning, which not only accurately extracts the image visual representations by fusing the segmentation region features and object detection features, but has the ability of learning one-to-many mappings from image space to caption space. The experimental results demonstrate that our approach achieves better interpretability and generalization performance compared with the state-of-the-art diverse image captioning models. Bing Liu 0016, Jinfu Lu, Hao Liu 0065, Yong Zhou 0003, Dongping Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Diverse Image Captioning via Conditional Variational Autoencoder and Dual Contrastive LearningabstractDiverse image captioning has achieved substantial progress in recent years. However, the discriminability of generative models and the limitation of cross entropy loss are generally overlooked in the traditional diverse image captioning models, which seriously hurts both the diversity and accuracy of image captioning. In this article, aiming to improve diversity and accuracy simultaneously, we propose a novel Conditional Variational Autoencoder (DCL-CVAE) framework for diverse image captioning by seamlessly integrating sequential variational autoencoder with contrastive learning. In the encoding stage, we first build conditional variational autoencoders to separately learn the sequential latent spaces for a pair of captions. Then, we introduce contrastive learning in the sequential latent spaces to enhance the discriminability of latent representations for both image-caption pairs and mismatched pairs. In the decoding stage, we leverage the captions sampled from the pre-trained Long Short-Term Memory (LSTM), LSTM decoder as the negative examples and perform contrastive learning with the greedily sampled positive examples, which can restrain the generation of common words and phrases induced by the cross entropy loss. By virtue of dual constrastive learning, DCL-CVAE is capable of encouraging the discriminability and facilitating the diversity, while promoting the accuracy of the generated captions. Extensive experiments are conducted on the challenging MSCOCO dataset, showing that our proposed methods can achieve a better balance between accuracy and diversity compared to the state-of-the-art diverse image captioning models. Bing Liu 0016, Yong Zhou 0003, Rui Yao 0006, Zhiwen Shao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Multi-Modal LiDAR Point Cloud Semantic Segmentation with Salience Refinement and Boundary PerceptionabstractPoint cloud segmentation is essential for scene understanding, which provides advanced information for many applications, such as autonomous driving, robots, and virtual reality. To improve the accuracy and robustness of point cloud segmentation, many researchers have attempted to fuze camera images to complement the color and texture information. The common fusion strategy is the combination of convolutional operations with concatenation, element-wise addition or element-wise multiplication. However, conventional convolutional operators tend to confine the fusion of modal features within their receptive fields, which can be incomplete and limited. In addition, the inability of encoder–decoder segmentation networks to explicitly perceive segmentation boundary information results in semantic ambiguity and classification errors at object edges. These errors are further amplified in point cloud segmentation tasks, significantly affecting the accuracy of point cloud segmentation. To address the above issues, we propose a novel self-attention multi-modal fusion semantic segmentation network for point cloud semantic segmentation. Firstly, to effectively fuze different modal features, we propose a self-cross fusion module (SCF), which models long-range modality dependencies and transfers complementary image information to the point cloud to fully leverage the modality-specific advantages. Secondly, we design the salience refinement module (SR), which calculates the importance of channels in the feature maps and global descriptors to enhance the representation capability of salient modal features. Finally, we propose the local-aware anisotropy loss measure the element-level importance in the data and explicitly provide boundary information for the model, which alleviates the inherent semantic ambiguity problem in segmentation networks. Extensive experiments on two benchmark datasets demonstrate that our proposed method surpasses current state-of-the-art methods. Yong Zhou 0003, Zeming Xie, Jiaqi Zhao 0001, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Modeling and Analysis of Finite-Scale Clustered Backscatter Communication NetworksabstractBackscatter communication (BackCom) is an intriguing technology that enables devices to transmit information by reflecting environmental radio frequency signals while consuming ultra-low energy. Applying BackCom in the Internet of things (IoT) networks can effectively address the power-unsustainability issue of energy-constraint devices. Considering many practical IoT applications, networks are finite-scale and devices are needed to be deployed at hotspot regions organized in clusters to cooperate for specific tasks. This paper considers finite-scale clustered backscatter communication networks (F-CBackCom Nets). To ensure communications, this paper establishes a theoretic model to analyze the communication connectivity of F-CBackCom Nets. Different from prior studies analyzing the connectivity with a focus on the transmission pair located at the center of the network, this paper analyzes the connectivity of a transmission pair located in an arbitrary location, because the performance of transmission pairs potentially varies with their network location. Extensive simulations validate the accuracy of our analytical model. Our results show that the connectivity of a transmission pair can be affected by its network location. Our analytical model and results can offer beneficial implications for constructing F-CBackCom Nets. Qiu Wang 0001, Yong Zhou 0003, Hongning Dai, Guopeng Zhang, Muhammad Imran 0001, Nidal Nasser |
ICC | 2 |
| 2023 | Personalized Image Aesthetics Assessment with Attribute-guided Fine-grained Feature RepresentationabstractPersonalized image aesthetics assessment (PIAA) has gained increasing attention from researchers due to its ability to measure individual users' specific aesthetic experiences. However, most existing PIAA methods rely on holistic features or simplistic coding to characterize users' aesthetic preferences for images, and we believe that more rich explicit features are needed in modeling PIAA. Consequently, we propose an attribute-guided fine-grained feature-aware personalized image aesthetics assessment method, which can fully capture fine-grained features from multiple attributes to represent users' aesthetic preferences for images. To achieve this, we first build a fine-grained feature extraction (FFE) module to obtain the refined local features of image attributes to compensate for holistic features. The FFE module is then used to generate user-level features, which are combined with the image-level features to obtain user-preferred fine-grained feature representations. By training extensive users' PIAA tasks, the aesthetic distribution of most users can be transferred to the personalized scores of individual users. To enable our proposed model to learn more generalizable aesthetics among individual users, we incorporate the degree of dispersion between users' personalized scores and image aesthetic distribution as a coefficient in the loss function during model training. Experimental results on several PIAA databases show that our method outperforms existing mainstream PIAA methods, and can effectively infer users' personalized aesthetics of images. Hancheng Zhu, Zhiwen Shao, Yong Zhou 0003, Guangcheng Wang, Pengfei Chen 0003, Leida Li |
ACM Multimedia | 3 |
| 2023 | Modeling and analysis of directional energy harvesting and spectrum sharing communications in massive D2D networks
Qiu Wang 0001, Yong Zhou 0003, Hongning Dai |
Ad Hoc Networks | 2 |
| 2023 | Identity-invariant representation and transformer-style relation for micro-expression recognition
Zhiwen Shao, Feiran Li, Yong Zhou 0003, Hancheng Zhu, Rui Yao 0006 |
Appl. Intell. | 3 |
| 2023 | Info-FPN: An Informative Feature Pyramid Network for object detection in remote sensing images
Silin Chen, Jiaqi Zhao 0001, Yong Zhou 0003, Hanzheng Wang, Rui Yao 0006, Lixu Zhang, Yong Xue |
Expert Syst. Appl. | 3 |
| 2023 | Context-aware and part alignment for visible-infrared person re-identification
Jiaqi Zhao 0001, Hanzheng Wang, Yong Zhou 0003, Rui Yao 0006, Lixu Zhang, Abdulmotaleb El Saddik |
Image Vis. Comput. | 3 |
| 2023 | Visible and Infrared Object Tracking via Convolution-Transformer Network With Joint Multimodal Feature LearningabstractThe existing Transformer-based RGBT tracker mainly focus on the enhancement of features extracted by Convolutional Neural Network (CNN). The potential of Transformer in representation learning remains under-explored. In this paper, we propose a Convolution-Transformer network with joint multimodal feature learning, in which both representation learning and feature fusion leverage Transformer. Specifically, we use the multi-branch Convolution-Transformer feature extraction network to process the extraction task of local modality-independent features and global modality-shared features respectively. Several simplified Transformer encoder layers form the Transformer backbone network, which is more suitable for the real-time object tracking. Besides, we found that inter-modality correlation is an important factor for modality interactions and mutual exploitation. Therefore, we propose a Joint Multimodal Feature Learning (JMFL) module, which uses cross-attention to capture the dependencies of cross-modal and enhance multimodal fusion by bidirectional guidance of multimodal information. The proposed method is fully experimented on two large benchmark datasets and compared with some current well-performing methods. The experimental results show that the proposed method performs well in terms of tracking accuracy and speed. Jiazhu Qiu, Rui Yao 0006, Yong Zhou 0003, Peng Wang 0015, Yanning Zhang 0001, Hancheng Zhu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Adversarial learning-based skeleton synthesis with spatial-channel attention for robust gait recognition
Ying Chen 0005, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Dongjun Zhu |
Multim. Tools Appl. | 4 |
| 2023 | Unsupervised RGB-T object tracking with attentional multi-modal feature fusion
Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Jiaqi Zhao 0001, Zhiwen Shao |
Multim. Tools Appl. | 3 |
| 2023 | Class-imbalanced complementary-label learning via weighted loss
Meng Wei 0006, Yong Zhou 0003, Zhongnian Li, Xinzheng Xu |
Neural Networks | 2 |
| 2023 | Semi-supervised transformable architecture search for feature distillation
Man Zhang 0006, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006, Zhiwen Shao, Hancheng Zhu |
Pattern Anal. Appl. | 2 |
| 2023 | Facial Action Unit Detection via Adaptive Attention and RelationabstractFacial action unit (AU) detection is challenging due to the difficulty in capturing correlated information from subtle and dynamic AUs. Existing methods often resort to the localization of correlated regions of AUs, in which predefining local AU attentions by correlated facial landmarks often discards essential parts, or learning global attention maps often contains irrelevant areas. Furthermore, existing relational reasoning methods often employ common patterns for all AUs while ignoring the specific way of each AU. To tackle these limitations, we propose a novel adaptive attention and relation (AAR) framework for facial AU detection. Specifically, we propose an adaptive attention regression network to regress the global attention map of each AU under the constraint of attention predefinition and the guidance of AU detection, which is beneficial for capturing both specified dependencies by landmarks in strongly correlated regions and facial globally distributed dependencies in weakly correlated regions. Moreover, considering the diversity and dynamics of AUs, we propose an adaptive spatio-temporal graph convolutional network to simultaneously reason the independent pattern of each AU, the inter-dependencies among AUs, as well as the temporal dependencies. Extensive experiments show that our approach (i) achieves competitive performance on challenging benchmarks including BP4D, DISFA, and GFT in constrained scenarios and Aff-Wild2 in unconstrained scenarios, and (ii) can precisely learn the regional correlation distribution of each AU. Zhiwen Shao, Yong Zhou 0003, Jianfei Cai 0001, Hancheng Zhu, Rui Yao 0006 |
IEEE Trans. Image Process. | 2 |
| 2023 | Attention-guided Adversarial Attack for Video Object SegmentationabstractVideo Object Segmentation (VOS) methods have made many breakthroughs with the help of the continuous development and advancement of deep learning. However, the deep learning model is vulnerable to malicious adversarial attacks, which mislead the model to make wrong decisions by adding adversarial perturbation that humans cannot perceive to the input image. Threats to deep learning models remind us that video object segmentation methods are also vulnerable to attacks, thereby threatening their security. Therefore, we study adversarial attacks on the VOS task to better identify the vulnerabilities of the VOS method, which in turn provides an opportunity to improve its robustness. In this paper, we propose an attention-guided adversarial attack method, which uses spatial attention blocks to capture features with global dependencies to construct correlations between consecutive video frames, and performs multipath aggregation to effectively integrate spatial-temporal perturbation, thereby guiding the deconvolution network to generate adversarial examples with strong attack capability. Specifically, the class loss function is designed to enable the deconvolution network to better activate noise in other regions and suppress the activation related to the object class based on the enhanced feature map of the object class. At the same time, attentional feature loss is designed to enhance the transferability against attack. The experimental results on the DAVIS dataset show that the proposed attention-guided adversarial attack method can significantly reduce the segmentation accuracy of OSVOS, and the J & F mean on DAVIS 2016 can reach 73.6% drop rate. The generated adversarial examples are also highly transferable to other video object segmentation models. Rui Yao 0006, Ying Chen 0005, Yong Zhou 0003, Fuyuan Hu, Jiaqi Zhao 0001, Bing Liu 0016, Zhiwen Shao |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2023 | TextDCT: Arbitrary-Shaped Text Detection via Discrete Cosine Transform MaskabstractArbitrary-shaped scene text detection is a challenging task due to the variety of text changes in font, size, color, and orientation. Most existing regression based methods resort to regress the masks or contour points of text regions to model the text instances. However, regressing the complete masks requires high training complexity, and contour points are not sufficient to capture the details of highly curved texts. To tackle the above limitations, we propose a novel light-weight anchor-free text detection framework called TextDCT, which adopts the discrete cosine transform (DCT) to encode the text masks as compact vectors. Further, considering the imbalanced number of training samples among pyramid layers, we only employ a single-level head for top-down prediction. To model the multi-scale texts in a single-level head, we introduce a novel positive sampling strategy by treating the shrunk text region as positive samples, and design a feature awareness module (FAM) for spatial-awareness and scale-awareness by fusing rich contextual information and focusing on more significant features. Moreover, we propose a segmented non-maximum suppression (S-NMS) method that can filter low-quality mask regressions. Extensive experiments are conducted on four challenging datasets, which demonstrate our TextDCT obtains competitive performance on both accuracy and efficiency. Specifically, TextDCT achieves F-measure of 85.1 at 17.2 frames per second (FPS) and F-measure of 84.9 at 15.1 FPS for CTW1500 and Total-Text datasets, respectively. Zhiwen Shao, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Rui Yao 0006 |
IEEE Trans. Multim. | 3 |
| 2023 | Weakly Supervised Few-Shot Semantic Segmentation via Pseudo Mask Enhancement and Meta LearningabstractFew shot semantic segmentation has been proposed to enhance the generalization ability of traditional models with limited data. Previous works mainly focus on the supervised tasks, while limited amount of work is explored for the weakly supervised tasks. Weakly supervised semantic segmentation has become an active research area because weakly supervised labels effectively reduce the annotation cost of visual tasks. To this end, we propose a weakly supervised few-shot semantic segmentation model based on the meta learning framework, which utilizes prior knowledge and adjusts itself according to new tasks. Thereupon then, the proposed network is capable of both high efficiency and generalization ability to new tasks. In the pseudo mask generation stage, we develop a WRCAM method with the channel-spatial attention mechanism to refine the coverage size of targets in pseudo masks. In the few-shot semantic segmentation stage, the optimization based meta learning method is used to realize few-shot semantic segmentation by virtue of the refined pseudo masks. The experimental results show that the proposed method not only significantly outperforms weakly supervised SOTA methods, but also could be comparative to some supervised SOTA methods. Man Zhang 0006, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006, Zhiwen Shao, Hancheng Zhu |
IEEE Trans. Multim. | 2 |
| 2023 | Spatial-Channel Enhanced Transformer for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a challenging task in computer vision, aiming at matching people across images from visible and infrared modalities. The widely used VI-ReID framework consists of a convolution neural backbone network that extracts the visual features, and a feature embedding network to project heterogeneous features to the same feature space. However, many studies based on the existing pre-trained models neglect potential correlations between different locations and channels within a single sample during the feature extraction. Inspired by the success of the Transformer in computer vision, we extend it to enhance feature representation for VI-ReID. In this paper, we propose a discriminative feature learning network based on a visual Transformer (DFLN-ViT) for VI-ReID. Firstly, to capture long-term dependencies between different locations, we propose a spatial feature awareness module (SAM), which utilizes a single-layer Transformer with a novel patch-embedding strategy to encode location information. Secondly, to refine the representation at each channel, we design a channel feature enhancement module (CEM). The CEM treats the features of each channel as a sequence of Transformer inputs, taking advantage of the Transformer's ability to model long-term dependencies. Finally, we propose a Triplet-aided Hetero-Center (THC) loss to learn more discriminative feature representation by balancing the cross-modality distance and intra-modality distance of the center. The experimental results on two datasets show that our method can significantly improve the VI-ReID performance, outperforming most state-of-the-art methods. Jiaqi Zhao 0001, Hanzheng Wang, Yong Zhou 0003, Rui Yao 0006, Silin Chen, Abdulmotaleb El Saddik |
IEEE Trans. Multim. | 3 |
| 2023 | Learning Personalized Image Aesthetics From Subjective and Objective AttributesabstractDue to the widespread popularity of social media, researchers have developed a strong interest in learning the personalized image aesthetics of online users. Personalized image aesthetics assessment (PIAA) aims to study the aesthetic preferences of individual users for images, which should be affected by the properties of both users and images. Existing PIAA approaches usually use the generic aesthetics learned from images as a prior model and adapt it to PIAA models through a small number of data annotated by individual users. However, the prior model merely learns the objective attributes of images, which is agnostic to the subjective attributes of users, complicating efficient learning of the personalized image aesthetics of individual users. Therefore, we propose a personalized image aesthetics assessment method that integrates the subjective attributes of users and objective attributes of images simultaneously. To characterize these two attributes jointly, an attribute extraction module is introduced to learn users’ personality traits and image aesthetic attributes. Then, an aesthetic prior model is built from numerous individual users’ annotated data, which leverages the personality traits of users and the aesthetic attributes of rated images as prior knowledge to model both the image aesthetic distribution and users’ residual scores relative to generic aesthetics simultaneously. Finally, a PIAA model is obtained by fine-tuning the aesthetic prior model with an individual user’s annotated data. Experiments demonstrate that the proposed method is superior to existing PIAA methods in learning individual users’ personalized image aesthetics. Hancheng Zhu, Yong Zhou 0003, Leida Li, Yandong Guo |
IEEE Trans. Multim. | 2 |
| 2023 | Cross-Class Bias Rectification for Point Cloud Few-Shot SegmentationabstractThe point cloud is a densely distributed 3D (three-dimensional) data, and annotating the point cloud is a time-consuming and labor-intensive work. The existing semantics segmentation work adopts few-shot learning to reduce the dependence on labeling samples while improving the generalization of the model to new categories. Since point clouds are 3D structures with rich geometric features, even objects of the same category have feature differences that cannot be ignored. Therefore, a few samples (support set) used to train the model do not cover all the features of this category. There is a distribution difference between the support samples and the samples used to verify the model performance (query set). In this paper, we propose an efficient point cloud few-shot segmentation method based on prototypes for bias rectification. A prototype is a vector representation of a category in the metric space. To make the prototype representation of the support set closer to the query set features, we define a feature bias term and reduce the distribution distance between the two sets by fusing the support set features and the bias term. On this basis, we design a feature cross-reference module. By mining the co-occurring features of the support and query sets, it can generate a more representative prototype which captures the overall features of the point cloud. Extensive experiments on two challenging datasets demonstrate that our method outperforms the state-of-the-art method by an average of 3.31$\%$in several N-way K-shot tasks, and achieves approximately 200 times faster reasoning speed. Our code is available athttps://github.com/964918993/2CBR. Guanyu Zhu, Yong Zhou 0003, Rui Yao 0006, Hancheng Zhu |
IEEE Trans. Multim. | 2 |
| 2023 | PACM: Privacy-Preserving Authentication Scheme With on-Chain Certificate Management for VANETsabstractPrivacy-preserving authentication is designed to protect vehicular ad-hoc networks (VANETs) from illegitimate users and fake messages while maintaining the privacy of legitimate users’ identities. However, existing authentication schemes have disadvantages such as non-transparent certificate issuance and revocation, high identity authentication and certificate revocation overhead. In this paper, we propose an efficient privacy-preserving authentication scheme with on-chain certificate management (PACM) in VANETs, where the service manager (SM) of each domain serves as a node of the blockchain to build a distributed system. Specifically, based on elliptic curve cryptography (ECC) and exclusive-OR operations, we achieve secure and lightweight mutual authentication between vehicles and roadside units (RSUs) by regularly updated pseudonyms. Then, we adopt the blockchain to record the issuance and revocation of all certificates, which makes SM’s activities transparent. Moreover, we introduce the counting garbled bloom filter (CGBF) to enable fast query and revocation of certificates. Besides, we design a non-forgeable and non-repudiable billing mechanism based on the hash chain technology. Security analysis and experimental results show that PACM achieves stronger security with less overhead. Guohuai Sang, Yiliang Liu, Haiqin Wu, Yong Zhou 0003, Shunrong Jiang |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2023 | Distilled Meta-learning for Multi-Class Incremental LearningabstractMeta-learning approaches have recently achieved promising performance in multi-class incremental learning. However, meta-learners still suffer from catastrophic forgetting, i.e., they tend to forget the learned knowledge from the old tasks when they focus on rapidly adapting to the new classes of the current task. To solve this problem, we propose a novel distilled meta-learning (DML) framework for multi-class incremental learning that integrates seamlessly meta-learning with knowledge distillation in each incremental stage. Specifically, during inner-loop training, knowledge distillation is incorporated into the DML to overcome catastrophic forgetting. During outer-loop training, a meta-update rule is designed for the meta-learner to learn across tasks and quickly adapt to new tasks. By virtue of the bilevel optimization, our model is encouraged to reach a balance between the retention of old knowledge and the learning of new knowledge. Experimental results on four benchmark datasets demonstrate the effectiveness of our proposal and show that our method significantly outperforms other state-of-the-art incremental learning methods. Hao Liu 0065, Zhaoyu Yan, Bing Liu 0016, Jiaqi Zhao 0001, Yong Zhou 0003, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Cyclic Self-attention for Point Cloud RecognitionabstractPoint clouds provide a flexible geometric representation for computer vision research. However, the harsh demands for the number of input points and computer hardware are still significant challenges, which hinder their deployment in real applications. To address these challenges, we design a simple and effective module named cyclic self-attention module (CSAM). Specifically, three attention maps of the same input are obtained by cyclically pairing the feature maps, thus exploring the features sufficiently of the attention space of the original input. CSAM can adequately explore the correlation between points to obtain sufficient feature information despite the multiplicative decrease in inputs. Meanwhile, it can direct the computational power to the more essential features, relieving the burden on the computer hardware. We build a point cloud classification network by simply stacking CSAM called cyclic self-attention network (CSAN). We also propose a novel framework for point cloud semantic segmentation called full cyclic self-attention network (FCSAN). By adaptively fusing the original mapping features and the CSAM extracted features, it can better capture the context information of point clouds. Extensive experiments on several benchmark datasets show that our methods can achieve competitive performance in classification and segmentation tasks. Guanyu Zhu, Yong Zhou 0003, Rui Yao 0006, Hancheng Zhu, Jiaqi Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Query Integrity Meets Blockchain: A Privacy-Preserving Verification Framework for Outsourced Encrypted DataabstractCloud outsourcing provides flexible storage and computation services for data users in a low cost, but it brings many security threats as the cloud server may not be fully trusted. Previous secure outsourcing solutions mostly assume that the server is honest-but-curious while the adversary model of a malicious server that may return incorrect results is rarely explored. Moreover, with the increasing popularity of verifiable computations, existing verification schemes are yet not efficient and cannot cater to different scenarios in practice. In this paper, we propose a blockchain-based verifiable search framework in the adversarial cloud outsourcing context. When outsourcing the encrypted data to the cloud or Interplanetary File System (IPFS), we also store the encrypted data index in a decentralized blockchain (i.e., Ethereum in this paper) which is public and cannot be modified. Once a user is authorized, he/she can flexibly obtain the query results and efficiently check the query integrity via the pre-deployed smart contract, without the need of the data owner being online. Moreover, for user's privacy protection, we construct a stealth authorization scheme to deliver the access authorization without any identity disclosure. Finally, theoretical analysis and performance evaluation validate the security and efficiency of our proposed framework. Shunrong Jiang, Jianqing Liu, Yiliang Liu, Liangmin Wang 0001, Yong Zhou 0003 |
IEEE Trans. Serv. Comput. | 6 |
| 2022 | Show, Deconfound and Tell: Image Captioning with Causal InferenceabstractThe transformer-based encoder-decoder framework has shown remarkable performance in image captioning. However, most transformer-based captioning methods ever overlook two kinds of elusive confounders: the visual confounder and the linguistic confounder, which generally lead to harmful bias, induce the spurious correlations during training, and degrade the model generalization. In this paper, we first use Structural Causal Models (SCMs) to show how two confounders damage the image captioning. Then we apply the backdoor adjustment to propose a novel causal inference based image captioning (CIIC) framework, which consists of an interventional object detector (IOD) and an interventional transformer decoder (ITD) to jointly confront both confounders. In the encoding stage, the IOD is able to disentangle the region-based visual features by deconfounding the visual confounder. In the decoding stage, the ITD introduces causal intervention into the transformer decoder and deconfounds the visual and linguistic confounders simultaneously. Two modules collaborate with each other to alleviate the spurious correlations caused by the unobserved confounders. When tested on MSCOCO, our proposal significantly outperforms the state-of-the-art encoder-decoder models on Karpathy split and online test split. Code is published in https://github.com/CUMTGG/CIIC. Bing Liu 0016, Xu Yang 0021, Yong Zhou 0003, Rui Yao 0006, Zhiwen Shao, Jiaqi Zhao 0001 |
CVPR | 4 |
| 2022 | HiFi-SVC: Fast High Fidelity Cross-Domain Singing Voice ConversionabstractThis paper presents HiFi-SVC, a small cross-domain singing voice conversion model for generating high-fidelity 22.05 kHz singing voices. Building on state-of-the-art neural vocoder HiFi-GAN and a convolution-based module for modeling F0, HiFi-SVC can be trained end-to-end with either speech or singing data, achieving better voice similarity on two of the datasets than FastSVC while using slightly smaller number of parameters. We also propose a pitch adjustment method for improving conversion quality. Yong Zhou 0003, Xiangju Lu |
ICASSP | 1 |
| 2022 | Spatial hierarchy perception and hard samples metric learning for high-resolution remote sensing image object detection
Dongjun Zhu, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Ying Chen 0005 |
Appl. Intell. | 4 |
| 2022 | Edge-aware and spectral-spatial information aggregation network for multispectral image semantic segmentation
Di Zhang 0020, Jiaqi Zhao 0001, Yong Zhou 0003, Boyu Shi, Rui Yao 0006 |
Eng. Appl. Artif. Intell. | 4 |
| 2022 | Multi-granularity semantic alignment distillation learning for remote sensing image semantic segmentation
Di Zhang 0020, Yong Zhou 0003, Jiaqi Zhao 0001, Zhongyuan Yang, Rui Yao 0006, Huifang Ma |
Frontiers Comput. Sci. | 2 |
| 2022 | Multi-source collaborative enhanced for remote sensing images semantic segmentation
Jiaqi Zhao 0001, Di Zhang 0020, Boyu Shi, Yong Zhou 0003, Rui Yao 0006, Yong Xue |
Neurocomputing | 4 |
| 2022 | Point cloud recognition based on lightweight embeddable attention module
Guanyu Zhu, Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Man Zhang 0006 |
Neurocomputing | 2 |
| 2022 | Performance on Cluster Backscatter Communication Networks With Coupled InterferencesabstractThis article presents an analytical model to analyze the communication performance of cluster backscatter communication networks (CBackCom Nets) by considering their unique interferences. In CBackCom Nets, interferences are from both backscatter transmitters (BTs) and carrier emitters (CEs), i.e., RF signal emitters. Because BTs are distributed in clusters around CEs, interferences from BTs and interferences from CEs constitute coupled interferences. In addition, since BTs conduct backscatter communications by reflecting RF signals from CEs, interfering signals from BTs, and interfering signals from CEs are power-correlated, leading to the particularity and complexity of coupled interferences of CBackCom Nets. In contrast to previous studies that analyze the performance of CBackCom Nets ignoring coupled interferences, this article develops a novel interference analysis approach to analyze their coupled interferences, and then analyze performance, including coverage probability and spatial throughput of a cluster. Our numerical results show that our analytical model can obtain more accurate results than prior analytical models. In addition, our results reveal the relationship between the communication performance and multiple factors, such as the node density, the energy harvesting model, the interferences from CEs, and the cluster size, offering insightful implications for constructing and configuring CBackCom Nets. Qiu Wang 0001, Yong Zhou 0003, Hongning Dai, Guopeng Zhang, Wei Zhang 0001 |
IEEE Internet Things J. | 2 |
| 2022 | Personality modeling from image aesthetic attribute-aware graph representation learning
Hancheng Zhu, Yong Zhou 0003, Qiaoyue Li, Zhiwen Shao |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | A Semi-Supervised Image-to-Image Translation Framework for SAR-Optical Image MatchingabstractSynthetic Aperture Radar (SAR) and optical image matching aims to acquire correspondences from a certain pair of SAR and optical images. Recent advances in the image-to-image translation provided a way to simplify the SAR-optical image matching into the SAR-SAR or optical-optical image matchings. Existing image-to-image translations mainly focus on supervised or unsupervised learning. However, gathering sufficient amounts of aligned training data for supervised learning is challenging, while unsupervised learning cannot guarantee enough correct correspondences. In this work, we investigate the applicability of semi-supervised image-to-image translation for SAR-optical image matching such that both aligned and unaligned SAR-optical images could be used. To this end, we combine the benefits of both supervised and unsupervised well-known image-to-image translation methods, i.e., Pix2pix and CycleGAN, and propose a simple yet effective semi-supervised image-to-image translation framework. Through extensive experimental comparisons to baseline methods, we verify the effectiveness of the proposed framework in both semi-supervised and fully-supervised settings. Our codes are available at https://github.com/WenliangDu/Semi-I2I. Wen-Liang Du 0002, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao, Xiaolin Tian 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Semantic Segmentation of Remote-Sensing Images Based on Multiscale Feature Fusion and Attention RefinementabstractIn recent years, the automatic extraction of remote-sensing image information has attracted full attention. However, the particularity of remote-sensing images and the scarcity of data sets with label information have brought new challenges to existing methods. Therefore, we develop a lightweight semantic segmentation network based onmultiscale feature fusion (MFF) and attention refinement (MFFANet). Our network relies on three crucial modules for improved performance. The multiscale attention refinement module strengthens the representation ability of feature maps extracted by the deep residual network. The MFF module aggregates the information carried by the high-level and low-level features while restoring the image resolution. Furthermore, the boundary enhancement module captures boundary details to solve the semantic ambiguity problem. We achieve 83.5% mean intersection over union (MIoU) on the Urban Semantic 3-D (US3D) data set and 69.3% MIoU on the Vaihingen data set with only 8.2M parameters. Xin He 0024, Yong Zhou 0003, Jiaqi Zhao 0001, Man Zhang 0006, Rui Yao 0006, Bing Liu 0016 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Semisupervised Multiscale Generative Adversarial Network for Semantic Segmentation of Remote Sensing ImageabstractSemantic segmentation of remote sensing images based on deep neural networks has gained wide attention recently. Although many methods have achieved amazing performance, they need large amounts of labeled images to distinguish the differences in angle, color, size, and other aspects for small targets in remote sensing data sets. However, with a few labeled images, it is difficult to extract the key features of small targets. We propose a semisupervised multiscale generative adversarial network (GAN), which not only utilizes the multipath input and atrous spatial pyramid pooling (ASPP) module but leverages unlabeled images and semisupervised learning strategy to improve the performance of small target segmentation in semantic segmentation when labeled data amount is small. Experimental results show that our model outperforms state-of-the-art methods with insufficient labeled data. Bing Liu 0016, Yong Zhou 0003, Jiaqi Zhao 0001, Shixiong Xia, Yuancan Yang, Man Zhang 0006, Liu Ming Ming |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Few-Shot Object Detection via Context-Aware Aggregation for Remote Sensing ImagesabstractFew-shot object detection methods have made prodigious progress in recent years. However, these methods are designed for optical images at a single scale, which leads to significantly degraded detection performance due to object scale variation of remote sensing images. In this letter, we propose a few-shot object detection method for the problem of scale variation in remote sensing images. More specifically, our model contains two main components: a context-aware pixel aggregation (CPA) that allows the model to adapt to objects at different scales through different scale convolution and a context-aware feature aggregation (CFA) that enhances context awareness to obtain more semantic information through a graph convolution network (GCN). Experiments on the DIOR dataset demonstrate that our model can achieve a satisfying detection performance on remote sensing images, and our model performs significantly better than the state-of-the-art model. Yong Zhou 0003, Jiaqi Zhao 0001, Hancheng Zhu, Rui Yao 0006, Wen-Liang Du 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Fine-Grained Feature Enhancement for Object Detection in Remote Sensing ImagesabstractRecently, object detection in aerial images has ushered in a new challenge—a new benchmark for fine-grained object recognition in high-resolution remote sensing imagery called FAIR1M has been proposed. Fine-grained categories usually have smaller inter class differences and intra-class similarities, which is more difficult to classify with existing object detectors. To address this problem, we propose two enhanced strategies on the current two-stage object detection algorithm. The first strategy uses attention-based group feature enhancement called group enhance module (GEM). By extending and grouping feature channels, the model can improve the ability to extract various discriminative features. The second strategy is to emphasize the sub-saliency feature learning, avoiding the network only focusing on the most significant part of the feature and ignoring the other parts. Our method is easy to implement and effective, and experiments show that our method can improve the Oriented regions with convolutional neural networks features (R-CNN) by about 1.45 mAP on the FAIR1M benchmark. Yong Zhou 0003, Sifan Wang, Jiaqi Zhao 0001, Hancheng Zhu, Rui Yao 0006 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Efficient lightweight video person re-identification with online difference discrimination module
Cunyuan Gao, Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Fuyuan Hu |
Multim. Tools Appl. | 3 |
| 2022 | Survey for person re-identification based on coarse-to-fine feature learning
Minjie Liu, Jiaqi Zhao 0001, Yong Zhou 0003, Hancheng Zhu, Rui Yao 0006, Ying Chen 0005 |
Multim. Tools Appl. | 3 |
| 2022 | ResT-ReID: Transformer block-based residual learning for person re-identification
Ying Chen 0005, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Dongjun Zhu, Dongjingdian Liu |
Pattern Recognit. Lett. | 4 |
| 2022 | Learning image aesthetic subjectivity from attribute-aware relational reasoning network
Hancheng Zhu, Yong Zhou 0003, Rui Yao 0006, Guangcheng Wang, Yuzhe Yang 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | Spatial-Temporal Based Multihead Self-Attention for Remote Sensing Image Change DetectionabstractThe neural network-based remote sensing image change detection method faces a large amount of imaging interference and severe class imbalance problems under high-resolution conditions, which bring new challenges to the accuracy of the detection network. In this work, to address the imaging interference caused by different imaging angles and times, the siamese strategy and multi-head self-attention mechanism are used to reduce the imaging differences between the dual-temporal images and fully exploit the inter-temporal information. Secondly, a learnable multi-part feature learning module is used to adaptively exploit features from different scales to obtain more comprehensive features. Finally, a mixed loss function strategy is used to ensure that the network converges effectively and excludes the adverse interference of a large number of negative samples to the network. Extensive experiments show that our method outperforms numerous methods on LEVIR-CD, WHU, and DSIFN datasets. Yong Zhou 0003, Fengkai Wang, Jiaqi Zhao 0001, Rui Yao 0006, Silin Chen, Heping Ma |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | FVC-Dedup: A Secure Report Deduplication Scheme in a Fog-Assisted Vehicular Crowdsensing SystemabstractIt is observed that modern vehicles are becoming more and more powerful in computing, communications, and storage capacity. By interacting with other vehicles or with local infrastructures (i.e., fog) such as road-side units, vehicles and fog devices can collaboratively provide services like crowdsensing in an efficient and secure way. Unfortunately, it is hard to develop a secure and privacy-preserving crowdsensing report deduplication mechanism in such a system. In this article, we propose a scheme FVC-Dedup to address this challenge. Specifically, we develop cryptographic primitives to realize secure task allocation and guarantee the confidentiality of crowdsensing reports. During the report submission, we improve the message-lock encryption (MLE) scheme to realize privacy-preserving report deduplication and resist the fake duplicate attacks. Besides, we construct a novel signature scheme to achieve efficient signature aggregation and record the contributions of each participant fairly without knowing the crowdsensing data. The security analysis and performance evaluation demonstrate that FVC-Dedup can achieve secure and privacy-preserving report deduplication with moderate computing and communication overhead. Shunrong Jiang, Jianqing Liu, Yong Zhou 0003, Yuguang Fang |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2022 | Swin Transformer Embedding UNet for Remote Sensing Image Semantic SegmentationabstractGlobal context information is essential for the semantic segmentation of remote sensing (RS) images. However, most existing methods rely on a convolutional neural network (CNN), which is challenging to directly obtain the global context due to the locality of the convolution operation. Inspired by the Swin transformer with powerful global modeling capabilities, we propose a novel semantic segmentation framework for RS images called ST-U-shaped network (UNet), which embeds the Swin transformer into the classical CNN-based UNet. ST-UNet constitutes a novel dual encoder structure of the Swin transformer and CNN in parallel. First, we propose a spatial interaction module (SIM), which encodes spatial information in the Swin transformer block by establishing pixel-level correlation to enhance the feature representation ability of occluded objects. Second, we construct a feature compression module (FCM) to reduce the loss of detailed information and condense more small-scale features in patch token downsampling of the Swin transformer, which improves the segmentation accuracy of small-scale ground objects. Finally, as a bridge between dual encoders, a relational aggregation module (RAM) is designed to integrate global dependencies from the Swin transformer into the features from CNN hierarchically. Our ST-UNet brings significant improvement on the ISPRS-Vaihingen and Potsdam datasets, respectively. The code will be available athttps://github.com/XinnHe/ST-UNet. Xin He 0024, Yong Zhou 0003, Jiaqi Zhao 0001, Di Zhang 0020, Rui Yao 0006, Yong Xue |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | CLT-Det: Correlation Learning Based on Transformer for Detecting Dense Objects in Remote Sensing ImagesabstractChallenges still exist in the task of object detection in remote sensing images with densely distributed objects due to large variation in scale and neglect of the relative position and correlation. To address these issues, a Correlation Learning Detector based on Transformer (CLT-Det) is proposed for detecting dense objects in remote sensing images. A Transformer Attention Module (TAM) is designed to improve the densely packed objects’ model representation ability by learning pixel-wise attention with Transformer. To alleviate the semantic gap caused by variations in scale, a Feature Refinement Module (FRM) is proposed by improving the multi-scale feature pyramid. A Correlation Transformer Module (CTM) is proposed to extract correlation information and encodes position information of dense objects’ features on the classification branch for fully utilizing the position information and correlation among objects. Extensive experiments compared with several state-of-art methods on two challenging remote sensing datasets, namely DOTA and HRSC2016, demonstrate that the proposed CLT-Det achieves promising and competitive performance. Yong Zhou 0003, Silin Chen, Jiaqi Zhao 0001, Rui Yao 0006, Yong Xue, Abdulmotaleb El Saddik |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Privacy-Preserving Scheme With Account-Mapping and Noise-Adding for Energy Trading Based on Consortium BlockchainabstractThe maturity in information technology and new energy technologies enables participants to generate, buy, and sell energy in energy trading systems. Although applying blockchain technology to energy trading has solved some drawbacks in traditional centralized energy systems, the openness and transparency characteristics make the trading records stored on the blockchain vulnerable to data-mining attacks that may cause indispensable privacy leakage. Due to high efficiency and low overhead, noise-addition is an appropriate solution for privacy preservation. Nonetheless, recent research on noise-addition needs to generate massive accounts, which brings a certain amount of waste and inconvenience for later regulation and management. To avoid the aforementioned issues, this paper proposes a consortium blockchain-enabled scheme to ensure the privacy of data stored on the blockchain and resist linking attacks initiated by data mining algorithms. Our scheme utilizes a dynamic partition algorithm to leverage an account mapping algorithm and a virtual token algorithm. Specifically, the account mapping algorithm utilizes a dynamic account allocation method to hide the trading distribution of active users. Furthermore, the virtual token algorithm applies Laplace noise to hide the actual energy consumption of inactive users and curb excessive accounts generation. Finally, we formally demonstrate the privacy and effectiveness of our proposed scheme in security analysis and experiment evaluations. Shunrong Jiang, Yiliang Liu, Tao Jiang 0017, Yong Zhou 0003 |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2022 | Clustering Matters: Sphere Feature for Fully Unsupervised Person Re-identificationabstractIn person re-identification (Re-ID) , the data annotation cost of supervised learning, is huge and it cannot adapt well to complex situations. Therefore, compared with supervised deep learning methods, unsupervised methods are more in line with actual needs. In unsupervised learning, a key to solving Re-ID is to find a standard that can effectively distinguish the difference (distance) between the features of images belonging to different pedestrian identities. However, there are some differences in the images captured by different cameras (such as brightness, angle, etc.). It is well known that the training of neural networks is mainly based on the distance between features, while in unsupervised learning, especially in unsupervised learning methods based on hierarchical clustering, the distance between features plays a more important role in the clustering phase. We improve the accuracy of a deep learning method based on hierarchical clustering under fully unsupervised conditions, starting from both feature and distance metrics. First, we propose to use spherical features, by normalizing the images in the feature space, to weaken the structural differences (length) between features, while saving the feature differences (direction) between different identities. Then, we use the sum of squared errors (SSE) as a regularization term to balance different cluster states. We evaluate our method on four large-scale Re-ID datasets, and experiments show that our method achieves better results than the state-of-the-art unsupervised methods. Yong Zhou 0003, Jiaqi Zhao 0001, Ying Chen 0005, Rui Yao 0006, Bing Liu 0016, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Facial action unit detection via hybrid relational reasoning
Zhiwen Shao, Yong Zhou 0003, Bing Liu 0016, Hancheng Zhu, Wen-Liang Du 0002, Jiaqi Zhao 0001 |
Vis. Comput. | 2 |
| 2021 | Path Planning based on Multi-objective Topological MapabstractRecently, intelligent robot technology has developed rapidly. As an important topic, path planning has attracted more and more attention. However, the performance of multi-objective path planning is limited by the scale of the problem, and the results of path planning are generally not optimal. In this work, based on 12 problems defined in multimodal multiobjective path planning optimization contest for CEC 2021 [1], we present a multi-objective path planning optimization algorithm, which includes data preprocessing, multi-objective genetic evolution path planning algorithm, and non-dominated sorting algorithm with elite strategy. This algorithm flow can realize multi-objective path planning and generate the optimal solution set (i.e. Pareto solution set). We implement our algorithm flow to calculate the time complexity and space complexity indicators. Experimental results on problems indicate the proposed multi-objective path planning algorithm can solve the optimal solution set in time. Due to the constraints of the problem, the number of optimal solutions is different for different problems. We show the validity of our method with experiments for path planning. Finally, we visually present some experimental results intuitively. The code is released at https://github.com/zhangruihao/pathPlanning. Jiaqi Zhao 0001, Zhijie Jia, Yong Zhou 0003, Zeming Xie, Di Zhang 0020 |
CEC | 3 |
| 2021 | Multi-Objective Net Architecture Pruning for Remote Sensing ClassificationabstractRemote sensing image scene classification has achieved significant breakthroughs in recent years. However, due to the high complexity and expensive computation most of CNNs used in the field of remote sensing imagery scene classification, it has become a challenging task for extracting effective features at restricted hardware conditions. To solve this problem, we present a model compression method by means of evolutionary algorithms. Specifically, we compress the model by pruning filters and transform the compression of the CNN model into a multi-objective optimization problem based on classification accuracy and compression ratio by using the adaptive-BN-based evaluation method. Furthermore, the prior knowledge of ResNet-50 on ImageNet is introduced to reduce the instability of evolutionary algorithm as a result of random population initialization. Experiments are implemented on three datasets with two evolutionary algorithms, and results demonstrate that our method can achieve state-of-the-art performances. Jiaqi Zhao 0001, Chengrun Yang, Yong Zhou 0003, Zhujun Jiang, Ying Chen 0005 |
IGARSS | 3 |
| 2021 | Joint Attention Mechanism for Unsupervised Video Object Segmentation
Rui Yao 0006, Xin Xu 0009, Yong Zhou 0003, Jiaqi Zhao 0001 |
PRCV (1) | 3 |
| 2021 | Point cloud classification by dynamic graph CNN with adaptive feature fusionabstractAbstract The deep neural network has made the most advanced breakthrough in almost all 2D image tasks, so we consider the application of deep learning in 3D images. Point cloud data, as the most basic and important form of representation of 3D images, can accurately and intuitively show the real world. The authors propose a new network based on feature fusion to improve the point cloud classification and segmentation tasks. Our network mainly consists of three parts: global feature extractor, local feature extractor and adaptive feature fusion module. A multi‐scale transformation network is devised to guarantee the invariance of the transformation of the global feature, and a residual block is introduced to alleviate the problem of gradient disappearance to enhance the global feature extractor. Based on the edge convolution and multi‐layer perceptron, a local feature extractor is constructed. Finally, an adaptive feature‐fusion module is proposed to complete the fusion of global features and local features. Extensive experiments on point cloud classification and segmentation tasks are carried out to verify the effectiveness of the proposed method. The classification accuracy of the ModelNet40 is 93.6%, which is 4.4% higher than that of the PointNet. Similarly, the segmentation accuracy on the ShapeNet is 85.6%, which is higher than other methods. Yong Zhou 0003, Jiaqi Zhao 0001, Yiyun Man, Minjie Liu, Rui Yao 0006, Bing Liu 0016 |
IET Comput. Vis. | 2 |
| 2021 | AMC-Net: Attentive modality-consistent network for visible-infrared person re-identification
Hanzheng Wang, Jiaqi Zhao 0001, Yong Zhou 0003, Rui Yao 0006, Ying Chen 0005, Silin Chen |
Neurocomputing | 3 |
| 2021 | Unsupervised cross-domain person re-identification with self-attention and joint-flexible optimization
Haopeng Hou, Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Ying Chen 0005, Abdulmotaleb El Saddik |
Image Vis. Comput. | 2 |
| 2021 | A siamese pedestrian alignment network for person re-identification
Yong Zhou 0003, Jiaqi Zhao 0001, Meng Jian, Rui Yao 0006, Bing Liu 0016, Ying Chen 0005 |
Multim. Tools Appl. | 2 |
| 2021 | Video-based person re-identification by semi-supervised adaptive stepwise learning
Yong Zhou 0003, Jiaqi Zhao 0001, Ying Chen 0005, Rui Yao 0006 |
Pattern Anal. Appl. | 2 |
| 2021 | Semi-supervised blockwisely architecture search for efficient lightweight generative adversarial network
Man Zhang 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Shixiong Xia, Zizheng Huang |
Pattern Recognit. | 2 |
| 2021 | Multi-Stage Fusion and Multi-Source Attention Network for Multi-Modal Remote Sensing Image SegmentationabstractWith the rapid development of sensor technology, lots of remote sensing data have been collected. It effectively obtains good semantic segmentation performance by extracting feature maps based on multi-modal remote sensing images since extra modal data provides more information. How to make full use of multi-model remote sensing data for semantic segmentation is challenging. Toward this end, we propose a new network called Multi-Stage Fusion and Multi-Source Attention Network ((MS) 2 -Net) for multi-modal remote sensing data segmentation. The multi-stage fusion module fuses complementary information after calibrating the deviation information by filtering the noise from the multi-modal data. Besides, similar feature points are aggregated by the proposed multi-source attention for enhancing the discriminability of features with different modalities. The proposed model is evaluated on publicly available multi-modal remote sensing data sets, and results demonstrate the effectiveness of the proposed method. Jiaqi Zhao 0001, Yong Zhou 0003, Boyu Shi, Jingsong Yang, Di Zhang 0020, Rui Yao 0006 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2020 | Secure and Privacy-preserving Energy Trading Scheme based on BlockchainabstractThe large-scale integration of distributed energy resources has resulted in surgical changes in energy trading systems. Traditional centralized trading systems suffer from high management cost and low efficiency. The recent advance of blockchain technology has enabled the invention of distributed energy trading systems, which can overcome the limitations of centralized trading systems. However, the distributed energy trading systems also bring new security and privacy challenges. For instance, transactions on blockchain are publicly visible which can lead to privacy leakage of trading information. Moreover, user privacy can also be leaked during verification of the aggregated energy trading result. In this paper, we propose a privacy-preserving energy trading scheme based on blockchain to meet the security requirements for distributed energy trading. We adopt a stealth transmission approach based on blockchain to ensure data privacy and break the linkage between consumers and providers in the energy trading process. We also use the non-interactive zero-knowledge proof technology to achieve privacy-preserving and trustworthy trading result verification. Security analysis and evaluation results have demonstrated that the proposed scheme can effectively protect the data privacy for distributed energy trading systems. Shunrong Jiang, Hao Yue 0001, Yong Zhou 0003 |
GLOBECOM | 5 |
| 2020 | Vehicular Edge Computing Meets Cache: An Access Control Scheme for Content DeliveryabstractVehicular Edge Computing (VEC) is an integration of Mobile Edge Computing with traditional vehicular networks, which aims to shift computing, communication, and storage resources to the edge of networks and is more close to vehicles. Due to the high mobility of vehicles, connection interruption and network changes may frequently occur. To tackle this issue, cache-based content delivery is regarded as a promising solution to achieve efficient data sharing in VEC. However, privacy-preserving access control and fair incentive distribution are rarely taken into account in prior VEC-oriented studies. In this paper, we propose an efficient and secure access control scheme for providing cache-based content delivery in VEC. Specifically, we construct two layer access control to enable flexible access control and fair incentive distribution with the assistance of edge nodes. Moreover, we construct a group signature-based scheme to achieve anonymous authentication and conditional revocation. The performance analysis shows that our secure scheme has an acceptable effect on network performance. Shunrong Jiang, Jianqing Liu, Longxia Huang, Haiqin Wu, Yong Zhou 0003 |
ICC | 5 |
| 2020 | Person image synthesis through siamese generative adversarial network
Ying Chen 0005, Shixiong Xia, Jiaqi Zhao 0001, Meng Jian, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Dongjun Zhu |
Neurocomputing | 5 |
| 2020 | Multiobjective ResNet pruning by means of EMOAs for remote sensing scene classification
Xuning Liu, Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Bing Liu 0016 |
Neurocomputing | 2 |
| 2020 | Diverse sample generation with multi-branch conditional generative adversarial network for remote sensing objects detection
Dongjun Zhu, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Meng Jian, Qiang Niu, Rui Yao 0006, Ying Chen 0005 |
Neurocomputing | 4 |
| 2020 | Appearance and shape based image synthesis by conditional variational generative adversarial network
Ying Chen 0005, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Dongjun Zhu |
Knowl. Based Syst. | 4 |
| 2020 | Remote sensing image captioning via Variational Autoencoder and Reinforcement Learning
Xiangqing Shen, Bing Liu 0016, Yong Zhou 0003, Jiaqi Zhao 0001 |
Knowl. Based Syst. | 3 |
| 2020 | Remote sensing image caption generation via transformer and reinforcement learning
Xiangqing Shen, Bing Liu 0016, Yong Zhou 0003, Jiaqi Zhao 0001 |
Multim. Tools Appl. | 3 |
| 2020 | Fusion based feature reinforcement component for remote sensing image object detection
Dongjun Zhu, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Ying Chen 0005 |
Multim. Tools Appl. | 4 |
| 2020 | Adaptively weighted learning for twin support vector machines via Bregman divergences
Zhizheng Liang, Lei Zhang 0029, Jin Liu 0006, Yong Zhou 0003 |
Neural Comput. Appl. | 4 |
| 2020 | Locality preserving difference component analysis based on the Lq norm
Zhizheng Liang, Xue-wen Chen 0001, Lei Zhang 0029, Jin Liu 0006, Yong Zhou 0003 |
Pattern Anal. Appl. | 5 |
| 2020 | Correlation classifiers based on data perturbation: New formulations and algorithms
Zhizheng Liang, Xue-wen Chen 0001, Lei Zhang 0029, Jin Liu 0006, Yong Zhou 0003 |
Pattern Recognit. | 5 |
| 2020 | GAN-based person search via deep complementary classifier with center-constrained Triplet loss
Rui Yao 0006, Cunyuan Gao, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Fuyuan Hu |
Pattern Recognit. | 5 |
| 2020 | Video Object Segmentation and Tracking: A SurveyabstractObject segmentation and object tracking are fundamental research areas in the computer vision community. These two topics are difficult to handle some common challenges, such as occlusion, deformation, motion blur, scale variation, and more. The former contains heterogeneous object, interacting object, edge ambiguity, and shape complexity; the latter suffers from difficulties in handling fast motion, out-of-view, and real-time processing. Combining the two problems of Video Object Segmentation and Tracking (VOST) can overcome their respective difficulties and improve their performance. VOST can be widely applied to many practical applications such as video summarization, high definition video compression, human computer interaction, and autonomous vehicles. This survey aims to provide a comprehensive review of the state-of-the-art VOST methods, classify these methods into different categories, and identify new trends. First, we broadly categorize VOST methods into Video Object Segmentation (VOS) and Segmentation-based Object Tracking (SOT). Each category is further classified into various types based on the segmentation and tracking mechanism. Moreover, we present some representative VOS and SOT methods of each time node. Second, we provide a detailed discussion and overview of the technical characteristics of the different methods. Third, we summarize the characteristics of the related video dataset and provide a variety of evaluation metrics. Finally, we point out a set of interesting future works and draw our own conclusions. Rui Yao 0006, Guosheng Lin, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2019 | An Efficient LightGBM Model to Predict Protein Self-interacting Using Chebyshev Moments and Bi-gram
Zhaohui Zhan, Zhu-Hong You, Yong Zhou 0003, Kai Zheng 0020, Zhengwei Li 0001 |
ICIC (2) | 3 |
| 2019 | iQIYI Celebrity Video Identification ChallengeabstractWe held the iQIYI Celebrity Video Identification Challenge in ACMMULTIMEDIA 2019. The purpose was to encourage the research on video-based person identification. We released the iQIYI-VID-2019 dataset, which contains 200K videos of 10K celebrities. In this paper, we introduce the organization of the challenge, the dataset, the evaluation process, and the results. Yuanliu Liu, Peipei Shi, Yong Zhou 0003, Jianbin Jiang, Yin Fan, Tingwei Gao, Ganwen Wang, Xiangju Lu, Danming Xie |
ACM Multimedia | 5 |
| 2019 | Lightweight Video Object Segmentation Based on ConvGRU
Rui Yao 0006, Yikun Zhang 0001, Cunyuan Gao, Yong Zhou 0003, Jiaqi Zhao 0001, Lina Liang |
PRCV (2) | 4 |
| 2019 | A Siamese Pedestrian Alignment Network for Person Re-identification
Yong Zhou 0003, Jiaqi Zhao 0001, Meng Jian, Rui Yao 0006, Bing Liu 0016, Xuning Liu |
PRCV (1) | 2 |
| 2019 | Further study on the maximum number of bent components of vectorial functions
Sihem Mesnager, Fengrong Zhang, Chunming Tang 0001, Yong Zhou 0003 |
Des. Codes Cryptogr. | 4 |
| 2019 | Structure-aware person search with self-attention and online instance aggregation matching
Cunyuan Gao, Rui Yao 0006, Jiaqi Zhao 0001, Yong Zhou 0003, Fuyuan Hu, Leida Li |
Neurocomputing | 4 |
| 2019 | Saliency detection via double nuclear norm maximization and ensemble manifold regularization
Bing Liu 0016, Yong Zhou 0003, Peng Liu 0013, Shuang Li 0011, Xinqiu Fang |
Knowl. Based Syst. | 2 |
| 2019 | Siamese Convolutional Neural Networks for Remote Sensing Scene ClassificationabstractThe convolutional neural networks (CNNs) have shown powerful feature representation capability, which provides novel avenues to improve scene classification of remote sensing imagery. Although we can acquire large collections of satellite images, the lack of rich label information is still a major concern in the remote sensing field. In addition, remote sensing data sets have their own limitations, such as the small scale of scene classes and lack of image diversity. To mitigate the impact of the existing problems, a Siamese CNN, which combines the identification and verification models of CNNs, is proposed in this letter. A metric learning regularization term is explicitly imposed on the features learned through CNNs, which enforce the Siamese networks to be more robust. We carried out experiments on three widely used remote sensing data sets for performance evaluation. Experimental results show that our proposed method outperforms the existing methods. Xuning Liu, Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Bing Liu 0016 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2018 | Efficient Framework for Predicting ncRNA-Protein Interactions Based on Sequence Information by Deep Learning
Zhaohui Zhan, Zhu-Hong You, Yong Zhou 0003, Liping Li 0003, Zhengwei Li 0001 |
ICIC (2) | 3 |
| 2018 | Multiobjective sparse ensemble learning by means of evolutionary algorithms
Jiaqi Zhao 0001, Licheng Jiao, Shixiong Xia, Vitor Basto-Fernandes, Iryna Yevseyeva, Yong Zhou 0003, Michael T. M. Emmerich |
Decis. Support Syst. | 6 |
| 2018 | Spectral regression based marginal Fisher analysis dimensionality reduction algorithm
Bing Liu 0016, Yong Zhou 0003, Zhanguo Xia, Peng Liu 0013, Qiuyan Yan |
Neurocomputing | 2 |
| 2018 | An improved efficient rotation forest algorithm to predict the interactions among proteins
Lei Wang 0121, Zhu-Hong You, Shixiong Xia, Xing Chen 0001, Yong Zhou 0003, Feng Liu 0039 |
Soft Comput. | 6 |
| 2017 | Computational Methods for the Prediction of Drug-Target Interactions from Drug Fingerprints and Protein Sequences by Stacked Auto-Encoder Deep Neural Network
Lei Wang 0121, Zhu-Hong You, Xing Chen 0001, Shixiong Xia, Feng Liu 0039, Yong Zhou 0003 |
ISBRA | 7 |
| 2017 | Video stitching based on iterative hashing and dynamic seam-line with local context
Rui Yao 0006, Jinliang Sun, Yong Zhou 0003, Dai Chen |
Multim. Tools Appl. | 3 |
| 2016 | Robust lifelong visual tracking using compact binary feature with color attributes
Rui Yao 0006, Shixiong Xia, Yong Zhou 0003, Qiang Niu |
Neurocomputing | 3 |
| 2016 | Manifold regularized extreme learning machine
Bing Liu 0016, Shixiong Xia, Yong Zhou 0003 |
Neural Comput. Appl. | 4 |
| 2016 | Exploiting Spatial Structure from Parts for Adaptive Kernelized Correlation Filter TrackerabstractDecomposing target into several parts may improve the capability of tracking algorithm to deal with appearance variations such as occlusion and deformation. In this letter, we propose a part-based appearance model by exploiting spatial structure from parts. The model minimizes appearance and deformation cost simultaneously to predict the new position of object. Then, the optimization problem is divided into two parts. Kernelized correlation filter (KCF) is used for tracking the appearance of parts separately to speed up the proposed tracker. Meanwhile, the deformation cost is minimized by structural learning schema, which can reduce the label noise that caused by inaccuracy bounding box. Finally, minimum spanning tree and dynamic programming are employed to combine the score map of the appearance and deformation of parts, and to detect best new position of target. Experimental results on several challenge sequences show the efficiency and effectiveness of the proposed tracking algorithm. Rui Yao 0006, Shixiong Xia, Fumin Shen, Yong Zhou 0003, Qiang Niu |
IEEE Signal Process. Lett. | 4 |
| 2015 | Extreme spectral regression for efficient regularized subspace learning
Bing Liu 0016, Shixiong Xia, Yong Zhou 0003 |
Neurocomputing | 4 |
| 2015 | Robust tracking via online Max-Margin structural learning with approximate sparse intersection kernel
Rui Yao 0006, Shixiong Xia, Yong Zhou 0003 |
Neurocomputing | 3 |
| 2015 | Semi-supervised extreme learning machine with manifold and pairwise constraints regularization
Yong Zhou 0003, Beizuo Liu, Shixiong Xia, Bing Liu 0016 |
Neurocomputing | 1 |
| 2013 | Normalized discriminant analysis for dimensionality reduction
Zhizheng Liang, Shixiong Xia, Yong Zhou 0003 |
Neurocomputing | 3 |
| 2013 | Unsupervised non-parametric kernel learning algorithm
Bing Liu 0016, Shixiong Xia, Yong Zhou 0003 |
Knowl. Based Syst. | 3 |
| 2013 | Training Lp norm multiple kernel learning in the primal
Zhizheng Liang, Shixiong Xia, Yong Zhou 0003, Lei Zhang 0029 |
Neural Networks | 3 |
| 2013 | Feature extraction based on Lp-norm generalized principal component analysis
Zhizheng Liang, Shixiong Xia, Yong Zhou 0003, Lei Zhang 0029, Youfu Li 0001 |
Pattern Recognit. Lett. | 3 |
| 2011 | Blockwise projection matrix versus blockwise data on undersampled problems: Analysis, comparison and applications
Zhizheng Liang, Shixiong Xia, Yong Zhou 0003, Youfu Li 0001 |
Pattern Recognit. | 3 |