Gensheng Pei

dblp:243/3679 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0002-7677-7487ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 16 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Beyond Quadratic: Linear-Time Change Detection with RWKV
abstract
Existing paradigms for remote sensing change detection are caught in a trade-off: CNNs excel at efficiency but lack global context, while Transformers capture long-range dependencies at a prohibitive computational cost. This paper introduces ChangeRWKV, a new architecture that reconciles this conflict. By building upon the Receptance Weighted Key Value (RWKV) framework, our ChangeRWKV uniquely combines the parallelizable training of Transformers with the linear-time inference of RNNs. Our approach core features two key innovations: a hierarchical RWKV encoder that builds multi-resolution feature representation, and a novel Spatial-Temporal Fusion Module (STFM) engineered to resolve spatial misalignments across scales while distilling fine-grained temporal discrepancies. ChangeRWKV not only achieves state-of-the-art performance on the LEVIR-CD benchmark, with an 85.46% IoU and 92.16% F1 score, but does so while drastically reducing parameters and FLOPs compared to previous leading methods. This work demonstrates a new, efficient, and powerful paradigm for operational-scale change detection.
Gensheng Pei, Tao Chen 0012, Xia Yuan, Haofeng Zhang 0001, Xiangbo Shu, Yazhou Yao
AAAI2
2026 CAFCL: Class-aware flow-based contrastive learning for out-of-distribution detection
Shiyu Peng, Jingzhu Li, Mingtao Wei, Gensheng Pei, Yazhou Yao
Pattern Recognit.5
2026 DepMatch: Boosting Semi-Supervised Semantic Segmentation by Exploring Depth Difference Knowledge
abstract
Existing semi-supervised semantic segmentation (SSS) methods fail to explore the potential of depth information in unlabeled data, as they suffer from 1) inter-class depth similarity, and 2) intra-class depth discrepancy. To address these challenges, this paper proposes DepMatch, a simple yet effective approach that leverages depth difference knowledge to guide consistency learning. Specifically, a Class-wise Depth Disparity Perception (CDDP) module is designed to exploit depth difference information, driven by class prediction priors, facilitating robust feature learning. Depth-feature discrepancy set is first constructed and then reliable pixel pairs are selected for inter-class depth disparity knowledge distillation. Simultaneously, exponential normalization is applied to intra-category depth disparity for suppressing large outlier variations, and an entropy-based adaptive weight is derived to prioritize feature learning of high entropy areas. Moreover, we propose the Uncertain Logit Disparity Regulation (ULDR) module, which leverages the depth variations at class boundaries to promote the mutual regulation of uncertain pixel logit information, enhancing the model's spatial understanding. Experiments on five public benchmarks show that DepMatch can be seamlessly incorporated as a plug-and-play plugin into popular SSS frameworks, achieving significant performance improvements across various visual encoders. The source code and models are made available at https://github.com/NUST-Machine-Intelligence-Laboratory/DepMatch.
Jianjian Yin, Xiruo Jiang, Tao Chen 0012, Gensheng Pei, Yazhou Yao, Fumin Shen, Heng Tao Shen
IEEE Trans. Image Process.4
2025 EEG-Based Seizure Detection and Type Classification with Structured State Space Modeling and Graph Neural Networks
abstract
Accurate seizure detection and type classification from Electroencephalography (EEG) are crucial for epilepsy diagnosis and clinical treatment. The intricate nature of seizure dynamics presents a significant challenge in effectively extracting distinguishing features from multivariate EEG signals, mainly due to long-range temporal dynamics and complex spatial dependencies between electrodes. To tackle these challenges, we introduce TS-S4GNN, a two-stream graph neural network (GNN) with structured state space modeling. Specifically, we combine channel-independent 1D-convolutional neural network with structured state space models to capture local and longrange temporal dependencies. The complex spatial dependencies between electrodes are learned with two GNN layers in parallel, one with static graph structure constructed according to the distance between electrodes, one with dynamically evolving graph structures learned from data. We validated TS-S4GNN on TUSZv1.5.2, the largest public EEG corpus with seizure type annotations. Experiments demonstrate that TS-S4GNN achieves an AUROC score of 0.913 in seizure detection and a weighted F1-score of 0.781 in seizure type classification, surpassing the current state-of-the-art models. Ablation studies validate the effectiveness of multi-scale temporal modeling and the twostream GNN approach. This study provides new insights into spatiotemporal modeling of multivariate signals and can be easily extended to other relevant learning tasks. The code is available at https://github.com/XploreAI-Lab/TS-S4GNN.
Xiaoya Fan, Pengyuan Ganzhang, Gensheng Pei, Shi-chun Bao
BIBM4
2025 Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection
abstract
The CLIP model has demonstrated significant advancements in aligning visual and language modalities through large-scale pre-training on image-text pairs, enabling strong zero-shot classification and retrieval capabilities on various domains. However, CLIP’s training remains computationally intensive, with high demands on both data processing and memory. To address these challenges, recent masking strategies have emerged, focusing on the selective removal of image patches to improve training efficiency. Although effective, these methods often compromise key semantic information, resulting in suboptimal alignment between visual features and text descriptions. In this work, we present a concise yet effective approach called Patch Generation-to-Selection (CLIP-PGS) to enhance CLIP’s training efficiency while preserving critical semantic con tent. Our method introduces a gradual masking process in which a small set of candidate patches is first pre-selected as potential mask regions. Then, we apply Sobel edge detection across the entire image to generate an edge mask that prioritizes the retention of the primary object areas. Finally, similarity scores between the candidate mask patches and their neighboring patches are computed, with optimal transport normalization refining the selection process to ensure a balanced similarity matrix. Our approach, CLIP-PGS, sets new state-of-the-art results in zero-shot classification and retrieval tasks, achieving superior performance in robustness evaluation and language compositionality benchmarks.
Gensheng Pei, Tao Chen 0012, Xinhao Cai, Xiangbo Shu, Tianfei Zhou, Yazhou Yao
CVPR1
2025 Cycle-Consistent Learning for Joint Layout-to-Image Generation and Object Detection
Xinhao Cai, Qiuxia Lai, Gensheng Pei, Xiangbo Shu, Yazhou Yao, Wenguan Wang
ICCV3
2025 PASS-SAM: Integration of Segment Anything Model for Large-Scale Unsupervised Semantic Segmentation
abstract
Large-scale unsupervised semantic segmentation (LUSS) is a sophisticated process that aims to segment similar areas within an image without relying on labeled training data. While existing methodologies have made substantial progress in this area, there is ample scope for enhancement. We thus introduce the PASS-SAM model, a comprehensive solution that amalgamates the benefits of various models to improve segmentation performance. Specifically, we enhance a baseline model utilizing self-attention and external attention modules. In the fine-tuning phase, we make use of conditional random fields (CRF) and the segment anything model (SAM) to refine and retrain the baseline model. During inferencing, we employ a model ensemble to blend predictions from different models, thereby enhancing segmentation accuracy. This approach secured first place in the LUSS track of the Third Jittor Artificial Intelligence Challenge. Our model, which makes use of the Jittor framework, is publicly available at https://github.com/PGSmall/jittor-PGSmall-LUSS.
Gensheng Pei, Qiong Wang 0003
Comput. Vis. Media3
2025 ChangeTitans: Toward Remote Sensing Change Detection With Neural Memory
abstract
Remote sensing change detection is essential for environmental monitoring, urban planning, and related applications. However, current methods often struggle to capture long-range dependencies while maintaining computational efficiency. Although Transformers can effectively model global context, their quadratic complexity poses scalability challenges, and existing linear attention approaches frequently fail to capture intricate spatiotemporal relationships. Drawing inspiration from the recent success of Titans in language tasks, we present ChangeTitans, the Titans-based framework for remote sensing change detection. Specifically, we propose VTitans, the first Titans-based vision backbone that integrates neural memory with segmented local attention, thereby capturing long-range dependencies while mitigating computational overhead. Next, we present a hierarchical VTitans-Adapter to refine multi-scale features across different network layers. Finally, we introduce TS-CBAM, a two-stream fusion module leveraging cross-temporal attention to suppress pseudo-changes and enhance detection accuracy. Experimental evaluations on four benchmark datasets (LEVIR-CD, WHU-CD, LEVIR-CD+, and SYSU-CD) demonstrate that ChangeTitans achieves state-of-the-art results, attaining 84.36% IoU and 91.52% F1-score on LEVIR-CD, while remaining computationally competitive. Our code and model are available at https://github.com/ChangeTitans/ChangeTitans.
Gensheng Pei, Yazhou Yao, Tianfei Zhou, Lizhong Ding 0003, Fumin Shen
IEEE Trans. Geosci. Remote. Sens.2
2025 Semi-Supervised Semantic Segmentation With Multi-Constraint Consistency Learning
abstract
Consistency regularization has prevailed in semi-supervised semantic segmentation and achieved promising performance. However, existing methods typically concentrate on enhancing the Image-augmentation based Prediction consistency and optimizing the segmentation network as a whole, resulting in insufficient utilization of potential supervisory information. In this paper, we propose a Multi-Constraint Consistency Learning (MCCL) approach to facilitate the staged enhancement of the encoder and decoder. Specifically, we first design a feature knowledge alignment (FKA) strategy to promote the feature consistency learning of the encoder from image-augmentation. Our FKA encourages the encoder to derive consistent features for strongly and weakly augmented views from the perspectives of point-to-point alignment and prototype-based intra-class compactness. Moreover, we propose a self-adaptive intervention (SAI) module to increase the discrepancy of aligned intermediate feature representations, promoting Feature-perturbation based Prediction consistency learning. Self-adaptive feature masking and noise injection are designed in an instance-specific manner to perturb the features for robust learning of the decoder. Experimental results on Pascal VOC2012 and Cityscapes datasets demonstrate that our proposed MCCL achieves new state-of-the-art performance. The source code and models are made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/MCCL.
Jianjian Yin, Tao Chen 0012, Gensheng Pei, Huafeng Liu 0004, Yazhou Yao, Liqiang Nie, Xian-Sheng Hua 0001
IEEE Trans. Multim.3
2024 VideoMAC: Video Masked Autoencoders Meet ConvNets
abstract
Recently, the advancement of self-supervised learning techniques, like masked autoencoders (MAE), has greatly influenced visual representation learning for images and videos. Nevertheless, it is worth noting that the predomi-nant approaches in existing masked image / video modeling rely excessively on resource-intensive vision transformers (ViTs) as the feature encoder. In this paper, we propose a new approach termed as VideoMAC, which combines video masked autoencoders with resource-friendly Con-vNets. Specifically, VideoMAC employs symmetric masking on randomly sampled pairs of video frames. To prevent the issue of mask pattern dissipation, we utilize ConvNets which are implemented with sparse convolutional operators as en-coders. Simultaneously, we present a simple yet effective masked video modeling (MVM) approach, a dual encoder architecture comprising an online encoder and an exponential moving average target encoder, aimed to facilitate inter-frame reconstruction consistency in videos. Additionally, we demonstrate that VideoMAC, empowering classical (ResNet) / modern (ConvNeXt) convolutional encoders to harness the benefits of MVM, outperforms ViT-based approaches on downstream tasks, including video object segmentation (+5.2% /6.4% J&F), body part propagation (+6.3% /3.1% mIoU), and human pose tracking (+10.2% / 11.1% [email protected]).
Gensheng Pei, Tao Chen 0012, Xiruo Jiang, Huafeng Liu 0004, Zeren Sun, Yazhou Yao
CVPR1
2024 Knowledge Transfer with Simulated Inter-image Erasing for Weakly Supervised Semantic Segmentation
Tao Chen 0012, Xiruo Jiang, Gensheng Pei, Zeren Sun, Yucheng Wang 0013, Yazhou Yao
ECCV (42)3
2024 Relating CNN-Transformer Fusion Network for Remote Sensing Change Detection
abstract
While deep learning, particularly convolutional neural networks (CNNs), has revolutionized remote sensing (RS) change detection (CD), existing approaches often miss crucial features due to neglecting global context and incomplete change learning. Additionally, transformer networks struggle with low-level details. RCTNet addresses these limitations by introducing (1) an early fusion backbone to exploit both spatial and temporal features early on, (2) a Cross-Stage Aggregation (CSA) module for enhanced temporal representation, (3) a Multi-Scale Feature Fusion (MSF) module for enriched feature extraction in the decoder, and (4) an Efficient Self-deciphering Attention (ESA) module utilizing transformers to capture global information and fine-grained details for accurate change detection. Extensive experiments demonstrate RCTNet’s clear superiority over traditional RS image CD methods, showing significant improvement and an optimal balance between accuracy and computational cost. Our source codes and pre-trained models are available at: https://github.com/NUST-Machine-Intelligence-Laboratory/RCTNet.
Yuhao Gao, Gensheng Pei, Mengmeng Sheng, Zeren Sun, Tao Chen 0012, Yazhou Yao
ICME2
2024 Universal Organizer of Segment Anything Model for Unsupervised Semantic Segmentation
abstract
Unsupervised semantic segmentation (USS) aims to achieve high-quality segmentation without manual pixel-level annotations. Existing USS models provide coarse category classifi-cation for regions, but the results often have blurry and imprecise edges. Recently, a robust framework called the segment anything model (SAM) has been proven to deliver precise boundary object masks. Therefore, this paper proposes a universal organizer based on SAM, termed as UO-SAM, to enhance the mask quality of USS models. Specifically, using only the original image and the masks generated by the USS model, we extract visual features to obtain positional prompts for target objects. Then, we activate a local region optimizer that performs segmentation using SAM on a per-object basis. Finally, we employ a global region optimizer to incorporate global image information and refine the masks to obtain the final fine-grained masks. Compared to existing methods, our UO-SAM achieves state-of-the-art performance. Our codes are available at https://github.com/NUST-Machine-Intelligence-Laboratory/UO-SAM.
Gensheng Pei, Xinhao Cai, Qiong Wang 0003, Huafeng Liu 0004, Yazhou Yao
ICME2
2024 Enhancing Robustness in Learning with Noisy Labels: An Asymmetric Co-Training Approach
abstract
Label noise, an inevitable issue in various real-world datasets, tends to impair the performance of deep neural networks. A large body of literature focuses on symmetric co-training, aiming to enhance model robustness by exploiting interactions between models with distinct capabilities. However, the symmetric training processes employed in existing methods often culminate in model consensus, diminishing their efficacy in handling noisy labels. To this end, we propose an Asymmetric Co-Training (ACT) method to mitigate the detrimental effects of label noise. Specifically, we introduce an asymmetric training framework in which one model (i.e., RTM) is robustly trained with a selected subset of clean samples while the other (i.e., NTM) is conventionally trained using the entire training set. We propose two novel criteria based on agreement and discrepancy between models, establishing asymmetric sample selection and mining. Moreover, a metric, derived from the divergence between models, is devised to quantify label memorization, guiding our method in determining the optimal stopping point for sample mining. Finally, we propose to dynamically re-weight identified clean samples according to their reliability inferred from historical information. We additionally employ consistency regularization to achieve further performance improvement. Extensive experimental results on synthetic and real-world datasets demonstrate the effectiveness and superiority of our method.
Mengmeng Sheng, Zeren Sun, Gensheng Pei, Tao Chen 0012, Haonan Luo 0002, Yazhou Yao
ACM Multimedia3
2023 Hierarchical Graph Pattern Understanding for Zero-Shot Video Object Segmentation
abstract
The optical flow guidance strategy is ideal for obtaining motion information of objects in the video. It is widely utilized in video segmentation tasks. However, existing optical flow-based methods have a significant dependency on optical flow, which results in poor performance when the optical flow estimation fails for a particular scene. The temporal consistency provided by the optical flow could be effectively supplemented by modeling in a structural form. This paper proposes a new hierarchical graph neural network (GNN) architecture, dubbed hierarchical graph pattern understanding (HGPU), for zero-shot video object segmentation (ZS-VOS). Inspired by the strong ability of GNNs in capturing structural relations, HGPU innovatively leverages motion cues (i.e., optical flow) to enhance the high-order representations from the neighbors of target frames. Specifically, a hierarchical graph pattern encoder with message aggregation is introduced to acquire different levels of motion and appearance features in a sequential manner. Furthermore, a decoder is designed for hierarchically parsing and understanding the transformed multi-modal contexts to achieve more accurate and robust results. HGPU achieves state-of-the-art performance on four publicly available benchmarks (DAVIS-16, YouTube-Objects, Long-Videos and DAVIS-17). Code and pre-trained model can be found at https://github.com/NUST-Machine-Intelligence-Laboratory/HGPU.
Gensheng Pei, Fumin Shen, Yazhou Yao, Tao Chen 0012, Xian-Sheng Hua 0001, Heng Tao Shen
IEEE Trans. Image Process.1
2023 Hierarchical Co-Attention Propagation Network for Zero-Shot Video Object Segmentation
abstract
Zero-shot video object segmentation (ZS-VOS) aims to segment foreground objects in a video sequence without prior knowledge of these objects. However, existing ZS-VOS methods often struggle to distinguish between foreground and background or to keep track of the foreground in complex scenarios. The common practice of introducing motion information, such as optical flow, can lead to overreliance on optical flow estimation. To address these challenges, we propose an encoder-decoder-based hierarchical co-attention propagation network (HCPN) capable of tracking and segmenting objects. Specifically, our model is built upon multiple collaborative evolutions of the parallel co-attention module (PCM) and the cross co-attention module (CCM). PCM captures common foreground regions among adjacent appearance and motion features, while CCM further exploits and fuses cross-modal motion features returned by PCM. Our method is progressively trained to achieve hierarchical spatio-temporal feature propagation across the entire video. Experimental results demonstrate that our HCPN outperforms all previous methods on public benchmarks, showcasing its effectiveness for ZS-VOS. Code and pre-trained model can be found at https://github.com/NUST-Machine-Intelligence-Laboratory/HCPN.
Gensheng Pei, Yazhou Yao, Fumin Shen, Xingguo Huang, Heng Tao Shen
IEEE Trans. Image Process.1
2022 Hierarchical Feature Alignment Network for Unsupervised Video Object Segmentation
Gensheng Pei, Fumin Shen, Yazhou Yao, Guosen Xie, Zhenmin Tang, Jinhui Tang 0001
ECCV (34)1
2022 A Novel Multi-Sample Data Augmentation Method for Oriented Object Detection in Remote Sensing Images
abstract
Data augmentation is widely used in computer vision tasks for enhancing the diversity of training data. However, due to sample redundancy and lack of object background, it is challenging to apply traditional data augmentation techniques to oriented object detection in remote sensing images. In this work, we propose SSMup (specifically synthetic mineral oversampling with mosaic and mixup), a multi-sample data augmentation method, to improve object detection performance in remote sensing images. Our method integrates Mosaic, Mixup, and SSMOTE to enable even distribution of target objects in augmented samples. Moreover, it equips the augmented samples with rich background information. Compared to existing state-of-the-art methods, our proposed method can remarkably improve the detection and generalization performance in remote sensing images. Comprehensive experiments are provided to demonstrate the effectiveness of the proposed method.
Guhua Chen, Gensheng Pei, Tao Chen 0012, Zhenmin Tang
MMSP2
2022 Feature Difference Enhancement Fusion for Remote Sensing Image Change Detection
Gensheng Pei, Tao Chen 0012, Yazhou Yao
PRCV (3)2
2022 Feature Hierarchical Differentiation for Remote Sensing Image Change Detection
abstract
Change detection (CD) is the localization of pixel-level differentiation between images in a specific setting,i.e., same-spatial different-temporal scenario. For high-resolution remote sensing (HRS) images, CD models should guarantee detection accuracy for the changes of interest and filter background noise for other regions. To this end, we propose a time-specific model, dubbed feature hierarchical differentiation (FHD), to achieve change perception aimed at HRS images. Specifically, we present the time-specific features (TSF) module to acquire each temporal image’s specific changes efficiently. Subsequently, the time-specific features from multi-temporal HRS images are adaptively fused by our proposed hierarchical differentiation (HD) module. Our FHD is subjected to elaborate experiments on four CD datasets. Quantitative and qualitative results outperform existing state-of-the-art methods. The ablation study further demonstrates the effectiveness of the proposed modules. Code is available at https://github.com/ZSVOS/FHD.
Gensheng Pei
IEEE Geosci. Remote. Sens. Lett.1
2021 Feature-label dual-mapping for missing label-specific features learning
Yusheng Cheng, Yibin Wang 0002, Gensheng Pei
Soft Comput.4
2019 Multi-label learning with kernel extreme learning machine autoencoder
Yusheng Cheng, Dawei Zhao 0002, Yibin Wang 0002, Gensheng Pei
Knowl. Based Syst.4