Chengjie Ge

dblp:275/1242 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning Robust Event-Guided Representations for Person Re-Identification
Chengzhi Cao, Xueyang Fu, Senyan Xu, Chengjie Ge, Zhengjun Zha
Int. J. Comput. Vis.4
2026 SkyFind: A Large-Scale Benchmark Unveiling Referring Expression Comprehension for UAV
abstract
Uncrewed aerial vehicles (UAV) are increasingly deployed to assist humans in diverse tasks, where understanding human intentions is critical to effective collaboration. Referring expression comprehension (REC) links language to visual targets, allowing UAV to recognize human-intended targets of interest, thereby supporting subsequent actions. However, existing REC research is almost exclusively confined to ground-based scenarios, leaving aerial scenarios largely unexplored. In this paper, we formally define UAV-based REC as a new research problem and highlight its unique challenges, including abundant background interference, small target size, and complex referring relations. To enable systematic study, we introduce SkyFind, a large-scale dataset with one million high-quality target-expression pairs, providing a solid foundation. In addition, we propose AerialREC, a baseline framework that reduces background interference in UAV imagery by searching for a potential target region before localization. We establish benchmark results on SkyFind using ten representative REC methods and validate the effectiveness of the AerialREC framework.
Guanbo Wu, Xueyang Fu, Kean Liu, Xin Lu 0008, Chengjie Ge, Wei Zhai, Zhengjun Zha
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 Toward Better De-Raining Generalization via Rainy Characteristics Memorization and Replay
abstract
Current image de-raining methods primarily learn from a limited dataset, leading to inadequate performance in varied real-world rainy conditions. To tackle this, we introduce a new framework that enables networks to progressively expand their de-raining knowledge base by tapping into a growing pool of datasets, significantly boosting their adaptability. Drawing inspiration from the human brain's ability to continually absorb and generalize from ongoing experiences, our approach borrows the mechanism of the complementary learning system. Specifically, we first deploy generative adversarial networks (GANs) to capture and retain the unique features of new data, mirroring the hippocampus's role in learning and memory. Then, the de-raining network is trained with both existing and GAN-synthesized data, mimicking the process of hippocampal replay and interleaved learning. Furthermore, we employ knowledge distillation with the replayed data to replicate the synergy between the neocortex's activity patterns triggered by hippocampal replays and the preexisting neocortical knowledge. This comprehensive framework empowers the de-raining network to accumulate knowledge from various datasets, continually enhancing its performance on previously unseen rainy scenes. Our testing on three benchmark de-raining networks confirms the framework's effectiveness. It not only facilitates continual knowledge accumulation across six datasets but also surpasses state-of-the-art methods in generalizing to new real-world scenarios. Our code is available at https://github.com/wangkunyu241/CLGID.
Xueyang Fu, Chengzhi Cao, Chengjie Ge, Wei Zhai, Zhengjun Zha
IEEE Trans. Neural Networks Learn. Syst.4
2025 EventMamba: Enhancing Spatio-Temporal Locality with State Space Models for Event-Based Video Reconstruction
abstract
Leveraging its robust linear global modeling capability, Mamba has notably excelled in computer vision. Despite its success, existing Mamba-based vision models have overlooked the nuances of event-driven tasks, especially in video reconstruction. Event-based video reconstruction (EBVR) demands spatial translation invariance and close attention to local event relationships in the spatio-temporal domain. Unfortunately, conventional Mamba algorithms apply static window partitions and standard reshape scanning methods, leading to significant losses in local connectivity. To overcome these limitations, we introduce EventMamba—a specialized model designed for EBVR task. EventMamba innovates by incorporating random window offset (RWO) in the spatial domain, moving away from the restrictive fixed partitioning. Additionally, it features a new consistent traversal serialization approach in the spatio-temporal domain, which maintains the proximity of adjacent events both spatially and temporally. These enhancements enable EventMamba to retain Mamba’s robust modeling capabilities while significantly preserving the spatio-temporal locality of event data. Comprehensive testing on multiple datasets shows that EventMamba markedly enhances video reconstruction, drastically improving computation speed while delivering superior visual quality compared to Transformer-based methods.
Chengjie Ge, Xueyang Fu, Peng He 0004, Chengzhi Cao, Zhengjun Zha
AAAI1
2025 Efficient Test-time Adaptive Object Detection via Sensitivity-Guided Pruning
abstract
Continual test-time adaptive object detection (CTTA-OD) aims to online adapt a source pre-trained detector to everchanging environments during inference under continuous domain shifts. Most existing CTTA-OD methods prioritize effectiveness while overlooking computational efficiency, which is crucial for resource-constrained scenarios. In this paper, we propose an efficient CTTA-OD method via pruning. Our motivation stems from the observation that not all learned source features are beneficial; certain domain-sensitive feature channels can adversely affect target domain performance. Inspired by this, we introduce a sensitivity-guided channel pruning strategy that quantifies each channel based on its sensitivity to domain discrepancies at both image and instance levels. We apply weighted sparsity regularization to selectively suppress and prune these sensitive channels, focusing adaptation efforts on invariant ones. Additionally, we introduce a stochastic channel reactivation mechanism to restore pruned channels, enabling recovery of potentially useful features and mitigating the risks of early pruning. Extensive experiments on three benchmarks show that our method achieves superior adaptation performance while reducing computational overhead by 12% in FLOPs compared to the recent SOTA method.
Xueyang Fu, Xin Lu 0008, Chengjie Ge, Chengzhi Cao, Wei Zhai, Zhengjun Zha
CVPR4
2025 Event Denoising Based on Iterative Tree-Structured Information Aggregation
abstract
Event cameras play a crucial role in the visual field; however, they are susceptible to noise. Traditional denoising algorithms for event cameras often struggle to balance accuracy and speed. To address this issue, this paper proposes an algorithm named Event Denoising Based on Iterative Tree-Structured Information Aggregation (EDIST). Specifically, the proposed method first establishes connections in event streams using a spatiotemporal window to extract Relation Tree. Then, a pruning algorithm is employed to streamline the subsequent information aggregation process, followed by a multi-stage convolution module designed to process the Relation Tree and obtain aggregated features. Finally, these aggregated features are fed into a classification module to determine whether an event is noise. During the inference stage, an information storage reuse module is designed to enable iterative execution, thereby enhancing inference speed. Experimental results demonstrate that the proposed algorithm outperforms existing methods in both denoising accuracy and inference speed on the DVSNOISE20 dataset.
Yueyang Xu, Chengjie Ge, Xueyang Fu, Zhengjun Zha
ICIP2
2025 PAID: Pairwise Angular-Invariant Decomposition for Continual Test-Time Adaptation
abstract
Continual Test-Time Adaptation (CTTA) aims to online adapt a pre-trained model to changing environments during inference. Most existing methods focus on exploiting target data, while overlooking another crucial source of information, the pre-trained weights, which encode underutilized domain-invariant priors. This paper takes the geometric attributes of pre-trained weights as a starting point, systematically analyzing three key components: magnitude, absolute angle, and pairwise angular structure. We find that the pairwise angular structure remains stable across diverse corrupted domains and encodes domain-invariant semantic information, suggesting it should be preserved during adaptation. Based on this insight, we propose PAID (Pairwise Angular Invariant Decomposition), a prior-driven CTTA method that decomposes weight into magnitude and direction, and introduces a learnable orthogonal matrix via Householder reflections to globally rotate direction while preserving the pairwise angular structure. During adaptation, only the magnitudes and the orthogonal matrices are updated. PAID achieves consistent improvements over recent SOTA methods on four widely used CTTA benchmarks, demonstrating that preserving pairwise angular structure offers a simple yet effective principle for CTTA. Our code is available at https://github.com/wangkunyu241/PAID.
Xueyang Fu, Yuanfei Bao, Chengjie Ge, Chengzhi Cao, Wei Zhai, Zhengjun Zha
NeurIPS4
2025 Event-Based Video Reconstruction With Deep Spatial-Frequency Unfolding Network
abstract
Current event-based video reconstruction methods, limited to the spatial domain, face challenges in decoupling brightness and structural information, leading to exposure distortion, and in efficiently acquiring non-local information without relying on computationally expensive Transformer models. To address these issues, we propose the Deep Spatial-Frequency Unfolding Reconstruction Network (DSFURNet), which explores and utilizes knowledge in the frequency domain for event-based video reconstruction. Specifically, we construct a variational model and propose three regularization terms: a brightness regularization term approximated by Fourier amplitudes, a structural regularization term approximated by Fourier phases, and an initialization regularization term that converts event representations into initial video frames. Then, we design corresponding spatial-frequency domain approximation operators for each regularization term. Benefiting from the global nature of computations in the frequency domain, the designed approximation operators can integrate local spatial and global frequency information at a lower computational cost. Furthermore, we combine the learned knowledge of the three regularization terms and unfold the optimization algorithm into an iterative deep network. Through this approach, the pixel-level initialization regularization constraint and the frequency domain brightness and structural regularization constraints can continuously play a role during the testing process, achieving a gradual improvement in the quality of the reconstructed video frames. Compared to existing methods, our network significantly reduces the number of network parameters while improving evaluation metrics.
Chengjie Ge, Xueyang Fu, Zhengjun Zha
IEEE Trans. Image Process.1
2025 Domain-Separated Bottleneck Attention Fusion Framework for Multimodal Emotion Recognition
abstract
As a focal point of research in various fields, human body language understanding has long been a subject of intense interest. Within this realm, the exploration of emotion recognition through the analysis of facial expressions, voice patterns, and physiological signals holds significant practical value. Compared with unimodal approaches, multimodal emotion recognition models leverage complementary information from vision, acoustic, and language modalities to robust perceive the human sentiment attitudes. However, the heterogeneity among modality signals leads to significant domain shifts, posing challenges for achieving balanced fusion. In this article, we propose a Domain-Separated Bottleneck Attention (DBA) Fusion Framework for human multimodal emotion recognition with lower computational complexity. Specifically, we partition each modality into two distinct domains: the invariant/private domain. The invariant domain contains crucial shared information, while the private domain aims to capture modality-specific representations. For the decomposed features, we introduce two sets of bottleneck cross-attention modules to effectively utilize the complementarity between domains to reduce redundant information. In each module, we interweave two Fusion Adapter blocks into the Self-Attention Transformer backbone. Each Fusion Adapter block integrates a small group of latent tokens as bridges for inter-modal and inter-domain interactions, mitigating the adverse effects of modality distribution differences and lowering computational costs. Extensive experimental results demonstrate that our method outperforms State-of-the-Art (SOTA) approaches across three widely used benchmark datasets.
Peng He 0004, Jun Yu 0001, Chengjie Ge, Lei Wang 0203, Zhen Kan
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Neuromorphic Event Signal-Driven Network for Video De-raining
abstract
Convolutional neural networks-based video de-raining methods commonly rely on dense intensity frames captured by CMOS sensors. However, the limited temporal resolution of these sensors hinders the capture of dynamic rainfall information, limiting further improvement in de-raining performance. This study aims to overcome this issue by incorporating the neuromorphic event signal into the video de-raining to enhance the dynamic information perception. Specifically, we first utilize the dynamic information from the event signal as prior knowledge, and integrate it into existing de-raining objectives to better constrain the solution space. We then design an optimization algorithm to solve the objective, and construct a de-raining network with CNNs as the backbone architecture using a modular strategy to mimic the optimization process. To further explore the temporal correlation of the event signal, we incorporate a spiking self-attention module into our network. By leveraging the low latency and high temporal resolution of the event signal, along with the spatial and temporal representation capabilities of convolutional and spiking neural networks, our model captures more accurate dynamic information and significantly improves de-raining performance. For example, our network achieves a 1.24dB improvement on the SynHeavy25 dataset compared to the previous state-of-the-art method, while utilizing only 39% of the parameters.
Chengjie Ge, Xueyang Fu, Peng He 0004, Chengzhi Cao, Zhengjun Zha
AAAI1
2024 Towards Generalized UAV Object Detection: A Novel Perspective from Frequency Domain Disentanglement
Xueyang Fu, Chengjie Ge, Chengzhi Cao, Zhengjun Zha
Int. J. Comput. Vis.3
2022 Learning Dual Convolutional Dictionaries for Image De-raining
abstract
Rain removal is a vital and highly ill-posed low-level vision task. While currently existing deep convolutional neural networks (CNNs) based image de-raining methods have achieved remarkable results, they still possess apparent shortcomings: First, most of the CNNs based models are lack of interpretability. Second, these models are not embedded with physical structures of rain streaks and background images. Third, they omit useful information in the background images. These deficiencies result in unsatisfied de-raining results in some sophisticated scenarios. To solve the above problems, we propose a Deep Dual Convolutional Dictionary Learning Network (DDCDNet) for these specific tasks. We firstly propose a new dual dictionary learning objective function, and then unfold it into the form of neural networks to learn prior knowledge from the data automatically. This network tries to learn the rain-streaks layer and the clean background using two dictionary learning networks instead of merely predicting the rain-streaks layer like most of the de-raining methods. To further increase the interpretability and generalization capability, we add sparsity and adaptive dictionary to our network to generate dynamic dictionary for each image based on content. Experimental results reveal that our model possesses outstanding de-raining ability on both synthetic and real-world data sets in terms of PSNR and SSIM as well as visual appearance.
Chengjie Ge, Xueyang Fu, Zhengjun Zha
ACM Multimedia1